Medical examination report interpretation method and device based on large language model

By performing optical character recognition and information extraction on medical examination reports, generating structured text, and using a large language model for reasoning, the problem of inaccurate interpretation of medical examination reports in existing technologies is solved, and efficient and accurate medical advice generation is achieved.

CN120636710APending Publication Date: 2025-09-12ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510726094.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately and efficiently interpret medical examination reports, resulting in poor information transmission during diagnosis and treatment.

Method used

A method based on a large language model is used to extract text from medical examination report images through optical character recognition technology, perform information extraction and structural processing, generate structured text, and use the large language model for reasoning to generate medical recommendations.

Benefits of technology

It improves the accuracy and efficiency of interpreting medical examination reports, ensures the reliability of medical advice, and supports better diagnosis and treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636710A_ABST
    Figure CN120636710A_ABST
Patent Text Reader

Abstract

One or more embodiments of the invention provide a medical examination report interpretation method and device based on a large language model, and the method comprises the steps: obtaining a to-be-interpreted medical examination report image, and carrying out the optical character recognition of the medical examination report image, identifying a medical examination report text from the medical examination report image; performing information extraction on the medical examination report text to extract a target text related to a medical examination result from the medical examination report text, and organizing the target text into a structured text according to a preset structured text format; and inputting the structured text into a large language model, and reasoning based on the structured text by the large language model to generate a medical suggestion text corresponding to the medical examination result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a method and device for interpreting medical examination reports based on a large language model. Background Art

[0002] In healthcare settings, medical examinations are scientific methods used to assess health status, diagnose diseases, or monitor treatment effectiveness. These tests primarily fall into five categories: laboratory tests, imaging tests, functional tests, pathological tests, and special tests. The specific tests used are tailored to the patient's condition and needs. Laboratory tests analyze body fluids or secretions to assess physiological or pathological conditions, including blood tests, urine tests, stool tests, and microbiological tests. Imaging tests use various techniques to observe the internal structure of the human body, including X-rays, ultrasounds, computed tomography (CT), magnetic resonance imaging (MRI), and nuclear medicine tests (e.g., PET-CT). Functional tests assess the physiological function of organs or systems, including electrocardiograms (ECGs), pulmonary function tests, electroencephalograms (EEGs), and endoscopic tests. Pathological tests use tissue or cell samples to determine the nature of a disease, including biopsies and cytology. Special tests are specialized assessments tailored to specific diseases or populations, including endoscopy, genetic testing, and allergen testing.

[0003] A medical examination report is a formal document compiled and issued by a medical institution or professional based on the results of various medical examinations conducted on a patient. This type of report typically records specific medical examination results, and some reports also provide possible medical conclusions. A medical examination report may include the patient's personal information (e.g., name, gender, age), examination items and purpose, examination methods and procedures, presentation of examination results (usually listing test values, images, or other forms of results, often with normal reference ranges for comparison), analysis, and conclusions.

[0004] Whether it is for diagnosing diseases, formulating treatment plans or monitoring health status, accurately interpreting medical examination reports is crucial. Accordingly, how to accurately interpret medical examination reports has become a highly concerned issue. Summary of the Invention

[0005] One or more embodiments of the present application provide the following technical solutions:

[0006] This application provides a method for interpreting medical examination reports based on a large language model, the method comprising:

[0007] Acquiring a medical examination report image to be interpreted, and performing optical character recognition on the medical examination report image to identify the medical examination report text from the medical examination report image;

[0008] Performing information extraction on the medical examination report text to extract target text related to the medical examination result from the medical examination report text, and organizing the target text into a structured text according to a preset structured text format;

[0009] The structured text is input into a large language model, and the large language model performs reasoning based on the structured text to generate a medical advice text corresponding to the medical examination result.

[0010] The present application also provides a medical examination report interpretation device based on a large language model, the device comprising:

[0011] an optical character recognition module, which obtains a medical examination report image to be interpreted and performs optical character recognition on the medical examination report image to identify the medical examination report text from the medical examination report image;

[0012] An examination result extraction module extracts information from the medical examination report text to extract target text related to the medical examination result from the medical examination report text, and organizes the target text into a structured text according to a preset structured text format;

[0013] The examination result interpretation module inputs the structured text into a large language model, and the large language model performs reasoning based on the structured text to generate a medical advice text corresponding to the medical examination result.

[0014] The present application also provides an electronic device, comprising:

[0015] processor;

[0016] a memory for storing processor-executable instructions;

[0017] The processor implements the steps of any of the above methods by running the executable instructions.

[0018] The present application also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of any of the methods described above.

[0019] In the above technical solution, optical character recognition can first be performed on the medical examination report image to be interpreted to identify the medical examination report text from the medical examination report image. Then, information extraction can be performed on the medical examination report text to extract the target text related to the medical examination results from the medical examination report text, and the extracted target text can be organized into structured text according to a pre-set structured text format. Finally, the structured text can be input into a large language model, and the large language model can perform inference based on the structured text to generate a medical advice text corresponding to the medical examination results contained in the structured text.

[0020] Using this approach, the medical examination report text obtained through optical character recognition of the medical examination report image can be organized into structured text with a specific format. This allows for unified storage and verification of the medical examination results, making them more organized and efficient. This improves the efficiency of subsequent interpretation of the medical examination report based on these results. Furthermore, using a large language model to generate medical recommendations corresponding to the medical examination results, enabling interpretation of the medical examination report, ensures accuracy and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The following is a description of the accompanying drawings required for describing the exemplary embodiments, in which:

[0022] Figure 1 It is a schematic diagram of a medical examination report shown in an exemplary embodiment of the present application.

[0023] Figure 2 It is a schematic diagram of another medical examination report shown in an exemplary embodiment of the present application.

[0024] Figure 3 It is a schematic diagram of a medical examination report interpretation system shown in an exemplary embodiment of the present application.

[0025] Figure 4 This is a flowchart of a method for interpreting medical examination reports based on a large language model, shown as an exemplary embodiment of the present application.

[0026] Figure 5 It is a structural diagram of a device shown in an exemplary embodiment of the present application.

[0027] Figure 6 This is a block diagram of a medical examination report interpretation device based on a large language model, shown as an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0028] The exemplary embodiments will be described in detail herein, with examples thereof being illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different drawings represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of the present application. Instead, they are merely examples consistent with some aspects of one or more embodiments of the present application.

[0029] It should be noted that, in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this application. In some other embodiments, the method may include more or fewer steps than those described in this application. In addition, a single step described in this application may be broken down into multiple steps for description in other embodiments; and multiple steps described in this application may be combined into a single step for description in other embodiments.

[0030] A medical examination report is a formal document compiled and issued by a medical institution or professional based on the results of various medical examinations conducted on a patient. This type of report typically records specific medical examination results, and some reports also provide possible medical conclusions. A medical examination report may include the patient's personal information (e.g., name, gender, age), examination items and purpose, examination methods and procedures, presentation of examination results (usually listing specific test values, images, or other forms of results, often with normal reference ranges for comparison), analysis, and conclusions.

[0031] The analyses and conclusions contained in medical examination reports are usually preliminary and simple. In order to better diagnose diseases, formulate treatment plans or monitor health status, it is necessary to interpret medical examination reports more accurately.

[0032] Medical test reports often contain numerous technical terms, abbreviations, and complex test results. This information can be difficult for non-professionals to understand. Correct interpretation can help patients understand the specific meaning of medical test results. When patients can understand their medical test reports, they can better communicate with their doctors, which is important for improving treatment satisfaction and compliance.

[0033] By interpreting medical test reports, a person's health status can be assessed, determining whether they have any illnesses or potential health risks. Accurately interpreting medical test reports helps doctors develop personalized treatment plans based on the patient's specific circumstances, achieving the best possible treatment outcomes. Sometimes, medical test results can also provide early warning of diseases or health issues that haven't yet presented symptoms, making it possible to take preventative measures, reduce the risk of illness, or initiate intervention at an early stage.

[0034] In this application, a large language model can be used as a tool for interpreting medical examination reports. Specifically, the medical examination report to be interpreted can be input into the large language model, and the large language model can perform reasoning based on the medical examination report to generate medical advice corresponding to the medical examination report. Among them, medical advice can include medical conclusions (for example, diagnosed diseases or health problems, etc.), feasible treatment plans (for example, drug selection, dosage and frequency of use, whether surgery is needed and its type, specific arrangements for rehabilitation training, etc.), recommended follow-up examinations, disease prevention measures (for example, lifestyle adjustment suggestions, etc.), etc.

[0035] Large language models are deep learning models trained using large amounts of text data. They can be used to generate natural language text or understand its meaning. Large language models can handle a variety of natural language tasks, such as text classification, named entity recognition (NER), question answering, and conversation, and are a key path to artificial intelligence.

[0036] In the field of natural language processing (NLP), large-scale text datasets are often referred to as corpuses. Corpuses can contain a variety of text data, such as literary works, academic papers, legal documents, news reports, everyday conversations, emails, and online forum posts. By learning from the text data in a corpus, large language models can acquire and understand the patterns and regularities of natural language, enabling effective processing and generation of human language.

[0037] Large language models typically use the Transformer architecture, meaning they are deep learning models based on the Transformer architecture. Transformer-based deep learning models are a type of neural network model that excels in fields like natural language processing.

[0038] Transformer is a neural network model used for sequence-to-sequence modeling. Transformer does not rely on recursive structures and can parallelize training and inference, speeding up model processing. In deep learning models based on the Transformer architecture, a multi-layer Transformer encoder is typically used to extract features from the input sequence, and a Transformer decoder is used to convert the extracted features into an output sequence. At the same time, such models typically also use a self-attention mechanism to capture long-distance dependencies in the input sequence, as well as residual connections and regularization methods to accelerate training and improve model performance.

[0039] A pretrained model is a large language model pretrained on large amounts of unlabeled text data. Pretrained models are general-purpose models; they are not designed or optimized for specific tasks. To adapt pretrained models to specific application scenarios and task requirements, they require fine-tuning to improve their performance on specific tasks. The large language model that is ultimately put into use is typically a pretrained model that has been further fine-tuned, performing supervised learning on labeled text data. Pretraining and fine-tuning are complementary processes: pretraining enables the model to acquire broad language understanding capabilities, while fine-tuning makes the model more specialized and accurate for specific tasks.

[0040] In other words, the training process of a large language model can be divided into two stages: pre-training and fine-tuning. During the pre-training stage, unsupervised learning (e.g., self-supervised learning) can be used on large-scale, unlabeled text datasets (e.g., online encyclopedias, online articles, books, etc.). Specifically, the model can predict missing parts or the next word based on the context, learn statistical laws such as semantics and syntax, and language structure. Backpropagation and optimization algorithms (e.g., gradient descent) are used to minimize prediction losses, iteratively update model parameters, and gradually improve the model's understanding of language. During the fine-tuning phase, you can select corresponding supervised learning tasks (for example, text classification, named entity recognition, question-answering systems, dialogue systems, etc.) based on the specific application scenarios and task requirements, and prepare task-specific text datasets. You can then use the pre-trained model as the starting point for fine-tuning, and use supervised learning to fine-tune on the task-specific text dataset. Specifically, you can perform the task based on the text dataset, and minimize the loss used to measure the performance of the model in processing specific tasks through backpropagation and optimization algorithms (for example, gradient descent). The model parameters are iteratively updated to gradually improve the model's performance on specific tasks. In practical applications, fine-tuning can flexibly choose supervised learning, unsupervised learning, or semi-supervised learning methods based on the specific application scenarios and the type of available data.

[0041] The language comprehension ability learned by the large language model during the pre-training and fine-tuning stages enables the large language model to perform logical inference, knowledge reasoning, or problem-solving by understanding, analyzing, and integrating text information when faced with complex problems or tasks. This ability is usually referred to as the reasoning ability of the large language model.

[0042] In practical applications, the pre-trained large language model is usually called the base model of the large language model, and the fine-tuned large language model is called the serving model of the large language.

[0043] Large language models are typically guided or stimulated by prompts (also known as prompts) to perform specific tasks. A prompt can be an initial text or text fragment provided to the large language model, such as a sentence, a question, or a conversation, intended to guide or stimulate the model to produce the corresponding output. Prompts are a key tool for guiding model output and can be very simple or quite complex, including instructions, examples, and descriptions of the desired output format. Prompts explicitly tell the large language model what task it is expected to perform, such as answering a question, simulating a conversation, writing an article, or translating text. Prompts also provide the large language model with necessary background information and context, enabling it to understand the logic, style, theme, or stance it should follow when generating content. Prompts can also inspire the large language model to demonstrate its inherent knowledge or specific language abilities, such as explaining complex concepts, citing regulations, or imitating the writing style of a specific author.

[0044] Since large language models are primarily used to understand and generate human language based on text processing, prompts usually appear in the form of text. However, in practical applications, large language models can also accept other forms of input as prompts, such as images, audio, and even video, provided that the large language model is designed or trained to process multimodal data.

[0045] However, in actual applications, on the one hand, the medical examination reports provided to patients are usually paper documents, or screenshots showing the electronic version of the medical examination reports. On the other hand, different medical examination reports usually contain complex and diverse table information. For example, in Figure 1 In the blood test report shown in the figure, each test item and its result form a table with a header, in which the header information includes four column names: "Item", "Result", "Reference Range", and "Unit". A row of information represents one test item and its result. Figure 2 The CT examination report shown contains a single-column table, in which each row of information represents the examination site and method (i.e., "examination items"), imaging images (i.e., "key images"), imaging manifestations (i.e., "radiological findings"), and imaging diagnostic opinions (i.e., "radiological opinions"). On the other hand, the medical examination reports provided to patients by different hospitals usually differ in format and wording. These situations will affect the model effect of the large language model when processing medical examination reports, and may even cause hallucination problems in the large language model. Among them, the hallucination problem refers to the content generated by the large language model looking very reasonable and coherent, and sometimes even able to imitate human emotions and thinking patterns, creating an illusion of "understanding" the input content, but in fact these contents are inaccurate or misleading.

[0046] In the above circumstances, it is expected that the medical examination reports can be interpreted more accurately and quickly, and the quality and efficiency of medical examination report resolution can be improved.

[0047] One or more embodiments of the present application provide a technical solution for interpreting medical examination reports based on a large language model. In this technical solution, optical character recognition can first be performed on the medical examination report image to be interpreted to identify the medical examination report text from the medical examination report image. Then, information extraction can be performed on the medical examination report text to extract target text related to the medical examination results from the medical examination report text. The extracted target text is organized into structured text according to a pre-set structured text format. Finally, the structured text can be input into a large language model, which performs inference based on the structured text to generate medical advice text corresponding to the medical examination results contained in the structured text.

[0048] Using this approach, the medical examination report text obtained through optical character recognition of the medical examination report image can be organized into structured text with a specific format. This allows for unified storage and verification of the medical examination results, making them more organized and efficient. This improves the efficiency of subsequent interpretation of the medical examination report based on these results. Furthermore, using a large language model to generate medical recommendations corresponding to the medical examination results, enabling interpretation of the medical examination report, ensures accuracy and efficiency.

[0049] Please refer to Figure 3 , Figure 3 It is a schematic diagram of a medical examination report interpretation system shown in an exemplary embodiment of the present application.

[0050] like Figure 3 As shown, the medical examination report interpretation system may include a server and at least one client accessing the server via any type of wired or wireless network.

[0051] Among them, the above-mentioned server can correspond to a server including an independent physical host, or a server cluster composed of multiple independent physical hosts; or, it can correspond to a virtual server, cloud server, etc. carried by a host cluster.

[0052] The above-mentioned client can correspond to terminal devices such as smart phones, tablet computers, laptops, desktop computers, PCs (Personal Computers), PDAs (Personal Digital Assistants), wearable devices (such as smart glasses, smart watches, etc.), smart car devices or game consoles.

[0053] Users can use the medical examination report interpretation service provided by the medical examination report interpretation system through the client; the client and the server can implement user-oriented medical examination report interpretation services through data interaction between each other.

[0054] Specifically, the server can be equipped with a large language model, and the medical examination report interpretation system can be based on the large language model. The large language model can perform reasoning based on the medical examination report to be interpreted to generate medical advice corresponding to the medical examination report.

[0055] For example, the client can output a corresponding user interface to the user, allowing the user to perform operations such as uploading an electronic version of the medical examination report or a medical examination report in image format, or inputting text as auxiliary information, in order to upload the medical examination report to be interpreted to the above-mentioned medical examination report interpretation system, and then use the medical examination report interpretation service provided by the medical examination report interpretation system. The client can send the medical examination report input by the user to the server, which will generate corresponding medical advice in text form based on the medical examination report and output the medical advice to the user, that is, return the medical advice to the client, and the client will display the medical advice to the user through the user interface for the user to review, thereby realizing a user-oriented medical examination report interpretation service.

[0056] It should be noted that the above-mentioned server can also be equipped with a medical knowledge base and an information retrieval component. Among them, the medical knowledge base is an external knowledge base relative to the large language model carried on the server, that is, the data in the medical knowledge base is not the knowledge acquired by the large language model through learning during the training process, but serves as auxiliary information in the reasoning process of the large language model, used to assist the large language model in generating answers corresponding to questions. During the reasoning process of the large language model, the information retrieval component can perform information retrieval in the medical knowledge base based on the query text (usually called Query or Question) as a prompt, so as to assist the large language model in generating an answer text corresponding to the query text through the retrieved relevant information.

[0057] The server can also be equipped with other functional components or subsystems such as a prompt generation component. These components or subsystems can work in conjunction with the large language model installed on the server to jointly generate the answer text corresponding to the query text as a prompt.

[0058] Furthermore, the medical examination report interpretation system may have only one large language model, which can be used to perform medical question-answering tasks. Alternatively, the medical examination report interpretation system may have multiple large language models. These large language models, in addition to those used for medical question-answering tasks, may also include large language models used for information extraction tasks, large language models used for structured text generation tasks, large language models used for knowledge base retrieval tasks, and so on.

[0059] Please refer to Figure 4 , Figure 4 This is a flowchart of a method for interpreting a medical examination report based on a large language model, as shown in an exemplary embodiment of the present application. The method for interpreting a medical examination report based on a large language model can be applied to Figure 3 The medical examination report interpretation system shown.

[0060] like Figure 4 As shown, the above-mentioned method for interpreting medical examination reports based on a large language model may include the following steps:

[0061] Step 402: Acquire a medical examination report image to be interpreted, and perform optical character recognition on the medical examination report image to identify the medical examination report text from the medical examination report image.

[0062] In this embodiment, to facilitate the interpretation of medical examination reports using the medical examination report interpretation system, users holding a paper medical examination report can scan or photograph the paper document to obtain an image of the medical examination report. Alternatively, a screenshot of the electronic version of the medical examination report held by the user can be directly used as the image of the medical examination report. The user can further upload the image of the medical examination report to be interpreted to the system, allowing the system to interpret the medical examination report based on the image of the medical examination report.

[0063] In actual applications, the above-mentioned medical examination report interpretation system can also obtain the medical examination report image to be interpreted through other means, and this application does not impose any special restrictions on this.

[0064] When a medical examination report image is obtained for interpretation, optical character recognition (OCR) can be performed on the image to identify text from the medical examination report image and use the identified text as the medical examination report text. Optical character recognition is a technology that can convert printed or handwritten text from an image into electronic text that can be processed by a computer. For images containing text, an optical character recognition algorithm can be used to analyze the shape of the text in the image, automatically identifying and converting the text in the image into electronic text that can be processed by a computer.

[0065] In some embodiments, since medical examination reports typically contain tabular information, to facilitate subsequent data processing, when identifying the medical examination report text from a medical examination report image to be interpreted through optical character recognition, optical character recognition can be performed on the medical examination report image to identify the text displayed in the medical examination report image, and then the text identified from the medical examination report image can be further processed into text in a structured text format as the medical examination report text. For example, the text identified from the medical examination report can be further organized into a key-value pair set (which can be referred to as a second-type key-value pair set), and the second-type key-value pair set can be used as the medical examination report text.

[0066] It should be noted that the above second-category key-value pair set includes at least one second-category key-value pair. For any second-category key-value pair in the second-category key-value pair set, the key in the second-category key-value pair may be a medical examination report item category, and the value may be a medical examination report item.

[0067] In practical applications, the text identified from the medical examination report can be further processed into text in JSON (JavaScript Object Notation) format. JSON is a structured text format designed for storing and exchanging data. JSON organizes data using key-value pairs and supports nested structures.

[0068] As Figure 2 Taking the CT examination report image shown as an example, by performing optical character recognition on the CT examination report image, the text recognized from the CT examination report image can be as follows (content in double quotes):

[0069] "Name:

[0070] Gender: Female

[0071] Age: 60

[0072] Radiology Examination Number:

[0073] Physical examination number:

[0074] Inspection time: 2024-05-10

[0075] Examination items: Low-dose spiral CT of the chest

[0076] Key images:

[0077] Radiological findings:

[0078] The thorax was symmetrical, with the trachea centered. Lung windows revealed increased and thickened markings in both lungs. A mass was observed in the right upper lobe, approximately 40 x 35 mm in size and shallowly lobed. Multiple, scattered, rounded nodules were found in both lungs, ranging in diameter from 7 to 10 mm with well-defined borders and uniform density. The hilum of the lungs was normal in size and morphology. A mediastinal window revealed normal cardiac silhouettes and major vessels. No masses or significantly enlarged lymph nodes were observed in the mediastinum. There was no pleural effusion or pleural thickening.

[0079] Radiological opinion: (for reference by clinicians only)

[0080] Right lung mass

[0081] Multiple nodules in both lungs suggest metastasis; please combine clinical findings with further examination to identify the primary tumor.

[0082] Assuming that the key-value pair is represented in the form of (key, value) and the text content is represented in the form of "text", the second type of key-value pair set formed by further organizing the text recognized from the above CT examination report image can be shown as follows:

[0083] {("Name","None"),("Gender","Female"),("Age","60 years old"),("Radiology examination number","None"),("Physical examination number","None"),("Examination time","2024-05-10"),("Examination items","Chest low-dose spiral CT"),("Key images","None"),("Radiology findings","Both sides of the chest are symmetrical, and the trachea is centered. The lung window shows increased and thickened markings on both lungs, and the right upper lobe is seen The mass, approximately 40 x 35 mm, is shallowly lobed. Multiple, scattered, quasi-round nodules are present in both lungs, ranging in diameter from 7 to 10 mm, with clear borders and uniform density. The size and morphology of the bilateral hilums are normal. A mediastinal window reveals normal cardiac shadows and major blood vessels. No masses or significantly enlarged lymph nodes are observed in the mediastinum. There is no pleural effusion or pleural thickening. "Radiological opinion," "Right lung mass. Multiple nodules in both lungs suggest metastasis. Please combine clinical findings with further examination to confirm the primary tumor."

[0084] Step 404: extract information from the medical examination report text to extract target text related to the medical examination results from the medical examination report text, and organize the target text into structured text according to a preset structured text format.

[0085] In this embodiment, in order to improve the efficiency of interpreting medical examination reports, after obtaining the above-mentioned medical examination report text, information extraction (IE) can be performed on the medical examination report text to extract text related to the medical examination results (which can be called target text) from the medical examination report text. Among them, information extraction is a natural language processing technology that is generally intended to automatically extract structured and meaningful information from unstructured or semi-structured text data. Information extraction can identify and extract specific types of facts, relationships or data items and convert them into a structured format that is easier to analyze and apply. The specific tasks of information extraction can generally include named entity recognition (NER), relation extraction, event extraction, attribute extraction, etc., and in actual applications, the specific tasks of the information extraction performed can be set according to actual needs and circumstances. For example, in this application, the specific task of information extraction can be to extract text related to the medical examination results from the medical examination report text obtained by optical character recognition.

[0086] Once the target text related to the medical examination results has been extracted from the medical examination report, the target text can be further organized into formatted text according to a pre-set structured text format. This allows the medical examination results to be stored and verified in a unified structured format, facilitating subsequent interpretation based on the medical examination results.

[0087] In some embodiments, a large language model can be used to accurately and efficiently extract information. Specifically, the medical examination report text can be input into the large language model, which then performs information extraction on the medical examination report text, thereby extracting text related to the medical examination results from the medical examination report text.

[0088] In practical applications, a prompt text can be constructed based on the medical examination report text to guide the large language model in performing the information extraction task, and the constructed prompt text can be input into the large language model. In this case, the large language model can then perform information extraction on the medical examination report text under the guidance of the prompt text, thereby extracting text related to the medical examination results from the medical examination report text.

[0089] The aforementioned large language model can refer to the serving model of the large language model. In practical applications, the constructed large language model can be pre-trained using unsupervised learning on a large, unlabeled text dataset to obtain the base model of the large language model. Furthermore, the information extraction task can be used as a supervised learning task during fine-tuning, and a text dataset specific to the information extraction task can be prepared. The base model of the large language model can then be used as the starting point for fine-tuning, using supervised learning to fine-tune the data on the text dataset specific to the information extraction task to obtain the serving model of the large language model.

[0090] When constructing a text dataset specific to the information extraction task, we can first define the type of information to be extracted (for example, disease, description, location, part, quantitative description, grade, suggestion, confidence type, etc.), and then annotate the collected medical examination report texts (these medical examination report texts are training samples) according to the defined information types. We can also annotate the text representing specific information in each medical examination report text (i.e., the text in a medical examination report), and annotate each text representing specific information with the corresponding information type (the annotated information type is the label of the training sample). In this way, the annotated medical examination report texts can be used as a text dataset specific to the information extraction task, and used for supervised training of the above-mentioned large language model.

[0091] Accordingly, the output of the above-mentioned large language model may include not only text related to the medical examination results extracted from the above-mentioned medical examination report text, but also information types corresponding to the extracted text, so that the extracted text can be further organized into structured text accurately and efficiently according to the corresponding information type.

[0092] It should be noted that the output of the large language model may be text related to the medical examination results extracted from the medical examination report text. Subsequently, the structured text generation component in the medical examination report interpretation system may further organize the text output by the large language model into structured text according to the structured text format. Alternatively, the large language model may be fine-tuned so that it can perform the structured text generation task. Thus, after extracting the text related to the medical examination results from the medical examination report text, the large language model may continue to organize the extracted text into structured text according to the structured text format.

[0093] In some embodiments, to facilitate subsequent data processing, the structured text format described above can be used to indicate that text related to medical examination results is organized into key-value pairs (referred to as first-class key-value pairs). The key in the first-class key-value pair can be the medical examination result category, and the value can be the medical examination result itself.

[0094] For example, the above structured text format can be represented by the following text (content within double quotes):

[0095] "Disease: #Disease name or conclusion

[0096] Description: #Description related to the disease being analyzed

[0097] Location: #Description of the location related to the disease, such as left, right, bilateral, left lobe, right lobe, etc.

[0098] Part: #Organ or part related to the disease

[0099] Quantitative description: #such as diameter, range, size, long diameter, etc.

[0100] Level: #Disease corresponding status, such as kidney tumor Bosniak, prostate PI-RAD, breast BI-RAD, etc.

[0101] Recommendations: Disease-related recommendations in the examination conclusions, such as follow-up visits and recommended follow-up examinations

[0102] Confidence type: #Suggestive, Suspicious, Seen, Normal, Postoperative Change, Post-treatment Change, Recheck".

[0103] Assuming that the key-value pairs are represented in the form of (key, value) and the text content is represented in the form of "text", the above structured text format indicates that the text related to the medical examination results is organized into 8 first-class key-value pairs, namely ("disease", value 1), ("description", value 2), ("location", value 3), ("part", value 4), ("quantitative description", value 5), ("grade", value 6), ("recommendation", value 7), ("confidence type", value 8). The first-class key-value pair set composed of these 8 first-class key-value pairs is {("disease", value 1), ("description", value 2), ("location", value 3), ("part", value 4), ("quantitative description", value 5), ("grade", value 6), ("recommendation", value 7), ("confidence type", value 8)}. This first-class key-value pair set can be regarded as a structured text organized by the text related to the medical examination results. Among them, the key "disease" in the first-class key-value pair ("disease", value 1) can be used as the primary key of the structured text.

[0104] Assume that the text of a medical examination report obtained through optical character recognition is as follows (the content within double quotes):

[0105] "Painless Electronic Gastroscopy Report

[0106] Check the image:

[0107] Esophageal cardia, gastric fundus, and gastric body

[0108] Antrum and pyloric angle Duodenal bulb Descending duodenum

[0109] Diagnosis Description:

[0110] Insertion status: Smooth

[0111] Delivery site: descending duodenum

[0112] Esophagus: The mucosa is smooth, the vascular network is clear, and the dilation is good.

[0113] Cardia: Opens and closes well, with clear dentate line.

[0114] Gastric fundus: A 0.5*0.6cm submucosal bulge can be seen with a smooth surface, unclear mucus lake and a large amount of secretions.

[0115] Gastric body: smooth mucosa.

[0116] Gastric antrum: The mucosa is congested and edematous, with scattered erosion foci and frequent peristalsis.

[0117] Pyloric bulb: smooth mucosa.

[0118] Gastric angle: arc-shaped, with smooth mucosa.

[0119] Duodenal bulb: No abnormalities were found.

[0120] Descending duodenum: no abnormalities were found.

[0121] Microscopic diagnosis:

[0122] 1. Chronic non-atrophic gastritis with erosion (grade II)

[0123] 2. The nature of the submucosal bulge in the gastric fundus remains to be determined.

[0124] This report is for clinical reference only! ”

[0125] The structured text organized by the text related to the medical examination results extracted from the medical examination report text can include two first-class key-value pair sets; the two first-class key-value pair sets are: {("disease", "chronic non-atrophic gastritis with erosion"), ("description", "gastric antrum: mucosal congestion and edema, with scattered erosive foci and frequent peristalsis"), ("location", "none"), ("site", "stomach"), ("quantitative description", "none"), ("grade", "grade II"), ("construction "Discussion", "None"), ("Confidence Type", "Hint")}, {("Disease", "Submucosal protrusion of gastric fundus"), ("Description", "Gastric fundus: A 0.5*0.6cm submucosal protrusion can be seen, with a smooth surface, unclear mucus lake, and a large amount of secretions"), ("Location", "None"), ("Location", "Stomach"), ("Quantitative description", "0.5*0.6cm"), ("Grade", "None"), ("Suggestion", "Nature to be investigated"), ("Confidence Type", "Suspicious")}.

[0126] For another example, the above structured text format can be represented by the following text (content within double quotes):

[0127] "Project: #Name of the medical examination project

[0128] Results: #Specific results of medical examination items

[0129] Reference interval: #Normal reference interval for medical examination items

[0130] Unit: #The unit used for the results of medical examination items.

[0131] Assuming that the key-value pairs are represented in the form of (key, value) and the text content is represented in the form of "text", the above structured text format indicates that the text related to the result of a medical examination item is organized into four first-class key-value pairs, namely ("item", value 1), ("result", value 2), ("reference interval", value 3), and ("unit", value 4). The first-class key-value pair set composed of these four first-class key-value pairs is {("item", value 1), ("result", value 2), ("reference interval", value 3), ("unit", value 4)}. This first-class key-value pair set can be regarded as a structured text organized by text related to the result of a medical examination item, and the structured text organized by text related to the results of each medical examination item can be regarded as a structured text organized by text related to the medical examination results.

[0132] As Figure 1 Taking the blood routine examination report shown in FIG. 1 as an example, since the blood routine examination report contains the results of 21 medical examination items in total, the structured text organized by the text related to the medical examination results in the blood routine examination report can include 21 first-class key-value pair sets; these 21 first-class key-value pair sets are: {("item", "white blood cell count"), ("result", "11.4"), ("reference interval", "3.5-9.5"), ("unit", "10 9 / L”)}, {(“Item”, “Neutrophil Percentage”), (“Result”, “28.0”), (“Reference Interval”, “40.0-70.0”), (“Unit”, “%”)}, and so on.

[0133] In practical applications, schemas can be used to format structured text. Schemas primarily define the structure, type, and constraints of data. They describe how data should be organized, including field names, data types (such as string, integer, Boolean, etc.), required or not, default values, and possible value ranges. The primary purpose of a schema is to ensure data consistency and validity, and it is often used to verify that data conforms to specific structural requirements. For example, JSONSchema is used to validate JSON documents.

[0134] For text, a text schema typically refers to a system of rules that define and constrain the structure, format, or semantics of a piece of text. It can be used to describe the meaning, data type, organization, and validation rules of fields within a text, aiming to ensure a consistent understanding and parsing of text across different systems. Therefore, by writing a schema, you can configure the format of this structured text. For example, a schema can represent the structured text shown in the example above.

[0135] Step 406: Input the structured text into a large language model, and the large language model performs reasoning based on the structured text to generate a medical advice text corresponding to the medical examination result.

[0136] In this embodiment, once the structured text organized from the target text is obtained, the structured text can be input into a large language model. Since the structured text is organized from the target text related to the medical examination results, the large language model can perform reasoning based on the structured text to generate medical advice text corresponding to the medical examination results, thereby achieving accurate and efficient interpretation of medical examination reports.

[0137] Furthermore, in order to enable the user who provides the medical examination report to obtain and view the specific content of the interpretation of the medical examination report, the medical advice text output by the large language model can also be output to the user who inputs the above-mentioned medical examination report image.

[0138] In practical applications, a query text (the query text is referred to as a prompt) can be constructed based on the structured text composed of the target text to guide the large language model in performing a medical question-answering task. This query text can then be input into the large language model. In this case, the large language model can perform reasoning based on the query text to answer the question posed by the query text, thereby generating a corresponding answer text. In this case, the answer text is the medical advice text corresponding to the medical examination results contained in the structured text.

[0139] The aforementioned large language model can refer to the serving model of the large language model. In practical applications, the constructed large language model can be pre-trained using unsupervised learning on a large, unlabeled text dataset to obtain the base model of the large language model. Furthermore, medical question-answering tasks can be used as supervised learning tasks for fine-tuning, and a text dataset specific to medical question-answering tasks can be prepared. The base model of the large language model can then be used as the starting point for fine-tuning, using supervised learning to fine-tune the model on a text dataset specific to medical question-answering tasks to obtain the serving model of the large language model.

[0140] When constructing a text dataset specific to medical question-answering tasks, we can annotate the collected prompt texts based on structured text (these structured texts are known as training samples), labeling each prompt text with the medical advice text corresponding to the structured text within it (the annotated medical advice text is known as the label of the training sample). This annotated prompt text can then serve as a text dataset specific to medical question-answering tasks, used for supervised training of the aforementioned large language model.

[0141] In some embodiments, in order to facilitate the user to compare the medical examination report and its medical advice, so that the user can better understand the corresponding medical advice, when the above-mentioned medical advice text is output to the above-mentioned user, the above-mentioned structured text organized by the above-mentioned target text can be output to the user together with the medical advice text.

[0142] In some embodiments, to further improve the accuracy of medical examination report interpretation, before inputting the structured text organized from the target text into the large language model to generate the corresponding medical advice text using the large language model, medical abnormality identification can be performed on the medical examination results corresponding to the first set of key-value pairs based on pre-set medical abnormality identification rules to obtain a medical abnormality identification result corresponding to the medical examination result. Subsequently, the medical abnormality identification result can also be organized into key-value pairs and added to the first set of key-value pairs, so that the medical abnormality identification result can serve as auxiliary information when the large language model generates the medical advice text corresponding to the medical examination result.

[0143] For example, in medical examination results, for each numerical result (for example, the result of white blood cell count, the result of neutrophil percentage, etc.), a corresponding reference interval is usually provided; taking one of the results as an example, if the result is within the reference interval, the medical abnormality identification result corresponding to the result can be determined as normal, if the result is higher than the reference interval, the medical abnormality identification result corresponding to the result can be determined as high, if the result is lower than the reference interval, the medical abnormality identification result corresponding to the result can be determined as low. For each result indicating negative / positive, if the result is negative, the medical abnormality identification result corresponding to the result can be determined as normal, if the result is positive, the medical abnormality identification result corresponding to the result can be determined as positive. For results that do not have a reference interval or contain multiple reference intervals and cannot be judged whether they are abnormal, the medical abnormality identification result corresponding to the result can be determined as undeterminable or undeterminable under multiple conditions.

[0144] As Figure 1 Taking the blood test report shown in the figure as an example, since the result of the white blood cell count is higher than the reference interval, the first type of key-value pair corresponding to the result of the white blood cell count can be updated to {("item","white blood cell count"), ("result","11.4"), ("reference interval","3.5-9.5"), ("unit","10 9 / L”), (“abnormal type”, “high”)}; since the result of neutrophil percentage is lower than the reference interval, the first type of key-value pair corresponding to the result of neutrophil percentage can be updated to {(“item”, “neutrophil percentage”), (“result”, “28.0”), (“reference interval”, “40.0-70.0”), (“unit”, “%”), (“abnormal type”, “low”)}; and so on.

[0145] In some embodiments, the medical examination report interpretation system primarily relies on the knowledge acquired by its large language model during training by learning from static corpus when interpreting medical examination reports. However, due to the limitations of this knowledge, the system may experience hallucinations when answering complex or specific questions. This reliance on static corpus limits the system's adaptability and response accuracy.

[0146] In order to improve the adaptability and response accuracy of the above-mentioned medical examination report interpretation system, the RAG (Retrieval-Augmented Generation) method can be adopted to combine information retrieval and model generation, so that when answering questions raised by users, the system no longer relies solely on the knowledge acquired by the large language model through learning static corpus during the training process, but can first perform information retrieval in the external knowledge base based on the question, and then understand and answer the question based on the retrieved relevant documents, and generate the corresponding answer. That is, the external knowledge base can be combined with the large language model, and during the model generation process, relevant information can be retrieved from the external knowledge base in real time to assist the large language model in making more accurate and comprehensive answers or decisions. Since the context of the retrieved information and the question is taken into account during the model generation process, it can be ensured that the generated content not only meets actual needs, but is also accurate, reliable, coherent, and natural.

[0147] Based on the above, before inputting the structured text into the large language model to generate the corresponding medical advice text using the large language model, the correlation between the text contained in the structured text and each medical knowledge text in the medical knowledge base can be calculated. Based on the calculated correlation, medical knowledge text related to the structured text can be retrieved from the medical knowledge base. The text contained in the structured text used for retrieval in the medical knowledge base can specifically be text such as examination item names and predicted disease names extracted from the text contained in the structured text. The retrieved medical knowledge text related to the examination item names and predicted disease names can be used to assist in analyzing the medical examination results contained in the structured text. For example, assuming the retrieved medical knowledge text is "When the result of examination item X is high, disease Y may be present," then if the result of examination item X contained in the medical examination result is high, it can be inferred that disease Y may be present.

[0148] In practical applications, a medical knowledge text in the above-mentioned medical knowledge base can be a sentence, a paragraph or a document, and this application does not impose any special restrictions on this. The medical knowledge base can be a document database for directly storing various medical knowledge texts. Alternatively, the medical knowledge base can also be a graph database, in which a node in the graph database can represent keywords such as the name of an examination item, the name of a predicted disease, etc. extracted from a medical knowledge text, and the medical knowledge text itself can be used as an attribute of the node. The specific form of the medical database can be selected according to the retrieval requirements and actual conditions, and this application does not impose any special restrictions on this.

[0149] When the above-mentioned medical knowledge text related to the above-mentioned structured text is obtained, the structured text and the related medical knowledge text can be input into the above-mentioned large language model, and the large language model performs reasoning based on the structured text and the related medical knowledge text to generate the above-mentioned medical advice text corresponding to the above-mentioned medical examination results contained in the structured text with the assistance of the relevant medical knowledge text.

[0150] It should be noted that the above-mentioned medical knowledge base can be dynamically updated, that is, the medical knowledge texts therein can be updated with the latest medical progress.

[0151] In some embodiments, in order to reduce the complexity of calculating the association between texts and improve efficiency, the association between texts can be converted into calculating the similarity between embedding vectors corresponding to the texts. Among them, embedding processing is used to map high-dimensional sparse data to low-dimensional dense vector space. For example, words can be converted into vector representations, so that the similarity between words can be measured in mathematical space. Specifically, on the one hand, for each medical knowledge text in the above-mentioned medical knowledge base, each medical knowledge text can be embedded in advance to generate an embedding vector corresponding to each medical knowledge text (which can be called a second embedding vector), and an index can be constructed based on the second embedding vector corresponding to each medical knowledge text. On the other hand, when the above-mentioned structured text is obtained, the structured text can be embedded to generate an embedding vector corresponding to the structured text (which can be called a first embedding vector); subsequently, the similarity between the first embedding vector and each second embedding vector can be calculated, and based on the calculated similarity, the medical knowledge text related to the structured text can be obtained from the medical knowledge base.

[0152] In one example, when obtaining medical knowledge text related to the above-mentioned structured text from the above-mentioned medical knowledge base based on the calculated similarity, it is possible to first determine a second embedding vector whose similarity to the first embedding vector corresponding to the structured text is greater than a preset threshold, and then obtain the medical knowledge text corresponding to the determined second embedding vector from the medical knowledge text contained in the medical knowledge base as the medical knowledge text related to the structured text.

[0153] In another example, when obtaining medical knowledge text related to the above-mentioned structured text from the above-mentioned medical knowledge base based on the calculated similarity, it is possible to first determine a preset number (i.e., Top K, K is the preset number) of first embedding vectors with the largest similarity to the first embedding vector corresponding to the structured text, and then obtain the medical knowledge text corresponding to the determined first embedding vector from the medical knowledge texts contained in the medical knowledge base as the medical knowledge text related to the structured text.

[0154] In some embodiments, a large language model that can provide knowledge base retrieval services can also be used to calculate the association between the structured text and each medical knowledge text in the medical knowledge base. Based on the calculated association, medical knowledge texts related to the structured text can be obtained from the medical knowledge base. For example, the structured text and each medical knowledge text in the medical knowledge base can be input into the large language model, which can generate a first embedding vector corresponding to the structured text and a second embedding vector corresponding to each medical knowledge text. The similarity between the first embedding vector and each second embedding vector can then be calculated. Based on the calculated similarity, medical knowledge texts related to the structured text can be obtained from the medical knowledge base.

[0155] It should be noted that the large language model used to perform question-answering tasks in the medical field, the large language model used to perform information extraction tasks, the large language model used to perform structured text generation tasks, the large language model used to perform knowledge base retrieval tasks, etc. can be different large language models or the same large language model, such as Figure 4 The large language model in the illustrated embodiment may also refer to different services provided by the same large language model, which may be specifically configured according to actual needs, and this application does not impose any special restrictions on this.

[0156] In the technical solution provided by one or more embodiments of the present application, optical character recognition can first be performed on the medical examination report image to be interpreted to identify the medical examination report text from the medical examination report image. Then, information extraction can be performed on the medical examination report text to extract the target text related to the medical examination results from the medical examination report text, and the extracted target text can be organized into structured text according to a pre-set structured text format. Finally, the structured text can be input into a large language model, and the large language model can perform inference based on the structured text to generate a medical advice text corresponding to the medical examination results contained in the structured text.

[0157] Using this approach, the medical examination report text obtained through optical character recognition of the medical examination report image can be organized into structured text with a specific format. This allows for unified storage and verification of the medical examination results, making them more organized and efficient. This improves the efficiency of subsequent interpretation of the medical examination report based on these results. Furthermore, using a large language model to generate medical recommendations corresponding to the medical examination results, enabling interpretation of the medical examination report, ensures accuracy and efficiency.

[0158] Corresponding to the aforementioned method embodiments, the present application also provides device embodiments.

[0159] Please refer to Figure 5 , Figure 5 : This is a structural diagram of a device shown in an exemplary embodiment of the present application. At the hardware level, the device includes a processor 502, an internal bus 504, a network interface 506, a memory 508, and a non-volatile memory 510, and of course may also include other required hardware. One or more embodiments of the present application can be implemented based on software, such as the processor 502 reading the corresponding computer program from the non-volatile memory 510 into the memory 508 and then running it. Of course, in addition to software implementation, one or more embodiments of the present application do not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic module, and can also be hardware or logic devices.

[0160] Please refer to Figure 6 , Figure 6 This is a block diagram of a medical examination report interpretation device based on a large language model, shown as an exemplary embodiment of the present application.

[0161] The above-mentioned medical examination report interpretation device based on large language model can be applied to Figure 5 The device shown in the figure is used to implement the technical solution of this application. The device includes:

[0162] The optical character recognition module 602 acquires a medical examination report image to be interpreted and performs optical character recognition on the medical examination report image to identify the medical examination report text from the medical examination report image;

[0163] The examination result extraction module 604 performs information extraction on the medical examination report text to extract target text related to the medical examination result from the medical examination report text, and organizes the target text into a structured text according to a preset structured text format;

[0164] The examination result interpretation module 606 inputs the structured text into a large language model, and the large language model performs reasoning based on the structured text to generate a medical advice text corresponding to the medical examination result.

[0165] In some embodiments, the apparatus further comprises:

[0166] a correlation calculation module, which calculates the correlation between the text contained in the structured text and each medical knowledge text in the medical knowledge base, and based on the correlation, obtains medical knowledge text related to the structured text from the medical knowledge base;

[0167] Inputting the structured text into a large language model, and having the large language model perform reasoning based on the structured text to generate a medical advice text corresponding to the medical examination result, includes:

[0168] The structured text and the medical knowledge text are input into a large language model, and the large language model performs reasoning based on the structured text and the medical knowledge text to generate a medical advice text corresponding to the medical examination result.

[0169] In some embodiments, calculating the relevance between the text contained in the structured text and each medical knowledge text in the medical knowledge base, and obtaining medical knowledge text related to the structured text from the medical knowledge base based on the relevance, includes:

[0170] Performing embedding processing on the text contained in the structured text to generate a first embedding vector corresponding to the structured text;

[0171] Obtaining a second embedding vector corresponding to each medical knowledge text in the medical knowledge base;

[0172] The similarity between the first embedding vector and each second embedding vector is calculated, and based on the similarity, medical knowledge text related to the structured text is obtained from the medical knowledge base.

[0173] In some embodiments, the structured text format is used to indicate that text related to medical examination results is organized into a first-category key-value pair set; wherein the key in each first-category key-value pair contained in the first-category key-value pair set is the medical examination result category, and the value is the medical examination result.

[0174] In some embodiments, the apparatus further comprises:

[0175] The abnormality recognition module performs medical abnormality recognition on the medical examination results corresponding to the first type of key-value pair set based on preset medical abnormality recognition rules before inputting the structured text into the large language model, and adds the key-value pairs organized by the medical abnormality recognition results to the first type of key-value pair set.

[0176] In some embodiments, performing optical character recognition on the medical examination report image to identify the medical examination report text from the medical examination report image includes:

[0177] Optical character recognition is performed on the medical examination report image to identify the text displayed in the medical examination report image, and the identified text is organized into a second-category key-value pair set as the medical examination report text; wherein the key of each second-category key-value pair contained in the second-category key-value pair set is the medical examination report item category, and the value is the medical examination report item.

[0178] In some embodiments, extracting information from the medical examination report text to extract target text related to the medical examination results from the medical examination report text, and organizing the target text into structured text according to a preset structured text format, includes:

[0179] The medical examination report text is input into a large language model, and the large language model extracts information from the medical examination report text to extract target text related to the medical examination results from the medical examination report text, and organizes the target text into structured text according to a preset structured text format.

[0180] In some embodiments, the apparatus further comprises:

[0181] The interpretation output module outputs the structured text and the medical advice text to the user.

[0182] For the device embodiments, they basically correspond to the method embodiments, so for relevant details, please refer to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separate, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the technical solution of this application.

[0183] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or physical devices, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.

[0184] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0185] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0186] Computer-readable media include permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0187] It should be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0188] The above description is of specific embodiments of the present application. Other embodiments are within the scope of this application. In some cases, the actions or steps described in this application can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0189] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a," "the," and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. The term "and / or" refers to and includes any or all possible combinations of one or more of the associated listed items.

[0190] The terms "one embodiment," "some embodiments," "example," "specific example," or "one implementation" used in one or more embodiments of the present application mean that the specific features or characteristics described in conjunction with the embodiment are included in at least one embodiment of the present application. The schematic descriptions of these terms do not necessarily refer to the same embodiment. Moreover, the specific features or characteristics described can be combined in an appropriate manner in one or more embodiments of the present application. In addition, different embodiments and specific features or characteristics in different embodiments can be combined without conflict.

[0191] It should be understood that although the terms first, second, third, etc. may be used to describe various information in one or more embodiments of the present application, these information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0192] The above description is merely a preferred embodiment of one or more embodiments of the present application and is not intended to limit one or more embodiments of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present application shall be included in the scope of protection of one or more embodiments of the present application.

[0193] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

Claims

1. A method for interpreting medical examination reports based on a large language model, the method comprising: Acquiring a medical examination report image to be interpreted, and performing optical character recognition on the medical examination report image to identify the medical examination report text from the medical examination report image; Performing information extraction on the medical examination report text to extract target text related to the medical examination result from the medical examination report text, and organizing the target text into a structured text according to a preset structured text format; The structured text is input into a large language model, and the large language model performs reasoning based on the structured text to generate a medical advice text corresponding to the medical examination result.

2. The method according to claim 1, further comprising: Calculating the relevance between the text contained in the structured text and each medical knowledge text in the medical knowledge base, and based on the relevance, obtaining medical knowledge text related to the structured text from the medical knowledge base; Inputting the structured text into a large language model, and having the large language model perform reasoning based on the structured text to generate a medical advice text corresponding to the medical examination result, includes: The structured text and the medical knowledge text are input into a large language model, and the large language model performs reasoning based on the structured text and the medical knowledge text to generate a medical advice text corresponding to the medical examination result.

3. The method according to claim 2, wherein calculating the degree of association between the text contained in the structured text and each medical knowledge text in the medical knowledge base, and obtaining medical knowledge text related to the structured text from the medical knowledge base based on the degree of association, comprises: Performing embedding processing on the text contained in the structured text to generate a first embedding vector corresponding to the structured text; Obtaining a second embedding vector corresponding to each medical knowledge text in the medical knowledge base; The similarity between the first embedding vector and each second embedding vector is calculated, and based on the similarity, medical knowledge text related to the structured text is obtained from the medical knowledge base.

4. The method according to claim 1, wherein the structured text format is used to indicate that the text related to the medical examination results is organized into a first type key-value pair set; wherein, The key of each first-category key-value pair contained in the first-category key-value pair set is the medical examination result category, and the value is the medical examination result.

5. The method according to claim 4, before inputting the structured text into the large language model, the method further comprises: Based on preset medical abnormality identification rules, medical abnormality identification is performed on the medical examination results corresponding to the first type of key-value pair set, and the key-value pairs organized by the medical abnormality identification results are added to the first type of key-value pair set.

6. The method according to claim 1, wherein performing optical character recognition on the medical examination report image to identify the medical examination report text from the medical examination report image comprises: Optical character recognition is performed on the medical examination report image to identify the text displayed in the medical examination report image, and the identified text is organized into a second-category key-value pair set as the medical examination report text; wherein the key of each second-category key-value pair contained in the second-category key-value pair set is the medical examination report item category, and the value is the medical examination report item.

7. The method according to claim 1, wherein the step of extracting information from the medical examination report text to extract target text related to the medical examination results from the medical examination report text and organizing the target text into structured text according to a preset structured text format comprises: The medical examination report text is input into a large language model, and the large language model extracts information from the medical examination report text to extract target text related to the medical examination results from the medical examination report text, and organizes the target text into structured text according to a preset structured text format.

8. The method according to claim 1, further comprising: The structured text and the medical advice text are output to the user.

9. A device for interpreting medical examination reports based on a large language model, the device comprising: an optical character recognition module, which obtains a medical examination report image to be interpreted and performs optical character recognition on the medical examination report image to identify the medical examination report text from the medical examination report image; An examination result extraction module extracts information from the medical examination report text to extract target text related to the medical examination result from the medical examination report text, and organizes the target text into a structured text according to a preset structured text format; The examination result interpretation module inputs the structured text into a large language model, and the large language model performs reasoning based on the structured text to generate a medical advice text corresponding to the medical examination result.

10. An electronic device comprising: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 8 by running the executable instructions.

11. A computer-readable storage medium having computer instructions stored thereon, wherein when the instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Multi-level interpretability method and system based on medical vision multi-modal large model, terminal and storage medium

    CN122287927A

  • A multi-level explainability method, system, terminal and storage medium based on a medical visual multi-modal large model

    CN122287927B