Medical report interpretation method and related device
By combining optical character recognition technology and multimodal large models, the problem of interpreting complex medical reports by single modal data is solved, and a fast and accurate interpretation solution for medical reports is provided, especially the recognition of complex medical terms and handwritten words, improving the accuracy and efficiency of interpretation.
Patent Information
- Application Number
- CN202510620063.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-19
AI Technical Summary
The prior art medical report interpretation methods that rely on single modal data cannot accurately identify complex medical terms and handwritten text, resulting in misdiagnosis or misdiagnosis, especially in the case of atypical lesions and poor image quality, which is limited in interpretation effects.
Combining optical character recognition technology and multimodal big models, by identifying text information in medical reports, using multimodal big models for semantic understanding, generating report summary and personalized suggestions, including cross-validation of text correction models and algorithm models to improve accuracy.
It realizes rapid, accurate and comprehensive interpretation of medical reports, and can effectively identify complex medical terms and handwritten words, improving the accuracy and efficiency of interpretation of medical reports.
Smart Images

Figure CN120510987A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a medical report interpretation method and related devices. Background Art
[0002] With the continuous advancement of information technology, the medical industry has accumulated a vast amount of medical data, such as electronic medical records and imaging reports. This explosive growth in data has driven the accumulation of medical knowledge and the advancement of intelligent medical applications, but it has also brought with it the complexity and diversity of medical information. Automatic interpretation of medical reports containing medical data can improve medical efficiency, assist doctors in making quick decisions, and enhance the patient experience, making it a critical issue in the medical field.
[0003] Currently, medical report interpretation typically relies on the analysis of single-modality data and relatively simple deep learning models, such as convolutional neural networks, which process a single type of imaging data to identify and classify specific lesions contained therein. However, interpreting medical reports based on single-modality data, such as using text-based deep learning models to understand and analyze text information while ignoring the interaction of image information; or using image-based deep learning models to understand and analyze image information while ignoring the assistance and interpretation of text information, limits the model's interpretation effectiveness. When faced with complex situations such as atypical lesions and poor quality single-modality data, accurate medical report interpretation is often impossible, and may even lead to missed or misdiagnosed lesions.
[0004] Therefore, how to improve the accuracy of interpreting medical reports has become a problem that needs to be solved. Summary of the Invention
[0005] Based on the above problems, the present application provides a medical report interpretation method and related devices, which can improve the accuracy of interpreting medical reports.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] In a first aspect, an embodiment of the present application provides a method for interpreting a medical report, the method comprising:
[0008] Obtaining a medical report; the medical report including at least one of text information, image information, and table information;
[0009] recognizing text information in the medical report by optical character recognition technology to obtain a text recognition result;
[0010] Recognize the medical report based on the text recognition result using a multimodal large model to obtain a first recognition result;
[0011] The multimodal large model performs semantic understanding based on the first recognition result and the target prompt to generate report interpretation content; the report interpretation content includes a report summary and personalized suggestions.
[0012] Optionally, after the medical report is identified based on the text recognition result using the multimodal large model and a first recognition result is obtained, the method further includes:
[0013] Based on the first recognition result, obtaining a revised recognition result by revising a model; the revised model includes a text model and an algorithm model;
[0014] The multimodal large model performs semantic understanding based on the first recognition result and the target prompt, and generates a report interpretation content, specifically:
[0015] The multimodal large model performs semantic understanding based on the modified recognition results and target prompts to generate report interpretation content.
[0016] Optionally, obtaining a revised recognition result by revising the model based on the first recognition result includes:
[0017] Correcting the first recognition result by using the text model in the correction model to obtain a text correction result;
[0018] Correcting the text correction result based on preset rules by using the algorithm model in the correction model to obtain an algorithm correction result;
[0019] The text correction result and the algorithm correction result are cross-validated to obtain a corrected recognition result.
[0020] Optionally, the correcting the first recognition result by using a text model in the correction model to obtain a text correction result includes:
[0021] By correcting a text model that adopts a self-attention mechanism in the model, the first recognition result is corrected based on the context information in the first recognition result to obtain a text correction result.
[0022] Optionally, the preset rules include symbol correction rules and logic judgment rules.
[0023] In a second aspect, an embodiment of the present application provides a medical report interpretation device, the device comprising:
[0024] An acquisition module, configured to acquire a medical report; the medical report includes at least one of text information, image information, and table information;
[0025] an optical character recognition model, configured to recognize text information in the medical report using optical character recognition technology to obtain a text recognition result;
[0026] A multimodal large model is used to identify the medical report based on the text recognition result to obtain a first recognition result; perform semantic understanding based on the first recognition result and the target prompt to generate report interpretation content; the report interpretation content includes a report summary and personalized suggestions.
[0027] Optionally, the medical report interpretation device further comprises: a correction model;
[0028] The correction model is used to correct the first recognition result to obtain a corrected recognition result; the correction model includes a text model and an algorithm model;
[0029] The multimodal large model is specifically used to perform semantic understanding based on the corrected recognition results and target prompts to generate report interpretation content.
[0030] Optionally, the correction model includes a text model, an algorithm model, and a verification unit:
[0031] The text model is used to correct the first recognition result to obtain a text correction result;
[0032] The algorithm model is used to correct the text correction result based on preset rules to obtain an algorithm correction result;
[0033] The verification unit is used to cross-verify the text correction result and the algorithm correction result to obtain a corrected recognition result.
[0034] In a third aspect, an embodiment of the present application provides a computer device, comprising: a system memory, at least one processor, and a bus connecting the system memory and the processor;
[0035] The system memory is used to store program code and transmit the program code to the processor;
[0036] The processor is used to execute the steps of the medical report interpretation method described in any embodiment of the first aspect according to the program code.
[0037] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on a medical report interpretation device, the medical report interpretation device executes the steps of the medical report interpretation method described in any embodiment of the first aspect.
[0038] Compared with the existing technology, this application has the following beneficial effects:
[0039] The embodiment of the present application provides a method for interpreting a medical report, which includes: first, obtaining a medical report; the medical report includes at least one of text information, image information, and table information; then, using optical character recognition technology to identify the text information in the medical report to obtain a text recognition result; using a multimodal large model, the medical report is identified based on the text recognition result to obtain a first recognition result; the multimodal large model performs semantic understanding based on the first recognition result and the target prompt to generate report interpretation content; the report interpretation content includes a report summary and personalized suggestions. Thus, by combining optical character recognition technology with a multimodal large model, it is possible to use optical character recognition technology to efficiently recognize text, and to use the multimodal large model to perform in-depth analysis of images, tables, and complex contexts, making up for the defects of optical character recognition technology in accurately recognizing complex medical terms and handwritten text, and relying on single modality data to accurately interpret medical reports, providing a fast, accurate, and comprehensive medical report interpretation solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0041] Figure 1 A flow chart of a medical report interpretation method provided in an embodiment of the present application;
[0042] Figure 2 A schematic diagram of a medical report provided in an embodiment of the present application;
[0043] Figure 3 A schematic diagram of a report interpretation content provided in an embodiment of the present application;
[0044] Figure 4 A schematic diagram of a first recognition result provided in an embodiment of the present application;
[0045] Figure 5 A schematic diagram of a corrected recognition result provided in an embodiment of the present application;
[0046] Figure 6 A schematic diagram of a medical report interpretation device provided in an embodiment of the present application;
[0047] Figure 7 A structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] The medical report interpretation method and related devices provided in this application can be used in the field of artificial intelligence. The above is only an example and does not limit the application field of the medical report interpretation method and related devices provided in this application.
[0049] The terms "first", "second", "third" and "fourth" in the specification, claims and drawings of this application are used to distinguish different objects rather than to limit a specific order.
[0050] In the embodiments of this application, words such as "as an example" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described in the embodiments of this application as "as an example" or "for example" should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "as an example" or "for example" is intended to present the relevant concepts in a concrete manner.
[0051] The terms used in the implementation section of this application are only used to explain the specific embodiments of this application and are not intended to limit this application.
[0052] As mentioned above, current medical report interpretation typically relies on the analysis of single-modality data and relatively simple deep learning models, such as convolutional neural networks. During training, these models use large amounts of labeled data, such as the shape, size, and density of tumors, to learn how to identify different lesion characteristics. By processing a single type of imaging data, these models can identify and classify specific lesions contained within, such as standardized lesions like tumors and nodules. However, based on single-modality data, complete or multidimensional interpretation of medical reports is often impossible. This is especially true for more complex medical reports, such as those with atypical lesion presentations or poor image quality. Accurate interpretation of medical reports can be challenging and may even lead to missed or misdiagnosed lesions.
[0053] Optical character recognition (OCR) technology is also widely used in medical report interpretation, rapidly extracting textual information and generating reports, enabling doctors to quickly review and analyze patient records. However, OCR technology struggles to accurately interpret low-quality reports, unstructured data, and medical terminology.
[0054] In view of this, an embodiment of the present application provides a method for interpreting a medical report. First, a medical report is obtained; the medical report includes at least one of text information, image information, and table information; then, the text information in the medical report is identified through optical character recognition technology to obtain a text recognition result; the medical report is identified based on the text recognition result through a multimodal large model to obtain a first recognition result; the multimodal large model performs semantic understanding based on the first recognition result and the target prompt to generate report interpretation content; the report interpretation content includes a report summary and personalized suggestions.
[0055] Therefore, the combination of optical character recognition technology and multimodal large models can not only use optical character recognition technology to efficiently recognize text, but also use multimodal large models to conduct in-depth analysis of images, tables and complex contexts. This makes up for the defects of optical character recognition technology in accurately recognizing complex medical terms and handwritten text, and the difficulty in accurately interpreting medical reports when relying on single modality data, and provides a fast, accurate and comprehensive medical report interpretation solution.
[0056] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0057] See also Figure 1 , which is a flow chart of a medical report interpretation method provided in an embodiment of the present application, the method comprising:
[0058] S101: Obtain medical report.
[0059] The medical report includes at least one of text information, image information and table information, and the medical report includes at least text information. For example, the medical report can be as follows: Figure 2 The test report shown, and / or medical images obtained through imaging techniques such as CT, MRI, X-ray or ultrasound, etc.
[0060] S102: Recognize text information in the medical report using optical character recognition technology to obtain a first recognition result.
[0061] Optical Character Recognition (OCR) technology is a technology that converts text content in images into computer-readable text.
[0062] As an example, the medical report can be preprocessed, such as adjusting contrast, binarization, denoising, and tilt correction, so as to improve the accuracy of the first recognition result; then, each individual character is located and separated from the preprocessed medical report, and the shape of each character and features such as stroke direction and intersection are analyzed. Using a pattern recognition algorithm such as a neural network or a support vector machine, the extracted features are compared with the character templates in the database to determine the most likely character and obtain the first recognition result, such as Figure 3 shown.
[0063] The use of OCR technology can quickly extract text content from medical reports, especially for reports with clear text and standard format. It can quickly identify text information from images, reducing the conversion time from image to text, and can quickly obtain basic data from medical reports, saving time for subsequent multimodal large models to make corrections and interpret medical reports based on the first recognition results, thereby improving the efficiency of medical report interpretation; while complex diagnostic analysis, chart analysis, and image recognition tasks are handed over to the multimodal large model for processing, avoiding repeated calculations and further improving the efficiency of medical report interpretation.
[0064] S103: Using the multimodal large model, the medical report is recognized based on the text recognition result to obtain a first recognition result.
[0065] A multimodal large model is an artificial intelligence model that can process and integrate data from multiple different data types or "modalities". It can simultaneously process multiple types of inputs such as images, text, and audio, and extract information from them to achieve more complex and comprehensive task understanding.
[0066] Due to its complex architecture and processing power requirements, multimodal large models may face slow processing speeds and high computing resource consumption when processing complex data such as medical reports. In the embodiments of the present application, OCR technology is used as a pre-step to process the text information in medical reports, and a multimodal large model is used to further recognize the medical reports based on the text recognition results. This can significantly improve the speed of the multimodal large model in recognizing medical reports and reduce the computing resources it needs to occupy during operation.
[0067] S104: The multimodal large model performs semantic understanding based on the first recognition result and the target prompt, and generates a report interpretation content.
[0068] The multimodal large model can not only process text information, but also integrate diverse information such as image information and table information. The multimodal large model combines technologies such as natural language processing and computer vision. It can identify structured elements such as tables in medical reports, unstructured elements such as doctor's notes, and image information such as CT scan images, so as to understand the structure and contextual information of medical reports. It also combines the tables and images contained in the medical reports to perform more complex visual and semantic analysis of the medical reports, and then perform semantic understanding of the first recognition result based on prompts to generate more accurate report interpretation content. Among them, the report interpretation content can include report summaries and personalized suggestions.
[0069] Exemplarily, the target prompt can be a pre-set fixed prompt, for example, "Please analyze the provided test report and identify all abnormal test items that are positive or exceed the reference value range. For each abnormal item, please analyze in detail the potential causes that may cause the abnormality and provide preliminary medical advice based on these findings"; it can also be obtained by adjusting the pre-set initial prompt through the first recognition result, so as to more effectively guide the multimodal large model to generate an answer that meets the expectations. For example, the initial prompt is "Please analyze the provided test report and identify all abnormal test items that are positive or exceed the reference value range. For each For abnormal items, please analyze in detail the potential causes that may lead to the abnormality and provide preliminary medical advice based on these findings. The adjusted target prompt is "Based on the latest blood test report provided by [patient name] (date: March 6, 2025), please do the following: 1. List all abnormal test items that are positive or exceed the normal reference range; 2. For each abnormal item, analyze its possible cause in detail; 3. Based on the above analysis, provide specific preliminary medical advice for each abnormal item; 4. Please organize your answers in the following format: [Abnormal Item Name] - [Cause Analysis] - [Suggestion]".
[0070] For example, a multimodal large model is used to perform semantic understanding based on the first recognition result and the target prompt to generate report interpretation content, which may specifically be:
[0071] S41: Generate personalized suggestions by performing semantic understanding based on the first recognition result and the first target prompt through a multimodal large model.
[0072] For example, the first target prompt might read, "Analyze abnormal test items that exceed the reference range and the potential causes of these abnormalities, and provide preliminary medical advice based on these findings." The multimodal large model can integrate medical data such as the patient's medical history, physical signs, and laboratory results to identify potential abnormalities in medical reports, provide decision support, and generate personalized recommendations for diagnosis or treatment. For example, the multimodal large model can derive diagnostic and treatment recommendations based on the lesion area and related text descriptions in the medical report image.
[0073] S42: Generate a report summary based on the first recognition result and personalized suggestions through the multimodal large model.
[0074] As an example, based on the first recognition result, more critical medical data such as the names and results of items whose corresponding symbols do not indicate "normal" can be extracted from the medical report, and / or treatment recommendations such as recommended drug types and dosages in personalized suggestions can be summarized to generate a report summary to help doctors quickly understand the patient's condition and treatment direction.
[0075] S43: Generate report interpretation content based on the report summary and personalized suggestions.
[0076] As an example, you can combine report summaries and personalized recommendations to generate report interpretation content. Figure 3 As shown, the report summary can be the tabular content of the test report, and the personalized suggestions can be suggestions such as "According to the table data, we can see that the result value of low-density lipoprotein is 3.74, which is higher than the standard value of the normal range. This may indicate that the patient has dyslipidemia and requires further examination and treatment. The result value of apolipoprotein A1 is 0.9, which is lower than the standard value of the normal range. This is usually related to an increased risk of cardiovascular disease. It is recommended that the doctor further evaluate the patient's overall health status and carry out appropriate treatment."
[0077] In an embodiment of the present application, first, a medical report is obtained; the medical report includes at least one of text information, image information, and table information; then, the text information in the medical report is recognized by optical character recognition technology to obtain a text recognition result; then, the medical report is recognized based on the text recognition result by a multimodal large model to obtain a first recognition result; finally, the multimodal large model performs semantic understanding based on the first recognition result and the target prompt to generate report interpretation content; the report interpretation content includes a report summary and personalized suggestions. Therefore, by combining optical character recognition technology with a multimodal large model, it is possible to use optical character recognition technology to efficiently recognize text, and to use a multimodal large model to perform in-depth analysis of images, tables, and complex contexts, making up for the defects of optical character recognition technology in accurately recognizing complex medical terms and handwritten text, and relying on single modality data to accurately interpret medical reports, and providing a fast, accurate, and comprehensive medical report interpretation solution.
[0078] In some embodiments, in order to make the recognition results ultimately used in the semantic understanding process of the multimodal large model more accurate, after obtaining the first recognition result, a revised recognition result can be obtained based on the first recognition result through a revised model. Among them, the revised model can include a text model and an algorithm model, which can correct errors in the first recognition result from both text content and logical rules through information such as contextual information and visual information, thereby improving the accuracy of the recognition result. In this case, the multimodal large model can perform semantic understanding based on the revised recognition result and the target prompt, and generate a report interpretation content.
[0079] Specifically, based on the first recognition result, a revised recognition result is obtained by revising the model, which may be:
[0080] S31: Correcting the first recognition result by using the text model in the correction model to obtain a text correction result.
[0081] Medical reports may contain complex medical terms, handwritten text, or illegible characters that are difficult to accurately recognize using OCR technology, resulting in character or vocabulary level recognition errors in the first recognition result. In an embodiment of the present application, the text model can be a Transformer architecture language model that uses a self-attention mechanism. It performs well in context understanding and semantic analysis. By training with a medical terminology library, it can accurately correct complex medical terms, handwritten text, and other difficult-to-recognize characters. Therefore, the text model uses context information and / or visual information for reasoning, and can further correct the first recognition result to obtain a text correction result.
[0082] Exemplarily, the text model may be a model such as BioBERT or ClinicalBERT that is good at correcting local errors in short texts.
[0083] S32: Correcting the text correction result based on preset rules through the algorithm model in the correction model to obtain an algorithm correction result.
[0084] The algorithmic model can be a rule-based error correction model or a deep learning model, which can correct deviations caused by misunderstanding of visual information or reasoning errors, and provide further rule-based corrections to the text correction results.
[0085] As an example, the preset rules may include symbol correction rules and logical judgment rules. Symbol correction rules are rules for correcting misplaced symbols. For example, the symbol correction rules can be used to compare the result with the reference value and correct the symbol if the comparison result does not match the symbol. Logic judgment rules are rules for detecting whether there are obvious anomalies in the text correction results. For example, the logic judgment rules can be used to check whether the value conforms to medical common sense.
[0086] See also Figure 4 Based on the symbol correction rules, the numerical values of each "result" item in the text correction result are compared with the "reference value". It can be found that the result corresponding to the project name "apolipoprotein B" is "1.03", and the corresponding reference value is "0.6-1.1". The comparison result is that the "result" is within the range of the "reference value". However, the symbol corresponding to "apolipoprotein B" is "↑" which means it is too high, that is, the comparison result does not match the symbol. In this case, the symbol can be corrected, that is, "↑" can be deleted, or the symbol indicating a high level (such as "↑") can be corrected to a symbol indicating normal level (such as "-"). Figure 5 shown.
[0087] In addition, based on logical judgment rules, it is also possible to check whether the numerical values, units, names and other contents in the text correction results conform to medical common sense. For example, the algorithm model can be trained in advance using medical data including project codes, project names, results, units and reference values. When there are obvious abnormalities that do not conform to medical common sense, such as the reference value corresponding to "triglycerides" in the text correction results is "038-188", the trained algorithm model can further correct the text correction results based on medical common sense, thereby further improving the accuracy of the recognition results.
[0088] S33: Cross-validate the text correction result and the algorithm correction result to obtain a corrected recognition result.
[0089] For example, a consistency check can be performed on the text correction results and the algorithm correction results. If there are discrepancies between the two, the final corrected recognition result can be obtained through a fusion strategy such as voting or stacking generalization, using pre-stored expert knowledge or the weights of the text correction results and the algorithm correction results. For example, for context understanding and vocabulary correction, a higher weight can be given to the text correction results; while for numerical comparison and symbol correction, a higher weight can be given to the algorithm correction results.
[0090] Therefore, the text model and the algorithm model work together to make personalized corrections to medical reports by combining the text correction results and the algorithm correction results. For example, for handwritten text reports, OCR technology may incorrectly recognize certain letters or symbols, but the text model can correct them through context, and the algorithm model can provide further corrections based on preset rules, so that the final corrected recognition result has higher accuracy, thereby improving the robustness and stability of the medical report interpretation device.
[0091] In addition, medical reports contain a large amount of data and information, which may contain noise such as repeated information. These noises can be identified and filtered through text models, and the logical relationship between data can be verified through algorithm models, and data that does not conform to the rules can be proposed, thereby reducing the interference of noise on the correction of recognition results.
[0092] See also Figure 6 , this figure is a schematic diagram of a medical report interpretation device provided in an embodiment of the present application, the device including: an acquisition module 601, an optical character recognition model 602 and a multimodal large model 603.
[0093] An acquisition module 601 is configured to acquire a medical report; the medical report includes at least one of text information, image information, and table information;
[0094] An optical character recognition model 602 is configured to recognize text information in a medical report using optical character recognition technology to obtain a first recognition result;
[0095] The multimodal large model 603 is used to identify the medical report based on the text recognition result to obtain a first recognition result; perform semantic understanding based on the first recognition result and the target prompt to generate report interpretation content; the report interpretation content includes a report summary and personalized suggestions.
[0096] Therefore, the combination of optical character recognition technology and multimodal large models can not only use optical character recognition technology to efficiently recognize text, but also use multimodal large models to conduct in-depth analysis of images, tables and complex contexts. This makes up for the defects of optical character recognition technology in accurately recognizing complex medical terms and handwritten text, and the difficulty in accurately interpreting medical reports when relying on single modality data, and provides a fast, accurate and comprehensive medical report interpretation solution.
[0097] Optionally, in some embodiments, the medical report interpretation device also includes: a correction model, used to correct the first recognition result to obtain a corrected recognition result; the correction model includes a text model and an algorithm model; a multimodal large model 603, specifically used to perform semantic understanding based on the corrected recognition result and the target prompt to generate report interpretation content.
[0098] Optionally, the correction model includes a text model, an algorithm model and a verification unit; wherein the text model is used to correct the first recognition result to obtain a text correction result; the algorithm model is used to correct the text correction result based on preset rules to obtain an algorithm correction result; the verification unit is used to cross-validate the text correction result and the algorithm correction result to obtain a corrected recognition result.
[0099] Optionally, the text model is specifically a text model that adopts a self-attention mechanism, which is used to correct the first recognition result based on context information in the first recognition result to obtain a text correction result.
[0100] See also Figure 7 , this figure is a structural diagram of a computer device provided in an embodiment of the present application. The computer device 01 includes a system memory 08, at least one processor 03, and a bus 04 connecting different system components (including the system memory 08 and the processor 03).
[0101] System memory 08 : used to store program codes and transmit the program codes to processor 03 .
[0102] Processor 03: used to execute the steps of the above-mentioned medical report interpretation method according to the instructions in the program code.
[0103] Bus 04 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0104] The computer device 01 typically includes a variety of computer system readable media, which can be any available media that can be accessed by the computer device 01, including volatile and non-volatile media, removable and non-removable media.
[0105] System memory 08 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 09 and / or cache memory 10. Computer device 01 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 11 may be used to read and write non-removable, non-volatile magnetic media ( Figure 7 Not shown, often called a "hard drive"). Although Figure 7 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 04 via one or more data medium interfaces. The memory 08 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present invention.
[0106] A program / utility 12 having a set (at least one) of program modules 13 may be stored, for example, in memory 08. Such program modules 13 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. The program modules 13 generally implement the functions and / or methods of the embodiments described herein.
[0107] The computer device 01 may also communicate with one or more external devices 02 (e.g., a keyboard, a pointing device, a display 07, etc.), one or more devices that enable a user to interact with the computer device 01, and / or any device that enables the computer device 01 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface 06. Furthermore, the computer device 01 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 05. Figure 7 As shown, the network adapter 05 communicates with other modules of the computer device 01 via the bus 04. Figure 7Not shown, other hardware and / or software modules may be used in conjunction with the computer device 01, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0108] The processor 03 executes various functional applications and data processing by running the programs stored in the system memory 08, such as implementing the medical report interpretation method provided in the embodiment of the present application.
[0109] In addition, the present application also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on a medical report interpretation device, the medical report interpretation device executes the steps of the above-mentioned medical report interpretation method.
[0110] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The device and storage medium embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components indicated as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.
[0111] The above is merely one specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for interpreting a medical report, characterized in that: The method comprises: Obtaining a medical report; the medical report including at least one of text information, image information, and table information; recognizing text information in the medical report by optical character recognition technology to obtain a text recognition result; Recognize the medical report based on the text recognition result using a multimodal large model to obtain a first recognition result; The multimodal large model performs semantic understanding based on the first recognition result and the target prompt to generate report interpretation content; the report interpretation content includes a report summary and personalized suggestions.
2. The method according to claim 1, characterized in that After the medical report is identified based on the text recognition result using the multimodal large model and a first recognition result is obtained, the method further includes: Based on the first recognition result, obtaining a revised recognition result by revising a model; the revised model includes a text model and an algorithm model; The multimodal large model performs semantic understanding based on the first recognition result and the target prompt, and generates a report interpretation content, specifically: The multimodal large model performs semantic understanding based on the modified recognition results and target prompts to generate report interpretation content.
3. The method according to claim 2, characterized in that The step of obtaining a revised recognition result by revising the model based on the first recognition result includes: Correcting the first recognition result by using the text model in the correction model to obtain a text correction result; Correcting the text correction result based on preset rules by using the algorithm model in the correction model to obtain an algorithm correction result; The text correction result and the algorithm correction result are cross-validated to obtain a corrected recognition result.
4. The method according to claim 3, characterized in that The step of correcting the first recognition result by using the text model in the correction model to obtain a text correction result includes: By correcting a text model that adopts a self-attention mechanism in the model, the first recognition result is corrected based on the context information in the first recognition result to obtain a text correction result.
5. The method according to claim 3, characterized in that The preset rules include sign correction rules and logic judgment rules.
6. A medical report interpretation device, characterized in that: The device comprises: An acquisition module, configured to acquire a medical report; the medical report includes at least one of text information, image information, and table information; an optical character recognition model, configured to recognize text information in the medical report using optical character recognition technology to obtain a text recognition result; A multimodal large model is used to identify the medical report based on the text recognition result to obtain a first recognition result; perform semantic understanding based on the first recognition result and the target prompt to generate report interpretation content; the report interpretation content includes a report summary and personalized suggestions.
7. The device according to claim 6, characterized in that The medical report interpretation device further includes: a correction model; The correction model is used to correct the first recognition result to obtain a corrected recognition result; the correction model includes a text model and an algorithm model; The multimodal large model is specifically used to perform semantic understanding based on the corrected recognition results and target prompts to generate report interpretation content.
8. The device according to claim 7, characterized in that The correction model includes a text model, an algorithm model and a verification unit: The text model is used to correct the first recognition result to obtain a text correction result; The algorithm model is used to correct the text correction result based on preset rules to obtain an algorithm correction result; The verification unit is used to cross-verify the text correction result and the algorithm correction result to obtain a corrected recognition result.
9. A computer device, characterized in that: The device includes: a system memory, at least one processor, and a bus connecting the system memory and the processor; The system memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the steps of the medical report interpretation method according to any one of claims 1 to 5 according to the program code.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program. When the computer program is executed on a medical report interpretation device, the medical report interpretation device executes the steps of the medical report interpretation method according to any one of claims 1 to 5.
Citation Information
Cited By
Medical record file management and intelligent report interpretation system based on multi-mode AI
CN121662253A
A Medical Record Management and Intelligent Report Interpretation System Based on Multimodal AI
CN121662253B