Apparatus and method for generating medical image interpretation report

A template-driven LLM system addresses the inefficiencies of existing medical imaging reporting by automating keyword extraction and report generation, ensuring rapid, accurate, and consistent cancer diagnosis documentation.

WO2026024058A1PCT designated stage Publication Date: 2026-01-29INHA UNIV RES & BUSINESS FOUNDATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/010807
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2025-07-22
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing medical imaging report generation systems are time-consuming, inconsistent, and lack standardization, leading to increased workload and errors, particularly in the context of cancer diagnosis, where rapid and accurate reporting is crucial.

Method used

A system utilizing a Large Language Model (LLM) to generate medical imaging reports through a template-based approach, automatically extracting keywords and generating findings, conclusions, and recommendations, tailored to cancer types, with standardized input methods and data sets for fine-tuning.

Benefits of technology

Ensures rapid, accurate, and consistent medical imaging reports, reducing doctor workload and improving the quality and consistency of medical documentation, enhancing clinician understanding and patient care.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025010807_29012026_PF_FP_ABST
    Figure KR2025010807_29012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an apparatus and a method for generating a medical image interpretation report, the apparatus comprising: a memory for storing at least one instruction for generating a medical image interpretation report by using a large language model (LLM); and a processor configured to perform an operation according to the instruction, wherein the processor can provide, through the interpretation report generation model, a template for calculating tumor, node, metastasis (TNM) for cancer, extract a keyword from a template input result when information is input to the provided template, input the extracted keyword as a prompt to the interpretation report generation model to automatically generate an interpretation report including findings, conclusion, and recommendation, complete the generated interpretation report, and use a combination of the keywords and the generated interpretation report as a dataset for fine tuning.
Need to check novelty before this filing date? Find Prior Art

Description

Medical image interpretation report generation device and method

[0001] The present disclosure relates to a device and method for generating a medical image interpretation text using a Large Language Model (LLM), and more specifically, to a device and method for generating an interpretation text that automatically generates an interpretation text using LLM by extracting keywords of the interpretation text through a template.

[0002] A report is a formal medical report written by a doctor after analyzing the results of a patient's examination, especially the results of imaging tests such as X-rays, CT scans, and MRIs. The report plays a crucial role in diagnosing and planning a patient's treatment, and facilitates communication among medical professionals. The report includes basic patient information such as name, age, gender, and patient ID; specific information about the test, such as the date, type, and site of examination; clinical findings, and test results. Test results refer to findings observed during imaging or examinations, and may include, for example, abnormalities in organs or tissues, the presence of tumors, or inflammation. Furthermore, the report may include the doctor's comprehensive assessment of the test results, including opinions and recommendations regarding the possibility of a specific disease or the need for further testing.

[0003] While reports are essential to the clinical process, writing them is time-consuming. Doctors often struggle with their busy schedules, requiring them to treat patients, perform surgeries, and conduct research, leaving them little time to dedicate to writing them. Furthermore, hospitals often require specific formats and procedures for writing reports, which, while essential for standardization, impose additional burdens on doctors. Furthermore, analyzing medical images and data and creating accurate reports based on these data is a complex process, especially when processing large amounts of information. Furthermore, reports serve as a means of communication with other medical professionals, so they must be clear, concise, and free of misunderstandings. This requires extra care when writing them.

[0004] Because the accuracy and speed of medical report writing can directly impact patient treatment outcomes, efforts are needed to leverage artificial intelligence (AI) to improve the efficiency of medical report writing and reduce errors. Meanwhile, conventional medical report creation processes are limited to AI solutions providing guidance on specific items.

[0005] The number of cancer imaging tests has been rapidly increasing both domestically and internationally. Consequently, many interpreters are experiencing burnout due to the burden of workload. Excessive workload and accumulated fatigue can lead to errors in interpretation and difficulty maintaining consistent accuracy in interpretation reports. Consequently, there is a risk of delays in the preparation of interpretation reports for expensive medical imaging tests essential for cancer diagnosis, or the production of reports lacking specialized expertise. Therefore, a system to address these issues is essential.

[0006] Medical imaging tests used to diagnose cancer contain a wealth of information about the disease, requiring accurate analysis and high-quality interpretation reports. Previously, medical imaging reports were written in a free-form format, relying on the expertise of individual specialists. This resulted in inconsistent formats and content. Furthermore, reports lacking expertise often made it difficult for clinicians to understand and interpret the medical imaging information.

[0007] Furthermore, most existing medical imaging diagnosis and report writing systems are foreign platforms. Their complex and multi-step process makes interpretation time consuming. This makes them unsuitable for the Korean healthcare environment, where a large number of reports must be produced within a certain timeframe. Furthermore, they increase workload and can discourage their use. Furthermore, they operate only in limited countries, meaning they are not available in Korea or are not covered by health insurance, incurring additional costs for users. Consequently, they are not widely adopted.

[0008] The purpose of the embodiment disclosed in the present disclosure is to accurately and quickly create a medical imaging report for a cancer patient using a template and a large language model (LLM)-based report generation model.

[0009] In addition, the embodiment disclosed in the present disclosure provides a template capable of calculating TNM (Tumor, Node, Metastasis) for 56 types of cancer, and the purpose is to automatically complete the template when medical staff inputs information into the template.

[0010] In addition, the embodiment disclosed in the present disclosure aims to provide a process for automatically extracting keywords based on input data, clearly defining a list of extracted keywords in advance, and automatically generating a reading text (including Findings, Conclusion, and Recommendation) by utilizing the extracted keywords.

[0011] In addition, the embodiment disclosed in the present disclosure aims to utilize keyword combinations and automatically generated reading sentences as a dataset for fine-tuning.

[0012] In addition, the purpose of the embodiment disclosed in the present disclosure is to standardize the contents of a reading text to create a template, and when a conclusion is entered into the created template, the opinion part is automatically created through a reading text creation model.

[0013] In addition, the embodiment disclosed in the present disclosure aims to provide doctors with a rapidly and automatically generated reading text to assist in providing information on a disease and establishing a treatment plan.

[0014] Furthermore, the embodiments disclosed herein create various templates based on the target cancer type, and once a conclusion is written for each template, the opinion section of the report is written using a statement generation model. Furthermore, the embodiments aim to train an artificial intelligence model using a data set collected for opinions matching the conclusion section, so that when the conclusion section is written in a report, the opinion section is automatically written.

[0015] In addition, the purpose of the embodiment disclosed in the present disclosure is to set up a data set based on clinical knowledge for each template and to write sentences connected through a conclusion, and to use a reading sentence generation model within the system to write conclusions and opinions through templates of various cancer types and different formats.

[0016] In addition, in the embodiment disclosed in the present disclosure, the writing method for reports differs from country to country, and the purpose is to enable the writing of a report by tuning the format of the reporter for each country to enable the writing of a report in a concise or long narrative manner depending on the writing method.

[0017] In addition, the purpose of the embodiment disclosed in the present disclosure is to input information after image reading using a standardized input method, and to upload it to a standardized list, and to link the conclusion and opinion sections and write it through a data set.

[0018] In addition, the purpose of the embodiment disclosed in the present disclosure is to provide a method in which a conclusion section is written and then an opinion section is written to complete the reading text.

[0019] However, the problems to be solved according to one embodiment are not limited to those mentioned above.

[0020] The medical image interpretation text generation device according to the present disclosure for achieving the above-described technical task comprises: a memory storing at least one command for generating a medical image interpretation text using an LLM (Large Language Model); And a processor for performing an operation according to the above command, wherein the processor collects a learning data set including a reading text, a reading text template, and a conclusion and opinion data set of the reading text to train a reading text generation model, and fine-tunes the trained reading text generation model with at least one of the learning data and synthetic data, and the processor provides a template for calculating TMN (Tumor, Node, Metastasis) for cancer through the reading text generation model, and when information is input to the provided template, extracts keywords from the template input result, and inputs the extracted keywords as a prompt to the reading text generation model to automatically generate a reading text including findings, conclusions, and recommendations, and completes the generated reading text, and utilizes a combination of the keywords and the generated reading text as a data set for fine tuning.

[0021] At this time, among the TMN, T (Tumor) may include information on the size and depth of invasion of the tumor, N (Node) may include information on the presence or absence of nearby lymph node metastasis and its extent, and M (Metastasis) may include information on the presence or absence of distant metastasis.

[0022] In addition, the processor can configure the detailed items of the template based on the characteristics of the organ according to the type of cancer in the body through the reading text generation model, and extract the stage and keywords based on the morphology, location, size, tumor invasion, and TNM of the lesion when reading the cancerous tumor.

[0023] Additionally, the template can extract keywords through at least one of the user's voice, touch, and mouse click, and calculate a cancer stage based on TNM.

[0024] Additionally, the processor can train the reading text generation model with a structured learning data set that pairs conclusions and opinions, and generate a reading text including a detailed description and observed content based on the keywords through the reading text generation model.

[0025] In addition, the processor inputs reading data into the reading text generation model or extracts a conclusion from a template in which the information is input and links with the LLM to generate a finding for the conclusion, and the finding may include a reading finding including at least one of an abnormal finding found in an imaging examination and a sign of a specific disease.

[0026] In addition, the processor tunes the LLM (Large Language Model) linked to the reading text generation model according to a mode based on a template applied differently depending on the cancer type, and the mode may include the type of cancer tumor.

[0027] In addition, when information is input after image reading using a standardized input method, the processor can upload the input image reading information to a standardized list, connect the conclusion and opinion portions of the opinion template input to the opinion generation model, and create an opinion based on the learning data set through the opinion generation model.

[0028] Additionally, the above learning data set may be a data set including a diagnosis corresponding to the findings of the conclusion and grounds and details matching the diagnosis.

[0029] In addition, the processor can summarize a conclusion from information after reading the input image through the reading text generation model, extract findings and diagnostic information from the summarized text through the reading text generation model, write an opinion using evidence and detailed information matched to the extracted findings and diagnostic information through the reading text generation model, recognize keywords included in the conclusion through the reading text generation model, and generate an opinion of the reading text by applying an opinion template matched in advance to the recognized keywords through the reading text generation model.

[0030] According to the present disclosure, it provides the effect of ensuring consistency and expertise in medical imaging examination results by generating a highly complete professional medical report for cancer tumor diagnosis and evaluation.

[0031] In addition, according to the present disclosure, the generated interpretation text provides high-quality medical image information to clinicians and specialists participating in multidisciplinary treatment, and provides a service that enables clinicians to easily explain cancer tumor imaging test results to patients.

[0032] In addition, according to the present disclosure, it provides the effect of ensuring that medical staff in charge of medical image interpretation work can create efficient cancer tumor diagnosis interpretation reports and increase the clinician's understanding of image findings.

[0033] In addition, according to the present disclosure, by standardizing the contents of the medical report and creating a template, doctors can write the medical report consistently and quickly, thereby providing the effect of improving the quality and consistency of medical documents.

[0034] In addition, according to the present disclosure, the burden on doctors is reduced by automatically creating the opinion section through an LLM-based reading text generation model, thereby enabling doctors to devote more time to treating patients.

[0035] In addition, the present disclosure provides the effect of enabling efficient diagnosis and treatment planning by quickly providing information on a disease through a rapidly and automatically generated reading.

[0036] In addition, according to the present disclosure, various types of templates are created according to the target cancer type, and conclusions and opinions suitable for each template are automatically created, thereby providing the effect of providing specialized and detailed interpretations for various cancer types.

[0037] In addition, according to the present disclosure, it helps to quickly and accurately create medical imaging interpretation reports for cancer patients, reduces the excessive workload of interpreting doctors, and smoothly provides high-quality information to clinical medical staff, thereby enabling the provision of high-quality medical services to cancer patients who have undergone medical imaging examinations.

[0038] In addition, the present disclosure provides the effect of increasing clinicians' understanding of the interpretation results and laying the foundation for providing a service that can explain the interpretation results in detail to patients.

[0039] In addition, the present disclosure supports different report writing methods for each country, enabling the writing of reports that meet the standards and requirements of each country, which has the effect of increasing compatibility in an international medical environment.

[0040] In addition, according to the present disclosure, the process of writing a reading report is made more efficient and accurate by inputting post-reading information through a standardized input method and automatically writing the conclusion and opinion sections through a data set.

[0041] In addition, the present disclosure provides the effect of increasing the work efficiency of doctors by automating the process of writing a medical report, and allowing them to focus on more important treatment and research by relieving them of repetitive tasks.

[0042] In addition, according to the present disclosure, by training an artificial intelligence model with a collected data set, the conclusion section is automatically written when writing a report, thereby providing the effect of continuously improving the accuracy and efficiency of writing a report.

[0043] The effects of the present invention are not limited to the effects described above, and should be understood to include all effects that can be inferred from the detailed description of the present invention or the composition of the invention described in the claims.

[0044] Figure 1 is a drawing showing a medical image interpretation text generation system using LLM according to an embodiment.

[0045] Fig. 2 is a block diagram showing a medical image interpretation document generation device according to an embodiment.

[0046] FIG. 3 is a diagram showing an interface of a template-based reading system for diagnosing rectal cancer according to an embodiment.

[0047] FIG. 4 is a diagram showing the configuration of a command set stored in memory according to an embodiment.

[0048] FIG. 5 is a diagram showing an output interface of a reading text generated by a reading text generation model according to an embodiment.

[0049] Figure 6 is a diagram showing a UI configuration that automatically disables or excludes items that do not apply among TNM staging stages depending on the type of cancer tumor in an embodiment.

[0050] FIG. 7 is a drawing showing an interface in which, when a user selects a specific value among TNM stages in an embodiment, a description text (Text) and a reference figure or image (Fig or Reference) for the stage are output together.

[0051] Fig. 8 is a diagram showing a process for generating a reading text according to an embodiment.

[0052] Figure 9 is a diagram showing the process of generating opinions on a reading text according to an embodiment.

[0053] Hereinafter, the embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Regardless of the drawing numbers, identical or similar components will be given the same reference numbers, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably only for the convenience of writing the specification, and do not in themselves have distinct meanings or roles. In addition, when describing the embodiments disclosed in this specification, if it is determined that a specific description of a related known technology may obscure the gist of the embodiments disclosed in this specification, a detailed description thereof will be omitted. In addition, the attached drawings are only intended to facilitate easy understanding of the embodiments disclosed in this specification, and the technical ideas disclosed in this specification are not limited by the attached drawings, and should be understood to include all modifications, equivalents, and substitutes included in the spirit and technical scope of the present invention.

[0054] Terms that include ordinal numbers, such as first, second, etc., may be used to describe various components, but the components are not limited by these terms. These terms are used solely to distinguish one component from another.

[0055] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.

[0056] In this application, terms such as “include” or “have” are intended to specify the presence of a feature, number, step, operation, component, part or combination thereof described in the specification, but should be understood not to exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof.

[0057] In this specification, the term "unit" includes a unit realized by hardware, a unit realized by software, and a unit realized using both. Furthermore, a single unit may be realized using two or more pieces of hardware, and two or more units may be realized by a single piece of hardware.

[0058] Some of the operations or functions described herein as being performed by a terminal, apparatus, or device may instead be performed by a server connected to the terminal, apparatus, or device. Similarly, some of the operations or functions described herein as being performed by a server may also be performed by a terminal, apparatus, or device connected to the server.

[0059] Hereinafter, the present invention will be described in detail with reference to the attached drawings.

[0060] Figure 1 is a drawing showing a medical image interpretation text generation system using LLM according to an embodiment.

[0061] Referring to FIG. 1, a medical image interpretation text generation system using an LLM-based interpretation text generation model according to an embodiment may be configured to include a medical image interpretation text generation device (100) and a user terminal (200).

[0062] In an embodiment, a medical image interpretation text generation device (100) extracts keywords of an interpretation text through a template when creating a cancer tumor medical image report, automatically generates an interpretation text using the interpretation text generation model, and completes the generated interpretation text.

[0063] In the embodiment, when a conclusion is automatically completed based on a template upon completion of the medical report, a professional opinion is generated using the medical report generation model. The medical image medical report generation device (100) offers higher efficiency than existing companies in that it can create a medical image report with a simple click.

[0064] The above template refers to a standardized form or input form used when creating a cancer oncology medical imaging report. The template includes customized configurations for each cancer type. Furthermore, the template details vary depending on the characteristics of the cancer, depending on the body part and organ. For example, template items for lung cancer and breast cancer may differ.

[0065] In this embodiment, the medical image interpretation report generation device (100) performs automated keyword extraction. In this embodiment, a template capable of generating TMNs for 56 types of cancer is provided, and the template is completed by medical staff entering information into the template. Thereafter, the medical image interpretation report generation device (100) automatically extracts keywords from the template input results.

[0066] In an embodiment, a medical image interpretation text generation device (100) receives morphology, location, size, tumor invasion, TNM stage (T, N, M), etc. as input and automatically extracts related keywords. In an embodiment, a user can input data using voice input, touch, mouse click, etc., and the TNM stage is automatically calculated based on the input information.

[0067] In an embodiment, keywords extracted based on a template may include keywords related to basic tumor characteristics, tumor invasion, TNM staging information, tumor deposits, and additional findings. For example, keywords related to basic tumor characteristics may include morphology such as polypoid, semicircular, circular, solid, mucinous, and mixed, and may include keywords indicating location and size information.

[0068] Keywords related to tumor invasion & spread may include at least one of invasion direction, MRF invasion, EMVI (Extramural Vascular Invasion), and depth of invasion.

[0069] In the embodiment, the direction of invasion may include whether there is invasion of surrounding organs (e.g., anterior invasion, posterior wall invasion), and the presence of MRF invasion (Mesorectal Fascia Involvement) and EMVI (Extramural Vascular Invasion) may include keywords such as no invasion (No) and presence of invasion (Yes). The depth of invasion may include keywords indicating whether there is invasion of the mucosa, muscularis propria, or serosal membrane.

[0070] Keywords for TNM Staging Information may include keywords for T Stage (tumor size and depth), N Stage (N Stage, whether there is lymph node metastasis), M Stage (M Stage, whether there is distant metastasis and its location), and distant metastatic site.

[0071] In the embodiment, the T stage may include T1, T2, T3, T4a, T4b, etc., and the N stage (N Stage, whether there is lymph node metastasis) may include N0 (no metastasis), N1a, N1b, N1c, N2a, N2b, etc. The M stage (M Stage, whether there is distant metastasis and location) may include M0 (no distant metastasis), M1a (metastasis to a single organ), M1b (metastasis to two or more organs), etc. The distant metastatic site may include the liver, lung, peritoneum, etc.

[0072] Keywords related to Tumor Deposits & Additional Findings may include keywords related to tumor deposits (whether present in the mesorectum), invasion of surrounding organs, lymph node status, and circumferential resection margin (CRM).

[0073] Keywords related to invasion of surrounding organs include keywords regarding invasion of surrounding organs such as the bladder, prostate, and vagina, and keywords related to lymph node status may include the number and size of enlarged lymph nodes. Circumferential Resection Margin (CRM) status may include positive / negative.

[0074] The Final Staging & Impression keywords may include related keywords such as the final staging decision (Overall Stage, Stage I, II, III, IV), the key summary that the above-mentioned reading text generation model will comprehensively derive (Findings), the clinical significance and diagnosis result (Conclusion), and recommendations for additional tests and treatment directions (surgery, chemotherapy, radiation therapy, etc.).

[0075] In addition, the medical image interpretation report generation device (100) can automatically generate an interpretation report using the interpretation report generation model based on extracted keywords, and can write an interpretation report based on a structured dataset in which conclusions and opinions are paired. In the embodiment, the template is a tool that allows medical staff to input cancer tumor information in a standardized manner and automatically extract keywords to support the writing of an interpretation report.

[0076] The medical image interpretation report generation device (100) according to the embodiment can provide a high-level opinion report by utilizing an LLM-based interpretation report generation model beyond a simple algorithm, and can maximize the ability to understand and predict complex data patterns, thereby providing more reliable reports and professional analysis results to the user.

[0077] In this embodiment, the medical image interpretation report generation device (100) provides convenience to users by allowing them to view the entire interpretation report, including keywords extracted using a template, at a glance during medical image interpretation. Furthermore, unlike overseas interpretation platforms that divide various information into multiple stages and describe them in detail, the device can assist in creating interpretation reports more easily through a user-friendly interface.

[0078] Additionally, the medical image interpretation report generation device (100) is driven to quickly and accurately generate interpretation reports based on medical image examination results using algorithms and LLM technology, thereby generating high-quality interpretation reports for a large volume of medical image examinations. This allows interpretation doctors to easily create expert interpretation reports and plays a crucial role in accurately conveying medical image information, thereby enabling clinicians to provide high-quality medical services.

[0079] In an embodiment, a medical image interpretation report generation device (100) standardizes the interpretation report content to generate a template, and when a conclusion is written in the generated template, the opinion part of the interpretation report is automatically created through the interpretation report generation model.

[0080] In addition, the medical image interpretation report generation device (100) provides doctors with interpretation reports that are automatically and quickly created, thereby assisting in establishing information about the disease and a treatment plan. In an embodiment, various types of templates are created according to the target cancer type, and after a conclusion is written for each template, the opinion section is created using the interpretation report generation model, and a data set collected for opinions matching the conclusion section can be used as learning data to train the interpretation report generation model. In addition, in an embodiment, the medical image interpretation report generation device (100) automatically creates the opinion section when the conclusion section is written when creating an interpretation report through a model.

[0081] In addition, in the embodiment, the medical image interpretation text generation device (100) sets a learning data set based on clinical knowledge collected for each template. In the embodiment, the learning data set matches the conclusion of the interpretation text and includes sentences connected to the conclusion, and the medical image interpretation text generation device (100) can write an opinion using the sentences included in the learning data set. In addition, the medical image interpretation text generation device (100) uses an interpretation text generation model within the system to write a conclusion and an opinion using templates of various cancer types and different formats.

[0082] In addition, the medical image interpretation report generation device (100) enables the interpretation report to be created by tuning the reporter format for each country in a way that the report is written concisely or in a long descriptive manner, depending on the method, as the writing method differs from country to country.

[0083] In addition, the medical image interpretation report generation device (100) inputs information after image interpretation using a standardized input method into the model, uploads the input information to a standardized list, connects the conclusion and opinion sections, and then creates an opinion based on the learning data set. Through this, the medical image interpretation report generation device (100) provides a method in which the conclusion section is written and then the opinion section is written, thereby completing the interpretation report.

[0084] A medical image interpretation report generation device (100) according to an embodiment enables rapid interpretation report generation along with efficient image analysis, and provides an accurate and consistent analysis method to generate standardized results.

[0085] Furthermore, the medical image interpretation report generation device (100) maximizes expertise in interpreting cancer medical images, providing highly reliable information and enhancing the understanding of clinicians, thereby preventing medical errors. Furthermore, it assists the work of medical image interpretation specialists, reducing fatigue and maintaining homeostasis, thereby improving the overall quality of healthcare.

[0086] Fig. 2 is a block diagram showing a medical image interpretation document generation device according to an embodiment.

[0087] In the embodiment, the medical image interpretation report generation device (100) may be configured as a server. A server is a computing system that provides services to other computers or devices in a computer network or stores and manages data. The server accepts requests from other computers or devices, called clients, and provides responses or data in response to those requests. The configuration of the medical image interpretation report generation device (100) illustrated in FIG. 2 is merely a simplified example.

[0088] The communication module (110) can be configured regardless of the communication mode, such as wired or wireless, and can be configured with various communication networks, such as a personal area network (PAN) and a wide area network (WAN). In addition, the communication module (110) can operate based on the well-known World Wide Web (WWW), and can also utilize a wireless transmission technology used for short-distance communication, such as infrared (IrDA: Infrared Data Association) or Bluetooth. For example, the communication module (110) can be responsible for transmitting and receiving data required to perform a technique according to an embodiment of the present disclosure.

[0089] The memory (120) may refer to any type of storage medium. For example, the memory (120) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., an SD or XD memory, etc.), a RAM (Random Access Memory), a SRAM (Static Random Access Memory), a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a magnetic memory, a magnetic disk, and an optical disk. Such a memory (120) may also constitute a database as illustrated in FIG. 1.

[0090] The memory (120) can store at least one instruction that can be executed by the processor (130). In addition, the memory (120) can store any type of information generated or determined by the processor (130) and any type of information received by the server (200). In addition, the memory (120) stores various types of modules, instruction sets, or models.

[0091] The processor (130) may perform technical features according to embodiments of the present disclosure, which will be described later, by executing at least one instruction stored in the memory (120). In one embodiment, the processor (130) may be configured with at least one core and may include a processor for data analysis and / or processing, such as a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), or a tensor processing unit (TPU) of a computer device.

[0092] This processor (130) can train a neural network or model designed using machine learning or deep learning methods. To this end, the processor (130) can perform calculations for neural network training, such as processing input data for training, extracting features from the input data, calculating errors, and updating the weights of the neural network using backpropagation. In addition, the processor (130) can also perform inference for a predetermined purpose using a model implemented using an artificial neural network method.

[0093] In the embodiment, the medical image interpretation report generated may be composed of 1) clinical information, 2) conclusion, 3) opinion, and 4) clinical recommendation. In the embodiment, the processor (130) extracts keywords of the interpretation report through a self-developed template when creating a cancer tumor medical image report, and inputs the extracted keywords into the interpretation report generation model to automatically generate the interpretation report.

[0094] In an embodiment, the cancer tumor template can be reconfigured to fit the specific characteristics of various cancer types in the body. Furthermore, when interpreting a cancer tumor, if at least one of the following is selected: morphology, location, size, tumor invasion, T (Tumor), N (Node), and M (Metastasis), a keyword with the stage entered is automatically extracted. In an embodiment, T may be configured to represent the size and depth of the tumor, N may be configured to represent the presence and extent of nearby lymph node metastasis, and M may be configured to represent the presence or absence of distant metastasis.

[0095] In an embodiment, the template allows the user to extract keywords through at least one of voice, touch, and mouse click, and at this time, a cancer stage based on 'TNM' is automatically calculated.

[0096] In an embodiment, a user can input information into a template using voice input, touch, or mouse click. For example, the user can say "3 cm tumor, no lymph node metastasis" or select the information using touch or the mouse. Keywords are then extracted based on the input information. In an embodiment, the template automatically extracts relevant keywords based on the information entered by the user (e.g., tumor size, location, invasion, etc.). For example, relevant keywords are extracted, such as "T2, N0, M0" (3 cm tumor, no lymph node metastasis, no distant metastasis).

[0097] Thereafter, the processor (130) analyzes the input keywords and information to automatically calculate the TNM stage (T, N, M values). For example, the T value (Tumor, tumor size and depth of invasion) is determined (e.g., T2) based on the tumor size (e.g., 3 cm) and the extent of invasion. The N value (Node, presence and extent of lymph node metastasis) checks whether there is lymph node metastasis (e.g., N0). The M value (Metastasis, presence or absence of remote metastasis) determines whether there is remote metastasis (e.g., M0). Thereafter, the TNM stage is automatically reflected in the template, and based on this, the processor (130) prepares to generate a reading text through the reading text generation model.

[0098] In an embodiment, when a keyword is extracted using the template, the processor (130) automatically generates a reading text based on the extracted keyword through the reading text generation model. In an embodiment, when a user inputs the morphology (Morphology), location (Location), size (Size), tumor invasion (Tumor Invasion), TNM stage (T, N, M), etc. of a lesion through the template, the processor (130) automatically extracts keywords related to the morphology (Morphology), location (Location), size (Size), tumor invasion (Tumor Invasion), and TNM stage (T, N, M) of the input lesion through the template.

[0099] Thereafter, the processor (130) uses the extracted keywords as input values ​​(Prompt) of the reading text generation model, so that the reading text generation model automatically generates an initial reading text including Findings, Conclusion, and Recommendation. For example, when the keyword "T2N0M0, tumor size 3cm, no invasion" is input, the processor (130) creates a reading text including the cancer status, diagnosis, and recommendations through the reading text generation model.

[0100] Thereafter, the processor (130) reviews the generated reading text and, if necessary, performs processes such as reflecting additional keywords, refining sentences, and supplementing expressions to complete the final reading text. The completed reading text is provided to medical staff and may ultimately be included in a medical imaging report.

[0101] For example, if a user inputs “3 cm sized tumor, no lymph node metastasis, no distant metastasis” into the template, the keyword “T2N0M0” is automatically extracted and automatically input into the readout generation module.

[0102] Subsequently, during the diagnostic report generation process, the diagnostic report generation model generates the following: "A 3 cm nodule was observed in the left upper lobe of the lung. No lymph node metastases or distant metastases were identified. Additional imaging tests and follow-up observation are recommended." Subsequently, through the diagnostic report supplementation and finalization process, the verdict can be completed as follows: "The likelihood of malignancy is low at the current stage, and regular follow-up examinations are recommended."

[0103] The above-described reading text generation model according to the embodiment is a model developed based on a structured dataset that pairs conclusions and opinions, and receives summarized keywords as input and provides a reading text including detailed descriptions and observed contents based on the keywords.

[0104] Fig. 3 shows an interface (30) of a template-based reading system for diagnosing rectal cancer according to an embodiment.

[0105] And, Fig. 3(a) shows a rectal cancer interpretation template. The rectal cancer interpretation template is a structured input form in which the user can select at least one of the morphology (Morphology), location (Location), size (Size), degree of invasion (Tumor Invasion), and whether or not metastasis (TNM Staging) of the cancer. The main input items may include Morphology (shape): Polypoid, Semicircular, Circular, etc., Location (location): location in the rectum (e.g., several cm from the anal verge), Size (size): maximum tumor length (mm), Tumor Invasion (degree of tumor invasion): how far the tumor has invaded, Tumor Deposits (whether or not tumor deposits are confirmed in the mesorectum (fat tissue), etc. In addition, the TNM stage (Stage) is selected. T (Stage) is the size and depth of invasion of the tumor (T1 to T4a / b), N (Stage) is whether there is lymph node metastasis (N0 to N3), and M (Stage) is whether there is distant metastasis (M0, M1a, M1b, etc.).

[0106] Figure 3 (b) illustrates an automatically generated diagnostic report. In an embodiment, the automatically generated diagnostic report (Findings, Conclusion, Recommendations, etc.) is output based on the values ​​selected by the user from the template. In an embodiment, the diagnostic report may include Tumor Invasion, TNM Stage, Overall Stage, Impression, and Recommendation. In the embodiment, Tumor Invasion describes how the tumor invades the surrounding tissue, and TNM Stage is the result of the stage classification according to the selected TNM stage.

[0107] Overall Stage (comprehensive staging) is the final stage (Stage III) determined based on TNM. Impression (findings) is the comprehensive meaning and clinical interpretation of the stage. Recommendation (recommendation) provides information suggesting endoscopic examination, additional biopsy, and treatment direction.

[0108] In addition, the processor (130) utilizes the extracted keyword combination and the automatically generated reading text as a fine-tuning dataset. In the embodiment, keywords extracted through a template (e.g., "T2N0M0, rectal cancer, mucosal invasion, no lymph node metastasis") and the reading text generated by the reading text generation model based on the keywords (e.g., "A tumor measuring 3 cm in size was observed in the left rectum, and lymph node metastasis and distant metastasis were not confirmed. The TNM stage is T2N0M0, and clinically determined to be Stage I. Additional follow-up observation is required.") are extracted as input and output.

[0109] In the embodiment, keywords extracted through a template (e.g., "T2N0M0, rectal cancer, mucosal invasion, no lymph node metastasis") and the keywords are set as input (Prompt), and the keywords (e.g., "T2N0M0, rectal cancer, mucosal invasion, no lymph node metastasis") and the sentence generated by the sentence generation model based on the keywords are set as output (Response), and then a learning dataset in the form of input-output pairs (Prompt-Response) is configured.

[0110] Thereafter, the processor (130) structures the dataset. In an embodiment, input data (X) can be generated by a keyword combination generated through a template, and output data (Y) is a reading text (including Findings, Conclusion, and Recommendation) generated by the reading text generation model.

[0111] For example, a learning dataset may include a dataset where the input keyword is 'T2N0M0, rectal cancer, mucosal invasion' and the output is 'A 3 cm sized tumor was observed in the left rectum, and no lymph node metastasis or distant metastasis was confirmed. The TNM stage is T2N0M0, and clinically determined to be Stage I. Further follow-up observation is required.'

[0112] Additionally, if the keyword input of the learning data set is 'T3N1M0, rectal cancer, lymph node metastasis,' the output can be set to the reading text, 'A 4 cm-sized tumor was identified in the right rectal wall, and metastasis was observed in the surrounding lymph nodes. The TNM stage is T3N1M0, and it is determined to be Stage III. Radiation and chemotherapy should be considered.'

[0113] In an embodiment, the processor (130) utilizes the constructed keyword-reading dataset to retrain the LLM, optimizing the model to learn various TNM combinations and clinical implications to generate more sophisticated readings. Furthermore, transfer learning techniques can be applied to add medical reading generation capabilities to the existing LLM. Furthermore, the processor (130) can input new patient data to evaluate the model's reading generation accuracy, and have actual medical staff review the results to verify clinical accuracy, sentence naturalness, and appropriateness of expression.

[0114] In the embodiment, a learning dataset is constructed by configuring extracted keywords and generated sentences as input-output pairs (Prompt-Response), and by fine-tuning the sentence generation model based on this, a more sophisticated customized model can be developed.

[0115] In addition, the processor (130) trains the reading text generation model based on a structured learning dataset that pairs conclusions and findings, and when keywords for reading text generation are extracted from the template, the processor receives them as input values ​​of the reading text generation model to automatically generate reading text.

[0116] To this end, the processor (130) constructs a structured learning dataset. In an embodiment, the processor (130) constructs a learning dataset by pairing the conclusion and the findings, which are the main components of the reading text. For example, the learning dataset may be composed of 'Conclusion 1: TNM stage is T2N0M0, and clinically determined to be Stage I. - Finding 1: A 3 cm sized tumor is observed in the left rectum, and lymph node metastasis and distant metastasis are not confirmed.', 'Conclusion 2: TNM stage is T3N1M0, and determined to be Stage III. - Finding 2: A 4 cm sized tumor is confirmed in the right rectal wall, and surrounding lymph node metastasis is observed.'

[0117] In an embodiment, the processor (130) fine-tunes the above-described reading text generation model using the constructed conclusion-finding pairs as learning data. At this time, the input data (X) are summarized keywords (TNM stage, tumor size, location, metastasis, etc.), and the output data (Y) may include a reading text in natural language generated based on the keywords. Here, the reading text may include findings and conclusions. Thereafter, keywords are extracted and input through a template. When the user selects the tumor type, size, invasion, TNM stage, etc. from the template, the template automatically extracts summarized keywords based on the information. For example, keywords such as 'T2N0M0, rectal cancer, invasion of the mucosal layer, no lymph node metastasis' may be extracted.

[0118] Thereafter, the processor (130) generates a detailed readout using the readout generation model. In an embodiment, the extracted summary keywords are input (prompted) into the readout generation model, and the readout generation model automatically generates a readout including detailed descriptions and observations by utilizing the learned conclusion-finding dataset. For example, if the input keywords are 'T2N0M0, rectal cancer, mucosal invasion, no lymph node metastasis', the findings (Findings) of the readout can be generated as 'A 3 cm sized tumor was observed in the left rectum, and lymph node metastasis and distant metastasis were not confirmed.' The conclusion can be generated as 'The TNM stage is T2N0M0, and clinically determined to be Stage I. Additional follow-up observation is required.'

[0119] In an embodiment, the processor (130) trains the above-mentioned reading text generation model using a structured learning dataset consisting of conclusion-opinion pairs, and automatically generates a reading text containing detailed descriptions and observations by receiving summary keywords extracted from the template as input. This enables the creation of automated, consistent, and clinically meaningful reading texts.

[0120] In one embodiment, the processor (130) automatically extracts keywords and calculates TNM staging when medical staff selects required information using an input template. Based on this, the LLM can automatically generate and complete a reading report.

[0121] In an embodiment, the final interpretation report can be reviewed and modified by a radiologist, and the processor (130) adds this modification information to a database and trains the interpretation report generation model with the added information. In an embodiment, the interpretation report generation model is trained, thereby enabling the construction of an automatic interpretation report generation system tailored to each individual radiologist.

[0122] FIG. 4 is a diagram showing the configuration of a command set stored in memory according to an embodiment.

[0123] Referring to FIG. 4, the instruction set according to the embodiment may be configured to include a collection unit (121), a preprocessing unit (122), a learning unit (123), and a feedback unit (124). The term 'unit' used in this specification should be interpreted to include software, hardware, or a combination thereof, depending on the context in which the term is used. For example, the software may be machine language, firmware, embedded code, and application software. As another example, the hardware may be a circuit, a processor, a computer, an integrated circuit, an integrated circuit core, a sensor, a MEMS (Micro-Electro-Mechanical System), a passive device, or a combination thereof.

[0124] The collection unit (121) collects a series of data required for generating a medical image interpretation and a learning data set for learning a interpretation generation model. In an embodiment, the collection unit (121) may include a data set for interpretations, interpretations by cancer type, interpretation templates by cancer type according to the target cancer type, and findings matched with the conclusion portion as a learning data set. In an embodiment, the conclusion portion includes the final diagnosis or evaluation of the cancer type in a medical report or interpretation, and the finding includes grounds or details supporting the conclusion. Accordingly, the collection unit (121) may collect the conclusion and findings matched with the conclusion portion of the interpretation portion as a learning data set.

[0125] Table 1 is a table showing an example of a learning data set of conclusions and opinions matching the conclusions according to an embodiment.

[0126] ConclusionsBreast CancerHistological examination results, immunohistochemical staining results, genetic analysis resultsLung CancerCT and PET scan results, biopsy tissue examination results, chest X-ray resultsColorectal CancerColonoscopy results, pathological examination results, fecal occult blood test resultsProstate CancerPSA level, prostate biopsy results, MRI scan resultsLiver CancerLiver function test results, ultrasound results, CT / MRI scan results

[0127] In embodiments, the training data set may include evidence and details matching the conclusion findings and diagnosis. Specifically, if the conclusion is "breast cancer diagnosis," the training data set may include at least one of the following findings matching the breast cancer diagnosis: 1) histological examination results: detection of malignant tumor in breast tissue, 2) immunohistochemical staining results: HER2 positivity, ER negativity, PR negativity, and 3) genetic analysis results: confirmation of BRCA1 / 2 mutation.

[0128] Additionally, the training data set may include at least one of the following findings that match the diagnosis of non-small cell lung cancer when the conclusion is a diagnosis of non-small cell lung cancer: 1) CT and PET scan results: 3 cm mass found in the right upper lobe of the lung, 2) increased FDG uptake, 3) biopsy tissue examination results: confirmed non-small cell lung cancer (NSCLC), and 4) chest X-ray results: increased lung opacity.

[0129] Additionally, if the conclusion part of the learning data set is 'diagnosis of colon cancer', the findings matching the diagnosis of colon cancer may include at least one of the following: 1) colonoscopy result: tumor found in colon mucosa, 2) pathological test result: adenocarcinoma confirmed, and 3) fecal occult blood test result: positive.

[0130] Additionally, in the example, if the conclusion part of the learning data set is 'prostate cancer diagnosis', findings matching the prostate cancer diagnosis may include 1) PSA level: increased to 20 ng / mL, 2) prostate biopsy result: detection of cancer cells in prostate tissue, and 3) MRI scan result: confirmation of tumor inside the prostate.

[0131] Additionally, if the conclusion part of the learning data set is 'diagnosis of hepatocellular carcinoma', the findings matching the diagnosis of hepatocellular carcinoma may include at least one of 1) liver function test results: increased AST and ALT levels, 2) ultrasound test results: detection of a tumor in the right lobe of the liver, and 3) CT / MRI scan results: confirmation of the size and location of a liver tumor.

[0132] Additionally, in the embodiment, the collection unit (121) can collect medical report forms for each country as a learning data set. To this end, the collection unit (121) can collect medical report forms by searching public databases provided by health authorities or related organizations in each country.

[0133] For example, in the United States, the collection unit (121) can collect data provided by the NIH (National Institutes of Health) or the CDC (Centers for Disease Control and Prevention) as a learning data set, and collaborate with various hospitals and medical institutions to directly obtain medical report forms used by them. Furthermore, medical report forms are extracted from academic papers or research materials. For this purpose, the collection unit (121) can use academic search engines such as PubMed and Google Scholar.

[0134] Additionally, in the embodiment, the collection unit (121) automatically collects publicly available medical report forms using web scraping technology. This is effective when the data volume is large and provided in a structured format. Furthermore, publicly available PDF, DOC, and XLS files can be directly downloaded to collect medical report forms, and form data can be automatically collected using APIs provided by health authorities or medical institutions.

[0135] Additionally, the collection unit (121) collects cancer diagnosis templates for each type as a learning data set. In one embodiment, the collection unit (121) can collaborate with various medical institutions and hospitals to obtain cancer diagnosis templates for their use. In one embodiment, collaboration can secure templates used in actual diagnosis, and cancer diagnosis templates are collected by utilizing databases provided by government or public health agencies.

[0136] For example, you can collect data provided by the National Cancer Institute (NCI) in the US or the National Health Service (NHS) in the UK. Additionally, you can extract cancer report templates from academic papers, research reports, and medical journals to collect training data corresponding to each cancer type. For this purpose, you can utilize academic search engines such as PubMed and Google Scholar.

[0137] In the embodiment, the collection unit (121) uses web scraping technology to automatically collect cancer diagnosis templates from public websites, and downloads and collects public diagnosis template files in PDF, DOC, and XLS formats. In addition, cancer diagnosis templates are automatically collected through APIs provided by medical institutions or data providers.

[0138] In the embodiment, the collection unit (121) converts the collected data into a unified format (e.g., a text file). PDF or image files can be converted into text using OCR (Optical Character Recognition) technology, and the refined data is stored in a database in a structured form.

[0139] The collection unit (121) classifies the database by cancer type and stores the items of each diagnostic report template as database fields. Furthermore, the collection unit (121) adds metadata for at least one of the following: the source, creation date, and medical institution used for each diagnostic report template, thereby clarifying the source and characteristics of the data.

[0140] The preprocessing unit (122) preprocesses the collected learning data set to remove biased or discriminatory data from the collected artificial intelligence learning data set. In the embodiment, the preprocessing unit (122) preprocesses the collected learning data set and processes it into a form suitable for learning the above-mentioned reading text generation model.

[0141] For example, the preprocessing unit (122) may perform at least one of noise removal, outlier removal, and missing value processing. Furthermore, the preprocessing unit (122) may perform at least one of data normalization, outlier removal, and data scaling through data preprocessing to prevent the model from learning unnecessary patterns.

[0142] The learning unit (123) trains a deep learning neural network using the collected learning data set to implement a deep learning model. In an embodiment, the model stored in the learning unit (123) and used to learn data may include the aforementioned reading text generation model. Each module or model stored in the learning unit (123) may be in the form of an application executable by the processor (130).

[0143] In an embodiment, the learning unit (123) may additionally generate synthetic data for model learning using the learning data set collected from the collection unit (121). In an embodiment, the synthetic data is data artificially generated for model learning by imitating actual data including the learning data set.

[0144] In an embodiment, the learning unit (123) automatically generates data for various cancer types using a given template to generate synthetic data. In an embodiment, the cancer types include, but are not limited to, at least one of stomach cancer, colon cancer, liver cancer, pancreatic cancer, and breast cancer.

[0145] In an embodiment, the learning unit (123) can generate synthetic data using various generative artificial intelligence models such as GPT-4.

[0146] Additionally, the learning unit (123) can generate synthetic data by combining learning data without using a generative artificial intelligence model. To generate synthetic data, the learning unit (123) sets a data template.

[0147] In this embodiment, a template is set for each cancer type, each template including at least one of the following: patient ID, diagnosis, stage, treatment, and prognosis. Then, various findings are collected for each cancer type, and synthetic data for various cancer types are automatically generated using the given cancer template.

[0148] Additionally, the learning unit (123) can generate detailed information specific to each cancer type through generative artificial intelligence and organize it into a template to create synthetic data. The synthetic data generated in this example can be utilized for various purposes, including data analysis, model training, and research.

[0149] In an embodiment, the report generation model learns from a training data set containing conclusions and matching opinions, and then writes sentences connected to the conclusions into the opinions of the report. In the embodiment, the report generation model utilizes a Large Language Model (LLM) within the system to generate conclusions and opinions using templates for various cancer types and in different formats. Furthermore, since report writing methods vary by country, the report generation model according to the embodiment can generate concise or descriptive reports depending on the writing method of each country.

[0150] The above-mentioned reading text generation model can input reading data into the model or extract a conclusion from an input template and link it with the LLM to generate a finding on the conclusion.

[0151] That is, the above-mentioned reading text generation model standardizes the reading text collected as a learning data set to generate a template, and collects the generated template as input data of the reading text generation model.

[0152] Thereafter, the above-mentioned interpretation generation model analyzes the input template to extract a conclusion and, in conjunction with the LLM, generates a finding based on the conclusion. In an embodiment, the finding may include an abnormal finding found in an imaging examination and a finding that includes at least one sign of a specific disease.

[0153] For example, the above-mentioned text generation model collects texts for each cancer type from at least one of a public database, medical institution collaboration, or academic paper. The collected texts are then converted to text using OCR, duplicate data is removed, and the format is standardized as needed.

[0154] The above-mentioned diagnostic report generation model then defines rules for standardizing the structure and format of the report. For example, the content included in the report is categorized into at least one of "patient information," "diagnosis results," and "medical opinion." A standard template is then created based on the defined rules.

[0155] In this embodiment, a standard template organizes various elements of a text into a consistent format. The collected text data is then mapped to the standard template. This transforms texts in various formats into a consistent format. The model then prepares the standardized template as input data for the text generation model, enabling it to be used as input data. To achieve this, the template is converted into a format the model can process.

[0156] The above-mentioned text generation model then analyzes the standardized template to extract the conclusion. Natural language processing (NLP) technology can be used for this purpose. In an embodiment, at least one of text analysis, keyword extraction, and semantic analysis can be performed to extract the conclusion. The model then works with the LLM to generate findings. For example, the text generation model works with a large-scale language model (LLM, GPT-4).

[0157] In one embodiment, the LLM can generate opinions based on the conclusions. In this embodiment, the reading text generation model inputs the extracted conclusions into the LLM and generates opinions based on them.

[0158] The LLM provides at least one additional detail, analysis, or rationale for the conclusion. In the embodiment, the model integrates the generated conclusions and findings into a single report, reviews the report, and revises it if necessary. This ultimately produces a complete report.

[0159] Additionally, in an embodiment, the above-described text generation model can generate a template by reconfiguring the template into detailed items according to the characteristics of each organ, depending on the type of cancer. To this end, the text generation model defines the characteristics of each cancer (e.g., including at least one of breast cancer, lung cancer, and colon cancer). In an embodiment, the characteristics of each cancer may include at least one of the following: tumor location, general diagnostic items, and examination method.

[0160] Additionally, the above-mentioned text generation model identifies important items that should be included in the text for each cancer tumor, reflecting expert opinions. The text generation model then defines a common basic template. For example, it sets general items to be included in all texts. Furthermore, it defines detailed items tailored to the characteristics of each cancer tumor. For example, for breast cancer, at least one of tumor location, size, and HER2 status may be included. In the embodiment, the basic template is reconfigured by adding detailed items for each cancer tumor.

[0161] Additionally, in the embodiment, the above-mentioned reading text generation model can receive a cancer tumor type as input and output a template suitable for that type through an algorithm that automatically generates an appropriate template based on the type of cancer tumor. The algorithm for generating an appropriate template includes, but is not limited to, at least one of a rule-based algorithm, a machine learning algorithm, and a Named Entity Recognition algorithm.

[0162] Afterwards, the above-mentioned diagnostic text generation model can be trained using the collected diagnostic text data. This allows for the generation of templates suitable for each cancer type. Examples of templates generated by the model according to cancer type are shown in Tables 2 to 4. Specifically, Table 2 represents a breast cancer diagnostic text template, Table 3 represents a lung cancer diagnostic text template, and Table 4 represents a colon cancer diagnostic text template.

[0163] Patient Information:- Name:- Age:- Sex:- Examination Date:- Date:- Examination Type:- Type: Mammography, Ultrasound, MRI- Diagnostic Results:- Tumor Location: Left / Right Breast, Upper / Lower- Tumor Size: cm- Histological Type: Infiltrative, Non-Infiltrative- HER2 Status: Positive / Negative- ER / PR Status: Positive / Negative- Additional Findings:- Lymph Node Invasion- Distant Metastasis

[0164] Patient Information:- Name:- Age:- Sex:- Date of Examination:- Date:- Type of Examination:- Type: CT, PET, X-ray- Diagnostic Results:- Tumor Location: Left / Right Lung, Upper / Lower Lobe- Tumor Size: cm- Histological Type: Small Cell Carcinoma, Non-Small Cell Carcinoma- EGFR Status: Positive / Negative- ALK Mutation: Positive / Negative- Additional Findings:- Lymph Node Invasion- Distant Metastasis

[0165] - Patient Information:- Name:- Age:- Sex:- Date of Examination:- Date:- Type of Examination:- Type: Colonoscopy, CT, MRI- Diagnostic Results:- Tumor Location: Colon / Rectum- Tumor Size: cm- Histological Type: Adenocarcinoma- MSI Status: Positive / Negative- Additional Findings:- Lymph node involvement- Distant metastasis

[0166] In addition, in the embodiment, the reading text generation model is tuned by the LLM linked to the reading text generation model with each mode based on a template applied differently depending on the cancer type.

[0167] To achieve this, the above-mentioned diagnostic text generation model collects diverse diagnostic text data for various cancer types. For example, it secures sufficient data for each cancer type, such as breast cancer, lung cancer, and colon cancer, and then refines and standardizes the collected data.

[0168] Afterwards, the above-mentioned reading text generation model organizes the data according to the template format and performs labeling work as needed.

[0169] The above-described text generation model then defines a basic template for each cancer type. For example, a breast cancer template may include at least one of tumor location, size, and HER2 status. Furthermore, in one embodiment, the model completes the template by adding detailed information for each cancer type. For example, a lung cancer template may include at least one of tumor location, size, and EGFR status.

[0170] Next, the text generation model selects a large language model, such as GPT-4, and initially tunes the LLM using data specific to each cancer type. This step involves learning information specific to each cancer type.

[0171] Afterwards, the above-mentioned reading text generation model sets a different mode for each cancer type. For example, it is defined as at least one of the following modes: breast cancer mode, lung cancer mode, and colon cancer mode.

[0172] The above-mentioned text generation model then fine-tunes the LLM for each mode. To achieve this, the model can be retrained using cancer-specific templates. In the embodiment, templates for each cancer type are defined and the LLM is tuned based on these templates to generate texts specific to each cancer type, thereby generating more accurate and appropriate texts for each cancer type.

[0173] Additionally, the above-mentioned text generation model fine-tunes the LLM to determine a mode that suits the text generation formats of each country and institution. For example, the text generation model assigns country and institution-specific labels to the collected data. This ensures that the LLM can identify the country and institution to which each data belongs during training.

[0174] Additionally, the above-mentioned reading text generation model classifies data by country and institution, and creates sub-datasets suitable for each classification.

[0175] Next, the above-mentioned text generation model selects an LLM, such as GPT-4, and prepares a tokenizer suitable for the selected LLM. Tokenizers are tools used in natural language processing (NLP) models that process text data, breaking sentences into smaller units such as words, partial words, or characters. Tokenizers are used to convert text data into numerical form so that the model can understand and process the text.

[0176] Thereafter, the above-mentioned reading text generation model performs learning settings for fine-tuning the LLM. For example, the above-mentioned reading text generation model sets learning parameters. In an embodiment, the parameters may include at least one of a learning rate, batch size, and number of epochs.

[0177] In addition, in the embodiment, the above-described reading text generation model can generate an opinion through finding after learning a learning data set of a preset level or higher including synthetic data, or, if an opinion is filled in a template, can actively infer and confirm a finding through an already written opinion, and then write a finding.

[0178] To this end, the learning unit (123) can fine-tune the LLM using the collected learning data set and synthetic data to train the reading text generation model to generate reading texts of a quality higher than a preset level. For example, the learning unit (123) first collects and prepares the data set required for learning. The collected data includes both actual reading text data and synthetic data.

[0179] In the embodiment, the real data includes at least one of actually written reading data, such as a medical reading or a research paper, and the synthetic data includes artificial data generated based on the real data.

[0180] Additionally, in the embodiment, the learning unit (123) can use the collected data to create a small-scale language model (SLL). For example, the learning unit (123) refines the collected data to remove unnecessary elements and preprocesses the text to convert it into a form that the model can learn. Subsequently, a small-scale model, such as gpt2, can be selected and trained as a small-scale language model that generates readable text.

[0181] Additionally, the learning unit (123) converts the collected learning data set and synthetic data into a format suitable for learning. Then, an LLM is selected. In the embodiment, a large-scale pre-trained language model such as GPT-4 can be used. Then, the real data and synthetic data are combined to create a dataset for learning.

[0182] Thereafter, the learning unit (123) fine-tunes the LLM using the collected learning data and synthetic data. The LLM learns to generate reading sentences through fine-tuning. Thus, the reading sentence generation model according to the embodiment can generate opinions through finding, or, if opinions are filled in the template, can actively infer and confirm the finding through the already-written opinion, and then create the finding.

[0183] Additionally, the above-mentioned reading text generation model can customize the reading text according to the writer's style. For example, to enable customization for the user who writes the reading text, the model provides the user with an interface necessary for setting input data.

[0184] Specifically, the above-described reading text generation model can display synonyms of words included in the reading text, sentences explaining the words, and higher-level words of a specific word according to a user's selection, or convert them into other words, phrases, or sentences.

[0185] To this end, the above-mentioned text generation model can tokenize the generated text word by word and tag each tokenized word with a detailed description. In an embodiment, the detailed description may include at least one of the following: a definition, synonyms, superordinate category words, advanced words, and alternative vocabulary for each word.

[0186] In an embodiment, the above-described text generation model initially generates a text, and when a user modifies a portion of the initially generated text, other portions of the text can be modified based on the modified text. In an embodiment, text modifications by the user and the model can be repeated to generate a text customized for the user.

[0187] Additionally, the above-described text generation model can generate a single text in various versions, depending on the author's style. In embodiments, the versions include, but are not limited to, easy and difficult versions. The easy version of the text provided in the embodiments is generated using simple terms and explanations for easy patient understanding, and includes at least one of basic diagnostic information, treatment method, and prognosis. The difficult version according to the embodiments is created using specialized terminology and detailed pathological information for easy understanding by medical professionals. The difficult version may include at least one of pathological results, specific details of treatment methods, and a detailed explanation of prognosis.

[0188] Additionally, fine-tuning is performed for each country and institution. To achieve this, the model defines a function for fine-tuning LLMs for each country and institution, and fine-tunes the model using datasets for each country and institution. Then, the mode is set for each country and institution. In the embodiment, an algorithm is implemented to determine the appropriate mode based on the input data. This is to apply the appropriate reading format for each country and institution.

[0189] In addition, in the embodiment, the reading text generation model uploads the input information to a standardized list when information is input after image reading using a standardized input method, connects the conclusion and opinion parts of the reading text template input to LLM, and writes an opinion using a learning data set.

[0190] Fig. 5 shows an output interface (4) of a reading text generated by a reading text generation model according to an embodiment.

[0191] In this embodiment, the above-mentioned interpretation text generation model collects information entered by the physician or interpreter after image interpretation. This information may include various details, such as tumor size, location, and histological characteristics.

[0192] Thereafter, the above-mentioned reading text generation model designs a user interface (UI) (4) to receive input information in a standardized format. For example, the input data is structured using at least one of a checkbox, a drop-down menu, and a text field.

[0193] The above-mentioned reading text generation model then converts the input information into a standardized list. This is to ensure that the input data is stored in a consistent format.

[0194] Afterwards, the above-mentioned reading text generation model uploads the standardized information to a database or storage accessible to the model.

[0195] The above-mentioned reading text generation model then links the conclusion and opinion sections of the reading text template. In one embodiment, the linking process involves consistently aligning the conclusion and opinion sections of the reading text template using the input information.

[0196] To this end, the above-described text generation model loads a text template. For example, the text generation model can load a text template for a specific cancer type. Then, based on the input information, the conclusion and opinion sections are linked to the template, and the opinion is generated using a training data set. In an embodiment, the opinion is automatically generated using the training data set. For this purpose, LLM is used. In an embodiment, the text generation model can generate opinions through LLM to obtain a final text, and can combine the conclusion and opinion to generate a final text.

[0197] Additionally, in the embodiment, the above-described reading text generation model summarizes the conclusion, extracts findings and diagnostic information from the summarized text, and writes the basis and details matching the extracted findings and diagnostic information as opinions.

[0198] In one embodiment, a medical professional can write a conclusion after reading an image. For example, if a medical professional writes, "The patient was diagnosed with breast cancer. The tumor is 3 cm in size and HER2 positive," the text generation model applies a summarization algorithm to the written conclusion. For example, the text generation model uses natural language processing (NLP) technology to summarize the conclusion.

[0199] In this embodiment, the above-mentioned text generation model extracts key information from the conclusion through a summarization algorithm to generate a concise summary. Then, findings and diagnostic information are extracted from the summarized text.

[0200] For example, the above-mentioned text generation model extracts findings and diagnostic information from the summarized text. This can be accomplished using named entity recognition (NER) or other NLP techniques. The findings and diagnostic information are then identified. For example, the above-mentioned text generation model identifies key findings (e.g., tumor size and location) and diagnostic information (e.g., HER2 status) from the extracted text.

[0201] The above-mentioned reading text generation model then matches the evidence and details. In an embodiment, the matching of evidence and details can be accomplished through an algorithm that matches the evidence and details corresponding to the findings and diagnostic information. This matching is performed using a predefined database or rules.

[0202] The above-mentioned diagnostic text generation model then builds a database containing the evidence and details that can be matched to each finding and diagnostic information. For example, the diagnostic information "HER2 positive" includes details describing the HER2 status test.

[0203] The above-mentioned interpretation generation model then generates an opinion based on the extracted findings and diagnostic information, along with supporting evidence and detailed information. In the example, the opinion generated can provide specific explanations and additional information to support the conclusion.

[0204] Additionally, in the embodiment, the above-mentioned reading text generation model automatically generates opinions using the LLM model. The model can input extracted information and generate appropriate opinion sentences.

[0205] Below, a specific example for writing comments on the above-mentioned reading text generation model is described.

[0206] For example, if the input conclusion is "The patient was diagnosed with breast cancer. The tumor size is 3 cm and it is HER2 positive.", the above-mentioned reading text generation model extracts the summary as "Breast cancer, tumor size 3 cm, HER2 positive" and extracts findings and diagnostic information from the summary.

[0207] In the example above, the findings could be extracted as "tumor size 3 cm" and "HER2 positivity." The extracted findings are then matched with the evidence and details. In this example, the matched evidence could be "HER2 status plays an important role in the diagnosis and treatment decisions of breast cancer," and the details could be "The tumor size is 3 cm and classified as a medium-sized tumor."

[0208] Afterwards, the above-mentioned reading generation model can write the findings as "The tumor size is 3 cm and is classified as a medium-sized tumor. HER2 status plays an important role in the diagnosis and treatment decisions of breast cancer."

[0209] Additionally, the above-mentioned interpretation generation model recognizes keywords contained in the conclusion and applies a pre-matched opinion template to the recognized keywords to generate the interpretation. To this end, the interpretation generation model recognizes keywords in the written conclusion.

[0210] In this example, a physician or reader writes a conclusion after reading the images. For example, the conclusion might be, "The patient was diagnosed with breast cancer, the tumor is 3 cm in size, and HER2 positive."

[0211] The above-mentioned text generation model then uses natural language processing (NLP) technology to identify key keywords from the conclusions. For example, keywords such as "breast cancer," "3cm," and "HER2 positive" can be extracted.

[0212] The above-mentioned reading text generation model then matches keywords with opinion templates. In one embodiment, a database is constructed that predefines opinion templates matching each keyword. For example, a specific opinion template is associated with the keyword "breast cancer."

[0213] The above-mentioned interpretation generation model then selects an appropriate finding template based on the recognized keywords, using an algorithm that retrieves the finding template matching the keyword. Then, specific information extracted from the conclusion is substituted into variables in the finding template. For example, the "tumor size" variable could be replaced with 3 cm.

[0214] Thereafter, the above-mentioned reading text generation model constructs a final opinion using the substituted template, and utilizes the LLM model to generate natural and consistent opinion sentences based on the template.

[0215] Specifically, in an embodiment, if the conclusion of the reading is "The patient was diagnosed with breast cancer, the tumor size is 3 cm, and it is HER2 positive," the reading generation model recognizes the keywords as ["breast cancer", "3 cm", "HER2 positive"], and matches the recognized keywords with the finding template. In an embodiment, using a keyword-template database, it can match "breast cancer": "The patient was diagnosed with breast cancer. This is found in X% of breast cancer patients.", "HER2 positive": "HER2 positive breast cancer has a different prognosis.", "3 cm": "The tumor size is 3 cm, which is a medium-sized tumor."

[0216] A matching template might be ["The patient was diagnosed with breast cancer. This is found in X% of breast cancer patients.", "HER2-positive breast cancer has a different prognosis.", "The tumor is 3 cm in size, which is a medium-sized tumor."].

[0217] Afterwards, the above-mentioned sentence generation model loads a template matching each keyword to apply the opinion template. Then, variables within the template are replaced with actual data as needed. For example, "X% of breast cancer patients" → "15% of breast cancer patients." The model then generates a final opinion sentence based on the loaded template. In the embodiment, LLM is utilized to naturally connect sentences and generate a consistent opinion. In the example described above, the final opinion can be generated in a format including, "The patient was diagnosed with breast cancer. This is found in 15% of breast cancer patients. The tumor size is 3 cm, which is a medium-sized tumor. HER2-positive breast cancer has a different prognosis."

[0218] Referring back to FIG. 3, the feedback unit (124) evaluates the learned artificial neural network model and deep learning model. In an embodiment, the feedback unit (124) may evaluate the artificial neural network model through at least one of accuracy, precision, and recall. Accuracy is an index that measures how well the results predicted by the artificial neural network model match the actual results. Precision is an index that measures the proportion of actual positives among the results predicted as positive. Recall is an index that measures the proportion of actual positives predicted by the model as positive. In an embodiment, the feedback unit (124) may calculate the accuracy, precision, and recall of the artificial neural network model, and evaluate the artificial neural network model based on at least one of the calculated indexes.

[0219] In an embodiment, the feedback unit (124) can measure the accuracy of the artificial neural network model using an evaluation dataset. The evaluation dataset consists of data that the model did not use for training and is used to objectively evaluate the model's performance. In an embodiment, the feedback unit (124) executes the artificial neural network model using the evaluation dataset and compares the artificial neural network model's predicted value for each input data with the actual correct answer value of the corresponding data. Thereafter, the accuracy of the model's predictions can be measured based on the comparison results. For example, the accuracy in the feedback unit (124) can be calculated as the ratio of data correctly predicted by the model among the entire data.

[0220] In addition, the feedback unit (124) can calculate the F1 score, which is an index indicating the balance of precision and recall, which is an index calculated as the harmonic mean of precision and recall, evaluate the artificial neural network model based on the calculated F1 score, generate an AUC-ROC curve, which is an index that visualizes the performance of the classification model in a graph, and evaluate the artificial neural network model based on the generated AUC-ROC curve. In an embodiment, the feedback unit (124) can evaluate that the performance of the model is better as the area under the ROC curve (AUC) is closer to 1.

[0221] In addition, the feedback unit (124) can evaluate the interpretability of the artificial neural network model. In an embodiment, the feedback unit (124) evaluates the interpretability of the artificial neural network model through SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) methods. SHAP (SHapley Additive exPlanations) is a library that provides an interpretation of the results predicted by the model, and the feedback unit (124) extracts SHAP values ​​from the library. In an embodiment, the feedback unit (124) can predict how much the characteristic information input to the model influenced the model prediction through the SHAP value extraction.

[0222] The Local Interpretable Model-agnostic Explanations (LIME) method is a method for explaining model predictions for individual samples. In one embodiment, the feedback unit (124) uses the LIME method to approximate the sample as an interpretable model and calculate the importance of each characteristic. Furthermore, the feedback unit (124) can estimate the influence of each characteristic variable by analyzing the model's internal weights and bias values.

[0223] The feedback unit (124) performs improvement work when the fairness of the artificial neural network model is low or shows discrimination. In an embodiment, the feedback unit (124) collects additional data representing the specific group when the data for the specific group is insufficient by a certain level or more and performs a data preprocessing process. In an embodiment, the feedback unit (124) performs a data preprocessing process including data normalization, outlier removal, and data scaling to prevent the model from learning unnecessary patterns. In addition, in an embodiment, the feedback unit (124) can prevent discrimination or ensure fairness by adding specific conditions to the model learning algorithm.

[0224] In an embodiment, the feedback unit (124) evaluates the performance of the model by comparing the model's predicted results with actual results through confusion matrix analysis to ensure fairness. A confusion matrix is ​​a matrix that evaluates the classification performance of a model in supervised learning. The confusion matrix displays the classification results by comparing the model's predicted results with actual results. In an embodiment, the feedback unit (124) can evaluate the performance of the model by calculating the accuracy and misclassification rate for each class through confusion matrix analysis.

[0225] Additionally, in the embodiment, the feedback unit (124) enables the distribution of data to be confirmed through visual analysis of learning data. For example, in the case of image data, image samples for each class can be visualized to evaluate the diversity and fairness of the data.

[0226] In addition, the feedback unit (124) verifies the fairness and diversity of the learning data and improves the artificial neural network model through fairness verification and evaluation index calculation. In an embodiment, fairness verification is to check whether the artificial neural network model shows discrimination for specific data attributes with respect to the learning information. In an embodiment, the feedback unit (124) can check whether discrimination for specific attributes is present by comparing the number of samples for each attribute or evaluating the classification performance for each attribute.

[0227] Additionally, the feedback unit (124) calculates various indicators to evaluate the performance of the artificial neural network model. For example, the model's performance can be evaluated by calculating at least one indicator among accuracy, precision, recall, and F1-score. At this time, the fairness and diversity of the model can be evaluated by calculating indicators for each class.

[0228] In addition, the feedback unit (124) collects feedback on problems that occur when the artificial neural network model is used in an actual environment, and continuously improves the artificial neural network model by reflecting the collected feedback in the artificial neural network model.

[0229] In an embodiment, the interface (4) includes a plurality of input item fields, and each field can be distinguished based on the patient's pathological findings, the location, size, and stage of the lesion. In an embodiment, the interface (4) can indicate at least one of morphological classification, microscopic findings, location information, lesion size, degree of tumor invasion, invasion of the anal sphincter, stage classification, and metastasis. The morphological classification (Morphology) can be configured to allow selection of the macroscopic shape of the lesion, such as Polypoid, Semicircular, Circular, and Mixed. The microscopic findings (Microscopy) can be configured with fields indicating histological characteristics, such as Solid and Mucinous.

[0230] Location information includes whether it is left / right / bilateral, distance from the anus (cm), circumference in a clockwise direction, and relationship to peritoneal reflection. Lesion size can be entered as a numerical value such as maximum longitudinal diameter. Tumor invasion can be composed of multiple fields describing detailed invasion findings such as invasion of muscularis propria (MP), extramural vascular invasion (EMVI), serosal invasion, and pelvic wall invasion. Anal sphincter invasion includes invasion of the internal and external sphincters and upper anatomical structures (levator ani muscle).

[0231] The T-Stage / N-Stage field provides fields for selecting the T-stage (T1T4b) or N-stage (N0N2b) according to the TNM classification system. Metastasis includes items for the presence or absence of distant metastasis (M0, M1a, M1b, etc.). The interface provided in the example is structured in a table format so that medical professionals can intuitively enter or confirm each item, and can contribute to the establishment of a diagnosis and treatment plan based on comprehensive information such as the tumor's anatomical location, size, extent of invasion, and presence or absence of lymph node metastasis.

[0232] Figure 6 shows a UI (40) configuration that automatically disables or excludes items that do not apply among the TNM stages depending on the type of cancer tumor in an embodiment.

[0233] Referring to Fig. 6, depending on the type of cancer (e.g., rectal cancer, stomach cancer, etc.), the non-applicable sub-stages among the TNM staging items can be automatically distinguished and deactivated (grayed out) or removed from the user interface (40).

[0234] When a user selects a specific cancer type (e.g., Rectal Cancer or Stomach Cancer), the system automatically determines TNM sub-stages that do not apply to the cancer type based on predefined staging mapping information. These inapplicable stages are visually distinguished in the UI and can be displayed as grayed-out areas (unselected areas) or visually blocked. For example, for Rectal Cancer, T1a, T1b, T1c, T2aT2c, T3aT3b, etc. are not applicable and are therefore grayed out and unselectable. Similarly, for Stomach Cancer, T1c, T2aT2c, T3aT3b, etc. are also grayed out and unselectable.

[0235] Additionally, in embodiments, an X mark or shade may be displayed together to visually clearly inform the user that the corresponding stage of the disease is not selectable. This is to increase the accuracy of the cancer type-specific stage classification system and prevent user input errors.

[0236] In this embodiment, only valid TNM staging items are dynamically loaded or activated based on the selected cancer type, and non-valid TNM staging items are automatically removed or deactivated from the UI. For example, when selecting rectal cancer, only T1T4b, N0N2b, and M0~M1c can be activated, and when selecting stomach cancer, only T1aT4b, N0N3b, and M0~M1 can be activated.

[0237] In the example, the TNM staging criteria are enabled based on the AJCC / UICC guidelines and the latest staging criteria. This provides medical professionals with an accurate and standardized staging entry environment based on cancer type. In the example, only valid TNM staging criteria are automatically selected based on cancer type, preventing input errors and improving the user experience. Furthermore, accurate keyword input is encouraged for automatic text generation.

[0238] FIG. 7 shows an interface (50) in which, when a user selects a specific value among TNM stages in an embodiment, a description text (Text) and a reference figure or image (Fig or Reference) for the stage are output together.

[0239] Referring to FIG. 7, when selecting a TNM stage for rectal cancer, a user interface (UI) (50) that provides the meaning of the stage in visual and text form is illustrated.

[0240] In the embodiment, the user can select one of the T, N, and M weapon stages above, and depending on the selected combination, detailed description text (Text) and related video image or diagram (Fig or Reference) are displayed together in the central box area.

[0241] For example, if you select T3, the text "T3: Tumor invades through muscularis propria into perirectal tissues" will be displayed, if you select N2a, the description "N2a: 4-6 regional lymph nodes are positive" will be displayed, and if you select M0, the description "M0: No distant metastasis" will be displayed. In addition, the comprehensive stage (Stage IIIB) corresponding to the corresponding TNM combination (T3, N2a, M0) will also be displayed.

[0242] Additionally, the right area provides reference images (Figure or Reference) corresponding to the selected stage. These images can be image files based on the user's own dataset, representing pathological features such as tumor invasion depth, extent of lymph node metastasis, and invasion of surrounding organs, or they can be retrieved and output from an external reference database (e.g., a RAG-based medical image repository). This configuration allows users to simultaneously perform accurate interpretation and visual understanding of TNM staging, thereby enhancing the accuracy of staging and learning efficiency.

[0243] Hereinafter, let us examine Fig. 8. The method for generating a reading text illustrated in Fig. 8 can be performed by a medical image reading text generation device (100) including a processor (130).

[0244] Additionally, the process of generating the opinion of the reading text shown in FIG. 9 can also be performed by a medical image reading text generation device (100) including a processor (130).

[0245] Meanwhile, Fig. 8 is merely exemplary, and the spirit of the present invention is not limited to what is illustrated in Fig. 8. For example, each step may be configured in a different order than that illustrated in Fig. 3, at least one of the steps illustrated in Fig. 8 may not be performed, or one or more steps not illustrated in Fig. 3 may be additionally performed. Hereinafter, a method for generating a reading text will be described in order. Since the operation (function) of the method for generating a reading text according to the embodiment is essentially the same as the function of the system, any description overlapping with that of Figs. 1 to 4 will be omitted.

[0246] Fig. 8 is a diagram showing a process for generating a reading text according to an embodiment.

[0247] Referring to Figure 8, in step S10, a template for calculating TMN (Tumor, Node, Metastasis) for cancer is provided, and in step S20, when information is entered into the provided template, keywords are extracted from the template input result. In step S30, the extracted keywords are input as a prompt to automatically generate a report including Findings, Conclusion, and Recommendation, and in step S40, the generated report is completed.

[0248] Figure 9 is a diagram showing the process of generating opinions on a reading text according to an embodiment.

[0249] In step S100, the collected texts are standardized by the text generation device to create a template. The generated template is then input into the text generation model. In step S200, the text generation model analyzes the input template to extract a conclusion, which is then linked to the LLM in step S300. Subsequently, in step S400, a finding is generated based on the conclusion of the text. In some embodiments, the finding may include abnormalities found in imaging tests or signs of a specific disease.

[0250] The medical image interpretation report generation device according to the embodiment generates a high-quality, professional medical report for cancer diagnosis and evaluation, ensuring consistency and expertise in medical imaging test results. Furthermore, the generated report provides high-quality medical imaging information to clinicians and specialists participating in multidisciplinary care, and enables clinicians to provide services that facilitate easy-to-understand explanations of cancer imaging test results to patients.

[0251] Furthermore, it enables medical professionals responsible for interpreting medical images to efficiently generate cancer diagnosis reports, ensuring clinicians' understanding of imaging findings. Furthermore, by standardizing the content of reports and creating templates, it enables physicians to consistently and quickly create reports, thereby improving the quality and consistency of medical documentation.

[0252] In addition, the medical image interpretation report generation device and method according to the embodiment standardizes the interpretation report content to create a template, thereby enabling doctors to consistently and quickly create interpretation reports, thereby providing the effect of improving the quality and consistency of medical documents.

[0253] Furthermore, the embodiment reduces the workload of doctors by automatically generating the opinion section using a Large Language Model (LLM), allowing doctors to devote more time to patient care. Furthermore, the embodiment provides the effect of efficiently establishing diagnosis and treatment plans by quickly providing information about the disease through automatically generated reports. Furthermore, the embodiment creates various templates according to the target cancer type and automatically generates conclusions and opinions tailored to each template, providing the effect of providing specialized and detailed reports for various cancer types.

[0254] Furthermore, by supporting different reporting methods for each country through examples, the system allows for the creation of reports tailored to each country's standards and requirements, thereby enhancing compatibility within the international healthcare environment. Furthermore, the system allows for the entry of post-report information through a standardized input method, and automatically generates conclusions and findings using data sets, making the report writing process more efficient and accurate.

[0255] Furthermore, through examples, automating the process of writing reports can increase physicians' work efficiency, freeing them from repetitive tasks and allowing them to focus on more important tasks such as clinical practice and research. Furthermore, by training an AI model on the collected data set, the conclusion section of a report is automatically filled in when the opinion section is written, thereby continuously improving the accuracy and efficiency of report writing.

[0256] A model in this specification may refer to any form of computer program that operates based on a network function, an artificial neural network, and / or a neural network. Throughout this specification, the terms model, neural network, network function, and neural network may be used interchangeably. A neural network is a network in which one or more nodes are interconnected through one or more links to form input node and output node relationships within the neural network. The characteristics of a neural network can be determined based on the number of nodes and links within the neural network, the correlation between the nodes and links, and the weight value assigned to each link. A neural network may be composed of a set of one or more nodes. A subset of the nodes constituting the neural network may constitute a layer.

[0257] A deep neural network (DNN) may refer to a neural network that includes multiple hidden layers in addition to an input layer and an output layer. A deep neural network may include a convolutional neural network (CNN), a recurrent neural network (RNN), an autoencoder, a generative adversarial network (GAN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a Q network, a U network, a Siamese network, a generative adversarial network (GAN), a transformer, and the like. The description of the above-described deep neural network is merely an example, and the present disclosure is not limited thereto.

[0258] Neural networks can learn through at least one of the following methods: supervised learning, unsupervised learning, semi-supervised learning, self-supervised learning, or reinforcement learning. Neural network learning can be the process of applying knowledge to the neural network to perform a specific action.

[0259] Neural networks can be trained to minimize output errors. This process involves repeatedly inputting training data into the neural network, calculating the neural network output and target error for the training data, and backpropagating the neural network error from the output layer to the input layer to update the weights of each node in the neural network to reduce the error. In supervised learning, labeled data is used for each training data, while unsupervised learning uses unlabeled data. The amount of change in the connection weights of each updated node can be determined by the learning rate. The neural network's calculation of input data and backpropagation of errors can constitute a learning cycle (epoch). The learning rate can vary depending on the number of iterations in the neural network's training cycle. Additionally, to prevent overfitting, at least one of the following methods can be applied: increasing the training data, regularization, dropout that disables some nodes, and batch normalization layer.

[0260] In one embodiment, the model may borrow at least a portion of a transformer. The transformer may be composed of an encoder that encodes embedded data and a decoder that decodes the encoded data. The transformer may have a structure that receives a series of data and outputs a series of data of different types through encoding and decoding steps. In one embodiment, the series of data may be processed into a form operable by the transformer. The process of processing the series of data into a form operable by the transformer may include an embedding process. Expressions such as data tokens, embedding vectors, and embedding tokens may refer to data embedded in a form operable by the transformer.

[0261] To encode and decode a series of data, a transformer can utilize an attention algorithm to process the encoders and decoders within the transformer. An attention algorithm can refer to an algorithm that, for a given query, calculates the similarity for one or more keys, reflects this similarity in the values ​​corresponding to each key, and then weights and sums the values ​​to which the similarity is reflected to calculate an attention value.

[0262] Depending on how the query, key, and value are configured, various types of attention algorithms can be categorized. For example, if attention is obtained by setting the query, key, and value all to the same value, this could be a self-attention algorithm. If attention is obtained by reducing the dimensionality of the embedding vector to process a series of input data in parallel and then generating individual attention heads for each segmented embedding vector, this could be a multi-head attention algorithm.

[0263] In one embodiment, the transformer may be composed of modules that perform multiple multi-head self-attention algorithms or multi-head encoder-decoder algorithms. In one embodiment, the transformer may also include additional components other than attention algorithms, such as embedding, normalization, and softmax. Methods for constructing a transformer using attention algorithms may include methods disclosed in Vaswani et al., Attention Is All You Need, 2017 NIPS, which is incorporated herein by reference.

[0264] A transformer can be applied to various data domains, such as embedded natural language, segmented image data, and audio waveforms, to transform a series of input data into a series of output data. To transform data with various data domains into a series of data that can be input to a transformer, the transformer can embed the data. The transformer can process additional data that expresses the relative positional relationship or phase relationship between the series of input data. Alternatively, vectors expressing the relative positional relationship or phase relationship between the input data can be additionally reflected in the series of input data to embed the series of input data. In one example, the relative positional relationship between the series of input data may include, but is not limited to, word order within a natural language sentence, the relative positional relationship between each segmented image, and the time order of segmented audio waveforms. The process of adding information expressing the relative positional relationship or phase relationship between the series of input data may be referred to as positional encoding.

[0265] In one embodiment, the model may include, but is not limited to, at least one of a Recurrent Neural Network (RNN), a Long Short Term Memory (LSTM) network, a Deep Neural Network (DNN), a Convolutional Neural Network (CNN), and a Bidirectional Recurrent Deep Neural Network (BRDNN). In one embodiment, the model may be a model trained by a transfer learning method. Here, transfer learning refers to a learning method in which a pre-trained model having a first task is obtained by pre-training a large amount of unlabeled training data by a semi-supervised learning or self-learning method, and the pre-trained model is fine-tuned to be suitable for a second task, and the model is trained by supervised learning on labeled training data to implement the target model.

[0266] Meanwhile, the disclosed embodiments may be implemented in the form of a recording medium storing computer-executable instructions. The instructions may be stored in the form of program code, and when executed by a processor, may generate program modules to perform the operations of the disclosed embodiments. The recording medium may be implemented as a computer-readable recording medium.

[0267] Computer-readable storage media include all types of storage media that store instructions that can be deciphered by a computer. Examples include read-only memory (ROM), random access memory (RAM), magnetic tape, magnetic disks, flash memory, and optical data storage devices.

[0268] The disclosed embodiments have been described with reference to the attached drawings as described above. Those skilled in the art will understand that the present disclosure can be implemented in forms other than the disclosed embodiments without altering the technical spirit or essential features of the present disclosure. The disclosed embodiments are illustrative and should not be construed as limiting.

Claims

1. A memory storing at least one command for generating a medical image interpretation text using LLM (Large Language Model); and A processor comprising: a processor that performs an operation according to the above command; The above processor, A method of training a text generation model by collecting a training data set including a text, a text template, and a data set of conclusions and opinions of the text, and fine-tuning the trained text generation model using at least one of the training data and synthetic data, The above processor, Through the above-mentioned reading text generation model, a template for generating TMN (Tumor, Node, Metastasis) for cancer is provided. When information is entered into the template provided above, keywords are extracted from the template input results, By inputting the extracted keywords as prompts into the above-mentioned reading text generation model, a reading text including findings, conclusions, and recommendations is automatically generated. Complete the above generated reading text, The combination of the above keywords and the generated reading text are used as a dataset for fine tuning. Medical image interpretation report generation device.

2. In paragraph 1, Among the above TMN, T(Tumor) contains information on the size and depth of invasion of the tumor, Among the above TMN, N (Node) includes information on the presence or absence of metastasis to nearby lymph nodes and the extent thereof. Among the above TMNs, M (Metastasis) includes information on whether there is distant metastasis. Medical image interpretation report generation device.

3. In paragraph 1, The above processor, Through the above-mentioned reading text generation model, the detailed items of the above template are configured based on the characteristics of the organ according to the type of cancer in the body, When reading a cancer tumor, the stage and keywords are extracted based on the morphology, location, size, tumor invasion, and TNM of the lesion. Medical image interpretation report generation device.

4. In paragraph 3, The above template is, Extract keywords through at least one of the user's voice, touch, and mouse clicks, Calculating cancer staging based on TNM, Medical image interpretation report generation device.

5. In paragraph 1, The above processor, The above-mentioned reading text generation model is trained with a structured learning data set that pairs conclusions and opinions, Generating a reading text that includes a detailed description and observed content based on the keyword through the above reading text generation model. Medical image interpretation report generation device.

6. In paragraph 1, The processor inputs reading data into the reading text generation model or extracts a conclusion from a template in which the information is input, and links with the LLM to generate an opinion (Finding) on ​​the conclusion. The above findings include interpretation findings including at least one of the abnormal findings found in the imaging examination and signs of a specific disease. Medical image interpretation report generation device.

7. In paragraph 1, The above processor tunes the LLM (Large Language Model) linked to the reading text generation model according to the mode based on a template applied differently depending on the cancer type, The above mode includes the type of cancer tumor, Medical image interpretation report generation device.

8. In paragraph 1, The above processor, When information is entered after image interpretation using a standardized input method, the entered image interpretation information is uploaded to a standardized list. Connecting the conclusion and opinion parts of the text template input to the above text generation model, and creating an opinion based on the learning data set through the text generation model. Medical image interpretation report generation device.

9. In paragraph 5, The above learning data set is, A dataset containing diagnoses corresponding to the findings of the conclusion and the basis and details matching the diagnosis, Medical image interpretation report generation device.

10. In paragraph 1, The above processor, Through the above-mentioned reading text generation model, the conclusion is summarized from the information after reading the input image, Through the above-mentioned reading text generation model, findings and diagnostic information are extracted from the above-mentioned summarized text, Through the above-mentioned reading text generation model, the basis and details matching the extracted findings and diagnostic information are written as opinions, Through the above reading text generation model, keywords included in the conclusion are recognized, Through the above-mentioned reading text generation model, the opinion of the reading text is generated by applying an opinion template matched in advance to the recognized keyword. Medical image interpretation report generation device.

11. In a method for generating a medical image interpretation report, performed by a processor of a device, A step of training a reading text generation model by collecting a learning data set including a reading text, a reading text template, and a conclusion and opinion data set of the reading text; A step of fine-tuning the learned reading text generation model using at least one of learning data and synthetic data; A step of providing a template for generating TMN (Tumor, Node, Metastasis) for cancer through the above-mentioned reading text generation model; When information is entered into the template provided above, a step of extracting keywords from the template input result; A step of automatically generating a reading text including findings, conclusions, and recommendations by inputting the extracted keywords as prompts into the reading text generation model; A step of completing the generated reading text; and A method comprising a step of utilizing the combination of the above keywords and the generated reading text as a dataset for fine tuning.

Citation Information

Patent Citations

  • Apparatus and Method for Generating Interpretation Text of Medical Image

    KR101881862B1

  • Display device having driving transistor to drive electro-optical element

    KR1020240154448A

  • Method for detecting abnormal findings and generating interpretation text of medical image

    KR102375786B1

  • KR20230018929A