Automated generation of text data from biomedical images
The integration of computer vision and natural language processing models in radiology practices addresses human error by providing real-time tumor detection and report generation, enhancing diagnostic accuracy and reducing errors in radiology workflows.
Patent Information
- Application Number
- PCT/US2025/014976
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-13
- Filing Date
- 2025-02-07
- Publication Date
- 2025-08-14
AI Technical Summary
Current radiology practices are prone to observational and interpretative errors, leading to human errors that contribute significantly to patient morbidity and mortality, particularly in cancer diagnosis and treatment, due to the lack of systematic methods to leverage multimodal AI models for real-time error detection and correction.
Integration of computer vision and natural language processing models to automatically detect tumors and discrepancies in radiology reports, providing real-time alerts and generating draft reports to enhance diagnostic accuracy and reduce human oversight.
The integrated AI pipeline significantly reduces human errors by proactively alerting radiologists to potential omissions and discrepancies, improving diagnostic accuracy and workflow efficiency, thereby enhancing patient outcomes.
Smart Images

Figure US2025014976_14082025_PF_FP_ABST
Abstract
Description
AUTOMATED GENERATION OF TEXT DATAFROM BIOMEDICAL IMAGESCROSS REFERENCES TO RELATED APPLICATIONS
[0001] The present application claims priority to U.S Provisional Patent Application No. 63 / 551,483, titled “Artificial Intelligence for Analyzing Biomedical Images and Reports,” filed February 8, 2024, and to U.S. Provisional Patent Application No. 63 / 564,910, titled “Automated Generation of Text Data from Biomedical Images,” filed March 13, 2024, each of which is incorporated by reference in its entirety.BACKGROUND
[0002] A computer system may apply one or more models on the input dataset to generate an output dataset.SUMMARY
[0003] Aspects of the present disclosure are directed to systems and methods for handling biomedical text and image data. One or more processors coupled with memory, can obtain data for a subject with a condition. The data can include (i) a biomedical image having a region of interest (ROI) corresponding to a first location associated with the condition in the subject and (ii) a report having text identifying a second location associated with the condition in the subject. The one or more processors can determine the first location corresponding to the ROI within the biomedical image using a first machine learning (ML) model. The one or more processors can identify the second location from the text of the report using a second ML model. The one or more processors can detect a deviation in the data for the subject responsive to the first location not corresponding to the second location. The one or more processors can store, using one or more data structures, an association between the data and an indication of the deviation.
[0004] In some embodiments, the one or more processors can provide, via an interface, an output including information associated with at least one of the biomedical image or the report, responsive to detecting the deviation. In some embodiments, the one ormore processors can receive, via the interface, feedback data identifying a third location associated with the condition in the subject, to update at least one of the first ML model, the second ML model, or the report.
[0005] In some embodiments, the one or more processors can detect a lack of any deviation in second data for a second subject responsive to a third location corresponding to a fourth location. The third location can be determined from a second biomedical image using the first ML model and the fourth location can be identified from text of a second ML model using the second ML model. In some embodiments, the one or more processors can store, using one or more second data structures, a second association between the second data and a second indication of the lack of any deviation.
[0006] In some embodiments, the one or more processors can identify second data comprising: (i) the first location corresponding to the ROI within the biomedical image determined using the first ML model and (ii) a plurality of second reports each having respective text identifying a respective location associated with the condition in the subject. In some embodiments, the one or more processors can provide the second data as input to a third ML model to generate a third report identifying a third location associated with the condition in the subject.
[0007] In some embodiments, the one or more processors can provide, via an interface, one or more user interface elements to accept or reject the third report. In some embodiments, the one or more processors can receive, via the one or more user interface elements of the interface, an indication of acceptance of the third report. In some embodiments, the one or more processors can identify the third report as the report to include in the data, responsive to receipt of the indication of acceptance.
[0008] In some embodiments, the one or more processors can provide, via an interface, the third report and one or more user elements to modify the third report. In some embodiments, the one or more processors can receive, via the one or more user interface elements of the interface, a modification to at least a portion of the third report. In some embodiments, the one or more processors can use the modification to at least the portion of the third report to update the third ML model.
[0009] In some embodiments, the one or more processors can receive, via an interface, the report generated by input from a user to identify the second location associated with the condition in the subject. In some embodiments, the one or more processors can generate, in accordance with a template defining locations, a first data structure identifying the first location in the subject. In some embodiments, the one or more processors can generate, in accordance with the template, a second data structure identifying the second location in the subject. In some embodiments, the one or more processors can detect the deviation by comparing the first data structure and the second data structure.
[0010] In some embodiments, the first ML model can include an image segmentation model established using a training dataset comprising a plurality of examples. Each example of the plurality of examples can identify: (i) a respective biomedical image having a respective ROI corresponding to a location associated with the condition in a respective subject and (ii) a respective annotation identifying the respective ROI within the respective biomedical image. The second ML model can include a natural language processing (NLP) model configured to extract the second location from the text of the report as associated with the condition within the subject. In some embodiments, the condition can include at least one of a tumor, a lesion, an infection, an inflammation, or cell damage in an anatomical site corresponding to at least one of the first location or the second location in the subject.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The foregoing and other objects, aspects, features, and advantages of the disclosure will become more apparent and better understood by referring to the following description taken in conjunction with the accompanying drawings, in which:
[0012] FIG. 1 is an example of three-dimensional lesion identifications after automatic linking of adjacent two-dimensional boxes;
[0013] FIG. 2 depicts a block diagram of a system for managing biomedical datasets comprising text and images, in accordance with an illustrative embodiment;
[0014] FIG. 3 depicts a block diagram of a process for obtaining biomedical datasets and determining locations of conditions from biomedical images in the system for managing biomedical datasets, in accordance with an illustrative embodiment;
[0015] FIG. 4 depicts a block diagram of a process for generating text reports using subject information and locations of conditions in the system for managing biomedical datasets, in accordance with an illustrative embodiment;
[0016] FIG. 5 depicts a block diagram of a process for comparing locations of conditions determined from biomedical images and from text reports in the system for managing biomedical datasets, in accordance with an illustrative embodiment;
[0017] FIG. 6 depicts a flow diagram of a process for detecting deviations between locations of conditions determined from biomedical images and from text reports manually generated by clinicians, in accordance with an illustrative embodiment;
[0018] FIG. 7 depicts a flow diagram of a process for automatically generating text reports from locations of conditions determined using biomedical images, in accordance with an illustrative embodiment;
[0019] FIG. 8 depicts a flow diagram of a method of managing biomedical datasets comprising text and images in accordance with an illustrative embodiment; and]0 20] FIG. 9 depicts a block diagram of a server system and a client computer system, in accordance with one or more implementations.DETAILED DESCRIPTION
[0021] Following below are more detailed descriptions of various concepts related to, and embodiments of, systems and methods for managing biomedical datasets comprising text and images. It should be appreciated that various concepts introduced above and discussed in greater detail below may be implemented in any of numerous ways, as the disclosed concepts are not limited to any particular manner of implementation. Examples of specific implementations and applications are provided primarily for illustrative purposes.
[0022] Section A describes a computer vision approach to detect and highlight lesions and tumors in radiology images.
[0023] Section B describes enhancing radiologic diagnosis through integrated computer vision and large language models.
[0024] Section C describes a radiology artificial intelligence assistant tool.
[0025] Section D describes systems and methods for handling biomedical text and image data.
[0026] Section E describes a network environment and computing environment, which may be useful for practicing various embodiments described herein.A. Computer Vision Approach to Detect and Highlight Lesions and Tumors in Radiology Images
[0027] The application of artificial intelligence (Al) to diagnostic radiology promises to revolutionize patient care by reducing detection errors, increasing accuracy, and improving the non-invasive characterization of disease. In cancer imaging, tumor segmentation represents a more accurate quantification of disease burden as compared with conventional linear measurement used by most radiologists in standard care and oncology clinical trials. The focus is to implement an Al model that automatically detects tumors, highlighting their locations on the radiology images and providing tumor volumes, which can be used by radiologists for improved diagnostic accuracy and precision.
[0028] Medical errors contribute to around 250,000 mortalities annually in the United States, thereby making them the third leading cause of death following heart disease and cancer. Radiology, despite its advancements, is particularly vulnerable to observational and interpretative errors, which can lead to dire patient consequences, especially in cases of overlooked cancer diagnoses in CT and MRI scans. Given the pivotal role radiologists play in cancer diagnosis and treatment response assessment, the increased diagnostic accuracy provided by this Al model is expected to improve patient outcomes.
[0029] Using computer vision techniques, tumors and their granular location may be detected. The tumor may be subsequently marked on the image to support the radiologist’s diagnostic assessment and provide automated volumetric measurements of tumor burden. Subsequently, the extracted information by the model can be used in the other Al-based models, for example to identify tumor genomic profiles and guide optimal therapy.
[0030] A vast collection of oncologic imaging data may be leveraged. Using computer vision models to identify tumors in MR brain images, the goal is to improve the model to detect and segment the tumor and its granular location. In the next step, detected tumors may be co-registered with the SRI24 Atlas, to detect the exact location and volume of the tumor in the brain. Different computer vision models, including Vision Transformers, CNN, segmentation, and semi-supervised methods, may be implemented to achieve this objective. Once validated, these findings and methodologies may be shared via publications and direct code dissemination.(0031 ] Image segmentation and detecting tumors based on two-dimensional images has been implemented using other approaches. However, in real life, there may be several images for the targeted organ, and if a computer vision model is to be implemented, just detecting a tumor based on a single radiology image may not be sufficient. Besides that, detecting the location of the tumor with high granularity is another challenge. Very few centers have both computer vision expertise and access to large databases of oncologic radiology images. In this present disclosure, a robust computer vision model may be built to obtain MRI images, detect tumor existence in the image, segment the tumor location, coregister that with the Brain atlas, and report the location of the tumor with high granularity (inside the organ, for example the frontal lobe).(0032] The specific aims may include: (1) implementing, testing and improving tumor detection models for CT and MRI; and (2) implementing a model that detects the location of the detected tumor by co-registering radiology images and the SRI24 brain atlas with high granularity. This means that once the tumor is detected, it analyzes whole images and tells which part of the organ the tumor is.
[0033] Radiologists currently produce reports based solely on their interpretation of images. This can lead to overlooked detections, which subsequently influence treatment plans. This may bring forth a novel Al-driven pipeline designed to pinpoint potential tumors and their locations in radiologist reports. This approach ensures enhanced accuracy and reliability in diagnosis since targeted parts of the images may be pinpointed for the radiologists. In the next step, it can be used to generate automated reports as well.
[0034] The pipeline may implement a simpler and faster diagnosis, a significant decrease in the number of undetected tumors and metastasis sites by radiologists, leading to decreased stress for medical professionals and potentially life-saving interventions for patients. The pipeline that obtains radiology images as soon as they are ready and processes them may be implemented, so the radiologist can look at both raw and processed images (e.g., as seen in FIG. 1), as well as metadata.
[0035] This computer vision model harnesses data-mined annotations from radiology images, deploying a self-supervised training paradigm to formulate tumor detection and segmentation models that are tailored for local radiology devices. This methodology triumphantly accomplishes a brain MRI tumor detection Fl score of 0.954 and a tumor segmentation average Dice score of 87%, all without any new manually curated training data. This model output for three-dimensional lesion identifications after automatic linking of adjacent two-dimensional boxes. The goal is to improve the model to be able to perform faster, such that it can be used in real time or close to real time (currently it takes several minutes for each MRI), and outputs the granular location of the tumor based on the images as a metadata.
[0036] A tumor may be detected, by detecting its granular location by co-registering the image with brain atlas images. The steps within the pipeline may include: a computer vision model detecting the existence of the tumor in radiology images; and detecting the exact location of the tumor by co-registering the image, with brain atlas images.
[0037] This model integration and application pipeline may be tested with radiologists, obtaining user validation and incorporating their valuable feedback to refine the application. The completed model pipeline may be used for implementation, for example, inhospitals nationwide, helping improve the diagnosis and care of the nearly 2 million new patients diagnosed with cancer annually in the United States, a number comparable to the population of Nebraska. With its potential to significantly enhance diagnostic accuracy, this application is poised to bring transformative change to healthcare.10038] In conclusion, the CV technologies may be amalgamated, resulting in a realtime enhanced intelligence tool that acts as an assistant for radiologists. The pioneering effort aims to drastically reduce human errors by proactively pinpointing radiologists of probable tumors and their exact location, as well as providing metadata for the tumor.B. Enhancing Radiologic Diagnosis Through Integrated Computer Vision and Large Language Models
[0039] Radiology reports occasionally contain human errors stemming from the interpretation of the corresponding images, with a prevalent error rate of 3%-5%. Yet, no systematized method exists to leverage multimodal Al models to detect such errors during clinical interpretation, enabling the radiologist to correct such mistakes in real-time.
[0040] Medical errors contribute to around 250,000 mortalities annually in the United States, thereby making them the third leading cause of death following heart disease and cancer. Radiology, despite its advancements, is particularly vulnerable to observational and interpretative errors, which can lead to dire patient consequences, especially in cases of overlooked cancer diagnoses in CT and MRI scans. Given the pivotal role radiologists play in cancer diagnosis and treatment response assessment, the increased diagnostic accuracy provided by this Al model is expected to improve patient outcomes.
[0041] Free-text radiology reports in oncology can be effectively processed by advanced language models to spot metastatic trends in real-time. Complemented by computer vision techniques, the aim may be to detect tumors and discrepancies in the radiology report, alerting radiologists promptly so these errors can be rectified.
[0042] The training of NLP and computer vision models for medical reports and images have relied on manual annotation by domain experts, a time-consuming and expensive process, but the number of annotations for radiology reports has increased at arapid rate at MSK. Furthermore, with recent advancements in BERT and dedicated language models trained in the medical domain, NLP has increased in accuracy for different types of tasks, such as classification and automated annotation. Using NLP in radiology reports to extract information can reduce the manual annotation effort and cost while increasing the availability of data for research. Similarly, the image annotations may be leveraged to train the computer vision models.
[0043] Very few centers have both the expertise in NLP, computer vision, and access to large databases of oncologic radiology images and reports. Metastatic patterns may be extracted from oncologic CT reports generated over a 10-year period, including data from nearly 100,000 patients. The CV-based approach may be used to detect and segment tumors in radiology images. Presented herein are the NLP and CV tools developed for processing radiology reports and images concurrently and presenting potentially missed lesions for a second review by the interpreting radiologist.
[0044] There may be multiple aims. First, to enhance and validate the language model to accurately detect cancer lesion data in CT and PET / CT reports. This phase may further the goals, targeting specific site optimization. Second, to enhance and validate the existing tumor detection models for CT and MRI. Third, develop a comprehensive machine learning pipeline harnessing the strengths of both the NLP and computer vision models. This integrated system may process images and reports simultaneously, notifying radiologists in real-time about potential discrepancies.
[0045] Radiologists produce reports based solely on their interpretation of images, without a second evaluation due to time and cost constraints. This can lead to overlooked detections, which subsequently influence treatment plans. The project brings forth a novel multimodal Al-driven pipeline designed to pinpoint potential inaccuracies in radiologist reports. This not only raises real-time alerts for a supplementary computer-aided examination but also acts as a collaborative tool that highlights potentially missed tumor lesions. The human-centric approach ensures a comprehensive solution, enhancing accuracy and reliability in results.[0046| This approach merges cutting-edge tumor detection computer vision models with advanced language models. This integration aims to identify and flag potential oversight in cancer lesion detection during the radiological process. The prime goal is to diminish human oversight in radiology report interpretations, facilitated by an intuitive and interactive application.
[0047] NLP, a branch of artificial intelligence that enables machines to comprehend, analyze, and generate human language, may be used. The NLP model preprocesses and classifies radiology reports. It leverages extensive language models to recognize complex patterns in the language used within these reports. This enables the detection of possible metastases in several sites based on the radiology reports, achieving an accuracy of over 95% in almost all instances (Table 1).
[0048] Table 1. NLP model performance in detecting metastasis in different sites, based on CT radiology notes:
[0049] The model reviews the radiologist’s report for overlooked tumor lesions and raises real-time alerts for a secondary, computer-aided review.
[0050] Under the CV approach, the computer vision model harnesses data-mined annotations from radiology images, deploying a self-supervised training paradigm to formulate tumor detection and segmentation models that are tailored for local radiology devices. This methodology triumphantly accomplishes a brain MRI tumor detection Fl score of 0.954 and a tumor segmentation average Dice score of 87%, all without any new manually curated training data.[0051 | Under the integrated approach: the strengths of the NLP and CV models can be combined. The user-friendly application can provide a more comprehensive and accurate review of both radiology reports and images. The steps within this pipeline may include: (1) the radiologist views the images and generates a draft report; (2) the radiologist initiates the Al application via a button press; and (3) the application processes the radiology report via the NLP model while the CV model examines the corresponding images.
[0052] If the system identifies a potential missed tumor, it alerts the radiologist in real-time with a pop-up window, prompting a closer review and offering a chance for immediate correction. If no discrepancies are identified, a pop-up window validates the report, confirming the accuracy of the radiologist’s interpretation.|0053| This model integration and application pipeline may be tested with these radiologists, obtaining user validation and incorporating their valuable feedback to refine the application. The completed model pipeline may be used for implementation in hospitals nationwide, helping improve the diagnosis and care of the nearly 2 million new patients diagnosed with cancer annually in the United States, a number comparable to the population of Nebraska. With its potential to significantly enhance diagnostic accuracy, the application is poised to bring transformative change to healthcare.
[0054] In conclusion, the NLP and CV technologies may be amalgamated, resulting in a real-time enhanced intelligence tool that acts as a secondary assessor for radiology reports. The effort may drastically reduce human errors by proactively alerting radiologists of probable omissions.C. Radiology Artificial Intelligence Assistant Tool|0055] Presented herein is a pipeline, integrating deep learning and language models to improve efficiency and accuracy in a radiologist’s workflow as well as generate robust structured data abstracting a patient’s MRIs. A universally compatible pipeline can be integrated across major clinical platforms and shared with various institutions, marking a pivotal step towards reducing radiologist human error.
[0056] Radiology imaging is important for disease diagnosis and treatment response monitoring. Despite the advent of computer-aided detection tools decades ago and recent advances in artificial intelligence (Al), most radiologists still interpret images and generate reports manually and without assistance from Al. This purely manual process leads to physician burnout and diagnostic errors that cause patient morbidity. A multimodal Al pipeline can abstract tumor lesions within radiology images, leveraging this information and the patient’s prior reports to generate a draft report for the radiologist. This approach may improve workflow efficiency for the radiologist and decrease diagnostic errors.
[0057] Improving the efficiency and accuracy of radiologist workflow may allow for better patient outcomes, reducing medical errors and allowing a quicker time to treatment by generating MRI reports in a shorter timeline. By building an integrated Al pipeline to support radiologists’ workflows, time to treatment and errors can be reduced, and a robust MRI database can be generated for further research. Within that pipeline, the computer vision techniques can be used to effectively detect tumors and generative Al techniques can efficiently generate draft radiology reports. The data engineering structure / pipeline and the tuning / building of the three model types may include computer vision, natural language processing and generative Al models.
[0058] Very few institutions have the expertise in NLP, computer vision models, generative Al models, and access to large databases of oncologic radiology images and reports. NLP and CV tools can be used for processing radiology images and previous reports to automatically prepare a draft report for radiologist review. The final report can be automatically analyzed to refine the NLP and CV models over time.
[0059] Radiologists produce reports based solely on their interpretation of images, typically without computer assistance or second evaluation due to time and cost constraints. This can lead to physician burnout and overlooked detections that can cause patient morbidity. This multimodal Al-driven pipeline can automatically generate a draft of the report based on the radiology images and previous reports. This approach can streamline the process for the radiologists and, by using the collaboration of Al and humans, decrease the risk of perceptual mistakes. The human-in-the-loop approach can ensure a comprehensive solution to enhance diagnostic speed, accuracy, and reliability. The automated pipeline canshorten radiology report turnaround time and improve diagnostic accuracy, translating to improved quality and timeliness of patient care.|0060] The pipeline can use computer vision (CV), natural language processing (NLP), generative artificial intelligence (Al), among others. The goal may have a sample prototype to begin working on change management and user adoption within the organization. The CV model can segment and detect tumor locations in the magnetic resonance images (MRIs). Mapping and image re-reregistering can be used to extract the exact location and size of the tumor. The generative Al can create a draft report from the locations identified using the computer vision. The NLP model extracts tumor locations from radiologist reports.
[0061] In conclusion, the advanced NLP and CV technologies can be integrated, resulting in a real-time artificial intelligence tool that acts as a secondary interpretation of radiology scans. Human errors can be reduced by proactively alerting radiologists to probable errors and omissions from their radiology reports.D. Systems and Methods for Handling Biomedical Text and Image Data
[0062] Referring now to FIG. 2, depicted is a block diagram of a system 100 for managing biomedical datasets comprising text and images. In overview, the system 100 can include at least one data processing system 105, at least one imaging device 110, at least one user device 115, and at least one database 120, communicatively coupled with one another via at least one network 125. The data processing system 105 can include at least one dataset indexer 130, at least one segment detector 135, at least one report generator 140, at least one text parser 145, at least one data comparator 150, at least one output handler 155, at least one computer vision (CV) model 160, at least one generative model 165, and at least one natural language processing (NLP) model 170, among others. Each of the components in the system 100 as detailed herein may be implemented using hardware (e.g., one or more processors coupled with memory), or a combination of hardware and software as detailed herein in Section B.
[0063] In further detail, the data processing system 105 may (sometimes herein generally referred to as a computing system or a server) be any computing device, includingone or more processors coupled with memory and software and capable of performing the various processes and tasks described herein. The data processing system 105 can be in communication with the imaging device 110, the user device 115, the database 120, and other devices via the network 125. The data processing system 105 may be situated, located, or otherwise associated with at least one server group. The server group may correspond to a data center, a branch office, or a site at which one or more servers corresponding to the data processing system 105 are situated. The data processing system 105 can perform or implement any of the functionalities detailed herein in conjunction with Sections A-C.
[0064] On the data processing system 105, the dataset indexer 130 can identify a biomedical data including a biomedical image and a text report for a subject with a condition. The segment detector 135 can use the CV model 160 to determine a location corresponding to the condition within the biomedical image. The report generator 140 can use the generative model 165 to generate a new text report using the location determined using the CV model 160. The text parser 145 can use the NLP model 170 to identify the location corresponding to the condition from the text of the report. The data comparator 150 can determine whether the location determined from the biomedical image corresponds to the location identified from the text report. The output handler 155 can provide information based on the comparison between the location determined from the biomedical image corresponds to the location identified from the text report.10065] The CV model 160 can be any type of machine learning algorithm or model to determine locations (e.g., image pixels or anatomical sites) corresponding to regions of interest (ROIs) from biomedical images. The CV model 160 can be maintained on the data processing system 105 or another computing device. The CV model 160 can be, for example, a deep learning artificial neural network (ANN), such as an encoder-decoder model with a convolutional neural network architecture. In general, the CV model 160 can have a biomedical image in any modality from a subject as an input, an identification of the ROI within the biomedical image as an output, and a set of weights relating the input to the output, among others. The CV model 160 may have been initialized, trained, and established using a training dataset in accordance with learning techniques (e.g., supervised or semi-supervised). The training dataset can include or identify a set of examples. Each example can include arespective biomedical image and an annotation defining the ROI within the respective biomedical image. In some embodiments, the annotation can define the anatomical site associated with the ROI.
[0066] The generative model 165 can be a generative artificial intelligence (Al), such as a large language model (LLM), to produce or generate text reports from input data. The generative model 165 can be maintained on the data processing system 105 or another computing device. The generative model 165 can be, for example, comprised of a transformer network (e.g., bidirectional encoder representations from transformers (BERT) and generative pre-trained transformer (GPT)) to process sequential data. In general, the generative model 165 can have inputs and outputs related by a set of weights. The input can include a data structure identifying a location of a condition in a subject, previously generated reports, and profile information associated with the subject as input. The output can include a newly generated report. The set of weights in the generative model 165 can be in accordance with the transformer network architecture, and can include an encoder and decoder, each with self-attention and feedforward mechanisms. The generative model 165 can be initialized, trained, and established using a corpus of text. The generative model 165 can learn the likelihood of a sequence of texts from the corpus of text. In some embodiments, the generative model 165 may have been pre-trained using a generalized corpus of text and then fine-tuned with a knowledge-specific corpus of text (e.g., previously created text reports).
[0067] The NLP model 170 can extract or identify the location corresponding to a condition within a subject from a text report. The NLP model 170 can be maintained on the data processing system 105 or another computing device. The NLP model 170 can be any machine learning model to extract information from data and can include, for example, named entity recognition (NER), dependency parsing, relation extraction, and template filling, among others. The NLP model 170 can include the text report as input, the location corresponding to the condition as the output, and a set of weights, among others. The set of weights in the NLP model 170 can be in accordance with any number of model architectures, such as neural networks (e.g., recurrent neural networks), hidden Markov models, transformer networks, Bayes classifiers, or condition random fields, among others.
[0068] The imaging device 110 (sometimes herein generally referred to as an imaging device or an image acquirer) may be any device to acquire biomedical images of subjects. The biomedical can be a tomogram acquired in accordance with a tomographic imaging technique, such as a magnetic resonance imaging (MRI) scanner, a nuclear magnetic resonance (NMR) scanner, an X-ray computed tomography (CT) scanner, an ultrasound imaging scanner, a positron emission tomography (PET) scanner, and a photoacoustic spectroscopy scanner, among others. Although primarily discussed in terms of a tomogram and MRI images in particular, other imaging modalities besides those listed above may be supported by the data processing system 105 for acquiring the biomedical image. The imaging device 110 can be in communication with the data processing system 105 and the user device 115 to provide acquired biomedical images.
[0069] The user device 115 (sometimes herein referred to as an operator device or a clinician device) can be any computing device comprising one or more processors coupled with memory and software and capable of performing the various processes and tasks described herein. The user device 115 can be associated with an entity (e.g., a clinician) examining the subject or biomedical images from the subject. The user device 115 can be in communication with the data processing system 105 and the imaging device 110 to exchange data. The user device 115 can display biomedical images acquired from the imaging device 110. The user device 115 can be used to input to create text reports for the biomedical images. The user device 115 can be used to provide feedback to the data processing system 105.
[0070] Referring now to FIG. 3, among others, depicted is a block diagram of a process 200 for obtaining biomedical datasets and determining locations of conditions from biomedical images in the system 100 for managing biomedical datasets. The process 200 can include or correspond to operations performed in the system 100 to obtain a dataset for subjects and determine locations of conditions from a biomedical image in the dataset. Under the process 200, the dataset indexer 130 executing on the data processing system 105 can retrieve, identify, or otherwise obtain at least one dataset 205 for at least one subject 210. The subject 210 may be a human or animal subject, among others. The subject 210 may have, may be at risk of, or may be afflicted with at least one condition on at least a site 215.The condition may include, for example, a tumor (e.g., from a type of cancer, such as breast cancer, bladder cancer, cervical cancer, colorectal cancer, kidney cancer, liver cancer, lung cancer, lymphoma, ovarian cancer, prostate cancer, skin cancer, or thyroid cancer), an infection (e.g., bacterial, viral, or fungal), an inflammation (e.g., acute or allergic), or other cell damage (e.g., necrosis, apoptosis, stress, radiation, or lesion), among others. The site 215 can correspond to an anatomical location, such as an organ (e.g., a brain, lung, heart, kidney, breast, prostate, ovary, pancreas, stomach, esophagus, bone, or epidermis) or a specific portion of the organ (e.g., frontal lobe, parietal lobe, occipital lobe, and temporal lobe of the brain or renal capsule, renal cortex, renal medulla, renal pelvis, and particular arteries of the kidney) of the subject 210.[00711 The dataset 205 can identify or include at least one biomedical image 220 (sometimes herein referred to as a tomogram or an image) of the subject 210. The biomedical image 220 can be obtained, produced, or otherwise acquired via the imaging device 110 and can be received by the dataset indexer 130 from the imaging device 110. The imaging device 110 can perform a scan of a section or a volume of the subject 210 using any number of imaging modalities or techniques (e.g., MRI or CT) to generate the biomedical image 220. The volume can include multiple organs in the subject 210. The biomedical image 220 can include a set of two-dimensional cross-sections (e.g., a front, a sagittal, a transverse, or an oblique plane) acquired from the three-dimensional volume. In some embodiments, the biomedical image 220 may be part of a video acquired of the sample over time. For example, the biomedical image 220 may correspond to a single frame of the video acquired of the sample over time at a frame rate. The biomedical image 220 may be maintained using one or more files in accordance with a format (e.g., single-file or multi-file DICOM format).
[0072] The biomedical image 220 can have at least one region of interest (RO I) 225 (also referred to herein as a structure of interest (SOI) or feature of interest (FOI)). The ROI 225 can correspond to at least a portion of an anatomical location (e.g., an organ) associated with the condition (e.g., tumor or lesion) within the subject 210. The ROI 225 can correspond to an area, section, or part of the biomedical image 220 that corresponds to the presence of the condition in the scanned section or volume of the subject 210 from which thebiomedical image 220 is acquired. For example, for a CT scan of the brain, the ROI 225 may correspond to a portion of the biomedical image 220 depicting a tumorous growth in a front lobe of the brain of a human subject. For an MRI image of the breast, the ROI 225 may correspond to a portion of the biomedical image 220 depicting a lesion in a lobule of the breast of a human subject.
[0073] The dataset 205 can identify or include at least one report 230. The report 230 can have or include text identifying at least a portion of an anatomical location (e.g., an organ) associated with the condition (e.g., tumor or lesion) within the subject 210. The report 230 can be composed, created, or otherwise generated via the user device 115. For instance, a clinician examining the biomedical image 220 can input or enter the text via an interface (e.g., graphical user interface) presented on the user device 115 to draft a radiology report corresponding to the report 230 about the subject 210. The report 230 generated via the user device 115 can identify an anatomical location corresponding or not corresponding to the actual site 215 in the subject 210. The report 230 can be retrieved, identified, or otherwise received by the dataset indexer 130 from the user device 115. The text of the report 230 may be unstructured (e.g., free-form text) or structured (e.g., populated in field-value format in accordance with a template). The report 230 may be maintained using one or more files in accordance with a format (e.g., a document format). In some embodiments, the dataset 205 may not have or lack the report 230 generated via the user device 115.
[0074] The segment detector 135 executing on the data processing system 105 can apply or use the CV model 160 to detect, identify, or determine at least one location 235 A corresponding to the ROI 225 in the biomedical image 220. The location 235 A can identify a portion of the biomedical image 220 corresponding to the ROI 225 associated with the condition in the subject 210. In some embodiments, the location 235 A can identify an anatomical site within the subject 210 associated with the condition. For example, the location 235 A can be mapped or registered by the text parser 145 to a model of an organ (e.g., brain atlas) associated with the anatomical site. The model may be a computer representation (e.g., stored and maintained on the data processing system 105) of the organ and used to co-register locations identified from the biomedical images with a standardized anatomical reference within the associated organ. The model may define locations in theorgan using a coordinate system (e.g., x, y, and z), to which the location 235A (or the biomedical image 220) may be mapped or registered. The model may be used to compare the locations derived from the images with locations extracted from text reports. To determine, the segment detector 135 can feed or apply the biomedical image 220 into the CV model 160. The segment detector 135 can process the biomedical image 220 in accordance with the set of weights of the CV model 160. From processing using the set of weights of the CV model 160, the segment detector 135 can output, produce, or otherwise generate the location 235 A. In some embodiments, the segment detector 135 can output, produce, or otherwise generate a segmented biomedical image. The segmented biomedical image can include a mark-up (e.g., highlight) of the location 235 A corresponding to the ROI 225 within the biomedical image 220.
[0075] With the determination of the location 235 A, the segment detector 135 can output, produce, or otherwise generate at least one data structure 240A. The data structure 240 A can be a standardized format for identifying the location 235 A with respect to subjects (e.g., the subject 210). For example, the data structure 240 A can be a table or matrix identifying anatomical sites for human subjects, one of which corresponds to the location with the condition (e.g., the tumor or lesion). The standardized format can facilitate comparison of data structures (e.g., the data structure 240A) with one another. The segment detector 135 can generate the data structure 240 A in accordance with a template defining anatomical locations within the subject 210. The template can specify, identify, or otherwise define the format for the data structure 240A. The data structure 240A can identify the location 235 A determined using the CV model 160.
[0076] Referring now to FIG. 4, among others, depicted is a block diagram of a process 300 for generating text reports using subject information and locations of conditions in the system 100 for managing biomedical datasets. The process 300 can include or correspond to operations performed in the system 100 to generate text reports using generative Al and interface with user devices to apply feedback to the text reports. Under the process 300, the report generator 140 executing on the data processing system 105 can retrieve, obtain, or otherwise identify a set of reports 230’ A-N (hereinafter generally referred to as reports 230’) from the database 120. Each report 230’ may have been previouslygenerated for the subject 210 in a similar manner as the report 230. For instance, each report 230’ may be composed by a clinician examining a respective, previously acquired biomedical image from the subject 210. The clinician can use an interface (e.g., graphical user interface) presented on the user device 115 to compose the report 230’. The report 230’ generated by the clinician can also identify an anatomical location corresponding or not corresponding to the actual site 215 in the subject 210. Once generated, the set of reports 230’ can be stored and maintained on the database 120.
[0077] In addition, the report generator 140 can retrieve, obtain, or otherwise identify profile information 305 associated with the subject 210 from the database 120. The profile information 305 can identify or include various information associated with the subject 210. For instance, the profile information 305 can include a patient identifier, an identifier of the condition, and traits (e.g., age, gender, background, and location), among others. The profile information 305 can be stored and maintained on the database 120. The report generator 140 can also retrieve, receive, or otherwise identify the location 235 A determined using the biomedical image 220. In some embodiments, the report generator 140 can process or parse the data structure 240A to extract, determine, or otherwise identify the location 235 A. In some embodiments, the report generator 140 can identify a dataset including the location 235A, the reports 230’, or the profile information 305, among others.
[0078] With the identifications, the report generator 140 can use the generative model 165 to produce, output, or otherwise generate at least one report 230”. In some embodiments, the report generator 140 can provide the location 235A, the reports 230’, and the profile information 305 as input to the generative model 165. For example, the report generator 140 can generate an input prompt in accordance with a template using the location 235 A, the reports 230’, and the profile information 305, and provide the input prompt as input into the generative model 165. The report generator 140 can process the input (e.g., the location 235A, the reports 230’, and the profile information 305) according to the set of weights of the generative model 165. From processing, the report generator 140 can output, produce, or otherwise generate the report 230”. The report 230” can be of a similar form as the reports 230 or 230’. The report 230” can have or include text identifying at least a portion of an anatomical location (e.g., an organ) associated with the condition (e.g., tumor or lesion)within the subject 210. The report 230” generated via the user device 115 can identify an anatomical location corresponding or not corresponding to the actual site 215 in the subject 210. The text of the report 230” may be unstructured (e.g., free-form text) or structured (e.g., populated in field-value format in accordance with a template). The report 230” may be maintained using one or more files in accordance with a format (e.g., a document format).
[0079] The report generator 140 can transmit, send, or otherwise provide the report 230” generated using the generative model 165 to the user device 115 for presentation via at least one interface 310. The biomedical image 220 (or a marked-up version of the biomedical image) can be also be provided with the report 230”. The interface 310 can include a set of user interface elements to accept, reject, or modify the report 230”. For instance, the clinician can use the user interface 310 to view the biomedical image 220 and inspect the contents of the report 230” for correctness. Using the interactions with the user interface elements of the interface 310, the user device 115 can create, produce, or otherwise generate at least one feedback 315. When the interactions indicate acceptance, the user device 115 can generate the feedback 315 to indicate acceptance of the report 230”. When the interactions indicate rejection, the user device 115 can generate the feedback 315 to indicate rejection of the report 230”. When the interactions include modifications to the report 230”, the user device 115 can generate the feedback 315 to include or identify the modifications. Upon entry, the user device 115 can return, send, or otherwise provide the feedback 315 to the report generator 140.
[0080] The report generator 140 can retrieve, identify, or otherwise identify feedback 315 from the user device 115. Upon receipt, the report generator 140 can process or parse the feedback 315. When the feedback 315 indicates acceptance, the report generator 140 can determine that the report 230” generated by the generative model 165 is correct. The report generator 140 can also use or identify the report 230” to include in the dataset 205 (e.g., to be used as the report 230). When the feedback 315 indicates rejection, the report generator 140 can refrain from using the report 230” generated by the generative model 165. In some embodiments, the report generator 140 can provide an indication of the rejection to the generative model 165 to regenerate the report 230”. The report generator 140 can repeat the process detailed herein with respect to the re-generated report 230”. When the feedback 315indicates modification, the report generator 140 can apply the modification to the report 230” to edit, alter, or otherwise change the text of the report 230”. Once applied, the report generator 140 can use or identify the report 230” to include in the dataset 205 (e.g., to be used as the report 230). In some embodiments, the report generator 140 can provide the modification to the report 230” to the generative model 165 to regenerate the report 230”. The report generator 140 can repeat the process detailed herein with respect to the regenerated report 230”.
[0081] Referring now to FIG. 5, among others, depicted is a block diagram of a process 400 for comparing locations of conditions determined from biomedical images and from text reports in the system 100 for managing biomedical datasets. The process 400 can include or correspond to operations performed in the system 100 to detect deviations in locations determined using text reports and images of subjects and provide output including information about the text reports and images. Under the process 400, the text parser 145 executing on the data processing system 105 can use the NLP model 170 to extract or identify at least one location 235B from the text of the report 230. The report 230 can be manually generated (e.g., by a clinician examining the biomedical image) or automatically generated (e.g., using the generative model 165). The location 235B can identify an anatomical site within the subject 210 associated with the condition. In some embodiments, the location 235B can be mapped or registered by the text parser 145 to a model of an organ (e.g., brain atlas) associated with the anatomical site. The model may be a computer representation of the organ, and used to co-register locations identified from reports with a standardized anatomical reference within the associated organ. The model may define locations in the organ using a coordinate system (e.g., x, y, and z), to which the location 235B may be mapped or registered. To identify, the text parser 145 can feed or apply the biomedical image 220 into the NLP model 170. The text parser 145 can process the report 230 in accordance with the set of weights of the NLP model 170. From processing using the set of weights of the NLP model 170, the text parser 145 can output, produce, or otherwise generate the location 235B.
[0082] With the determination of the location 235B, the text parser 145 can output, produce, or otherwise generate at least one data structure 240B. The data structure 240B canbe a standardized format for identifying the location 235B with respect to subjects (e.g., the subject 210). For example, the data structure 240B can be a table or matrix identifying anatomical sites for human subjects, one of which corresponds to the location with the condition (e.g., the tumor or lesion). The standardized format can facilitate comparison of data structures (e.g., the data structure 240B) with one another. The text parser 145 can generate the data structure 240B in accordance with a template defining anatomical locations within the subject 210. The template can specify, identify, or otherwise define the format for the data structure 240B. The data structure 240B can identify the location 235B determined using the NLP model 170.
[0083] The data comparator 150 executing on the data processing system 105 can determine, identify, or detect whether there is at least one deviation in the dataset 205 based on a comparison between the location 235A and the location 235B. In some embodiments, the data comparator 150 can compare the data structure 240 A and the data structure 240B to determine whether there is a deviation between the biomedical image 220 and the report 230 of the dataset 205. From comparing, the data comparator 150 can determine whether the location 235 A corresponds to (e.g., matches) the location 235B. In some embodiments, the data comparator 150 can use the model for the organ to compare the location 235 A and the location 235B in the computer representation of the organ. In some embodiments, the data comparator 150 can determine whether the location 235 A corresponds to (e.g., matches) the location 235B based on a threshold distance. The threshold distance may define a value for the distance between the two locations at which the locations are determined to be corresponding. The threshold distance may have any range, for example, between 0.1 mm and 5 cm. The data comparator 150 may calculate or determine a distance between the location 235A and the location 235B. With the determination, the data comparator 150 may compare the distance with the distance threshold. When the distance is greater than the distance threshold, the data comparator 150 may determine that the location 235 A does not correspond to the location 235B. Conversely, when the distance is less than or equal to the distance threshold, the data comparator 150 may determine that the location 235 A corresponds to the location 235B.
[0084] Based on the determination, the data comparator 150 can produce, output, or otherwise generate at least one indication 405. When the location 235 A does not correspond to the location 235B, the data comparator 150 can detect the occurrence of the deviation in the dataset 205. The data comparator 150 can generate the indication 405 to indicate the occurrence of the deviation (or discrepancy). The data comparator 150 can also store and maintain an association between the dataset 205 and the indication 405 of the occurrence of the deviation using one or more data structures (e.g., table, matrix, linked list, array, binary tree, heap, hash table, or another object) on the database 120. Conversely, when the location 235 A corresponds to the location 235B, the data comparator 150 can detect an absence of the deviation in the dataset 205. The data comparator 150 can generate the indication 405 to indicate the absence of the deviation. The data comparator 150 can also store and maintain an association between the dataset 205 and the indication 405 of the absence of the deviation using one or more data structures on the database 120.|0085] The output handler 155 executing on the data processing system 105 can create, produce, or otherwise generate at least one output 410. The output 410 can include information associated with the biomedical image 220 or the report 230, among others. The biomedical image 220 can include or identify at least one mark-up 415 identifying the ROI 225 therein, determined using the CV model 160. The output 410 can include or identify the indication 405 of the occurrence or the absence of the deviation. In some embodiments, the output handler 155 can generate the output 410, with the detection of the occurrence of the deviation in the dataset 205. The output 410 may serve as an alert to notify discrepancies between the location 235B identified in the report 230 versus the location 235 A determined from the CV model 160. In some embodiments, the output handler 155 can generate the output 410, independent of the detection of the occurrence of the deviation.
[0086] With the generation, the output handler 155 can send, transmit, or otherwise provide the output 410 to the user device 115 for presentation via the interface 310. The interface 310 can include a set of user interface elements to provide feedback on the information on the output 410, such as acceptance or modification of the report 230 or the locations 235A or 235B associated with the condition in the subject 210. For instance, the clinician can use the user interface 310 to view the biomedical image 220 and inspect thecontents of the report 230 to investigate the occurrence of the deviation. Using the interactions with the user interface elements of the interface 310, the user device 115 can create, produce, or otherwise generate at least one feedback 415. When the interactions with the interface 310 indicate acceptance (e.g., if there is no deviation), the user device 115 can generate the feedback 415 to indicate acceptance of the output 410. When the interactions include modifications to the output (e.g., to the report 230 or the location 235A or 235B to correct the deviation), the user device 115 can generate the feedback 415 to include or identify the modifications. Upon entry, the user device 115 can return, send, or otherwise provide the feedback 415 to the output handler 155.
[0087] The output handler 155 can retrieve, identify, or otherwise identify feedback 415 from the user device 115. Upon receipt, the output handler 155 can process or parse the feedback 415. When the feedback 415 indicates acceptance, the output handler 155 can determine that the output 410 is correct. When the feedback 415 indicates modification of the report 230 (or the location 235B), the output handler 155 can determine that the report 230 (or the location 235B) is incorrect. The output handler 155 can also apply the modification to edit, alter, or otherwise change the report 230 (e.g., with the correct anatomical location). When the feedback 415 indicates modification of the location 235A, the output handler 155 can determine that the location 235 A determined using the biomedical image 220 is incorrect. In some embodiments, when the feedback 415 indicates modification to the output 410, the output handler 155 can modify, change, or otherwise update the weights of at least one of the models, such as the CV model 160, the generative model 165, or the NLP model 170. The updating can be in accordance with the network architecture for the model. In some embodiments, the output handler 155 can store and maintain the dataset 205 (including the modifications to the report 230) on the database 120.
[0088] In this manner, the data processing system 105 can process data from different modalities in the form of reports 230 (e.g., in text) and biomedical images 220 (e.g., an image), without the reliance on having separate systems to process the data. The data processing system 105 can detect deviations between the locations 235 A and 235B determined from the biomedical image 220 and the report 230 respectively. The data processing system 105 can also use feedback 315 or 415 to correct or emend the text of thereport 230, thereby improving the accuracy and quality of the datasets 205 stored and maintained on the database 120. By improving the accuracy and quality of datasets 205, the data processing system 105 can reduce the consumption of computing resources (e.g., processing and memory), which would have otherwise been spent on providing faulty data and manually correcting locations 235A and 235B. By also being able to process data of different modalities, the data processing system 105 can reduce the consumption of network bandwidth from data transfers that would have been used to send the data in different modalities to separate computing systems for processing. The data processing system 105 can also use the generative model 165 to create newly generated reports 230’ to reduce the amount of time and effort exerted by users in creating reports 230 about the subjects 210. The reduction in time and effort can also improve the quality of human-computer interactions (HCI) between the user and the user device 115 and the overall system 100.
[0089] Furthermore, the functionalities of the data processing system 105 can be used to also improve clinical outcomes. For example, the subject 210 may be diagnosed with or evaluated for a cancer (or another condition), and a scan of the subject 210 at the anatomical site 215 associated with the cancer may be taken for further examination. The scan may be used to create the biomedical image 220 of the anatomical site 215 to be used by a radiologist to evaluate the state of the cancer within the subject 210. In this scenario, there may be a risk of observational and interpretative errors, leading to missed diagnoses and incorrect treatment plans. The CV model 160 may be used to identify a granular location (e.g., the location 235 A) of the tumor depicted the biomedical image 220, and the identified location in turn may be mapped to the anatomical site 215. The generative model 165 may be used to automatically generate draft radiology reports (e.g., the report 230”) based on the detected tumor locations. This can significantly reduce the time radiologists spend on report writing and allows them to focus on reviewing and validating the reports. By automating the detection and reporting processes, the data processing system 105 can decrease the turnaround time for radiology reports.
[0090] In addition, with respect to report manually drafted by the radiologist (e.g., the report 230), the NLP model 170 may be used to extract the location of the tumor (e.g., the location 235B) as indicated in the report. With the extraction, the data processing system 105can determine whether there are any discrepancies or deviations between the location indicated in the report and the location determined from the CV model 160. If there are deviations, an alert may be provided to the radiologist to notify the radiologist. The alert may serve to notify of potential errors and to ensure that the error be corrected immediately, improving diagnostic accuracy. The radiologist can also take these discrepancies to update the treatment plan (e.g., adjusting dosage and target area for radiotherapy). The data processing system 105 can thus act as a secondary assessor for radiology reports, providing a comprehensive review of both the scanned images and the text reports, thus reducing the risk of human oversight. The radiologist (or the user) can also provide feedback 315 or 415 in response to outputs from the models. The feedback 315 or 415 can be incorporated to continuously update and improve the machine learning and generative models, thereby enhancing the reliability and increasing accuracy of the final reports. The data processing system 105 can thus improve clinical outcomes by enabling quicker and more accurate diagnoses.[0091 | Referring now to FIG. 6, among others, depicted is a flow diagram of a process 500 for detecting deviations between locations of conditions determined from biomedical images and from text reports manually generated by clinicians. The process 500 can be implemented or performed using any components detailed herein, such as the system 100 or a server system 800. Under the process 500, a patient may undergo a scan (e.g., a magnetic resonance imaging (MRI)) (502) to produce a biomedical image 504 (e.g., an MRI image). The biomedical image may be assigned to a clinician (e.g., a radiologist) for review (506). A computing system can use a computer vision model to identify tumor locations within the biomedical imager (508). The computing system can use a mapping to generate structured data (510) to output a data structure 512 to identify tumor locations.[00921 In conjunction, the clinician can review the biomedical image (514) and can create a draft report 516. The computing system can use a natural language processing (NLP) model to extract tumor locations from the draft report (518) to output a data structure 520 to identify tumor locations. The computing system can compare the data structures 512 and 520 for discrepancies (522). The computing system can use mapping to mark the biomedical image with discrepancies (524) to generate a marked-up image 526 with discrepancies. Theclinician can review for additional discrepancies and can override upon selection (528). The computing system can generate a final report (530).
[0093] Referring now to FIG. 7, among others, depicted is a flow diagram of a process 600 for automatically generating text reports from locations of conditions determined using biomedical images. The process 600 can be implemented or performed using any components detailed herein, such as the system 100 or the server system 800. Under the process 600, a patient may undergo a scan (e.g., a magnetic resonance imaging (MRI)) to produce a biomedical image (e.g., an MRI image) (602). A computing system can use a computer vision model to identify tumor locations within the biomedical imager (604). The computing system can use a mapping to generate structured data (606) to output a data structure 608 to identify tumor locations. The computing system can also generate a marked- up biomedical image 610.
[0094] The computing system can identify previous reports for the patient (612). The computing system can use a generative artificial intelligence (Al) model (614) to produce a draft report 616. A radiologist may review the marked-up biomedical image (618) to produce a modified report 620. The computing system can use a natural language processing (NLP) model to extract locations from the report (622) to output a data structure 624 to identify tumor locations. The computing system can compare the data structures 608 and 624 for improving various models (626).
[0095] Referring now to FIG. 8, among others, depicts a flow diagram of a method 700 of managing biomedical datasets comprising text and images. The method 700 can be implemented or performed using any components detailed herein, such as the system 100 or the server system 800. Under the method 700, a computing system can identify a biomedical image (705). The computing system can determine a location of a region of interest (ROI) from a biomedical image (710). The computing system can identify previous reports (715). The computing system can generate a report (720). The computing system can identify a location of the ROI from text of a report (725). The computing system can detect whether is a deviation between the location of ROI identified from the biomedical image and the location of the ROI from the text of the report (730). If there is a correspondence between the location of ROI identified from the biomedical image and the location of the ROI fromthe text of the report, the computing system can detect an occurrence of the deviation (735). Otherwise, if there is no correspondence, the computing system can detect an absence of the deviation (740). The computing system can store an association (745). The computing system can provide an output about the report to a user device (750). The computing system can receive feedback from a user device (755). The computing system can update a report or models (760).E. Computing and Network Environment
[0096] Various operations described herein can be implemented on computer systems. FIG. 9 shows a simplified block diagram of a representative server system 800, client computing system 814, and network 826 usable to implement certain embodiments of the present disclosure. In various embodiments, the server system 800 or similar systems can implement services or servers described herein or portions thereof. Client computing system 814 or similar systems can implement clients, described herein. The systems 100 described herein can be similar to the server system 800. The server system 800 can have a modular design that incorporates a number of modules 802 (e.g., blades in a blade server embodiment); while the two modules 802 are shown, any number can be provided. Each module 802 can include processing unit(s) 804 and local storage 806.
[0097] The processing unit(s) 804 can include a single processor, which can have one or more cores, or multiple processors. In some embodiments, the processing unit(s) 804 can include a general-purpose primary processor as well as one or more special-purpose coprocessors, such as graphics processors, digital signal processors, or the like. In some embodiments, some or all the processing units 804 can be implemented using customized circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In other embodiments, the processing unit(s) 804 can execute instructions stored in local storage 806. Any type of processors in any combination can be included in the processing unit(s) 804.
[0098] Local storage 806 can include volatile storage media (e.g., DRAM, SRAM, SDRAM, or the like) and / or non-volatile storage media (e.g., magnetic or optical disk, flashmemory, or the like). Storage media incorporated in local storage 806 can be fixed, removable, or upgradeable as desired. Local storage 806 can be physically or logically divided into various subunits, such as a system memory, a read-only memory (ROM), and a permanent storage device. The system memory can be a read-and-write memory device or a volatile read-and-write memory, such as dynamic random-access memory. The system memory can store some or all of the instructions and data that the processing unit(s) 804 need at runtime. The ROM can store static data and instructions that are needed by the processing unit(s) 804. The permanent storage device can be a non-volatile read-and-write memory device that can store instructions and data even when the module 802 is powered down. The term “storage medium” as used herein, includes any medium in which data can be stored indefinitely (subject to overwriting, electrical disturbance, power loss, or the like) and does not include carrier waves and transitory electronic signals propagating wirelessly or over wired connections.
[0099] In some embodiments, local storage 806 can store one or more software programs to be executed by the processing unit(s) 804, such as an operating system and / or programs implementing various server functions, such as functions of the system 100 or any other system described herein, or any other server(s) associated with system 100 or any other system described herein.
[0100] Software” refers generally to sequences of instructions that, when executed by the processing unit(s) 804, cause the server system 800 (or portions thereof) to perform various operations, thus defining one or more specific machine embodiments that execute and perform the operations of the software programs. The instructions can be stored as firmware residing in read-only memory and / or program code stored in non-volatile storage media that can be read into volatile working memory for execution by the processing unit(s) 804. Software can be implemented as a single program or a collection of separate programs or program modules that interact as desired. From local storage 806 (or non-local storage described below), the processing unit(s) 804 can retrieve program instructions to execute and data to process in order to execute various operations described above.
[0101] In some the server systems 800, multiple modules 802 can be interconnected via a bus or other interconnect 808, forming a local area network that supportscommunication between the modules 802 and other components of the server system 800.The interconnect 808 can be implemented using various technologies including server racks, hubs, routers, etc.
[0102] A wide area network (WAN) interface 810 can provide data communication capability between the local area network (interconnect 808) and the network 826, such as the Internet. Technologies can be used, including wired (e.g., Ethernet, IEEE 802.3 standards) and / or wireless technologies (e.g., Wi-Fi, IEEE 802.11 standards).
[0103] In some embodiments, local storage 806 is intended to provide working memory for the processing unit(s) 804, providing fast access to programs and / or data to be processed while reducing traffic on the interconnect 808. Storage for larger quantities of data can be provided on the local area network by one or more mass storage subsystems 812 that can be connected to the interconnect 808. The mass storage subsystem 812 can be based on magnetic, optical, semiconductor, or other data storage media. Direct attached storage, storage area networks, network-attached storage, and the like can be used. Any data stores or other collections of data described herein as being produced, consumed, or maintained by a service or server can be stored in the mass storage subsystem 812. In some embodiments, additional data storage resources may be accessible via the WAN interface 810 (potentially with increased latency).
[0104] The server system 800 can operate in response to requests received via the WAN interface 810. For example, one of the the modules 802 can implement a supervisory function and assign discrete tasks to other the modules 802 in response to received requests. Work allocation techniques can be used. As requests are processed, results can be returned to the requester via the WAN interface 810. Such operation can generally be automated. Further, in some embodiments, the WAN interface 810 can connect multiple server systems 800 to each other, providing scalable systems capable of managing high volumes of activity. Other techniques for managing server systems and server farms (collections of server systems that cooperate) can be used, including dynamic resource allocation and reallocation.
[0105] The server system 800 can interact with various user-owned or user-operated devices via a wide-area network such as the Internet. An example of a user-operated deviceis shown in FIG. 8 as the client computing system 814. The client computing system 814 can be implemented, for example, as a consumer device such as a smartphone, other mobile phone, tablet computer, wearable computing device (e.g., smart watch, eyeglasses), desktop computer, laptop computer, and so on.
[0106] For example, the client computing system 814 can communicate via the WAN interface 810. The client computing system 814 can include computer components, such as processing unit(s) 816, storage device 818, network interface 820, user input device 822, and user output device 837. The client computing system 814 can be a computing device implemented in a variety of form factors, such as a desktop computer, laptop computer, tablet computer, smartphone, other mobile computing device, wearable computing device, or the like.
[0107] The processing unit(s) 816 and storage device 818 can be similar to the processing unit(s) 804 and local storage 806 described above. Suitable devices can be selected based on the demands to be placed on the client computing system 814; for example, the client computing system 814 can be implemented as a “thin” client with limited processing capability or as a high-powered computing device. The client computing system 814 can be provisioned with program code executable by the processing unit(s) 816 to enable various interactions with the server system 800.
[0108] The network interface 820 can provide a connection to the network 826, such as a wide area network (e.g., the Internet), to which the WAN interface 810 of the server system 800 is also connected. In various embodiments, the network interface 820 can include a wired interface (e.g., Ethernet) and / or a wireless interface implementing various RF data communication standards, such as Wi-Fi, Bluetooth, or cellular data network standards (e.g., 3G, 4G, LTE, etc ).
[0109] The user input device 822 can include any device (or devices) via which a user can provide signals to the client computing system 814; the client computing system 814 can interpret the signals as indicative of particular user requests or information. In various embodiments, the user input device 822 can include any or all of a keyboard, touch pad,touch screen, mouse or other pointing device, scroll wheel, click wheel, dial, button, switch, keypad, microphone, and so on.
[0110] The user output device 837 can include any device via which the client computing system 814 can provide information to a user. For example, the user output device 837 can include display-to-display images generated by or delivered to the client computing system 814. The display can incorporate various image generation technologies, e.g., a liquid crystal display (LCD), light-emitting diode (LED), including organic lightemitting diodes (OLED), projection system, cathode ray tube (CRT), or the like, together with supporting electronics (e.g., digital-to-analog or analog-to-digital converters, signal processors, or the like). Some embodiments can include a device such as a touchscreen that functions as both input and output device. In some embodiments, other the user output devices 837 can be provided in addition to or instead of a display. Examples include indicator lights, speakers, tactile “display” devices, printers, and so on.
[0111] Some embodiments include electronic components, such as microprocessors, storage, and memory that store computer program instructions in a computer readable storage medium. Many of the features described in this specification can be implemented as processes that are specified as a set of program instructions encoded on a computer readable storage medium. When these program instructions are executed by one or more processing units, they cause the processing unit(s) to perform various operations indicated in the program instructions. Examples of program instructions or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter. Through suitable programming, the processing unit(s) 804 and 816 can provide various functionality for the server system 800 and the client computing system 814, including any of the functionality described herein as being performed by a server or client, or other functionality.
[0112] It will be appreciated that the server system 800 and the client computing system 814 are illustrative and that variations and modifications are possible. Computer systems used in connection with embodiments of the present disclosure can have other capabilities not specifically described here. Further, while the server system 800 and theclient computing system 814 are described with reference to particular blocks, it is to be understood that these blocks are defined for convenience of description and are not intended to imply a particular physical arrangement of component parts. For instance, different blocks can be but need not be located in the same facility, in the same server rack, or on the same motherboard. Further, the blocks need not correspond to physically distinct components. Blocks can be configured to perform various operations, e.g., by programming a processor or providing appropriate control circuitry, and various blocks might or might not be reconfigurable depending on how the initial configuration is obtained. Embodiments of the present disclosure can be realized in a variety of apparatus including electronic devices implemented using any combination of circuitry and software.IOU3| While the disclosure has been described with respect to specific embodiments, one skilled in the art will recognize that numerous modifications are possible. Embodiments of the disclosure can be realized using a variety of computer systems and communication technologies, including, but not limited to, specific examples described herein. Embodiments of the present disclosure can be realized using any combination of dedicated components, programmable processors, and / or other programmable devices. The various processes described herein can be implemented on the same processor or different processors in any combination. Where components are described as being configured to perform certain operations, such configuration can be accomplished, e.g., by designing electronic circuits to perform the operation, by programming programmable electronic circuits (such as microprocessors) to perform the operation, or any combination thereof. Further, while the embodiments described above may refer to specific hardware and software components, those skilled in the art will appreciate that different combinations of hardware and / or software components may also be used and that particular operations described as being implemented in hardware might also be implemented in software or vice versa.[0114| Computer programs incorporating various features of the present disclosure may be encoded and stored on various computer readable storage media; suitable media includes magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD), flash memory, and other non-transitory media. Computer readable media encoded with the program code may be packaged with a compatible electronic device,or the program code may be provided separately from electronic devices (e.g., via Internet download or as a separately packaged computer-readable storage medium).
[0115] Thus, although the disclosure has been described with respect to specific embodiments, it will be appreciated that the disclosure is intended to cover all modifications and equivalents within the scope of the following claims.
Claims
WHAT IS CLAIMED IS:
1. A method of handling biomedical text and image data, comprising: obtaining, by one or more processors, data for a subject with a condition, the data comprising (i) a biomedical image having a region of interest (RO I) corresponding to a first location associated with the condition in the subject and (ii) a report having text identifying a second location associated with the condition in the subject; determining, by the one or more processors, the first location corresponding to the ROI within the biomedical image using a first machine learning (ML) model; identifying, by the one or more processors, the second location from the text of the report using a second ML model; detecting, by the one or more processors, a deviation in the data for the subject responsive to the first location not corresponding to the second location; and storing, by the one or more processors, using one or more data structures, an association between the data and an indication of the deviation.
2. The method of claim 1, further comprising: providing, by the one or more processors, via an interface, an output including information associated with at least one of the biomedical image or the report, responsive to detecting the deviation; and receiving, by the one or more processors, via the interface, feedback data identifying a third location associated with the condition in the subject, to update at least one of the first ML model, the second ML model, or the report.
3. The method of claim 1, further comprising: detecting by the one or more processors, a lack of any deviation in second data for a second subject responsive to a third location corresponding to a fourth location, wherein the third location is determined from a second biomedical image using the first MLmodel, and wherein the fourth location is identified from text of a second ML model using the second ML model; and-36-storing, by the one or more processors, using one or more second data structures, a second association between the second data and a second indication of the lack of any deviation.
4. The method of claim 1, further comprising: identifying, by the one or more processors, second data comprising: (i) the first location corresponding to the ROI within the biomedical image determined using the first ML model and (ii) a plurality of second reports each having respective text identifying a respective location associated with the condition in the subject; and providing, by the one or more processors, the second data as input to a third ML model to generate a third report identifying a third location associated with the condition in the subject.
5. The method of claim 4, further comprising: providing, by the one or more processors, via an interface, one or more user interface elements to accept or reject the third report; and receiving, by the one or more processors, via the one or more user interface elements of the interface, an indication of acceptance of the third report, and wherein obtaining the data further comprises identifying the third report as the report to include in the data, responsive to receipt of the indication of acceptance.
6. The method of claim 4, further comprising: providing, by the one or more processors, via an interface, the third report and one or more user elements to modify the third report; and receiving, by the one or more processors, via the one or more user interface elements of the interface, a modification to at least a portion of the third report, and using, by the one or more processors, the modification to at least the portion of the third report to update the third ML model.-37-7. The method of claim 1, wherein identifying the data further comprises receiving, via an interface, the report generated by input from a user to identify the second location associated with the condition in the subject.
8. The method of claim 1, wherein determining the first location further comprising generating, in accordance with a template defining locations, a first data structure identifying the first location in the subject, wherein identifying the second location further comprises generating, in accordance with the template, a second data structure identifying the second location in the subject, and wherein detecting the deviation further comprises comparing the first data structure and the second data structure.
9. The method of claim 1, wherein the first ML model comprises an image segmentation model established using a training dataset comprising a plurality of examples, each example of the plurality of examples identifying: (i) a respective biomedical image having a respective ROI corresponding to a location associated with the condition in a respective subject and (ii) a respective annotation identifying the respective ROI within the respective biomedical image; and wherein the second ML model further comprises a natural language processing (NLP) model configured to extract the second location from the text of the report as associated with the condition within the subject.
10. The method of claim 1, wherein the condition further comprises at least one of a tumor, a lesion, an infection, an inflammation, or cell damage in an anatomical site corresponding to at least one of the first location or the second location in the subject.I L A system for handling biomedical text and image data, comprising: one or more processors coupled with memory, configured to: obtain data for a subject with a condition, the data comprising (i) a biomedical image having a region of interest (ROI) corresponding to a first location associated with thecondition in the subject and (ii) a report having text identifying a second location associated with the condition in the subject; determine the first location corresponding to the ROI within the biomedical image using a first machine learning (ML) model; identify the second location from the text of the report using a second ML model; detect a deviation in the data for the subject responsive to the first location not corresponding to the second location; and store, using one or more data structures, an association between the data and an indication of the deviation.
12. The system of claim 11, wherein the one or more processors are further configured to: provide, via an interface, an output including information associated with at least one of the biomedical image or the report, responsive to detecting the deviation; and receive, via the interface, feedback data identifying a third location associated with the condition in the subject, to update at least one of the first ML model, the second ML model, or the report.
13. The system of claim 11, wherein the one or more processors are further configured to detect a lack of any deviation in second data for a second subject responsive to a third location corresponding to a fourth location, wherein the third location is determined from a second biomedical image using the first ML model and wherein the fourth location is identified from text of a second ML model using the second ML model; and store, using one or more second data structures, a second association between the second data and a second indication of the lack of any deviation.
14. The system of claim 11, wherein the one or more processors are further configured to identify second data comprising: (i) the first location corresponding to the ROI within the biomedical image determined using the first ML model and (ii) a plurality of second reports each having respective text identifying a respective location associated with the condition in the subject; andprovide the second data as input to a third ML model to generate a third report identifying a third location associated with the condition in the subject.
15. The system of claim 14, wherein the one or more processors are further configured to provide, via an interface, one or more user interface elements to accept or reject the third report; and receive, via the one or more user interface elements of the interface, an indication of acceptance of the third report, and identify the third report as the report to include in the data, responsive to receipt of the indication of acceptance.
16. The system of claim 14, wherein the one or more processors are further configured to: provide, via an interface, the third report and one or more user elements to modify the third report; and receive, via the one or more user interface elements of the interface, a modification to at least a portion of the third report, and use the modification to at least the portion of the third report to update the third ML model.
17. The system of claim 11, wherein the one or more processors are further configured to receive, via an interface, the report generated by input from a user to identify the second location associated with the condition in the subject.
18. The system of claim 11, wherein the one or more processors are further configured to: generate, in accordance with a template defining locations, a first data structure identifying the first location in the subject, generate, in accordance with the template, a second data structure identifying the second location in the subject, and detect the deviation by comparing the first data structure and the second data structure.
19. The system of claim 11, wherein the first ML model comprises an image segmentation model established using a training dataset comprising a plurality of examples, each example of the plurality of examples identifying: (i) a respective biomedical image having a respective ROI corresponding to a location associated with the condition in a respective subject and (ii) a respective annotation identifying the respective ROI within the respective biomedical image; and wherein the second ML model further comprises a natural language processing (NLP) model configured to extract the second location from the text of the report as associated with the condition within the subject.
20. The system of claim 11, wherein the condition further comprises at least one of a tumor, a lesion, an infection, an inflammation, or cell damage in an anatomical site corresponding to at least one of the first location or the second location in the subject.-41-
Citation Information
Patent Citations
System and method for using three dimensional infrared imaging for libraries of standardized medical imagery
US20100191541A1
Systems and methods for improved analysis and generation of medical imaging reports
US20200043600A1
Medical report labeling system and method for use therewith
US20210118552A1