Methods and systems for performing real-time radiology
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- WHITERABBIT AI INC
- Filing Date
- 2025-07-30
- Publication Date
- 2026-03-19
AI Technical Summary
Breast cancer screening is hampered by poor patient experience, inconsistent turnaround times, varying radiologist performance, and unclear pricing, leading to inconsistent standards of care.
A method and system utilizing artificial intelligence to stratify radiological workflows by classifying medical images into categories, directing them to appropriate radiologists for evaluation, and providing real-time or near-real-time diagnostic results.
Improves patient experience by reducing delays and inconsistencies in breast cancer screening, enhancing radiologist performance, and standardizing care through efficient, accurate, and transparent diagnostic processes.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] cross reference
[0001] This application is a division of U.S. Provisional Patent Application No. 62 / 958,848, filed January 9, 2020. No. 59, which is incorporated herein by reference in its entirety. [Background technology]
[0002]
[0002] Breast cancer is the most common cancer among women in the United States, with 250,000 cases reported in 2017 alone. There have been over 100 new diagnoses in the United States. Approximately 1 in 8 women will be diagnosed with breast cancer at some point in their lifetime. Despite improvements in treatment, more than 40,000 women die from breast cancer each year in the United States. Great strides have been made in reducing breast cancer mortality, in part due to widespread uptake of screening mammography. Breast cancer screening can help identify early-stage cancers, which have a much better prognosis and lower treatment costs compared with late-stage cancers. This difference can be very significant: women with localized breast cancer have a 5-year survival rate of nearly 99%, while women with metastatic breast cancer have a 5-year survival rate of 27%.
[0003]
[0003] Despite these proven benefits, screening mammography Uptake is hampered in part by poor patient experience, including long delays in obtaining appointments, unclear pricing, long wait times to receive results, and confusing reports. Furthermore, problems arising from a lack of pricing transparency are compounded by wide variations in costs across healthcare facilities. Similarly, turnaround times for receiving results are inconsistent across healthcare facilities. Additionally, significant variations in radiologist performance mean that patients experience widely varying standards of care depending on location and income. Summary of the Invention
[0004]
[0004] The present disclosure relates to a method for further screening and analyzing medical image data using artificial intelligence.
[0003] Methods and systems are provided for performing radiological evaluations of subjects by stratifying them into individual radiological workflows for diagnostic and / or evaluation. Such subjects may include subjects with disease (e.g., cancer) and subjects without disease (e.g., cancer). Screening may be directed toward cancer, such as breast cancer. Stratification may be based on disease-related assessments or other assessments (e.g., estimated difficulty of the case).
[0005] In one aspect, the present disclosure provides a method for (a) acquiring at least one image of a body part of a subject. (b) using a trained algorithm to classify at least one image or a derivative thereof into one of a plurality of categories, wherein the classifying step comprises applying an image processing algorithm to the at least one image or a derivative thereof; (c) upon classifying the at least one image or a derivative thereof in (b), (i) sending the at least one image or a derivative thereof to a first radiologist for radiological evaluation if the at least one image is classified into a first category of the plurality of categories, or (ii) sending the at least one image or a derivative thereof to a second radiologist for radiological evaluation if the at least one image is classified into a second category of the plurality of categories; and (d) receiving a radiological evaluation of the subject from the first radiologist or the second radiologist based at least in part on the radiological analysis of the at least one image or a derivative thereof.
[0006] In some embodiments, (b) comprises generating at least one image or a derivative thereof. , classifying the at least one image or a derivative thereof as normal, equivocal, or suspicious. In some embodiments, the method further includes sending the at least one image or a derivative thereof to a classifier based on the classification of the at least one image or a derivative thereof in (b). In some embodiments, (c) includes sending the at least one image or a derivative thereof to a first radiologist of the first plurality of radiologists or a second radiologist of the second plurality of radiologists for radiological evaluation. In some embodiments, the at least one image or a derivative thereof is a medical image.
[0007] In some embodiments, the trained algorithm is In some embodiments, the trained algorithm is configured to classify at least one image, or a derivative thereof, as normal, equivocal, or suspicious with at least about 80% sensitivity. In some embodiments, the trained algorithm is configured to classify at least one image, or a derivative thereof, as normal, equivocal, or suspicious with at least about 80% specificity. In some embodiments, the trained algorithm is configured to classify at least one image, or a derivative thereof, as normal, equivocal, or suspicious with at least about 80% positive predictive value. In some embodiments, the trained algorithm is configured to classify at least one image, or a derivative thereof, as normal, equivocal, or suspicious with at least about 80% negative predictive value.
[0008] In some embodiments, the trained machine learning algorithm determines whether the tissue contains abnormal tissue. , configured to identify at least one region of the at least one image or a derivative thereof that is suspected of containing abnormal tissue.
[0009] In some embodiments, the trained algorithm is The derivatives are classified as normal, equivocal, or suspicious for indicating cancer. In some embodiments, the cancer is breast cancer. In some embodiments, at least one image or a derivative thereof is a three-dimensional image of the subject's breast. In some embodiments, the trained machine learning algorithm is trained using at least about 100 independent training samples including images indicating or suspected of indicating cancer.
[0010] In some embodiments, the trained algorithm determines whether a cancer is indicative of cancer or is suspected to be indicative of cancer. The algorithm is trained using a first plurality of independent training samples including suspected positive images and a second plurality of independent training samples including negative images that do not represent or are not suspected of representing cancer. In some embodiments, the trained algorithm comprises a supervised machine learning algorithm. In some embodiments, the supervised machine learning algorithm comprises a deep learning algorithm, a support vector machine (SVM), a neural network, or a random forest.
[0011] In some embodiments, the method further comprises monitoring the subject. and wherein the monitoring step includes evaluating images of the region of the subject's body at a plurality of time points, the evaluating step being based at least in part on classifying at least one image, or a derivative thereof, at each of the plurality of time points as normal, equivocal, or suspicious. In some embodiments, differences in the evaluation of the images of the subject's body at the plurality of time points are indicative of one or more clinical indicators selected from the group consisting of: (i) a diagnosis of the subject, (ii) a prognosis of the subject, and (iii) the effectiveness or ineffectiveness of a course of treatment for the subject.
[0012] In some embodiments, (c) comprises (i) at least one image or a derivative thereof. sending the images to a first radiologist of a first set of radiologists for radiological evaluation, the first radiologist determining at least in part whether at least one image is classified as suspicious; (ii) sending at least one image or a derivative thereof to a second radiologist of the second set of radiologists for radiological evaluation and producing a screening result based at least in part on whether the at least one image is classified as equivocal; or (iii) sending at least one image or a derivative thereof to a third radiologist of the third set of radiologists for radiological evaluation and producing a screening result based at least in part on whether the at least one image is classified as normal. In some embodiments, (c) further includes, if at least one image is classified as equivocal, sending at least one image or a derivative thereof to a first radiologist of the first set of radiologists for radiological evaluation and producing a screening result. In some embodiments, (c) further includes, if at least one image is classified as equivocal, sending at least one image or a derivative thereof to a second radiologist of the second set of radiologists for radiological evaluation and producing a screening result. In some embodiments, (c) further comprises, if the at least one image is classified as normal, sending the at least one image, or a derivative thereof, to a third radiologist of the third set of radiologists for radiological evaluation to produce a screening result. In some embodiments, the screening result for the subject is produced in the same clinic visit as the step of acquiring the at least one image, or a derivative thereof. In some embodiments, the first set of radiologists is located at an on-site clinic, and the at least one image, or a derivative thereof, is acquired at the on-site clinic.
[0013] In some embodiments, the second set of radiologists includes radiologists, The radiologist is trained to classify at least one image, or a derivative thereof, as normal or suspicious with greater accuracy than the trained algorithm. In some embodiments, a third set of radiologists is located remotely from the on-site clinic, and at least one image is acquired at the on-site clinic. In some embodiments, a third radiologist of the third set of radiologists performs a radiological evaluation of at least one image, or a derivative thereof, of a batch comprising a plurality of images, the batch being selected to improve efficiency of the radiological evaluation.
[0014] In some embodiments, the method further comprises: and generating a diagnostic result for the subject based at least in part on the step of acquiring the at least one image. In some embodiments, the diagnostic result for the subject is generated in the same clinic visit as the step of acquiring the at least one image. In some embodiments, the diagnostic result for the subject is generated within about one hour of the step of acquiring the at least one image.
[0015] In some embodiments, at least one image or a derivative thereof is The image is sent to the first radiologist, the second radiologist, or a third radiologist based at least in part on additional characteristics of the body part, in some embodiments, the additional characteristics include anatomy, tissue characteristics (e.g., tissue density or physical properties), the presence of foreign bodies (e.g., implants), the type of finding, a medical condition (e.g., predicted by an algorithm, such as a machine learning algorithm), or a combination thereof.
[0016] In some embodiments, at least one image or a derivative thereof is a first radiation to the first radiologist, the second radiologist, or the third radiologist based at least in part on additional characteristics of the radiologist, the second radiologist, or the third radiologist (e.g., the first radiologist, the second radiologist, or the third radiologist's personal ability to perform a radiological assessment of the at least one image or derivative thereof).
[0017] In some embodiments, (c) comprises generating at least one image or a derivative thereof. generating an alert based at least in part on sending the at least one image or a derivative thereof to a first radiologist or a second radiologist. In some embodiments, the method further includes sending the alert to the subject or the subject's clinical caregiver. In some embodiments, the method further includes sending the alert to the subject through a patient mobile application. In some embodiments, the alert is generated in real time with (b) or near real time with (b).
[0018] In some embodiments, the step of applying an image processing algorithm comprises at least The method further comprises identifying a region of interest within one of the images or a derivative thereof, and labeling the region of interest to produce at least one labeled image. In some embodiments, the method further comprises storing the at least one labeled image in a database. In some embodiments, the method further comprises storing the at least one image or one or more of its derivatives and the classification in a database. In some embodiments, the method further comprises generating a presentation of the at least one image based at least in part on the at least one image or one or more of its derivatives and the classification. In some embodiments, the method further comprises storing the presentation in a database.
[0019] In some embodiments, (c) is synchronized in real time with (b) or near real time with (b). In some embodiments, the at least one image comprises multiple images acquired from the subject, the multiple images being acquired using different modalities or at different points in time. In some embodiments, the classifying comprises processing clinical health data of the subject.
[0020] In another aspect, the present disclosure provides a method for processing at least one image of a body part of a subject. The present invention provides a computer system for: a database configured to store at least one image of a body part of a subject; and one or more computer processors operably coupled to the database, the one or more computer processors performing a classifying step: (a) using a trained algorithm to classify at least one image, or a derivative thereof, into one of a plurality of categories, the classifying step including applying an image processing algorithm to the at least one image, or a derivative thereof; and (b) upon classifying the at least one image, or a derivative thereof, in (a), (i) classifying the at least one image, or a derivative thereof, into one of a plurality of categories. and one or more computer processors individually or collectively programmed to: (i) send at least the image or a derivative thereof to a first radiologist for radiological evaluation if the image is classified into a first category of the plurality of categories; or (ii) send at least one image or a derivative thereof to a second radiologist for radiological evaluation if at least one image is classified into a second category of the plurality of categories; and (c) receive a radiological evaluation of the subject from the first radiologist or the second radiologist based at least in part on the radiological analysis of the at least one image or a derivative thereof.
[0021] In some embodiments, (a) comprises: In some embodiments, the one or more computer processors are individually or collectively programmed to further perform the step of sending at least one image or a derivative thereof to a classifier based on the classification in (a) of the at least one image or derivative thereof. In some embodiments, (b) includes sending the at least one image or derivative thereof to a first radiologist of the first plurality of radiologists or a second radiologist of the second plurality of radiologists for radiological evaluation. In some embodiments, the at least one image or derivative thereof is a medical image.
[0022] In some embodiments, the trained algorithm is In some embodiments, the trained algorithm is configured to classify at least one image, or a derivative thereof, as normal, equivocal, or suspicious with at least about 80% sensitivity. In some embodiments, the trained algorithm is configured to classify at least one image, or a derivative thereof, as normal, equivocal, or suspicious with at least about 80% specificity. In some embodiments, the trained algorithm is configured to classify at least one image, or a derivative thereof, as normal, equivocal, or suspicious with at least about 80% positive predictive value. In some embodiments, the trained algorithm is configured to classify at least one image, or a derivative thereof, as normal, equivocal, or suspicious with at least about 80% negative predictive value. In some embodiments, the trained machine learning algorithm is configured to identify at least one region of at least one image, or a derivative thereof, that contains abnormal tissue or is suspected of containing abnormal tissue.
[0023] In some embodiments, the trained algorithm is The derivatives are classified as normal, equivocal, or suspicious for indicating cancer. In some embodiments, the cancer is breast cancer. In some embodiments, at least one image or a derivative thereof is a three-dimensional image of the subject's breast. In some embodiments, the trained machine learning algorithm is trained using at least about 100 independent training samples including images indicating or suspected of indicating cancer.
[0024] In some embodiments, the trained algorithm determines whether a cancer is indicative or suspected to be indicative of cancer. The algorithm is trained using a first plurality of independent training samples including suspected positive images and a second plurality of independent training samples including negative images that do not represent or are not suspected of representing cancer. In some embodiments, the trained algorithm comprises a supervised machine learning algorithm. In some embodiments, the supervised machine learning algorithm comprises a deep learning algorithm, a support vector machine (SVM), a neural network, or a random forest.
[0025] In some embodiments, the one or more computer processors and (iii) a diagnostic step of evaluating the subject's bodily region, the diagnostic step comprising evaluating images of the subject's bodily region at a plurality of time points, the evaluating step being based at least in part on a classification of at least one image, or a derivative thereof, at each of the plurality of time points as normal, equivocal, or suspicious. In some embodiments, differences in the evaluation of the images of the subject's body at the plurality of time points are indicative of one or more clinical indicators selected from the group consisting of: (i) a diagnosis of the subject, (ii) a prognosis of the subject, and (iii) the effectiveness or ineffectiveness of a course of treatment for the subject.
[0026] In some embodiments, (b) comprises (i) at least one image or a derivative thereof. (ii) sending the at least one image or derivative thereof to a first radiologist of the first set of radiologists for radiological evaluation and generating a screening result based at least in part on whether the at least one image or derivative thereof is classified as suspicious; and (iii) sending the at least one image or derivative thereof to a second radiologist of the second set of radiologists for radiological evaluation and generating a screening result based at least in part on whether the at least one image or derivative thereof is classified as equivocal. and (iii) sending at least one image or a derivative thereof to a third radiologist of the third set of radiologists for radiological evaluation to produce a screening result based at least in part on whether the at least one image or a derivative thereof is classified as normal. In some embodiments, (b) further comprises, if the at least one image is classified as suspicious, sending the at least one image or a derivative thereof to a first radiologist of the first set of radiologists for radiological evaluation to produce a screening result. In some embodiments, (b) further comprises, if the at least one image is classified as equivocal, sending the at least one image or a derivative thereof to a second radiologist of the second set of radiologists for radiological evaluation to produce a screening result. In some embodiments, (b) further comprises, if the at least one image is classified as normal, sending the at least one image or a derivative thereof to a third radiologist of the third set of radiologists for radiological evaluation to produce a screening result. In some embodiments, the subject screening results are produced in the same clinic visit as the step of acquiring the at least one image, hi some embodiments, a first set of radiologists is located at the on-site clinic, and the at least one image is acquired at the on-site clinic.
[0027] In some embodiments, the second set of radiologists includes radiologists, The radiologist is trained to classify at least one image, or a derivative thereof, as normal or suspicious with greater accuracy than the trained algorithm. In some embodiments, a third set of radiologists is located remotely from the on-site clinic, and at least one image is acquired at the on-site clinic. In some embodiments, a third radiologist of the third set of radiologists performs a radiological evaluation of at least one image, or a derivative thereof, of a batch comprising a plurality of images, the batch being selected to improve efficiency of the radiological evaluation.
[0028] In some embodiments, one or more computer processors The imaging devices are individually or collectively programmed to further obtain a diagnostic result for the subject from a diagnostic procedure performed on the subject based at least in part on the results of the imaging test. In some embodiments, the diagnostic result for the subject is produced in the same clinic visit as the step of acquiring the at least one image. In some embodiments, the diagnostic result for the subject is produced within about one hour of the step of acquiring the at least one image.
[0029] In some embodiments, at least one image or a derivative thereof is of the subject. The image is sent to the first radiologist, the second radiologist, or a third radiologist based at least in part on additional characteristics of the body part, in some embodiments, the additional characteristics include anatomy, tissue characteristics (e.g., tissue density or physical properties), the presence of foreign bodies (e.g., implants), the type of finding, a medical condition (e.g., predicted by an algorithm, such as a machine learning algorithm), or a combination thereof.
[0030] In some embodiments, at least one image or a derivative thereof is to the first radiologist, the second radiologist, or the third radiologist based at least in part on additional characteristics of the radiologist, the second radiologist, or the third radiologist (e.g., the first radiologist, the second radiologist, or the third radiologist's personal ability to perform a radiological assessment of the at least one image or derivative thereof).
[0031] In some embodiments, (b) comprises generating at least one image or a derivative thereof. generating an alert based at least in part on the step of sending the at least one image or a derivative thereof to a first radiologist or the step of sending the at least one image or a derivative thereof to a second radiologist; In some embodiments, the one or more computer processors are individually or collectively programmed to further transmit the alert to the subject or the subject's clinical caregiver. In some embodiments, the one or more computer processors are individually or collectively programmed to further transmit the alert to the subject through a patient mobile application. In some embodiments, the alert is generated in real time with (a) or near real time with (a).
[0032] In some embodiments, the step of applying an image processing algorithm comprises at least The method includes identifying a region of interest within one of the images or a derivative thereof, and labeling the region of interest to produce at least one labeled image. In some embodiments, the one or more computer processors are individually or collectively programmed to further store the at least one labeled image in a database. In some embodiments, the one or more computer processors are individually or collectively programmed to further store one or more of the at least one image or its derivatives and a classification in a database. In some embodiments, the one or more computer processors are individually or collectively programmed to further generate a presentation of the at least one image or its derivative based at least in part on the one or more of the at least one image and the classification. In some embodiments, the one or more computer processors are individually or collectively programmed to further store the presentation in a database.
[0033] In some embodiments, (b) is synchronized in real time with (a) or near real time with (a). In some embodiments, the at least one image comprises multiple images acquired from the subject, the multiple images being acquired using different modalities or at different points in time. In some embodiments, the classifying comprises processing clinical health data of the subject.
[0034] Another aspect of the present disclosure is a method for implementing a method for generating a program executed by one or more computer processors. When provided, a non-transitory computer-readable medium is provided that includes machine-executable code for implementing any of the methods described above or elsewhere herein.
[0035] Another aspect of the present disclosure is one or more computer processors and associated The present invention provides a system comprising an integrated computer memory including machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.
[0036]
[0036] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description. It will be clear that in the following detailed description, merely exemplary embodiments of the present disclosure are shown and described. As will be understood, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0037] Incorporation by Reference All publications, patents, and patent applications mentioned herein are the property of their respective owners. All such publications, patents, or patent applications are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, the present specification is intended to supersede and / or supersede such conflicting material. will be done.
[0038] The novel features of the invention are set forth with particularity in the appended claims. The features and advantages will be better understood by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "drawings" and "figures"): [Brief explanation of the drawings]
[0039] [Figure 1]
[0039] FIG. 1 is an exemplary workflow diagram of a method for sending a case for radiological review (eg, by a radiologist or radiographer) according to the disclosed embodiments. [Figure 2]
[0040] FIG. 1 illustrates an example method of using a triage engine configured to stratify subjects undergoing mammography screening by classifying the subject's mammography data into one of three different workflows: normal, equivocal, and suspicious, according to disclosed embodiments. [Figure 3A]
[0041] 1 is a diagram of an example user interface for a real-time radiology system including a view from the perspective of a mammography technician or technician assistant, according to a disclosed embodiment; [Figure 3B] FIG. 1 is an illustration of an example user interface for a real-time radiology system including a view from a radiologist's perspective, in accordance with a disclosed embodiment; [Figure 3C] FIG. 1 is an illustration of an example user interface for a real-time radiology system including a view from a billing officer's perspective, in accordance with a disclosed embodiment; [Figure 3D] 1 is a diagram of an example user interface for a real-time radiology system including a view from the perspective of an ultrasound technician or technician assistant, in accordance with a disclosed embodiment; [Figure 4]
[0042] FIG. 1 is a diagram of a computer system programmed or configured to implement the methods provided herein. [Figure 5]
[0043] 1 is an exemplary plot of the detection frequency of breast cancer tumors of various sizes (ranging from 2 mm to 29 mm) detected using a real-time radiology system, according to disclosed embodiments. [Figure 6]
[0044] 1 is an exemplary plot of positive predictive value (PPV1) versus callback rate from screening mammography, according to disclosed embodiments. [Figure 7]
[0045] 10 is an exemplary plot comparing batch (including control, BI-RADs, and density) interpretation time (left) and percentage improvement in interpretation time over the control (right) for a first set of radiologists, a second set of radiologists, and the total set of radiologists overall, in accordance with disclosed embodiments. [Figure 8]
[0046] 1 is a receiver operating characteristic (ROC) curve illustrating the performance of a DNN on a binary classification task evaluated on a test dataset, according to disclosed embodiments. [Figure 9]
[0047] FIG. 1 is a diagram of an example schematic of patient flow through a clinic using an AI-enabled real-time radiology system and a patient mobile application (app), according to disclosed embodiments. [Figure 10]
[0048] FIG. 1 is a diagram of a schematic example of an AI-assisted radiological evaluation workflow, according to the disclosed embodiments. [Figure 11]
[0049] FIG. 1 is a diagram of an example triage software system developed using machine learning for screening mammography to enable more timely report delivery and follow-up for suspicious cases (e.g., as occurs in a batch reading setting), according to disclosed embodiments. [Figure 12A]
[0050] FIG. 1 is an example of a composite 2D mammography (SM) image derived from a digital breast tomosynthesis (DBT) examination for Breast Imaging Reporting and Data System (BI-RADS) breast density category (A) almost entirely fatty, according to a disclosed embodiment. [Figure 12B]FIG. 1 is an example of a synthetic 2D mammography (SM) image derived from a DBT examination for BI-RADS breast density category (B) scattered areas of fibroglandular density, in accordance with disclosed embodiments. [Figure 12C] FIG. 1 is an example of a synthetic 2D mammography (SM) image derived from a DBT examination for BI-RADS breast density category (C) heterogeneously dense, in accordance with the disclosed embodiments. [Figure 12D] FIG. 1 is an example of a synthetic 2D mammography (SM) image derived from a DBT examination for a BI-RADS breast density category (D) extremely dense, in accordance with the disclosed embodiments. [Figure 13A]
[0051] 13A-13D are comparative images of the same breast under the same pressure, according to disclosed embodiments: Figure 13A is a full-field digital mammography (FFDM) image. [Figure 13B] Figure 13B is a composite 2D mammography (SM) image. [Figure 13C] FIG. 13C is a zoomed-in view of the FFDM image with the original site indicated by the white square. [Figure 13D] Figure 13D is a zoomed-in area of the SM image with the original area indicated by a white square. Figures 13C and 13D are intended to highlight the differences in texture and contrast that can occur between the two image types. [Figure 14A]
[0052] 1 is a confusion matrix for the Breast Imaging Reporting and Data System (BI-RADS) breast density task evaluated against a full-field digital mammography (FFDM) test set, according to disclosed embodiments. The number of test samples (trials) in each bin is shown in parentheses. [Figure 14B]1 shows a confusion matrix for a binary density task (BI-RADS C+D, which is high density, vs. BI-RADS A+B, which is non-high density) evaluated on a full-field digital mammography (FFDM) test set, according to disclosed embodiments. The number of test samples (tests) in each bin is shown in parentheses. [Figure 15A]
[0053] 1 is a confusion matrix for the Breast Imaging Reporting and Data System (BI-RADS) breast density task without adaptation evaluated against the Site1SM test set, with the number of test samples (trials) in each bin shown in parentheses, in accordance with disclosed embodiments. [Figure 15B] 1 is a confusion matrix for a binary density task (BI-RADS C+D, which is dense, versus BI-RADS A+B, which is non-dense) without adaptation, evaluated against the Site1SM test set, according to a disclosed embodiment. The number of test samples (trials) in each bin is shown in parentheses. [Figure 15C] 1 is a confusion matrix for the BI-RADS breast density task with adaptation by matrix calibration of 500 training samples, evaluated against the Site1SM test set, in accordance with a disclosed embodiment. The number of test samples (tests) in each bin is shown in parentheses. [Figure 15D] 1 is a diagram of a confusion matrix for a binary density task (dense vs. non-dense) with adaptation by matrix calibration of 500 training samples, evaluated against the Site1SM test set, in accordance with a disclosed embodiment. The number of test samples (tests) in each bin is shown in parentheses. [Figure 16A]
[0054] FIG. 1 is a confusion matrix for the Breast Imaging Reporting and Data System (BI-RADS) breast density task without adaptation evaluated against the Site2SM test set, in accordance with disclosed embodiments. [Figure 16B]FIG. 10 is a confusion matrix for a binary density task (BI-RADS C+D, which is dense, versus BI-RADS A+B, which is non-dense) without adaptation, evaluated against the Site2SM test set, in accordance with disclosed embodiments. [Figure 16C] FIG. 10 is a confusion matrix for the BI-RADS breast density task with adaptation by matrix calibration of 500 training samples evaluated against the Site2SM test set, in accordance with disclosed embodiments. [Figure 16D] 1 is a confusion matrix for a binary density task (dense vs. non-dense) with adaptation by matrix calibration of 500 training samples, evaluated against the Site2SM test set, according to a disclosed embodiment. The number of test samples (tests) in each bin is shown in parentheses. [Figure 17A]
[0055] FIG. 10 is a diagram of the impact of the amount of training data on the performance of adaptive methods, as measured by macroAUC, for the Site1 dataset, in accordance with disclosed embodiments. [Figure 17B] FIG. 10 is a diagram of the impact of the amount of training data on the performance of adaptive methods, as measured by linearly weighted Cohen's Kappa coefficient, for the Site1 dataset, in accordance with disclosed embodiments. [Figure 17C] FIG. 10 is a diagram of the impact of the amount of training data on the performance of adaptive methods, as measured by macroAUC, for the Site2SM dataset, in accordance with disclosed embodiments. [Figure 17D] FIG. 10 is a diagram of the impact of the amount of training data on the performance of adaptive methods, as measured by linearly weighted Cohen's Kappa coefficient, for the Site2SM dataset, in accordance with disclosed embodiments. [Figure 18]
[0056] FIG. 1 is a diagram of an example of a schematic of a real-time radiological assessment workflow. [Figure 19]
[0057] FIG. 1 is a diagram of an example of a schematic of a real-time radiological assessment workflow. [Figure 20]
[0058] FIG. 1 is a schematic example of an AI-assisted radiological evaluation workflow in a teleimaging setting. DETAILED DESCRIPTION OF THE INVENTION
[0040]
[0059] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that It will be apparent that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It will be understood that various alternatives may be employed to the embodiments of the invention described herein.
[0041]
[0060] As used in this specification and claims, unless expressly stated otherwise, Unless otherwise specified, the singular forms "a," "an," and "the" include the plural. For example, the term "a nucleic acid" includes "nucleic acids," including mixtures thereof.
[0042]
[0061] As used herein, the term "subject" generally refers to a testable or detectable Genetic genetic information refers to an entity or medium that contains genetic information. A subject can be a person, an individual, or a patient. A subject can be a vertebrate, such as a mammal. Non-limiting examples of mammals include humans, monkeys, livestock, game animals, rodents, and pets. A subject can be a person with cancer or suspected of having cancer. A subject can exhibit symptoms indicative of the subject's health or physiological state or condition, such as the subject's cancer (e.g., breast cancer). Alternatively, a subject can be asymptomatic with respect to such health or physiological state or condition.
[0043]
[0062] Breast cancer is the most common cancer among women in the United States, with 250,000 cases reported in 2017 alone. There have been over 100 new diagnoses in the United States. Approximately 1 in 8 women will be diagnosed with breast cancer at some point in their lifetime. Despite improvements in treatment, more than 40,000 women die from breast cancer each year in the United States. Great strides have been made in reducing breast cancer mortality, in part due to widespread uptake of screening mammography. Breast cancer screening can help identify early-stage cancers, which have a much better prognosis and lower treatment costs compared with late-stage cancers. This difference can be very significant: women with localized breast cancer have a 5-year survival rate of nearly 99%, while women with metastatic breast cancer have a 5-year survival rate of 27%.
[0044]
[0063] Despite these proven benefits, screening mammography Uptake is hampered in part by poor patient experience, including long delays in obtaining appointments, unclear pricing, long wait times to receive results, and confusing reports. Furthermore, problems arising from a lack of pricing transparency are compounded by wide variations in costs across healthcare facilities. Similarly, turnaround times for receiving results are inconsistent across healthcare facilities. Additionally, significant variations in radiologist performance mean that patients experience widely varying standards of care depending on location and income.
[0045]
[0064] The present disclosure uses artificial intelligence to further screen and analyze medical image data.
[0003] The present invention provides a method and system for performing real-time radiology of subjects by layering individual radiology workflows for diagnostic and / or evaluation. Such subjects may include subjects with and without cancer. Screening may be for cancer, such as breast cancer.
[0046]
[0065] FIG. 1 illustrates a method for performing a radiology procedure (e.g., by a radiologist, radiology specialist, or 1 illustrates an exemplary workflow of a method for sending a case for radiological review (by a radiologist or radiologist). In one aspect, the present disclosure provides a method 100 for processing at least one image of a part of a body of a subject. Method 100 may include acquiring an image of the part of the body of the subject (as per act 102). Then, method 100 may include classifying the image, or a derivative thereof, into one of a plurality of categories using a trained algorithm (as per act 104). For example, classifying may include applying an image processing algorithm to the image, or a derivative thereof. Next, method 100 may include determining whether the image has been classified into a first category or a second category of the plurality of categories (as per act 106). If the image has been classified into the first category, method 100 may include sending the image to a first radiologist for radiological evaluation (as per act 108). If the image is classified in the second category, method 100 may include sending the image to a second radiologist for radiological evaluation (as per operation 110). Method 100 may then include receiving a suggestion to examine the subject (e.g., from the first radiologist or the second radiologist, or from another radiologist or physician) based on the radiological evaluation of the image (as per operation 112).
[0047]
[0066] FIG. 2 shows mammography data for a subject, divided into normal, Illustrates an example of how a triage engine configured to stratify subjects undergoing mammography screening by classifying them into one of three different workflows: likely normal, possibly suspicious, and suspicious. First, a dataset containing the patient's electronic health record (EHR) and medical images is provided. The AI-based triage engine then processes the EHR and medical images to analyze the dataset and classify it as likely normal, possibly suspicious, or likely suspicious. The patient's dataset is then stratified into one of three workflows: likely normal, likely suspicious, and likely suspicious, based on the dataset's classification as likely normal, likely suspicious, or suspicious, respectively. Each patient's data set is processed through one of three workflows: a normal case workflow, an uncertain case workflow, and a suspicious case workflow. Each of the three workflows may include radiologist review or further AI-based analysis (e.g., by a trained algorithm). The normal case workflow may include an AI-based (optionally cloud-based) confirmation that the patient's data set is normal, which then completes routine screening. For example, a group of radiologists may review the normal case workflow cases in a high-volume, efficient manner. Alternatively, the normal case workflow may include an AI-based (optionally cloud-based) determination that the patient's data set is suspicious, which then orders a rapid radiologist review of the patient's data set. For example, a second group of radiologists may review the suspicious case workflow cases in a smaller number of less efficient ways (e.g., a radiologist performs a more detailed radiological evaluation). Similarly, the uncertain case and suspicious case workflows may also include a rapid radiologist review of the patient's data set. In some embodiments, different sets of radiologists are used to review different workflows, as described elsewhere herein. In some embodiments, the same set of radiologists is used to review different workflows (eg, at different points in time depending on the priority of the case for radiological evaluation).
[0048]
[0067] 3A-3D are diagrams illustrating a mammography system according to a disclosed embodiment. 3A and 3B show diagrams of example user interfaces for a real-time radiology system, including views from the perspectives of an assistant (FIG. 3A), a radiologist (FIG. 3B), a billing specialist (FIG. 3C), and an ultrasound technician or technician assistant (FIG. 3D). The views may include a heat map showing which areas have been identified as suspicious by the AI algorithm. The mammography technician or technician assistant may ask the patient several questions and evaluate the patient's responses to determine whether the patient is suitable for real-time radiology evaluation. The radiologist may read or interpret the patient's medical images (e.g., mammography images) according to the disclosed real-time radiology methods and systems. The billing specialist may estimate the diagnosis cost based on the patient's suitability for real-time radiology evaluation. The mammography / ultrasound technician or technician assistant may inform the patient to wait for the results of the real-time radiology evaluation. The user interface may provide a notification to the technician or technician assistant that the acquired image is of poor quality (e.g., generated by an AI-based algorithm) so that the technician or technician assistant can make corrections to the acquired image or repeat the image acquisition.
[0049] Medical Image Acquisition
[0068] Medical images can be acquired or derived from human subjects (e.g., patients). The medical images may be stored in a database, such as a computer server (e.g., a cloud-based server), a local server, a local computer, or a mobile device (such as a smartphone or tablet). The medical images may be obtained from subjects with cancer, from subjects suspected of having cancer, or from subjects without or not suspected of having cancer.
[0050]
[0069] Medical images may be obtained before and / or after treatment of subjects with cancer Medical images may be obtained from a subject during a treatment or treatment regime. Multiple sets of medical images may be obtained from a subject to monitor the effectiveness of treatment over time. Medical images may be obtained from subjects with known or suspected cancer (e.g., breast cancer) where a definitive positive or negative diagnosis is not available through clinical testing. Medical images may be obtained from subjects suspected of having cancer. Medical images may be obtained from subjects experiencing unexplained symptoms such as fatigue, nausea, weight loss, aches and pains, weakness, or bleeding. The medical images may be obtained from subjects who have the condition described. The medical images may be obtained from subjects who are at risk for developing cancer due to factors such as family history, age, hypertension or pre-hypertension, diabetes or pre-diabetes, overweight or obesity, environmental exposures, lifestyle risk factors (e.g., smoking, alcohol consumption, or drug use), or the presence of other risk factors.
[0051]
[0070] Medical imaging includes mammography scans, computed tomography (CT) scans, The medical images may be obtained using one or more imaging modalities, such as a CT scan, magnetic resonance imaging (MRI) scan, ultrasound scan, digital X-ray scan, positron emission tomography (PET) scan, PET-CT scan, nuclear medicine scan, thermography scan, ophthalmology scan, optical coherence tomography scan, electrocardiogram scan, endoscopic scan, fluoroscopy scan, bone densitometry scan, optical scan, or any combination thereof. The medical images may be preprocessed using image processing techniques or deep learning to enhance image characteristics (e.g., contrast, brightness, sharpness), remove noise or artifacts, filter frequency ranges, compress images to small file sizes, or sample or crop images. The medical images may be raw or reconstructed (e.g., to create a 3D volume from multiple 2D images). The images may be processed to compute maps correlated to tissue properties or functional behavior, such as in functional MRI (fMRI) or resting-state fMRI. The images may be overlaid with additional information, such as heat maps or fluid profiles. The image may be created from a composite of images from several scans of the same subject, or from several subjects.
[0052]
[0071] Pre-trained Algorithms
[0072] A dataset containing multiple medical images of one or more subject body parts is provided. Once acquired, the trained algorithm can be used to process the dataset and classify the images as normal, equivocal, or suspicious. For example, the trained algorithm can be used to determine regions of interest (ROIs) in multiple medical images of a subject and process the ROIs to classify the images as normal, equivocal, or suspicious. The trained algorithm may be configured to classify images as normal, equivocal, or suspicious with at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or greater than 99% accuracy for at least about 25, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, or more than about 500 independent samples.
[0053]
[0073] The trained algorithm may include a supervised machine learning algorithm. The algorithm may include a classification and regression tree (CART) algorithm. The supervised machine learning algorithm may include, for example, a random forest, a support vector machine (SVM), a neural network (e.g., a deep neural network (DNN)), or a deep learning algorithm. The trained algorithm may include an unsupervised machine learning algorithm.
[0054]
[0074] The trained algorithm accepts multiple input variables and generates a result based on the multiple input variables. The plurality of input variables may be configured to generate one or more output values based on the plurality of input variables. The plurality of input variables may include features extracted from one or more datasets including medical images of parts of the subject's body. For example, the input variables may include the number of potentially cancerous or suspicious regions of interest (ROIs) in a dataset of medical images. Potentially cancerous or suspicious regions of interest (ROIs) may be identified or extracted from the dataset of medical images using various image processing techniques, such as image segmentation. The input variables may also include several images from a 3D volume or slices in multiple visits over time. The input variables may also include clinical health data of the subject.
[0055]
[0075] In some embodiments, the clinical health data includes age, weight, height, body mass index (BMI), and The clinical health data may include one or more quantitative measurements of the subject, such as MI, blood pressure, heart rate, glucose level, etc. As another example, the clinical health data may include one or more categorical measurements, such as race, ethnicity, medication or other clinical treatment history, smoking history, alcohol consumption history, daily activity or exercise level, genetic test results, blood test results, imaging results, and screening results.
[0056]
[0076] The trained algorithm is applied to one or more images (e.g., radiology images). The trained algorithm may include one or more modules configured to perform image processing on the medical images to produce detection or segmentation of one or more images. The trained algorithm may include a classifier (e.g., a linear classifier, a logistic regression classifier, etc.) such that each of one or more output values includes one of a fixed number of possible values, thereby demonstrating the classification of a dataset including medical images by the classifier. The trained algorithm may include a binary classifier such that each of one or more output values includes one of two values (e.g., {0, 1}, {positive, negative}, {high risk, low risk}, or {suspicious, normal}), thereby demonstrating the classification of a dataset including medical images by the classifier. The trained algorithm may also include another type of classifier such that each of one or more output values includes one of three or more values (e.g., {0, 1, 2}, {positive, negative, or neutral}, {high risk, medium risk, or low risk}, or {suspicious, normal, or unclassifiable}), thereby demonstrating the classification of a dataset including medical images by the classifier. The output value may include a descriptive label, a numerical value, or a combination thereof. A portion of the output value may include a descriptive label. Such a descriptive label may provide an identification, indication, likelihood, or risk of a disease or disorder condition for the subject, and may include, for example, positive, negative, high risk, medium risk, low risk, suspicious, normal, or unclassifiable. Such a descriptive label may provide the identification of a follow-up diagnostic procedure or treatment for the subject, and may include, for example, a therapeutic intervention suitable for treating cancer or other conditions, the duration of the therapeutic intervention, and / or the administration method of the therapeutic intervention. Such a descriptive label may provide the identification of secondary clinical tests that may be appropriate to perform on the subject, and may include, for example, imaging tests, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, digital X-rays, positron emission tomography (PET) scans, PET-CT scans, or any combination thereof. As another example, such a descriptive label may provide a prognosis for the subject's cancer.As another example, such descriptive labels may provide a relative assessment of a subject's cancer (e.g., estimated stage or tumor burden). Some descriptive labels may be mapped to numerical values, for example, by mapping "positive" to 1 and "negative" to 0.
[0057]
[0077] Some of the output values may include numeric values, such as binary values, integers, or continuous values. Such binary output values may include, for example, {0,1}, {positive,negative}, or {high risk,low risk}. Such integer output values may include, for example, {0,1,2}. Such continuous output values may include, for example, probability values at least 0 and less than or equal to 1. Such continuous output values may include, for example, the center coordinates of an ROI. Such continuous output values may indicate the prognosis of cancer for the subject. Some numerical values may be used, for example, with 1 being "positive" , 0 may be mapped to a descriptive label by mapping 0 to "negative." An array or map of numerical values may be produced, such as a cancer probability map.
[0058]
[0078] Portions of the output values may be assigned based on one or more cutoff values. For example, if a dataset including medical images indicates that a subject has cancer (e.g., breast cancer) at least 50% of the time, then a binary classification of the dataset including medical images may assign an output value of "positive" or 1. For example, if a dataset including medical images indicates that a subject has cancer less than 50% of the time, then a binary classification of the dataset including medical images may assign an output value of "negative" or 0. In this case, a single cutoff value of 50% is used to classify the dataset including medical images into one of two possible binary output values. Examples of single cutoff values include about 1%, about 2%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, and about 99%.
[0059]
[0079] In another example, the subject may be at least about 50%, at least about 55%, or at least A classification of a dataset including medical images may assign an output value of "positive" or 1 if the dataset including medical images indicates that the subject has about a 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or more probability of having cancer. If the dataset including the medical images indicates that the subject has more than about 50%, more than about 55%, more than about 60%, more than about 65%, more than about 70%, more than about 75%, more than about 80%, more than about 85%, more than about 90%, more than about 91%, more than about 92%, more than about 93%, more than about 94%, more than about 95%, more than about 96%, more than about 97%, more than about 98%, or more than about 99% chance of having cancer, the classification of the sample may be assigned an output value of "positive" or 1.
[0060]
[0080] Subjects are less than about 50%, less than about 45%, less than about 40%, less than about 35%, and about 30% If the dataset including medical images indicates that the subject has less than about 50% or less, less than about 45% or less, less than about 40% or less, less than about 35% or less, less than about 30% or less, less than about 25% or less, less than about 20% or less, less than about 15% or less, less than about 10% or less, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1% chance of having cancer, the classification of the dataset including medical images may assign an output value of "negative" or 0 if the dataset including medical images indicates that the subject has less than about 50% or less, less than about 45%, less than about 40%, less than about 35%, less than about 30%, less than about 25%, less than about 20%, less than about 15%, less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1% chance of having cancer.
[0061]
[0081] A dataset containing medical images is classified as "positive," "negative," 1, or 0. If not, the classification of a dataset containing medical images may assign an output value of "middle" or 2. In this case, a set of two cutoff values is used to classify a dataset containing medical images into one of three possible output values. Example sets of cutoff values include {1%, 99%}, {2%, 98%}, {5%, 95%}, {10%, 90%}, {15%, 85%}, {20%, 80%}, {25%, 75%}, {30%, 70%}, {35%, 65%}, {40%, 60%}, and {45%, 55%}. Similarly, a set of n cutoff values may be used to classify a dataset containing medical images into one of n+1 possible output values, where n is any positive integer.
[0062]
[0082] A trained algorithm may be trained using multiple independent training samples. Each independent training sample may include a dataset including medical images from a subject, an associated dataset (e.g., labels or annotations) obtained by analyzing the medical images, and one or more known output values corresponding to the dataset including medical images (e.g., difficulty of interpreting the images, time taken to interpret the images, clinical diagnosis, prognosis, defect, efficacy of treatment for the subject's cancer). The independent training sample may include a dataset including medical images and associated datasets and outputs obtained or derived from multiple different subjects. The independent training sample may include a dataset including medical images and associated datasets and outputs obtained from the same subject at multiple different points in time (e.g., periodically, such as weekly, monthly, or yearly). The independent training sample may be associated with the presence of cancer or disease (e.g., a training sample including a dataset including medical images and associated datasets and outputs obtained or derived from multiple subjects known to have cancer or disease). The independent training sample may be associated with the absence of cancer or disease (e.g., a training sample including a dataset containing medical images and associated datasets and output obtained or derived from multiple subjects who are known not to have been previously diagnosed with cancer or who have received negative test results for cancer or disease).
[0063]
[0083] The trained algorithm may be at least about 50, at least about 100, at least about The system may be trained using 250, at least about 500, at least about 1,000, at least about 5,000, at least about 10,000, at least about 15,000, at least about 20,000, at least about 25,000, at least about 30,000, at least about 35,000, at least about 40,000, at least about 45,000, at least about 50,000, at least about 100,000, at least about 150,000, at least about 200,000, at least about 250,000, at least about 300,000, at least about 350,000, at least about 400,000, at least about 450,000, or at least about 500,000 independent training samples. The independent training samples may include datasets including medical images associated with the presence of a disease (e.g., cancer) and / or datasets including medical images associated with the absence of a disease (e.g., cancer). The trained algorithm may be trained using independent training samples associated with the presence of about 500,000 or less, about 450,000 or less, about 400,000 or less, about 350,000 or less, about 300,000 or less, about 250,000 or less, about 200,000 or less, about 150,000 or less, about 100,000 or less, about 50,000 or less, about 25,000 or less, about 10,000 or less, about 500 or less, about 250 or less, about 100 or less, or about 50 or less diseases (e.g., cancer). In some embodiments, the dataset including medical images is unrelated to the samples used to train the trained algorithm.
[0064]
[0084] The trained algorithm generates independent predictors associated with the presence of a disease (e.g., cancer). The model may be trained using a first number of training samples and a second number of independent training samples associated with the absence of a disease (e.g., cancer). The first number of independent training samples associated with the presence of a disease (e.g., cancer) may be less than or equal to the second number of independent training samples associated with the absence of a disease (e.g., cancer). The first number of independent training samples associated with the presence of a disease (e.g., cancer) may be equal to the second number of independent training samples associated with the absence of a disease (e.g., cancer). The first number of independent training samples associated with the presence of a disease (e.g., cancer) may be greater than the second number of independent training samples associated with the absence of a disease (e.g., cancer).
[0065]
[0085] The trained algorithm has at least about 50%, at least about 55%, About 60%, at least about 65%, at least about 70%, at least about 75%, at least at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or higher accuracy, at least about 50, at least about 100, at least about 250, at least about 500 , at least about 1,000, at least about 5,000, at least about 10,000, at least about 15,000, at least about 20,000, at least about 25,000, at least about 30,000, at least about 35,000, at least about 40,000, at least about 45,000, at least about 50,000, at least about 100,000, at least about 150,000, at least about 200,000, at least about 250,000, at least about 300,000, at least about 350,000, at least about 400,000, at least about 450,000, or at least about 500,000 independent test samples. The accuracy of classifying medical images by the trained algorithm may be calculated as the percentage of independent test samples that are correctly identified or classified as normal or suspicious (e.g., images from subjects known to have cancer or subjects with negative clinical tests for cancer).
[0066]
[0086] The trained algorithm should be able to accurately interpret medical images at least approximately 5%, at least approximately 10%, , at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or more. The PPV of classifying medical images using a trained algorithm may be calculated as the percentage of medical images identified or classified as suspicious that correspond to subjects with a truly abnormal condition (e.g., cancer).
[0067]
[0087] The trained algorithm should be able to accurately interpret medical images at least approximately 5%, at least approximately 10%, , at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or more negative predictive value (NPV). The NPV of classifying medical images using a trained algorithm may be calculated as the percentage of medical images that are identified or classified as normal, which corresponds to subjects who do not truly have an abnormal condition (e.g., cancer).
[0068]
[0088] The trained algorithm should be able to accurately interpret medical images at least approximately 5%, at least approximately 10%, , at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83% , at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999% or more clinical sensitivity. The clinical sensitivity of classifying medical images using a trained algorithm may be calculated as the percentage of medical images obtained from subjects known to have a condition (e.g., cancer) that are correctly identified or classified as suspicious for that condition.
[0069]
[0089] The trained algorithm should be able to accurately interpret medical images at least approximately 5%, at least approximately 10%, , at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, or at least The trained algorithm may be configured to classify medical images with a clinical specificity of at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or more. Clinical specificity for classifying medical images using the trained algorithm may be calculated as the percentage of medical images obtained from subjects without a condition (e.g., subjects with negative laboratory tests for cancer) that are correctly identified or classified as normal for that condition.
[0070]
[0090] The trained algorithm processes medical images with a mean score of at least about 0.50, at least about 0. The algorithm may be configured to classify data sets containing medical images with an Area-Under-Curve (AUC) of at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, at least about 0.99, or more. The AUC may be calculated as the integral of a receiver operating characteristic (ROC) curve (e.g., the area under the ROC curve) associated with the trained algorithm in classifying a dataset containing medical images as normal or suspicious.
[0071]
[0091] The trained algorithms were evaluated for their performance in identifying cancer, accuracy, PPV, and NPV. The trained algorithm may be adjusted or fine-tuned to improve one or more of the following: clinical sensitivity, clinical specificity, or AUC. The trained algorithm may be adjusted or fine-tuned by adjusting the parameters of the trained algorithm (e.g., a set of cutoff values used to classify a dataset including medical images as described elsewhere herein, or parameters or weights of a neural network). Trained Algorithm may be continually adjusted or fine-tuned during the training process or after the training process is complete.
[0072]
[0092] After the first pre-trained algorithm is trained, a subset of the inputs is used to generate high-quality classifications. For example, a subset of a plurality of features of a dataset including medical images may be identified as the most influential or most important to include for high-quality classification or cancer identification. The plurality of features of a dataset including medical images, or a subset thereof, may be ranked based on a classification metric that indicates the influence or importance of each individual feature toward high-quality classification or cancer identification. Such metrics can be used to reduce, possibly significantly, the number of input variables (e.g., predictor variables) that can be used to train a trained algorithm to a desired performance level (e.g., based on a desired minimum accuracy, PPV, NPV, clinical sensitivity, clinical specificity, AUC, or a combination thereof). For example, if training a trained algorithm with a plurality of variables that includes tens to hundreds of input variables to the trained algorithm results in greater than 99% classification accuracy, instead training the trained algorithm with only a selected subset of such most influential or most significant input variables from the plurality of variables, such as about 5 or less, about 10 or less, about 15 or less, about 20 or less, about 25 or less, about 30 or less, about 35 or less, about 40 or less, about 45 or less, about 50 or less, or about 100 or less, may result in a reduced, but still acceptable, classification accuracy (e.g., at least about 50%, at least about 55%, At least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.The subset may be selected by rank ordering the input variables across the plurality of input variables and selecting a predetermined number (e.g., about 5 or less, about 10 or less, about 15 or less, about 20 or less, about 25 or less, about 30 or less, about 35 or less, about 40 or less, about 45 or less, about 50 or less, or about 100 or less) of input variables having the best classification metrics.
[0073] Identifying or monitoring cancer
[0093] A trained algorithm is used to generate a data set containing multiple medical images of parts of the subject's body. After processing the dataset to classify the images as normal, equivocal, or suspicious, cancer may be identified or monitored in the subject. The identification may be based, at least in part, on the classification of the images as normal, equivocal, or suspicious, on a plurality of features extracted from the dataset including the medical images, and / or on clinical health data of the subject. The identification may be performed by a radiologist, multiple radiologists, or a trained algorithm.
[0074]
[0094] Cancer is at least about 50%, at least about 55%, at least about 60%, The accuracy of cancer identification within a subject may be at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or greater. Accuracy of cancer identification may be calculated as the percentage of subjects of an independent test (e.g., subjects known to have cancer or subjects with a negative clinical test result for cancer) that are correctly identified or classified as having or not having cancer.
[0075]
[0095] Cancer is at least about 5%, at least about 10%, at least about 15%, The PPV for identifying cancer may be calculated as the percentage of subjects of an independent test that are identified or classified as having cancer, which corresponds to the percentage of subjects who truly have cancer.
[0076]
[0096] Cancer is at least about 5%, at least about 10%, at least about 15%, either may be identified in a subject with a negative predictive value (NPV) of about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or more. The NPV of identifying cancer using a trained algorithm may be calculated as the percentage of subjects from an independent test who are identified or classified as not having cancer, which corresponds to subjects who truly do not have cancer.
[0077]
[0097] Cancer is at least about 5%, at least about 10%, at least about 15%, At least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, %, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or more. Clinical sensitivity for identifying cancer may be calculated as the percentage of subjects for which an independent test associated with the presence of cancer (e.g., subjects known to have cancer) are correctly identified or classified as having cancer.
[0078]
[0098] Cancer is at least about 5%, at least about 10%, at least about 15%, At least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, The clinical specificity for identifying cancer may be calculated as the percentage of subjects with an independent test associated with the absence of cancer (e.g., subjects with a negative clinical test result for cancer) that are correctly identified or classified as not having cancer.
[0079]
[0099] In some embodiments, the subject may be identified as being at risk for cancer. After identifying a subject as at risk for cancer, a clinical intervention may be selected for the subject based at least in part on the cancer for which the subject was identified as at risk. In some embodiments, the clinical intervention is selected from a plurality of clinical interventions (e.g., clinically indicated for various types of cancer).
[0080]
[0100] In some embodiments, the trained algorithm determines whether the subject is at least about 5% , at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more.
[0081]
[0101] The trained algorithm is designed to ensure that subjects are at least about 50%, at least about 55%, At least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about The risk of cancer may be determined with an accuracy of at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or even higher.
[0082]
[0102] If a subject is identified as having cancer, the subject may optionally receive therapeutic intervention. Therapeutic intervention may include prescribing an effective amount of a drug, further testing or evaluation of the cancer, further monitoring of the cancer, or a combination thereof. If the subject is currently being treated for cancer with a course of treatment, therapeutic intervention may include a subsequent, different course of treatment (e.g., increasing the effectiveness of the treatment because the current course of treatment is ineffective).
[0083]
[0103] Therapeutic interventions may include recommending secondary laboratory testing in subjects to confirm a cancer diagnosis. This secondary laboratory testing may include imaging, blood tests, computerized diagnostic imaging, and The scans may include a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest x-ray, a positron emission tomography (PET) scan, a PET-CT scan, or any combination thereof.
[0084]
[0104] Classification of images as normal, equivocal, or suspicious,datasets containing medical images Multiple features extracted from and / or a subject's clinical health data may be evaluated over a period of time to monitor a patient (e.g., a subject with cancer or a subject being treated for cancer). In some cases, the classification of a patient's medical images may change over the course of treatment. For example, features from a dataset of patients whose risk of cancer is reduced by effective treatment may shift to resemble the profile or distribution of healthy subjects (e.g., cancer-free subjects). Conversely, features from a dataset of patients whose risk of cancer is increased by ineffective treatment may shift to resemble the profile or distribution of subjects at higher risk of cancer or with more advanced cancer.
[0085]
[0105] The subject's cancer is monitored through a course of treatment to treat the subject's cancer. The monitoring may include assessing the cancer in the subject at two or more time points. The assessing may be based on at least a classification of the image as normal, equivocal, or suspicious, a plurality of features extracted from a dataset including medical images, and / or clinical health data of the subject determined at each of the two or more time points.
[0086]
[0106] In some embodiments, the classification of an image as normal, equivocal, or suspicious The differences in the features extracted from a dataset including medical images, and / or the subject's clinical health data determined between two or more time points may be indicative of one or more clinical indicators, such as (i) a diagnosis of cancer in the subject, (ii) a prognosis of cancer in the subject, (iii) an elevated risk of cancer in the subject, (iv) a reduced risk of cancer in the subject, (v) the effectiveness of a course of treatment for treating cancer in the subject, and (vi) the ineffectiveness of a course of treatment for treating cancer in the subject.
[0087]
[0107] In some embodiments, the classification of an image as normal, equivocal, or suspicious A difference in a plurality of features extracted from a dataset including medical images, and / or a subject's clinical health data determined between two or more time points may indicate a diagnosis of cancer in the subject. For example, if no cancer is detected in the subject at an earlier time point, but cancer is detected in the subject at a later time point, the difference indicates a diagnosis that the subject has cancer. A clinical action or decision may be made based on this indication of a diagnosis that the subject has cancer, such as prescribing a new therapeutic intervention for the subject. The clinical action or decision may include recommending a secondary laboratory test for the subject to confirm the cancer diagnosis. The secondary laboratory test may include an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest x-ray, a positron emission tomography (PET) scan, a PET-CT scan, or any combination thereof.
[0088]
[0108] In some embodiments, the classification of an image as normal, equivocal, or suspicious Differences in the features extracted from a dataset including medical images, and / or clinical health data of a subject determined between two or more time points may indicate a prognosis for the subject's cancer.
[0089]
[0109] In some embodiments, the classification of an image as normal, equivocal, or suspicious differences between multiple features extracted from datasets containing medical images, and / or The subject's clinical health data determined between the above time points may indicate that the subject has an elevated risk of cancer. For example, if cancer is detected in the subject at both the earlier and later time points and the difference is positive (e.g., increasing from the earlier time point to the later time point), the difference may indicate that the subject has an elevated risk of cancer. Clinical action or decision may be made based on this indication of an elevated cancer risk, such as prescribing the subject a new therapeutic intervention or switching therapeutic interventions (e.g., terminating the current treatment and prescribing a new treatment). Clinical action or decision may include recommending a secondary laboratory test for the subject to confirm the elevated cancer risk. The secondary laboratory test may include an imaging test, blood test, computed tomography (CT) scan, magnetic resonance imaging (MRI) scan, ultrasound scan, chest x-ray, positron emission tomography (PET) scan, PET-CT scan, or any combination thereof.
[0090]
[0110] In some embodiments, the classification of an image as normal, equivocal, or suspicious A difference in a cancer risk profile, features extracted from a dataset including medical images, and / or clinical health data of a subject determined between two or more time points may indicate a reduced risk of cancer for the subject. For example, if cancer is detected in the subject at both an earlier and a later time point and the difference is negative (e.g., decreasing from the earlier time point to the later time point), the difference may indicate a reduced risk of cancer for the subject. A clinical action or decision may be made based on this indication that the subject has a reduced risk of cancer (e.g., continuing or terminating a current therapeutic intervention). The clinical action or decision may include recommending a secondary laboratory test for the subject to confirm the reduced risk of cancer. The secondary laboratory test may include an imaging test, blood test, computed tomography (CT) scan, magnetic resonance imaging (MRI) scan, ultrasound scan, chest x-ray, positron emission tomography (PET) scan, PET-CT scan, or any combination thereof.
[0091]
[0111] In some embodiments, the classification of an image as normal, equivocal, or suspicious The difference between the features extracted from a dataset including medical images, and / or clinical health data of a subject determined between two or more time points may indicate the effectiveness of a course of treatment for treating the subject's cancer. For example, if cancer is detected in the subject at an earlier time point but not at a later time point, the difference indicates the effectiveness of the course of treatment for treating the subject's cancer. A clinical action or decision may be made based on this indication of the effectiveness of the course of treatment for treating the subject's cancer, such as continuing or terminating the subject's current therapeutic intervention. The clinical action or decision may include recommending a secondary laboratory test for the subject to confirm the effectiveness of the course of treatment for treating the cancer. The secondary laboratory test may include an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest x-ray, a positron emission tomography (PET) scan, a PET-CT scan, or any combination thereof.
[0092]
[0112] In some embodiments, the classification of an image as normal, equivocal, or suspicious Differences in features extracted from a dataset including medical images, and / or clinical health data of a subject determined between two or more time points may indicate the ineffectiveness of a course of treatment to treat cancer in the subject. For example, if cancer is detected in a subject at both an earlier and later time point, and the difference is positive or zero (e.g., increases or remains constant from the earlier to the later time point), and an effective treatment was indicated at the earlier time point, the difference may indicate the ineffectiveness of a course of treatment to treat cancer in the subject. A clinical action or decision may be made based on this indication of ineffectiveness of the course of treatment to treat the subject's cancer, for example, terminating the current therapeutic intervention and / or switching to (e.g., prescribing) a different new therapeutic intervention for the subject. The clinical action or decision may include recommending a secondary laboratory test for the subject to confirm the ineffectiveness of the course of treatment to treat the cancer. The secondary laboratory test may include an imaging test, blood test, computed tomography (CT) scan, magnetic resonance imaging (MRI) scan, ultrasound scan, chest x-ray, positron emission tomography (PET) scan, PET-CT scan, or any combination thereof.
[0093]
[0113] Printing disease reports
[0114] After cancer is identified or the risk of disease or cancer is increased in a subject After the subject has been monitored, a report may be electronically output that indicates the disease or cancer in the subject (e.g., identifies or provides an indication of the disease or cancer). The subject may not exhibit the disease or cancer (e.g., the disease or cancer is asymptomatic, such as due to complications). The report may be presented on a graphical user interface (GUI) of the user's electronic device. The user may be the subject, a caregiver, a doctor, a nurse, or another healthcare worker.
[0094]
[0115] The report may include: (i) a diagnosis of cancer in the subject; (ii) a prognosis of the subject's disease or cancer; The report may then include one or more clinical indicators, such as (iii) an increased risk of the subject's disease or cancer, (iv) a decreased risk of the subject's disease or cancer, (v) the effectiveness of a course of treatment for treating the subject's disease or cancer, (vi) the ineffectiveness of a course of treatment for treating the subject's disease or cancer, (vii) the location and / or level of suspicion of the disease or cancer, and (viii) a measure of the effectiveness of a proposed course of diagnosis for the disease or cancer. The report may include one or more clinical actions or decisions made based on such one or more clinical indicators. Such clinical actions or decisions may be directed to therapeutic intervention or further clinical evaluation or testing of the subject's disease or cancer.
[0095]
[0116] For example, a clinical indicator of a subject's disease or cancer diagnosis may be a new The clinical action may involve prescribing a therapeutic intervention. As another example, a clinical indication of an increased risk of disease or cancer in a subject may involve the clinical action of prescribing a new therapeutic intervention for the subject or switching therapeutic interventions (e.g., terminating a current treatment and prescribing a new treatment). As another example, a clinical indication of a decreased risk of disease or cancer in a subject may involve the clinical action of continuing or terminating a current therapeutic intervention for the subject. As another example, a clinical indication of the effectiveness of a course of treatment for treating disease or cancer in a subject may involve the clinical action of continuing or terminating a current therapeutic intervention for the subject. As another example, a clinical indication of the ineffectiveness of a course of treatment for treating disease or cancer in a subject may involve the clinical action of terminating a current therapeutic intervention for the subject and / or switching (e.g., prescribing) a different new therapeutic intervention. As another example, a clinical indication of the site of disease or cancer may involve the clinical action of prescribing a new diagnostic test, particularly any particular parameter of that test that may be the target of the indicator.
[0096] Computer Systems
[0117] The present disclosure also provides a computer system programmed to implement the methods of the present disclosure. FIG. 4 illustrates a computer that is programmed or configured to, for example, train and test a trained algorithm, process medical images using the trained algorithm to classify the images as normal, equivocal, or suspicious, identify or monitor cancer in a subject, and electronically output a report indicating cancer in the subject. 4 shows a computer system 401.
[0097]
[0118] The computer system 401 may be, for example, a computer system for training and testing a trained algorithm. The computer system 401 may coordinate various aspects of the analysis, calculation, and generation of the present disclosure, such as analyzing, computing, and generating a medical image using a trained algorithm to classify the image as normal, equivocal, or suspicious, identifying or monitoring cancer in a subject, and electronically outputting a report indicating cancer in a subject. The computer system 401 may be a user's electronic device or a computer system located remotely to the electronic device. The electronic device may be a mobile electronic device.
[0098]
[0119] The computer system 401 includes a central processing unit (CPU, herein referred to as a “processor”). The computer system 401 includes a CPU 405, which may be a single-core or multi-core processor, or may have multiple processors for parallel processing. The computer system 401 also includes memory or memory locations 410 (e.g., random access memory, read-only memory, flash memory), electronic storage 415 (e.g., hard disk), a communication interface 420 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 425, such as cache, other memory, data storage, and / or electronic display adapters. The memory 410, storage 415, interface 420, and peripheral devices 425 are in communication with the CPU 405 through a communication bus (solid lines), such as a motherboard. The storage 415 may be a data storage device (or data repository) for storing data. The computer system 401 may be operably coupled to a computer network ("network") 430 with the aid of the communication interface 420. The network 430 may be the Internet, an Internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet.
[0099]
[0120] In some cases, the network 430 may be a telecommunications and / or data network. Network 430 can include one or more computer servers that can enable distributed computing, such as cloud computing. For example, one or more computer servers can enable cloud computing on network 430 (the "cloud") to perform various aspects of the analysis, calculation, and generation of the present disclosure, such as training and testing trained algorithms, using the trained algorithms to process medical images and classify the images as normal, equivocal, or suspicious, identifying or monitoring cancer in a subject, and electronically outputting a report indicating cancer in a subject. Such cloud computing can be provided by, for example, Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM. The network 430 may be implemented by a cloud computing platform such as a cloud. The network 430, in some cases aided by the computer system 401, may implement a peer-to-peer network, which may allow devices coupled to the computer system 401 to act as clients or servers.
[0100]
[0121] CPU 405 may be one or more computer processors and / or The CPU 405 may include one or more graphics processing units (GPUs). The CPU 405 is capable of executing sequences of machine-readable instructions, which may be embodied as a program or software. The instructions may be stored in a memory location, such as memory 410. The instructions may be sent to the CPU 405, which may then be programmed or configured to implement the methods of the present disclosure. Examples of operations that may be performed include fetch, decoder, execute, and writeback.
[0101]
[0122] The CPU 405 may be part of a circuit such as an integrated circuit. One or more other components of 401 may be included in a circuit, which in some cases is an application specific integrated circuit (ASIC).
[0102]
[0123] The storage device 415 stores drivers, libraries, and stored programs. The storage device 415 may store files. The storage device 415 may store user data, such as, for example, user settings and user programs. The computer system 401 may include one or more additional data storage devices that are external to the computer system 401, such as, in some cases, located on a remote server that communicates with the computer system 401 over an intranet or the Internet.
[0103]
[0124] The computer system 401 communicates with one or more It is possible to communicate with a remote computer system. For example, computer system 401 can communicate with a user's remote computer system. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad®, a Samsung® Galaxy Tab), a phone, a smartphone (e.g., an Apple® iPhone®, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access computer system 401 via network 430.
[0104]
[0125] The methods described herein utilize electronic storage locations in computer system 401. The instructions may be implemented by machine (e.g., a computer processor) executable code stored on a computer, such as on memory 410 or electronic storage 415. Machine-executable or machine-readable code may be provided in the form of software. During use, the code is executable by processor 405. In some cases, the code may be retrieved from storage 415 and stored in memory 410 for easy access by processor 405. In some situations, electronic storage 415 may be omitted, and machine-executable instructions are stored in memory 410.
[0105]
[0126] The code may be used in a machine having a processor configured to execute the code. The code can be precompiled and configured for execution, or can be compiled during run-time. The code can be supplied in a programming language that can be selected so that the code can be executed in a precompiled or as-compiled manner.
[0106]
[0127] The systems and methods provided herein, such as computer system 401 Aspects may be embodied as programming. Various aspects of the technology may be thought of as a "product" or "article of manufacture," typically in the form of machine (or processor) executable code and / or associated data executed on or embodied in some type of machine-readable medium. The machine-executable code may be stored in electronic storage, such as memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. "Storage" type media may include any or all of the tangible memory of a computer, processor, etc., or its associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which may provide non-transitory storage for software programming at any time. Software All or portions of may sometimes be communicated over the Internet or various other telecommunications networks. Such communication may, for example, allow the software to be loaded from one computer or processor to another, e.g., from a management server or host computer to an application server computer platform. Thus, other types of media that may carry software elements include optical, electrical, and electromagnetic waves, such as those used between physical interfaces between local devices, through wired and optical landline networks, and on various air-links. Physical elements, such as wired or wireless links, optical links, that carry such waves may also be considered software-bearing media. As used herein, unless qualified as a non-transitory, tangible "storage" medium, the term computer or machine "readable medium" refers to any medium that participates in providing instructions to a processor for execution.
[0107]
[0128] Thus, machine-readable media such as computer-executable code may be considered tangible storage media. The tangible transmission media may take the form of, but may not be limited to, a medium, a carrier wave medium, or a physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, as shown in the drawings, any of the storage devices in any computer, such as those that may be used to implement a database, etc. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire, and optical fiber, including wiring that comprises a bus within a computer system. Carrier wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic medium, CD-ROMs, DVDs or DVD-ROMs, any other optical medium, punched card paper tape, any other physical storage medium with a pattern of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves that transport data or instructions, cables or links that transport such carrier waves, or any other medium from which a computer can read programming code and / or data. Many such forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0108]
[0129] The computer system 401 may include an electronic display 435 or may be The electronic display 435 can be in communication with a user interface (UI) 440 for providing, for example, visual displays indicating training and testing of a trained algorithm, visual displays of image data indicating classification as normal, equivocal, or suspicious, identification of a subject as having cancer, or an electronic report (e.g., a diagnostic or radiology report) indicating cancer in a subject. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0109]
[0130] The methods and systems of the present disclosure are implemented by one or more algorithms. The algorithms are implemented in software when executed by the central processing unit 405. The algorithms may, for example, train and test a trained algorithm, process medical images using the trained algorithm to classify the images as normal, equivocal, or suspicious, identify or monitor cancer in a subject, and electronically output a report indicating cancer in a subject.
[0110] Example
[0131] [Example]
[0111] Improving patient care with real-time radiology
[0132] The systems and methods of the present disclosure can be used to perform real-time radiology screening. The diagnostic workflow was performed on multiple patients. For example, on the first day of the real-time radiology clinic, patients received immediate results for normal cases, leaving them feeling reassured and calm.
[0112]
[0133] In another example, the day after the real-time radiology clinic, another patient was screened. A suspicious finding was received during the scan, and diagnostic follow-up for the suspicious finding was performed within three hours. The patient was told by the radiologist that the finding was benign and not suspected to be cancerous. The patient was very relieved and happy to avoid the anxiety of waiting for a final diagnosis. On average, anywhere in the United States, such a process can take two to eight weeks. Even in certain clinics with rapid workflows, the process can take one to two weeks without the assistance of real-time radiology.
[0113]
[0134] Another example: on another day in a real-time radiology clinic, A real-time radiology system detected a 3 mm breast cancer tumor that was confirmed as cancerous by biopsy five days later. FIG. 5 shows an exemplary plot of the detection frequency of breast cancer tumors of various sizes (ranging from 2 mm to 29 mm) detected by radiologists, according to disclosed embodiments. A real-time radiology system can provide life-saving clinical impact by reducing time to treatment. The cancer may continue to grow until the patient undergoes their next screening or diagnostic procedure, at which time removal or treatment may become more life-threatening, painful, expensive, and less successful.
[0114]
[0135] Another example of a real-time radiology clinic is when a patient is diagnosed with a suspected condition within an hour. The patient received a diagnostic follow-up procedure for the new findings. A biopsy was required, but because the patient was taking aspirin, the biopsy was completed the next office day. The biopsy confirmed the cancer detected by real-time radiology. The radiology workup period was reduced from 8 office days to 1 day, and the time to diagnosis was reduced from 1 month to 1 week.
[0115]
[0136] The clinical impact of real-time radiology systems is based on PPV1 and callbacks. PPV1 may be measured by screening mammography metrics such as callback rate. PPV1 generally refers to the percentage of consultations with abnormal initial interpretations by radiologists that result in a tissue diagnosis of cancer within one year. Callback rate generally refers to the percentage of consultations with abnormal initial interpretations (e.g., "recall rate"). Over a six-week span, the real-time radiology clinic processed 796 patient cases using AI-based analytics, of which 94 were flagged to be reviewed by a radiologist in real time. A total of four cases were diagnosed with cancer, of which three were confirmed as cancer (e.g., by biopsy).
[0116]
[0137] FIG. 6 illustrates a positive result from a screening mammogram in accordance with disclosed embodiments. An exemplary plot of predictive value (PPV1) versus callback rate is shown. A prospective study had a callback rate of 11.8% with a PPV1 of 3.2%. In contrast, intermediate radiologists had a callback rate of 11.6% with a PPV1 of 4.4%.
[0117]
[0138] FIG. 7 illustrates a first set of radiologists, radiologists' AI sorted batches for the second set and the entire total set of radiologists An example plot comparing interpretation time (including Bi-RADS Assessment and density) for reading images in a randomly shuffled batch (left) and the percentage improvement in interpretation time relative to a control reading randomly shuffled batches (right) is shown. This figure demonstrates that AI-driven workflows can improve radiologist productivity to a statistically significant degree (ranging from approximately 13% to 21%).
[0118]
[0139] [Example]
[0119] Classification of suspicious findings in screening mammography using deep neural networks
[0140] Deep learning can be applied to a variety of computer vision and image processing applications. For example, deep learning may be used to automatically learn image features relevant to a given task, and may be used for a variety of tasks ranging from classification and detection to segmentation. Computational models based on deep neural networks (DNNs) have been developed for radiology applications such as screening mammography and are used to identify suspicious and potentially abnormal or high-risk lesions, improving radiologist productivity. In some cases, deep learning models can match or even exceed human-level performance. Additionally, deep learning can be used to help improve the performance of general radiologists to approach that of breast imaging experts. For example, general radiologists typically have lower cancer detection rates and much higher recall rates than fellowship-trained breast radiologists.
[0120]
[0141] Deep learning can improve the accuracy of screening mammograms, including distinguishing between malignant and benign findings. DNN models can be used to interpret images, identifying missed cancers or reducing false positive callbacks, especially for non-expert readers.
[0121]
[0142] The DNN model is stored in a publicly accessible Digital Database. The dataset was trained using the Dataset for Screening Mammography (DDSM) (eng.usf.edu / cvprg / Mammography / Database.html). The DDSM contains 2,620 cases with over 10,000 digitized scan-film mammography images. The images were evenly divided into normal mammograms and mammograms with suspicious findings. Normal mammograms were confirmed over a 4-year follow-up period. Suspicious findings were further divided into biopsy-verified benign findings (51%) and biopsy-verified malignant findings (49%). All cases with apparently benign findings not followed up with biopsy as part of routine clinical care were excluded from the dataset. As a result, distinguishing benign from malignant findings in this dataset can be more challenging than in a typical clinical mammography screening scenario.
[0122]
[0143] The DDSM dataset consists of a subset containing the training dataset, a validation dataset, and The dataset was split into a training dataset and a test dataset. Using the training dataset, the DNN was trained to distinguish between benign or normal areas of the breast and malignant findings. The dataset contained annotations pointing out the location of tumors in the image, which can be crucial in guiding the deep learning process.
[0123]
[0144] The performance of the DNN for this binary classification task is evaluated using the receiver operating characteristic (RO C) The DNN model was evaluated against the test dataset through the use of the ROC curve (as shown in Figure 8). The DNN model distinguished between malignant and benign findings within the ROC curve (AUC) of 0.89. The DNN model was used to distinguish between malignant and benign findings with high accuracy, as indicated by the area. In contrast, radiologists can typically achieve 84.4% sensitivity and 90.8% specificity for the task of cancer detection in screening mammograms. The DNN model was used to distinguish between malignant and benign findings with 79.2% sensitivity and 80.0% specificity for the more challenging cases found in the DDSM dataset. The performance gap relative to radiologists is due in part to the relatively small size of the dataset and can be mitigated by incorporating a larger training dataset. Furthermore, the DNN model can be further configured to outperform typical radiologists in terms of accuracy, sensitivity, specificity, AUC, positive predictive value, negative predictive value, or a combination thereof.
[0124]
[0145] High-precision DNN models trained on limited public benchmark datasets Although the dataset is arguably more challenging than those in clinical settings, the DNN model was able to distinguish between malignant and benign findings with near-human-level performance.
[0125]
[0146] A similar DNN model is being developed in partnership with Washington University in St. Louis. The deep learning model may be trained using a clinical mammography dataset from the Joanne Knight Breast Health Center in St. Louis, MD. This dataset contains a large medical record database, encompassing over 100,000 patients, 4,000 biopsy-confirmed cancer cases, and over 400,000 imaging sessions consisting of 1.5 million images. The dataset may be manually or automatically labeled (e.g., by building annotations) to optimize the deep learning process. Because DNN performance improves significantly with the size of the training dataset, this unique, vast, and rich dataset may dramatically improve the sensitivity and specificity of DNN models compared to DNN models trained on DDSM data. Such highly accurate DNN models offer an opportunity for transformative improvements in breast cancer screening, ensuring all women have access to expert-level care.
[0126]
[0147] [Example]
[0127] Artificial Intelligence (AI)-Driven Radiology Clinic for Early Cancer Detection
[0148] Introduction
[0149] Breast cancer is the most common cancer among women in the United States, with over 250,000 new cases in 2017 alone. A diagnosis has been made. Approximately 1 in 8 women will be diagnosed with breast cancer at some point in their lifetime. Despite improvements in treatment, more than 40,000 women die from breast cancer each year in the United States. Widespread uptake of screening mammography has made great strides in reducing some breast cancer mortality rates (a 39% decrease since 1989). Breast cancer screening can help identify early-stage cancers, which have a much better prognosis and lower treatment costs compared with late-stage cancers. This difference can be very significant: women with localized breast cancer have a 5-year survival rate of nearly 99%, while women with metastatic breast cancer have a 5-year survival rate of 27%.
[0128]
[0150] Despite these proven benefits, only about half of women currently use Amex. Not receiving mammograms at the rates recommended by the American College of Radiology. This low mammography utilization can result in significant burdens for patients and for the health care system in the form of increased expenses and costs. Screening mammography uptake is hindered in part by poor patient experience, including long delays in obtaining an appointment, unclear pricing, long wait times to receive results, and confusing reports. Furthermore, the problems arising from a lack of pricing transparency are exacerbated by wide variations in costs across health care facilities. Similarly, the cost of receiving results can be difficult to justify. Transmission times are inconsistent between healthcare facilities.
[0129]
[0151] In addition, significant variability in radiologist performance can lead to patient location and attendance. Radiologists experience highly variable standards of care depending on their experience. For example, cancer detection rates are more than twice as high for radiologists in the 90th percentile compared with those in the 10th percentile. False-positive rates (e.g., the percentage of healthy patients incorrectly recalled for follow-up visits) vary even more significantly between these two groups. Aggregating all screening tests performed in the United States, approximately 96% of patients who are called back are false-positives. Given the significant societal and personal burden of cancer, often coupled with poor patient experience, inconsistent screening performance, and large cost variability, AI-based or AI-assisted screening methods can be developed to significantly improve this clinical accuracy of mammography screening.
[0130]
[0152] Innovations in artificial intelligence and software could lead to early detection of cancer AI can be leveraged to achieve significant improvements in health outcomes, including accurate detection. Such improvements can impact one or more steps in patient behavior, from cost transparency, appointment scheduling, patient care, radiology workflow, diagnostic accuracy, and result communication to follow-up. AI-driven networks of imaging centers can be developed to achieve high-quality service, timeliness, accuracy, and cost-effectiveness. In such clinics, women can immediately schedule mammograms and receive a cancer diagnosis before they leave in a single visit. By using the "real-time radiology" methods and systems disclosed herein, AI-driven clinics can transform the traditional two-visit screening-diagnosis paradigm into a single visit. Artificial intelligence may be used to customize clinical workflow for each patient using a triage engine and to adjust how screening tests are read to significantly improve radiologist accuracy (e.g., by reducing radiologist fatigue), thereby improving the accuracy of cancer detection. AI-based or AI-assisted approaches can be used to achieve additional improvements to the screening / diagnostic process, such as patient scheduling, improved adherence to screening guidelines through customer outreach, and timeliness of report delivery using patient-facing applications. Self-improving systems may use AI to create better clinics that generate data to improve AI-based systems.
[0131]
[0153] A key component of creating an AI-powered radiology network is success through patient acquisition. While other components of the system can streamline radiology workflow processes and provide an improved and streamlined experience for patients, patient recruitment and enrollment is critical to collecting enough data to train the AI-driven system for high performance.
[0132]
[0154] Furthermore, AI-powered clinics can improve the patient experience before they even arrive at the clinic. Improvements may reduce barriers to screening mammography. This may involve addressing two major barriers that limit uptake: (1) concerns about the cost of the examination and (2) lack of awareness of conveniently located clinics. When pricing and availability are highly uncertain, as in traditional clinics, there can be significant variation in prices and services, creating barriers to patient scheduling of appointments.
[0133]
[0155] AI-based scheduling to streamline the scheduling process and provide transparency to patients A user application may be developed for the user. The system may be configured to provide the user with a map of Kiku Clinics, as well as available appointment times. For those with health insurance, both 2D and 3D screening mammograms are free of charge. This, along with any potential costs that may be incurred, may be made clear to the patient at the time of scheduling. Assurances regarding the timeliness of results may also be provided to the patient, addressing potential sources of patient anxiety that may discourage them from scheduling an appointment.
[0134]
[0156] The application verifies the patient's insurance and schedules the appointment if necessary. The application may be configured to request work orders from a primary care physician (PCP) during the appointment process. The application may be configured to receive user input of pre-appointment forms to more efficiently process patients during their clinic visit. If the patient has remaining forms to complete before their appointment, the patient may be given a device upon checking into the clinic to complete the remaining forms. The application may be configured to facilitate electronic completion of such forms to reduce or eliminate the time-consuming and error-prone tasks of handwritten paper forms, as is currently done in the standard of care. By facilitating user input of paperwork in advance of the appointment date, the application provides patients with a more streamlined experience, with less time and resources allocated to on-site operational tasks.
[0135]
[0157] The patient's previous mammograms may also be obtained prior to the consultation. With images acquired in a mobile clinic, this process can occur transparently to the patient. By capturing previous images prior to the visit, a potential bottleneck to the rapid review of newly acquired images can be eliminated.
[0136]
[0158] After scheduling an appointment, the application will take action to improve attendance. The application may be configured to provide patients with reminders about upcoming appointments. The application may also be configured to provide patients with information about the appointment in advance to minimize anxiety and reduce time spent in the exam room explaining the procedure. Additionally, to build relationships with primary care physicians (PCPs), referring physicians may check whether their patients have scheduled mammogram appointments. This allows physicians to assess compliance and encourage patients who have not yet scheduled appointments in a timely manner as recommended by their physicians.
[0137]
[0159] Real-time Radiology System
[0160] The traditional breast cancer screening paradigm results in significant delays that cause anxiety for patients. This may include a 20-day waiting period. This could reduce the number of women who choose to receive this preventive care, putting them at risk of finding their cancer later, when the treatment is more difficult and more life-threatening. A typical patient visits a clinic for a screening mammogram and leaves after spending about 30 minutes in the clinic. The woman then waits 30 days for a phone call or letter informing her that there was a suspicious abnormality on her screening mammogram and that she should schedule a follow-up diagnostic appointment. The patient then waits another week for that appointment, during which time she may undergo additional imaging to determine whether a biopsy is needed.
[0138]
[0161] The current paradigm is to implement large-scale (e.g., >100 patients per day) Motivated by the volume of patients being screened, such imaging centers typically have a backlog of screening visits that require at least 1-2 days of reading before radiologists can process the screening mammograms performed on a given day. If any of these cases require diagnostic workup, the visits can vary widely in the length of the diagnostic visit (e.g., 20 minutes to 12 minutes). 0 minutes) and therefore often cannot be done immediately. Scheduling does not take this into account, resulting in long wait times for patients and poor workflow for technicians.
[0139]
[0162] Received an immediate, real-time reading of their screening mammogram Patients may experience less anxiety than those who do not until three weeks later. In contrast, women who received a false-positive screening result (a normal case flagged as suspicious) but received an immediate interpretation experienced levels of anxiety similar to those of women who received normal mammograms. Most of these women were not aware that they had an abnormal screen. However, women who are aware that they have an abnormal screen tend to seek further medical attention for breast-related concerns and other medical problems. Furthermore, women may be more satisfied with the screening process and may be more likely to comply with future screening recommendations if they know they will leave the mammography clinic with their mammogram results. Such increased patient satisfaction may improve member retention in health plans. Additionally, immediate interpretation of suspicious cases may decrease the time to breast cancer diagnosis, thereby improving patient care and outcomes.
[0140]
[0163] In some cases, clinics may be able to limit real-time Such clinics may offer a variety of screening services. Such clinics may schedule only a few patients at a given time so that patients can immediately follow up the screening procedure with a diagnostic consultation if the need arises. This procedure can be expensive, time-consuming, and unacceptable on a large scale, which still means that most women must wait several weeks for potentially life-changing results. Roughly 4 million women may encounter such an unpleasant screening process each year.
[0141]
[0164] Using the methods and systems of the present disclosure, an AI-based triage system: It may also be developed for screening mammography.
[0165] Screening examination images are received from clinical imaging systems These images may be processed by an AI-driven Triage Engine, which then stratifies patient cases into one of multiple workflows. For example, the multiple workflows may include two categories (e.g., normal and suspicious). As another example, the multiple workflows may include three categories (e.g., normal, equivocal, and suspicious). Each such category may then be handled by a different set of dedicated radiologists who specialize in performing their particular set of workflows.
[0142]
[0166] FIG. 9 illustrates an AI-enabled real-time radiology system in accordance with disclosed embodiments. and a patient mobile application (app). The patient begins by registering on the website or patient app. The patient then schedules a radiology screening appointment using the patient app. The patient then completes a pre-examination form using the patient app. The patient then arrives at the clinic for their screening visit. An AI-based radiology evaluation is then performed on the medical images obtained from the patient's screening visit. The patient's images and visit results are then provided to the patient through the patient app. The patient then reschedules their appointment using the patient app, if necessary or recommended. The screening visit process may proceed as before.
[0143]
[0167] FIG. 10 illustrates an AI-assisted radiological evaluation workflow in accordance with disclosed embodiments. A simplified example is shown. First, a dataset containing a patient's electronic health record (EHR) and medical images is provided. Next, an AI-based triage engine processes the EHR and medical images to analyze the dataset and classify the dataset as likely normal, possibly suspicious, or likely suspicious. Next, a workflow distribution module distributes the patient's dataset to one of three workflows: a normal workflow, an uncertain workflow, and a suspicious workflow, based on the dataset's classification as likely normal, possibly suspicious, or likely suspicious, respectively. Each of the three workflows may include radiologist review or further AI-based analysis (e.g., by a trained algorithm).
[0144]
[0168] The majority of mammography screening examinations can be classified in the normal category. By focusing a first set of radiologists solely on this workflow, the concept and value of “batch reading” and its associated productivity gains can be applied and expanded. Because the cases handled by this first set of radiologists are likely to be almost all normal, there may be fewer context switches and penalties for handling significantly unusual cases. In an AI-based system, reports may be automatically pre-populated, allowing radiologists to spend significantly more time interpreting images rather than writing reports. In the rare cases where a radiologist disagrees with the AI-assessed normal case and considers the case suspicious, the case may be handled as usual and the patient may be scheduled for a diagnostic consultation. These normal cases may be further subdivided into even more homogeneous batches to realize productivity improvements by grouping cases that the AI-based system has determined to be similar. For example, batching all AI-assessed dense breasts together, or batching cases that are visually similar based on AI-derived features.
[0145]
[0169] A small percentage of mammography screening visits are uncertain workflows Such sessions may be classified as "normal," but the AI system may not classify them as normal, but they may involve findings that do not meet the fully suspicious threshold. These may be highly complex cases requiring significantly more time per session of radiologist evaluation than cases in the normal or suspicious workflow. For this reason, it may be beneficial to focus a second set of distinct radiologists on this smaller set of tasks, which are less homogeneous and potentially have significantly more interpretation and reporting requirements. These radiologists, through years of experience or training, have greater expertise in interpreting these challenging cases. This specialization may be even more specific based on the category or characteristics identified by the AI. For example, one group of radiologists may perform better than another group at correctly assessing AI-determined tumor masses. Therefore, consultations so identified by the algorithm may be routed to this more suitable group of experts. In some cases, the second set of radiologists is the same as the first set of radiologists, but radiological evaluations of different sets of cases are performed at different times based on case priority. In some cases, the second set of radiologists is a subset of the first set of radiologists.
[0146]
[0170] The smallest but most important part of a mammography screening examination is screening for suspected The workflow for these cases may be classified as a "suspicious case" workflow. A third set of radiologists may be assigned to this role, effectively reading these cases as part of their "on-call" duties. Most of the radiologist's time may be spent performing scheduled diagnostic consultations. However, in downtime between consultations, the radiologist may be alerted to any suspicious cases so that the diagnosis can be verified as soon as possible. Such cases may be used to ensure that the patient begins a follow-up diagnostic consultation as soon as possible. In some cases, the third set of radiologists is the same as the first or second set of radiologists, but the radiological evaluation of different sets of cases occurs at different times based on case priority. In some cases, the third set of radiologists is a subset of the first or second set of radiologists.
[0147]
[0171] In some cases, the workflow involves analyzing medical images and providing a medical imaging radiologist The method may include applying an AI-based algorithm to determine the difficulty of performing a radiological evaluation, and then prioritizing medical images for radiological evaluation or assigning medical images to a set of radiologists (e.g., among a plurality of different sets of radiologists) based on the determined degree of difficulty. For example, less difficult cases (e.g., more "routine" cases) may be assigned to a set of radiologists with a relatively lower degree of skill or experience, while more difficult cases (e.g., more questionable or unconventional cases) may be assigned to a different set of radiologists with a relatively higher degree of skill or experience (specialized radiologists). For example, less difficult cases (e.g., more "routine" cases) may be assigned to a first set of radiologists with a relatively lower level of schedule availability, while more difficult cases (e.g., more questionable or unconventional cases) may be assigned to a different set of radiologists with a relatively higher level of schedule availability.
[0148]
[0172] In some cases, the degree of difficulty is determined by the estimated time required to fully evaluate the image. It may be measured by length (e.g., about 1 minute, about 2 minutes, about 3 minutes, about 4 minutes, about 5 minutes, about 6 minutes, about 7 minutes, about 8 minutes, about 9 minutes, about 10 minutes, about 15 minutes, about 20 minutes, about 25 minutes, about 30 minutes, about 40 minutes, about 50 minutes, about 60 minutes, or more than about 60 minutes). In some cases, the degree of difficulty may be measured by the estimated degree of agreement or consensus in radiological evaluations of medical images across multiple unrelated radiological evaluations (e.g., by different radiologists or by the same radiologist on different days). For example, the estimated degree of agreement or consensus in radiological evaluations may be about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or greater than about 99%. In some cases, the degree of difficulty may be measured by the desired level of radiologist education, experience, or expertise (e.g., less than about 1 year, about 1 year, between 1 and 2 years, about 2 years, between 2 and 3 years, about 3 years, between 3 and 4 years, about 4 years, between 4 and 5 years, about 5 years, between 5 and 6 years, about 6 years, between 6 and 7 years, about 7 years, between 7 and 8 years, about 8 years, between 8 and 9 years, about 9 years, between 9 and 10 years, about 10 years, or more than about 10 years). In some cases, the degree of difficulty may be measured by the estimated sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), or accuracy of the radiological assessment (e.g., about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or greater than about 99%).
[0149]
[0173] In some cases, the workflow analyzes medical images to identify categories of medical images. The workflow may include applying an AI-based algorithm to analyze medical images to determine a lesion type in the medical images, and then prioritizing the medical images for radiological evaluation or assigning the medical images to a set of radiologists (e.g., among a plurality of different sets of radiologists) based on the determined categorization of the medical images. For example, a set of cases with similar characteristics may be categorized together and assigned to the same radiologist or set of radiologists, thereby achieving reduced context switches and increased efficiency and accuracy. The similar characteristics may be based, for example, on the region of the body where the ROI is located, tissue density, BIRADS score, etc. In some cases, the workflow may include applying an AI-based algorithm to analyze medical images to determine a lesion type in the medical images, and then prioritizing the medical images for radiological evaluation based on the determined lesion type in the medical images. This may include prioritizing medical images for review or assigning medical images to a set of radiologists (e.g., among a plurality of different sets of radiologists).
[0150]
[0174] In some cases, workflows are handed down to radiologists through market-based systems. This workflow may include having cases self-assigned, whereby each case is evaluated by an AI-based algorithm to determine the appropriate price or cost for a radiology evaluation. Such price or cost may be a defined relative unit of value that is reimbursed to each radiologist upon completion of the radiology evaluation. For example, each radiology evaluation of a case may be priced based on defined characteristics (e.g., difficulty, length of consultation). In such a workflow, cases may not be assigned to radiologists, thereby avoiding the problem of radiologists choosing relatively routine or easy cases to earn a high reimbursement rate per case.
[0151]
[0175] In some cases, the workflow may involve aligning cases with the radiologist's assessed performance. This may include assigning radiologists roles based on their performance (e.g., the radiologist's previous sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, or efficiency in performing radiological assessments). Such performance may be determined or improved based on blinded assignment of control cases (e.g., positive or negative control cases) to radiologists to ensure quality control. For example, radiologists with better performance may be assigned high-volume cases or high-value or high-compensation cases. By defining such clear roles for a given radiologist (e.g., on any given day), each workflow can be individually optimized for task-specific needs. An AI-driven triage engine can enable real-time radiology delivery to patients at scale. The system may also enable dynamic case allocation based on expertise. For example, a fellowship-trained breast imager may be the most valuable imager in an uncertain case workflow where their exceptional experience can be leveraged. Additionally, we are able to perform cross-clinic screen interpretation across a network of clinics, ensuring efficient use of radiologist time regardless of any individual clinic's staffing or patient base.
[0152]
[0176] Reporting may be done as follows: Mammography Qual The Medical Quality Standards Act (MQSA) mandates that all patients receive a written lay summary of their mammography report directly. This report must be sent within 30 days of the mammogram. While verbal results are often used to guide care and ease anxiety, they must be supported by a written report. Reports can be mailed, sent electronically, or handed to the patient. Typically, clinics can deliver reports to patients using paper mail. AI-based clinics may deliver mammography reports electronically through a patient application. Source images may also be made available electronically so patients can easily retrieve and transfer information to other clinics. Patients in a real-time radiology workflow can receive their screening and diagnostic reports immediately before leaving the clinic.
[0153]
[0177] Timely reporting of screening results is crucial to patient satisfaction. Waiting longer than two weeks for results and not being able to contact someone to have questions answered have been identified as major reasons contributing to patient dissatisfaction (which may also result in lower screening rates in the future). This system can ensure that patients do not unexpectedly receive erroneous reports and that there is no uncertainty about when they will receive their results.
[0154]
[0178] AI-based systems may be continually trained in clinical practice. As systems operate, new data is constantly collected and used to further train and refine AI systems, thereby further improving the quality of care and enabling new improvements to the patient experience. Each patient encounter provides the system with annotated, possibly live, examples to add to the dataset. In particular, the workflow of real-time radiology systems facilitates prioritizing the capture of high-value cases. Identifying false positives and false negatives (unflagged but suspicious cases) can be important for improving system performance by providing challenging examples with high educational value. Even cases that are correctly classified (e.g., as ground truth for radiologist review) can provide useful feedback. Incorporating such examples into the training dataset can provide the system with a valuable source of information about uncertain calibrations, which ensures that the confidence values produced by AI-based systems are accurate. This can dramatically improve overall robustness and, therefore, trust in the system. By improving the end-to-end patient workflow and keeping radiologists in the loop, AI-based clinical systems can automatically discover the important information outlined above. The resulting system is constantly improving, providing consistently high quality patient care and radiologist support.
[0155]
[0179] AI-powered mammography screening clinics streamline the screening process "The patient application is designed to provide patients with high-quality service and accuracy. Patients can visit the clinic, be screened for cancer, receive any necessary follow-up care, and leave with their diagnosis, thereby completing the entire screening and diagnosis process in a single visit with rapid results. The patient application is designed to provide price transparency, hassle-free scheduling, error-free form filling, and instant delivery of reports and images, thereby improving the ease, stress, and efficiency of the patient screening process."
[0156]
[0180] Radiologists are assigned a triage engine to triage patients into normal and uncertain cases. Adopting a specialized set of workflows for both positive and suspicious cases (or alternative categorizations based on AI evaluation of images) can provide more accurate and more productive results. Clinicians may become more effective as the AI system learns and enhances their capabilities. AI-based or AI-assisted mammography may be performed on a large population scale at low cost and with high efficiency, thereby improving the cancer screening process and patient outcomes.
[0157]
[0181] [Example]
[0158] Real-time radiology in breast cancer screening mammography when combined with artificial intelligence techniques
[0182] Suspicious screening mammograms for immediate review by a radiologist A software system configured to prioritize mammograms has been developed, thereby reducing the time to diagnostic follow-up. The software system aims to significantly reduce patient anxiety as well as the overall time to treatment by shortening the review time for suspicious mammogram cases. Reducing the wait time, which is often up to approximately 2-4 weeks between the first and second evaluations, can be expected to increase the life expectancy of these patients who are actually positive for breast cancer. An additional potential benefit is that the software may reduce the likelihood of missing any cancer.
[0159]
[0183] Some studies have shown that screening can result in false positives (positives flagged as suspicious). As is typically the case (BIRADS 0), women who receive immediate follow-up may experience levels of anxiety similar to those experienced by women who receive a normal diagnosis. Many of these women may not even be aware that they have an abnormal screening result. Therefore, immediate follow-up care can alleviate the potential anxiety caused by a false-positive screening result.
[0160]
[0184] On the other hand, you may receive a false-positive screening result and have it removed days or weeks later. Women who are called back for a follow-up diagnostic consultation tend to seek further medical attention for breast-related concerns and other medical problems. Thus, women who are able to receive their final mammography results during the same clinic visit as their mammography scan may be more satisfied with their screening experience and more likely to have high compliance with future screening recommendations.
[0161]
[0185] However, many breast imaging centers do not offer immediate follow-up examinations. In some cases, the reader may be unable to communicate the findings. This can be attributed to several difficulties, including scheduling constraints, the timeliness of receiving previous evaluations from other institutions, and the lost productivity of reading each examination immediately after acquisition. Perhaps most importantly, reading several breast screening cases together greatly improves the reader's assessment accuracy. This requires waiting until a sufficiently large batch of cases has been collected before reading the examination, making it impossible to provide immediate results and follow-up examinations to patients when indicated.
[0162]
[0186] Machine learning-based methods have been used to identify suspicious areas in mammography and tomosynthesis images. A triage software system is developed using machine learning for screening mammography to enable more timely report delivery and follow-up for suspicious cases (e.g., as done in a batch reading setting) (as shown in Figure 11). Medical images are provided to a real-time radiology system for processing. The real-time radiology system's AI-based triage engine processes the medical images and classifies them as suspicious or non-suspicious (e.g., normal or routine). If the image is classified as suspicious by the AI-based triage engine, it is sent for immediate radiologist review (e.g., during the same visit as the initial screening appointment or on the same day). The immediate radiologist review may confirm the suspicious case (leading to an immediate diagnostic consultation being ordered) or overturn the suspicious case (leading to the next scheduled routine annual screening). If the image is classified as not suspicious (e.g., normal or routine) by the AI-based triage engine, it is sent for routine radiologist review, which may assess the case as suspicious (leading to a routine diagnostic consultation being ordered) or may confirm the case as not suspicious (leading to the next scheduled routine annual screening being performed).
[0163]
[0187] The software allows large breast screening clinics to identify abnormalities It allows patients with mammography results to have diagnostic follow-up imaging on the same day or within the same visit. Leveraging such rapid diagnostic follow-up imaging can pave the way for breast imaging clinics to deliver the highest accuracy with the highest level of service and significantly reduce patient anxiety.
[0164]
[0188] Such machine learning-based methods can be used to assess patients and The time to treatment of the true tumor is reduced so that patients have a higher probability of living longer than patients who do not receive same-day follow-up diagnostic evaluation.
[0165]
[0189] To evaluate suspicious findings on mammography and tomosynthesis images The machine learning-based approach contemplates several advantages and objectives, as follows: First, the time from the initial screening visit to the communication of diagnostic imaging results for breast cancer screening is reduced (possibly significantly), improving the likelihood of an accurate diagnosis. For example, such a diagnosis may be produced with increased sensitivity, specificity, positive predictive value, negative predictive value, area under the receiver operating characteristic (AUROC), or a combination thereof. Second, an approach combining radiologists with artificial intelligence may efficiently improve the speed and / or quality of the initial assessment. Third, a more advanced diagnostic visit (e.g., additional X-ray-based imaging, ultrasound imaging, another type of medical imaging, or a combination thereof) may be completed within a short period (e.g., within 60 minutes) after the patient receives their screening results. Fourth, such a method may advantageously lead to improved patient satisfaction due to more timely communication of results and follow-up imaging.
[0166]
[0190] method
[0191] Clinical workflows are optimized to provide a higher level of service to patients. As more patients and data are collected in the training dataset, the machine learning algorithm continually improves the accuracy (or sensitivity, specificity, positive predictive value, negative predictive value, AUROC, or a combination thereof) of its computer-aided diagnosis.
[0167]
[0192] Computer algorithms and software can reduce the risk of breast screening images Software can be developed to automatically classify abnormal and normal categories with high accuracy. Such software could enable large-scale breast screening clinics to provide same-day or same-visit diagnostic follow-up imaging to patients with abnormal initial screening results. This would also require evaluating changes to clinical operations, particularly how screening cases are read and how a second diagnostic evaluation can be performed within the first 60 minutes of the exam.
[0168]
[0193] Rapid screening methods are available for all breast screening clinics. Approximately 10% of patients undergoing screening have a suspicious result and are subsequently recommended for a diagnostic consultation on the same day or during the same visit. Rapid turnaround time for screening results and follow-up diagnostic consultations is possible through careful coordination between radiologists, clinical staff, and patients in a clinical setting. As more information is collected, machine learning trained on increasingly large training datasets will provide even greater levels of accuracy in detecting suspicious mammography scans.
[0169]
[0194] Once the screening exam acquisition is complete, the images are sent to the router and processed by the software. The results are received by the machine learning algorithm and quickly classified (e.g., within about one minute). If the screening is marked as probably normal by the machine learning algorithm, the patient completes their visit and exits the clinic as usual. However, if the screening is flagged as probably abnormal by the machine learning algorithm, the patient is asked to wait up to about 10 minutes while the case is quickly reviewed by a radiologist (as shown in Figure 11).
[0170]
[0195] A given clinic screens approximately 30 patients per day, and Assuming a 10% chance of misidentification, approximately 3 patients per day would typically be identified as positive by the machine learning algorithm and, after review by a radiologist, designated as suitable for real-time diagnostic follow-up (e.g., typically with additional tomosynthesis imaging and possibly ultrasound consultation).
[0171]
[0196] To demonstrate the effectiveness of real-time radiology methods and systems, several Several metrics are used: First, the change in time between the patient's initial screening encounter and communication of diagnostic imaging results under the conventional and proposed real-time workflows can be measured to capture both changes in the delay in when a case's screening is reviewed, as well as changes in logistics such as mailing documents and scheduling appointments.
[0172]
[0197] Second, real-time radiology models are based on the most recent data collected. It is continually evaluated (e.g., on a monthly basis). For example, the parameters of the computer vision algorithm are tweaked and changed to improve its accuracy for the immediately subsequent screening period (e.g., one month). The effectiveness of changes to the computer program is evaluated against a blind test dataset of several hundred representative visits and from preliminary results from the subsequent screening period.
[0173]
[0198] Third, patient satisfaction surveys are reviewed periodically and are short (e.g., approximately 60 minutes). The study will help determine how operational processes can be improved to better enable follow-up diagnostic consultation within the hospital.
[0174]
[0199] The following data is available via a real-time radiology workflow: For each patient undergoing a screening / diagnostic evaluation, the following may be collected: patient demographics (e.g., age, race, height, weight, socioeconomic background, smoking status, etc.), patient imaging data (e.g., obtained by mammography), patient results (e.g., BIRADS for screening and diagnostic visits and biopsy pathology results if applicable), timestamps of patient visit events, patient callback rates for batch reading and real-time cases, and radiologist interpretation times for screening and diagnostic cases.
[0175]
[0200] The methods and systems of the present disclosure can be used to provide potential benefits including: Real-time radiology can: detect tumors that might otherwise go unrecognized (or only be recognized once the tumor has progressed); reduce time to treatment; improve patient longevity due to recognition and treatment compared to traditional evaluation processes; reduce patient anxiety due to the elimination of waiting times between exams.
[0176]
[0201] [Example]
[0177] Multi-site investigation of a deep learning model of breast density for full-field digital mammography and digital breast tomosynthesis examinations
[0202] overview
[0203] Deep learning (DL) models show promise for mammographic breast density estimation, but training Performance can be hindered by limited data or potential imaging differences across clinics. Digital breast tomosynthesis (DBT) exams are increasingly standard for breast cancer screening and breast density assessment, but much more data is available for full-field digital mammography (FFDM) exams. For synthetic 2D mammography (SM) images derived from 3D DBT examinations, a breast density DL model was developed in a multisite setting using FFDM images and limited SM data. Breast Imaging Reporting A DL model was trained to predict breast density using FFDM images acquired from 2008 to 2017 (Site 1: 57,492 patients, 750,752 images) for a retrospective study. The FFDM model was evaluated against SM datasets from two institutions (Site 1: 3,842 patients, 14,472 images; Site 2: 7,557 patients, 63,973 images). Adaptive methods were explored to improve performance on the SM datasets, and the effect of dataset size on each adaptive method was considered. Statistical significance was assessed through the use of confidence intervals and estimated by bootstrapping. Even without indications, the model showed close agreement with the initial reporting radiologist for all three datasets (Site 1 FFDM: linearly weighted κw = 0.75, 95% confidence interval (CI): [0.74, 0.76]; Site 1 SM: κw = 0.71, CI: [0.64, 0.78]; Site 2 SM: κw = 0.72, CI: [0.70, 0.75]). With indications, the use of only 500 SM images improved performance for Site 2 (Site 1: κw = 0.72, CI: [0.66, 0.79]; Site 2: κw = 0.79, CI: [0.76, 0.81]). These results establish that the BI-RADS breast density DL model demonstrated high levels of performance on FFDM and SM images from two institutions using methods that required little or no SM images.
[0178]
[0204] Multisite surveys are, for example, those by Matthews et al., "A Multisite Study of a Breast Density Deep Learning To develop a deep learning model for breast density for full-field digital mammography and synthetic mammography, as described in "Model for Full-Field Digital Mammography and Synthetic Mammography," Radiology:Artificial Intelligence, doi.org / 10.1148 / ryai.2020200015. This document is incorporated herein by reference in its entirety.
[0179]
[0205] Introduction
[0206] Breast density is an important risk factor for breast cancer, and the denser the area, the more likely It may mask findings in mammograms, further reducing sensitivity. Some conditions require clinics to inform women of their density. Radiologists typically divide breast density into four categories: The Biomarker and Data System (BI-RADS) lexicon was used to assess breast density: almost entirely fatty, scattered areas of fibroglandular density, heterogeneously dense, and extremely dense (as shown in Figures 12A-12D). Unfortunately, radiologists exhibit intra- and inter-reader variability in assessing BI-RADS breast density, which can translate into differences in clinical care and estimated risk.
[0180]
[0207] Figures 12A to 12D show four Breast Imaging Reports. Breast Imaging and Data System (BI-RADS) breast density categories: (A) almost entirely fatty (Figure 12A), (B) scattered areas of fibroglandular density (Figure 12B), and (C) heterogeneously Examples of composite 2D mammography (SM) images derived from digital breast tomosynthesis (DBT) examinations are shown for dense (heterogeneously dense) (Figure 12C) and extremely dense (Figure 12D). Images are normalized to fit the grayscale intensity window found in the Digital Imaging and Communications in Medicine (DICOM) header, ranging from 0.0 to 1.0.
[0181]
[0208] Deep learning (DL) is used to analyze the image quality of film images and full-field digital mammograms. BI-RADS models may be employed to assess breast density on both FFDM and FFDM images, with some models demonstrating closer agreement with consensus predictions than individual radiologists. To realize the promise of using such DL models in clinical practice, two key challenges must be met. First, as breast cancer screening increasingly shifts to digital breast tomosynthesis (DBT) due to improved reader performance, DL models may need to accommodate DBT consultations. Figures 13A-13D illustrate the differences in image characteristics between 2D images for FFDM and DBT consultations. However, the relatively recent adoption of DBT at many institutions means that the datasets available to train DL models are often quite limited for DBT consultations compared to FFDM consultations. Second, DL models may need to provide consistent performance across sites, where differences in imaging technology, patient demographics, or assessment practices may affect model performance. In practice, this may need to be achieved with little or no additional data from each site.
[0182]
[0209] 13A-13D are full-field images of the same breast under the same pressure on a subject. Comparison of a digital FFDM (Fast Fibroblast Disease Modeling) image (Figure 13A) and a composite 2D mammography (SM) image (Figure 13B), as well as zoomed-in areas to highlight possible texture and contrast differences between the two image types, with the original areas indicated by white boxes for both the FFDM (Figure 13C) and SM (Figure 13D) images. Images are normalized to fit the grayscale intensity window found in the Digital Imaging and Communications in Medicine (DICOM) header, ranging from 0.0 to 1.0.
[0183]
[0210] This is the first report of both FFDM and DBT consultations at two institutions. A BI-RADS breast density DL model was developed that provided close agreement with performing radiologists. First, the DL model was trained to predict BI-RADS breast density using a large FFDM dataset from one institution. The model was then evaluated on a set of FFDM exams from the same institution and from separate institutions, as well as synthetic 2D mammography (SM) images (C-View, Hologic, Inc., Marlborough, MA) generated as part of a DBT exam. To improve performance on the two SM datasets, adaptive techniques requiring fewer SM images were explored.
[0184]
[0211] Materials and Methods
[0212] A retrospective study was conducted by the Institutional Review Board at each of the two sites where data were collected. The study was approved (Site 1: internal Institutional Review Board, Site 2: Western Institutional Review Board). Informed consent was waived, and all data were handled in accordance with the Health Insurance Portability and Accountability Act.
[0185]
[0213] The dataset was collected from two sites: Site 1, located in the Midwestern United States; Academic Medical Center, Site 2 is an outpatient radiology clinic located in Northern California. At Site 1, 191,493 mammography examinations were selected (FFDM: n = 187,627; SM: n = 3,866). Examinations were reviewed by 1 of 11 radiologists with breast imaging experience. At Site 2, 16,283 examinations were selected. Examinations were reviewed by 1 of 12 radiologists with breast imaging experience ranging from 9 to 41 years. Radiologists' BI-RADS breast density assessments were obtained from each site's mammography reporting software (Site 1: Magview version 7.1, Magview, Burtonsville, MD; Site 2: MRS version 7.2.0; MRS Systems Inc., Seattle, WA). To facilitate the development of our DL model, patients were randomly assigned to training (FFDM: 50,700, 88%; Site 1 SM: 3,169, 82%; Site 2 SM: 6,056, 80%), validation (FFDM: 1,832, 3%; Site 1 SM: 403, 10%; Site 2 SM: 757, 10%), or testing (FFDM: 4,960, 9%; Site 1 SM: 270, 7%; Site 2 SM: 744, 10%). All visits with BI-RADS breast density assessment were included. For the exam set, visits were required to have all four standard screening mammography images (medial and lateral oblique views and craniocaudal views of both breasts). The distribution of BI-RADS breast density assessments per set is shown in Table 1 (Site 1) and Table 2 (Site 2).
[0186] [Table 1]
[0187]
[0214] Table 1: Full-field digital mammography (FFDM) and Description of the training (Train), validation (Val), and test (Test) datasets for 2D mammography and synthetic 2D mammography (SM). The total number of patients, visits, and images is given for each dataset. The number of images in the four Breast Imaging Reporting and Data System (BI-RADS) breast density categories is also given.
[0188] [Table 2]
[0189]
[0215] Table 2: Site 2 synthetic 2D mammography (SM) training (Train), Description of the validation (Val) and test (Test) datasets. The total number of patients, visits, and images is given for each dataset. The number of images in the four Breast Imaging Reporting and Data System (BI-RADS) breast density categories is also given.
[0190]
[0216] The two sites represent different patient populations. The patient cohort from Site 1 Site 1 was 59% Caucasian (34,192 / 58,397), 23% African American (13,201 / 58,397), 3% Asian (1,630 / 58,397), and 1% Hispanic (757 / 58,397), while Site 2 was 58% Caucasian (4,350 / 7557), 1% African American (110 / 7557), 21% Asian (1,594 / 7557), and 7% Hispanic (522 / 7557).
[0191]
[0217] Deep Learning Model
[0218] DL models and training procedures include deep neural network models. The model was implemented using the Pytorch DL framework (pytorch.org, version 1.0). The base model architecture included a pre-activation Resnet-34 in which batch normalization layers were replaced with group normalization layers. The model was configured to process as input a single image corresponding to one of the views from a mammography examination and produce estimated probabilities that the image was a breast image belonging to each of the BI-RADS breast density categories.
[0192]
[0219] Deep learning (DL) models have a learning speed of 10 -4 and weight decay is 10 -3 A The model was trained using the full-field digital mammography (FFDM) dataset (shown in Table 1) using the dam optimizer. Weight decay was not applied to parameters belonging to the normalization layer. The input was resized to 416 × 320 pixels, and pixel intensity values were normalized to fit the grayscale intensity window found in the Digital Imaging and Communications in Medicine (DICOM) header, ranging from 0.0 to 1.0. Training was performed using mixed precision and gradient checkpointing with a batch size of 256 distributed across two NVIDIA GTX 1080 Ti graphics processing units (Santa Clara, CA). Each batch was sampled such that the probability of selecting a BI-RADS B or BI-RADS C sample was four times higher than the probability of selecting a BI-RADS A or BI-RADS D sample, roughly corresponding to the density distribution found in the United States. Horizontal and vertical flipping was employed for data augmentation. To obtain more frequent information about the training progress, epochs were capped at 100,000 samples, for a total training set size of over 672,000 samples. The model was trained for 100 such epochs. Results are reported for the epoch with the smallest cross-entropy loss on the validation set, which occurred after 93 epochs.
[0193]
[0220] The parameters for the vector and matrix calibration methods are The parameters were chosen by minimizing the cross-entropy loss function using the optimization method (scipy.org, version 1.1.0). The parameters were initialized so that the linear layers correspond to the identity transformation. Training was performed using a linear layer with the L2 norm of the gradients set to 10 -6 The training was stopped when the learning rate was less than 10 or when the number of iterations exceeded 500. Retraining the final fully connected layer for fine-tuning was -4 and weight decay is 10 -5 The fine-tuning and training from scratch were performed using the Adam optimizer. The batch size was set to 64. The fully connected layers were trained for 100 epochs from random initialization, and the results are reported for the epoch with the smallest validation cross-entropy loss. Training from scratch on a synthetic 2D mammography (SM) dataset was performed following the same procedure as for the base model. For fine-tuning and training from scratch, the size of the epochs was The size was set to the number of training samples.
[0194]
[0221] Domain Adaptation
[0222] A model trained on a dataset from one domain (the source domain) Domain adaptation is performed using a DL model to transfer that knowledge to a dataset from another domain (the target domain), which is usually of much smaller size. Features learned by DL models in previous layers can be general, i.e., domain and task agnostic. Depending on the similarity of the domains and tasks, it is possible to reuse deeper features learned from one domain for another domain or task. A model that can be directly applied to a new domain without modification is said to generalize.
[0195]
[0223] The DL model trained on the FFDM image (source domain) is used to model the SM image ( A technique was developed to reuse all features learned from the FFDM domain to adapt to the target domain. First, to calibrate the neural network, a small linear layer was added after the last fully connected layer. Two forms of linear layer were considered: (1) vector calibration, where the matrix is diagonal, and (2) matrix calibration, where the matrix is allowed to deform freely. Second, the last fully connected layer of the Resnet-34 model was retrained on samples from the target domain, a process called fine-tuning.
[0196]
[0224] To investigate the effect of the dataset size of the target domain, an adaptation technique The adaptation process was repeated for various SM training sets over a range of sizes. The adaptation process was repeated 10 times for each dataset size using different random samples of the training data. For each sample, training images were selected randomly without replacement from the full training set. As a baseline, a Resnet-34 model was trained from scratch, e.g., from random initialization, on the maximum number of training samples for each SM dataset.
[0197]
[0225] Statistical analysis
[0226] To obtain a consultation-level assessment, each image in the consultation is processed using a DL model. The resulting probabilities were averaged. Several performance metrics were calculated from these average probabilities for the four-class BI-RADS breast density task and the binary high-density (BI-RADS C+D) vs. non-high-density (BI-RADS A+B) task: (1) estimated accuracy based on agreement with the initial reporting radiologist, (2) area under the receiver operating characteristic curve (AUC), and (3) Cohen's kappa coefficient (scikit-learn.org, version 0.20.0). Confidence intervals were calculated using non-Studentized pivoted bootstrap of the test set for 8,000 random samples. For the four-class problem, macroAUC (the average of the four AUC values from one task vs. the other) and linearly weighted Cohen's kappa coefficient were reported. For the binary density task, predicted high-density and non-high-density probabilities were calculated by adding the predicted probabilities for the corresponding BI-RADS density category.
[0198]
[0227] result
[0228] The performance of the deep learning model for FFDM examinations was evaluated as follows: First, the trained model was evaluated on a large, held-out set of FFDM examinations from Site 1 (4,960 patients, 53,048 images, mean age: 56.9, age range: 23-97). In this case, the images were from the same institution and of the same type as those employed to train the model. The BI-RADS breast density distribution predicted by the DL model (A: 8.5%, B: 52.2%, C: 36.1%, D: 3.0%) was compared. The 2% distribution of initial reporting radiologists was similar (A: 9.3%, B: 52.0%, C: 34.6%, D: 4.0%). The DL model demonstrated close agreement with radiologists for the 4-class BI-RADS breast density task across a variety of performance measures (as shown in Table 3), including accuracy (82.2%, 95% confidence interval (CI): [81.6%, 82.9%]) and linearly weighted Cohen's kappa coefficient (κw = 0.75, CI: [0.74, 0.76]). A high level of agreement was also observed for the binary breast density task (accuracy = 91.1%, CI: [90.6%, 91.6%]; AUC = 0.971, CI: [0.968, 0.973]; κ = 0.81, CI: [0.80, 0.82]). As demonstrated by the confusion matrices shown in Figures 14A-14D, the DL model rarely deviated from multiple breast density categories (e.g., by calling extremely dense breast a scattered result; 0.03%, 4 / 13262), which is learned implicitly by the DL model without any explicit penalty for these types of large errors.
[0199]
[0229] Figures 14A-14B show full-field digital mammography (FFDM) Confusion matrices are shown for the Breast Imaging Reporting and Data System (BI-RADS) breast density task (Figure 14A) and for the binary density task (BI-RADS C+D, which is dense, vs. BI-RADS A+B, which is non-dense) evaluated against the test set. The number of test samples (visits) in each bin is shown in parentheses.
[0200] [Table 3]
[0201]
[0230] Table 3: 4-Class Breast Imaging Reporting and Performance of the disclosed deep learning model on the full-field digital mammography (FFDM) consultation test set for both the Biometric Data System (BI-RADS) breast density task and the binary density task (BI-RADS C+D, which is dense, versus BI-RADS A+B, which is non-dense). 95% confidence intervals are given in brackets. Results from other studies are evaluated for their respective test sets and are provided as a point of comparison.
[0202]
[0231] To place the results in the context of other studies, The performance of the deep learning model was evaluated against other large FFDM datasets obtained from academic centers and compared with commercially available breast density software (as shown in Table 3). The FFDM DL model appears to give comparable performance.
[0203]
[0232] The performance of the deep learning model for DBT consultations was evaluated as follows: Results were first reported for the Site 1 SM study set (270 patients, 1080 images, mean age: 54.6, age range: 28–72), which may occur between the two sites. This is to avoid any discrepancies that may arise from the model's adaptation. As shown in Table 4, even when performed without adaptation, the model showed close agreement with the radiologist who originally reported on the BI-RADS breast density task (accuracy = 79%, CI: [74%, 84%]; κw = 0.71, CI: [0.64, 0.78]). The DL model slightly underestimated breast density on SM images (as shown in Figures 15A–15D), producing a BI-RADS breast density distribution with more non-dense cases and fewer dense cases (A: 10.4%, B: 57.8%, C: 28.9%, D: 3.0%) than the radiologist (A: 8.9%, B: 49.6%, C: 35.9%, D: 5.6%). This bias may be due to the difference shown in Figure 13, where certain areas of the breast appear darker on SM images. Similar biases have been shown in other automated breast density estimation software
[33] . Agreement for the binary density task was also extremely high without adaptation (accuracy = 88%, CI: [84%, 92%]; kappa = 0.75, CI: [0.67, 0.83]; AUC = 0.97, CI: [0.96, 0.99]).
[0204] [Table 4]
[0205]
[0233] Table 4: Deep learning (DL) models trained on one dataset, 50 Performance of the disclosed method and system for adapting another model with a set of 0 synthetic 2D mammography (SM) images. The datasets are denoted as "MM" for the full-field digital mammography (FFDM) dataset, "C1" for the Site 1 SM dataset, and "C2" for the Site 2 SM dataset. For reference, the performance of a model trained from scratch on the FFDM dataset (672,000 training samples) and evaluated on its test set is also shown. The 95% confidence intervals calculated by bootstrapping for the test set are given in brackets.
[0206]
[0234] After adaptation by matrix calibration using 500 SM images, the density distribution is Although the radiologist distributions (A: 5.9%, B: 53.7%, C: 35.9%, D: 4.4%) were more similar, overall agreement was similar (accuracy = 80%, CI: [76%, 85%]; κw = 0.72, CI: [0.66, 0.79]). Accuracy for the two dense classes improved at the expense of the two non-dense classes (as shown in Figures 15A-D). A significant improvement was seen in the binary density task, with Cohen's kappa coefficient increasing from 0.75 [0.67, 0.83] to 0.82 [0.76, 0.90] (accuracy = 91%, CI: [88%, 95%]; AUC = 0.97, CI: [0.96, 0.99]).
[0207]
[0235] Figures 15A-15D show the adaptive SM test set evaluated for Site 1. Confusion matrices are shown for the Breast Imaging Reporting and Data System (BI-RADS) breast density task (Figure 15A), the binary density task (BI-RADS C+D, which is high density, vs. BI-RADS A+B, which is non-high density) without adaptation (Figure 15B), the BI-RADS breast density task with adaptation using matrix calibration of 500 training samples (Figure 15C), and the binary density task (high density vs. non-high density) with adaptation using matrix calibration of 500 training samples (Figure 15B). The number of test samples (visits) in each bin is shown in parentheses.
[0208]
[0236] Site 2 SM test set without adaptation (744 patients, 6192 images, mean age For the SM data (age range: 30–92, mean age: 55.2, mean age: 30–92), a high degree of agreement was again observed between the DL model and the radiologist who originally reported the data (accuracy = 76%, CI: [74%, 78%]; κw = 0.72, CI: [0.70, 0.75], as shown in Table 4). The BI-RADS breast density distribution predicted by the DL model (A: 5.7%, B: 48.8%, C: 36.4%, D: 9.1%) was more similar to that seen in the Site 1 dataset. The model may not be optimal for Site 2, where patient demographics differ, possibly due to prior learning from the Site 1 FFDM dataset. The predicted density distribution does not appear to be biased toward low density estimates as seen in Site 1 (as shown in Figures 16A–16D). This may suggest some differences in the SM images or their interpretation between the two sites. Agreement was particularly strong for the binary density task (Accuracy = 92%, CI: [91%, 93%]; Kappa = 0.84, CI: [0.81, 0.87]; AUC = 0.980, CI: [0.976, 0.986]). The very good performance on the Site2 dataset without adaptation demonstrates that the DL model can generalize well across sites.
[0209]
[0237] By adapting the matrix calibration on 500 training samples, S Performance on the BI-RADS breast density task for the ite2 SM dataset improved significantly (precision = 80, CI: [78, 82]; κw = 0.79, CI: [0.76, 0.81]). After adaptation, the predicted BI-RADS breast density distribution (A: 16.9%, B: 43.3%, C: 29.4%, D: 10.4%) was more similar to the radiologist's distribution (A: 15.3%, B: 42.2%, C: 30.2%, D: 12.3%). Adaptation may have helped adjust for the demographic distribution of breast density at this site. There was less improvement in the binary breast density task (accuracy = 92, CI: [91, 94]; kappa = 0.84, CI: [0.82, 0.87]; AUC = 0.983, CI: [0.978, 0.988]).
[0210]
[0238] Figures 16A-16D show the adaptive SM test set evaluated for Site 2. Confusion matrices are shown for the Breast Imaging Reporting and Data System (BI-RADS) breast density task (Figure 16A), the binary density task without adaptation (BI-RADS C+D, which is high density, vs. BI-RADS A+B, which is non-high density) (Figure 16B), the BI-RADS breast density task with adaptation using matrix calibration of 500 training samples (Figure 16C), and the binary density task (high density vs. non-high density) with adaptation using matrix calibration of 500 training samples (Figure 16B). The number of test samples (visits) in each bin is shown in parentheses.
[0211]
[0239] The relative performance of different adaptation methods depends on the number of training samples available for adaptation. It may depend on the number of training samples, and more training samples may be beneficial for methods with more parameters. 17A-17D show the effect of the amount of training data on the performance of the adaptive method as measured by macroAUC for the Site 1 dataset (FIG. 17A), the effect of the amount of training data on the performance of the adaptive method as measured by linearly weighted Cohen's Kappa coefficient for the Site 1 dataset (FIG. 17B), the effect of the amount of training data on the performance of the adaptive method as measured by macroAUC for the Site 2 SM dataset (FIG. 17C), and the effect of the amount of training data on the performance of the adaptive method as measured by linearly weighted Cohen's Kappa coefficient for the Site 2 SM dataset (FIG. 17D), according to disclosed embodiments. Results are reported for 10 random realizations of the training data for each dataset size (as described elsewhere herein) to explore uncertainties arising from the selection of the training data rather than uncertainties arising from the limited size of the test set, as was done in calculating the 95% confidence intervals. Each adaptation method has a range of sample numbers for which it provides the best performance, with the range corresponding to the number of parameters of the adaptation method (vector calibration: 4 + 4 = 8 parameters; matrix calibration: 4 × 4 + 4 = 20 parameters; fine-tuning: 512 × 4 + 4 = 2052 parameters). When the number of training samples was very small (e.g., less than 100 images), some adaptation methods negatively impacted performance. Even with the largest dataset size, the amount of training data was too limited for a Resnet-34 model trained from scratch on SM images to outperform a model adapted from FFDM.
[0212]
[0240] 17A-17D show the target values for the performance of the adapted model. The impact of the number of training samples in the domain is shown as measured by macroAUC (Figure 17A) and linearly weighted Cohen's kappa coefficient (Figure 17B) for the synthetic 2D mammography (SM) test set from Site 1, and as measured by macroAUC (Figure 17C) and linearly weighted Cohen's kappa coefficient (Figure 17D) for the SM test set from Site 2. Results are shown for vector calibration, matrix calibration, and retraining the final fully connected layer (fine-tuning). Error bars indicate the standard error of the mean calculated on 10 random samplings of the training data. Performance before adaptation (none) and training from scratch are shown as references. For the SM study from Site 1, full-field digital mammography (FFDM) performance served as an additional reference. Note that each graph is shown using its own full dynamic range to facilitate comparison of different adaptation methods for a given metric and dataset.
[0213]
[0241] Discussion
[0242] Breast Imaging Reporting and Data System Although BI-RADS breast density can be an important indicator of breast cancer risk and radiologist sensitivity, intra- and inter-reader variability can limit the usefulness of this measure. Deep learning (DL) models for estimating breast density may be constructed to reduce this variability while still achieving accurate assessments. However, such DL models have been demonstrated to be applicable to digital breast tomosynthesis (DBT) consultations and can be generalized across institutions, thereby demonstrating their suitability as useful clinical tools. To overcome the limited training data available for DBT consultations, the DL model was first trained on a large set of full-field digital mammography (FFDM) images. When evaluated against a held-out examination set of FFDM images, the model demonstrated close agreement between radiologist-reported BI-RADS breast density (κw = 0.75, 95% confidence interval (CI): [0.74, 0.76]). The model was then evaluated against two datasets of synthetic 2D mammography (SM) images generated as part of DBT consultations. FFDM data and SM from the same institution High levels of agreement were also found for the DBT dataset (Site 1: κw = 0.71, CI: [0.64, 0.78]), and the SM dataset from another institution (Site 2: κw = 0.72, CI: [0.70, 0.75]). The strong performance of the DL model demonstrates that it can be generalized to DBT consultations and data from different institutions. Further adaptation of the model to the SM dataset led to some improvement at Site 1 (κw = 0.72, CI: [0.66, 0.79]) and much larger improvement at Site 2 (κw = 0.79, CI: [0.76, 0.81]).
[0214]
[0243] The initial reporting radiologist's assessment was accepted as ground truth. When applied to a given dataset, the level of inter-reader variability between radiologists has a significant impact on the performance that can be achieved for a given dataset. For example, after adaptation, the performance obtained for the Site2 SM dataset was higher than that obtained for the FFDM dataset used to train the model. This is likely a result of the limited inter-reader variability for the Site2 SM dataset, as over 80% of the examinations were read by only two readers.
[0215]
[0244] In contrast to other methods, the BI-RADS breast density DL model The DL model was evaluated on SM images from multiple institutions and on data from multiple institutions. Furthermore, as discussed above, the DL model demonstrated comparable performance compared to other DL models and commercially available breast density software when evaluated on FFDM images (κw = 0.75, CI: [0.74, 0.76] vs. Lehman et al. 0.67, CI: [0.66, 0.68]; Volpara 0.57, CI: [0.55, 0.59]; Quantra 0.46, CI: [0.44, 0.47]) [19, 3]. For each method, results are reported for its individual test set, similar to the way our own results are reported.
[0216]
[0245] Other measures of breast density, such as volumetric breast density, are available through automated 3D tomosynthesis volumes. It may be estimated by software developed or inferred from DBT examination. A threshold can be chosen to convert such a measure to BI-RADS breast density, but this may result in a lower level of agreement than direct estimation of BI-RADS breast density (e.g., agreement between radiologist-assessed BI-RADS breast density and assessment derived from volumetric breast density was κw = 0.47). Here, BI-RADS breast density is estimated from 2D SM images instead of 3D tomosynthesis volumes, as this simplifies transfer learning from FFDM images and mirrors how breast radiologists assess density.
[0217]
[0246] In some cases, when a deep learning (DL) model is adapted to a new engine, Adjustments may be made for cross-institutional differences in patient demographics, patient demographics, or interpreting radiologists. This final adjustment may result in some inter-reader variability between the initial DL model and the adapted DL model, but this may be lower than the inter-reader variability if the model learns the consensus of radiologists in each group. As a result, the improved DL model performance observed after adaptation for the Site 2 SM dataset may be due to differences in patient demographics or radiologist evaluation practices compared to the FFDM dataset. The weaker improvement for the Site 1 SM dataset may be due to similarities in these same factors. Regarding the comparison of domain adaptation techniques as a function of the number of training samples, adjusting the number of parameters in the model based on the number of training samples may yield better performance than training a DL model trained from scratch.
[0218]
[0247] These results are published in Breast Imaging Reporting and The widespread use of the BI-RADS breast density deep learning (DL) model has demonstrated great potential for improving clinical care. The success of the DL model without adaptation indicates that the features learned by the model are broadly applicable to both full-field digital mammography (FFDM) and synthetic 2D mammography (SM) images from digital breast tomosynthesis (DBT) examinations, as well as across different readers and institutions. Therefore, the BI-RADS breast density DL model can be deployed to new sites and institutions without the additional effort of compiling large datasets and training models from scratch. The BI-RADS breast density DL model, which can generalize across sites and image types, can be used to provide faster, lower-cost, and more consistent estimates of women's breast density.
[0219]
[0248] [Example]
[0220] Real-time radiology for optimized radiology workflow
[0249] A machine learning-based classification system was developed and applied to a dataset containing medical images of subjects. The radiology interpretation task (e.g., among multiple different workflows) may be sorted, prioritized, enriched, or edited based on an analysis of the case data. Sorting, prioritizing, enriching, or editing cases for radiology evaluation may be performed based on medical image data (e.g., image data headers or database elements, instead of relying solely on metadata such as labels or annotation information). For example, medical images may be processed by one or more image processing algorithms. A machine learning-based radiology system enables advanced radiology workflows that convey faster and more accurate diagnoses by allowing medical image datasets to be stratified into various radiology evaluations based on their suitability for such evaluations. For example, the multiple different workflows may include radiology evaluations by multiple different sets of radiologists. The radiologists may be on-site or remote from the clinic where the patient's medical images are obtained.
[0221]
[0250] In some embodiments, the machine learning based classification system is The AI triage engine may be configured to sort or prioritize radiology interpretation tasks among a plurality of different workflows based on an analysis of the datasets that include images. For example, one set of datasets that include medical images may be prioritized for radiology evaluation over another set of datasets that include medical images based on a determination by the AI triage engine that the first set of datasets has a higher priority or urgency than the second set of datasets.
[0222]
[0251] In some embodiments, the real-time radiology system is an AI-enabled triadic radiology system. A page workflow is used to obtain medical images of a subject through a screening visit, and then AI is used to communicate radiological results (e.g., screening results and / or diagnostic results) to the patient within minutes (e.g., within about 5 minutes, about 10 minutes, about 15 minutes, about 30 minutes, about 45 minutes, about 60 minutes, about 90 minutes, about 2 hours, about 3 hours, about 4 hours, about 5 hours, about 6 hours, about 7 hours, or about 8 hours) after the medical images are obtained.
[0223]
[0252] In some embodiments, the real-time radiology system detects errors in the AI determinations. The system includes a real-time notification system for interacting with clinic staff about patient cases. The notification system is installed in various locations within the screening clinic (e.g., at clinic staff workstations). Users (e.g., physicians and clinic staff) are assigned roles and receive different notifications for each role. Notifications are sent to clinic staff about patient cases. The notification is triggered when an emergency situation is determined by a trained algorithm. For example, the notification may include both advisory information as well as authorized users entering information that may affect the patient's clinical workflow in real time during the visit. A physician (e.g., a treating physician or radiologist) is notified of emergency cases as they occur via real-time alerts and uses the information from the notification to provide a better diagnosis.
[0224]
[0253] In some embodiments, the real-time radiology system sends a notification to the patient. The patient mobile application (app) includes notifications for the patient's screening / diagnostic visit status, radiology evaluations performed on the patient's medical images, presentations constructed from the radiology evaluations, etc.
[0225]
[0254] In some embodiments, the real-time radiology system may include a system for predicting future extractions. The real-time radiology system includes a database configured to acquire, retrieve, and store datasets including medical images (e.g., radiology images), AI enrichment of the datasets (e.g., medical images labeled, annotated, or processed by AI, such as via image processing algorithms), screening results, diagnostic results, and presentations of the medical images and results. The real-time radiology system is configured to provide services to patients and their clinical caregivers (e.g., radiologists and clinic staff) to retrieve, access, and view the contents of the database. Services of the real-time radiology system may support building complex computational graphs from the stored datasets, including chaining several AI models.
[0226]
[0255] Figure 18 shows an example of a schematic of a real-time radiology evaluation workflow. A real-time radiology evaluation workflow may include acquiring images from a subject (e.g., via mammography). The images may be processed using the systems and methods of the present disclosure (e.g., including AI algorithms) to detect that the images correspond to suspicious cases. A clinician may be alerted that the subject is suitable for real-time radiology evaluation. While the subject waits in the clinic, the images are sent to a radiologist for radiology evaluation, and the results of the radiology evaluation are provided to the clinician for further review.
[0227]
[0256] Figure 19 shows another example of a real-time radiology evaluation workflow overview. Using the disclosed systems and methods (e.g., including AI algorithms), a subject's images are retrieved from a PACS database and analyzed. If the AI analysis indicates that a given subject (e.g., patient) does not have suspicious images, a patient coordinator is notified, who then informs the patient that the results will be received at home after a radiological evaluation is performed. If the AI analysis indicates that the patient has suspicious images, a technologist is notified, who then either (1) updates the medical history, notifies a radiologist to perform a radiological evaluation, and provides the results to the patient coordinator, or (2) notifies billing to process the patient's copay for a follow-up visit and notifies the patient coordinator. The patient coordinator may share the results with the patient and schedule a follow-up appointment, if necessary.
[0228]
[0257] In some embodiments, the real-time radiology assessment workflow comprises: (i) (ii) sending the image or a derivative thereof to a first radiologist of a first set of radiologists for radiological evaluation and generating a screening result based at least in part on whether the image is classified as suspicious; and (iii) sending the image or a derivative thereof to a second radiologist of a second set of radiologists for radiological evaluation and generating a screening result based at least in part on whether the image is classified as equivocal. or (iii) sending the image or a derivative thereof to a third radiologist of a third set of radiologists for radiological evaluation and producing a screening result based at least in part on whether the image is classified as normal.
[0229]
[0258] In some embodiments, the real-time radiology assessment workflow comprises at least If any one image is classified as suspicious, sending the image or a derivative thereof to a first radiologist of the first set of radiologists for radiological evaluation to produce a screening result. In some embodiments, the real-time radiology evaluation workflow includes, if the image is classified as equivocal, sending the image or a derivative thereof to a second radiologist of the second set of radiologists for radiological evaluation to produce a screening result. In some embodiments, the real-time radiology evaluation workflow includes, if the image is classified as normal, sending the image or a derivative thereof to a third radiologist of the third set of radiologists for radiological evaluation to produce a screening result.
[0230]
[0259] In some embodiments, the subject screening results are displayed as an image or a derivative thereof. In some embodiments, the first set of radiologists are located at an on-site clinic (e.g., the clinic where the images or their derivatives were acquired).
[0231]
[0260] In some embodiments, the second set of radiologists are radiologists (e.g., In some embodiments, the third set of radiologists includes a radiologist who is trained to classify images, or derivatives thereof, as normal or suspicious with greater accuracy than the trained algorithm. In some embodiments, the third set of radiologists is located remotely from the on-site clinic (e.g., the clinic where the images were acquired). In some embodiments, a third radiologist in the third set of radiologists performs radiological evaluation of images, or derivatives thereof, of a batch comprising multiple images (e.g., where the batch is selected to improve efficiency of the radiological evaluation).
[0232]
[0261] In some embodiments, the real-time radiology evaluation workflow comprises: and performing a diagnostic procedure based at least in part on the screening results to produce a diagnostic result for the subject. In some embodiments, the diagnostic result for the subject is produced in the same clinic visit as the step of acquiring the image. In some embodiments, the diagnostic result for the subject is produced within about one hour of the step of acquiring the image.
[0233]
[0262] In some embodiments, the image or a derivative thereof may include additional images of parts of the subject's body. In some embodiments, the additional characteristics include anatomy, tissue characteristics (e.g., tissue density or physical properties), the presence of foreign bodies (e.g., implants), the type of finding, a medical condition (e.g., predicted by an algorithm, such as a machine learning algorithm), or a combination thereof.
[0234]
[0263] In some embodiments, the image or a derivative thereof is provided by a first radiologist, a second radiologist, to the first radiologist, the second radiologist, or the third radiologist based at least in part on additional characteristics of the radiologist, or the third radiologist (e.g., the first radiologist, the second radiologist, or the third radiologist's personal ability to perform a radiological assessment of the at least one image or derivative thereof).
[0235]
[0264] In some embodiments, the real-time radiology review workflow comprises: generating an alert based at least in part on sending the image or a derivative thereof to a first radiologist, or sending the image or a derivative thereof to a second radiologist. In some embodiments, the real-time radiology review workflow includes sending the alert to the subject or the subject's clinical caregiver. In some embodiments, the real-time radiology review workflow includes sending the alert to the subject through a patient mobile application. In some embodiments, the alert is generated in real time with (b) or near real time with (b).
[0236]
[0265] In some embodiments, the real-time radiology system provides AI-driven remote imaging. The remote imaging platform includes an AI-based radiology work distributor that routes cases for physician review in real time or substantially real time with the acquisition of medical images. The remote imaging platform may be configured to perform AI-based profiling of image types and physicians to assign each case to one physician from among multiple physicians based on the individual physician's aptitude for handling, evaluating, or interpreting a given case's dataset. Radiologists may belong to a network of radiologists, each with a distinct set of radiology skills, expertise, and experience. The remote imaging platform may assign cases to physicians based on searching the network for physicians with a desired combination of skills, expertise, experience, and cost. The radiologist may be on-site or remote from the clinic where the patient's medical images are acquired. In some embodiments, a radiologist's expertise may be determined by comparing the radiologist's performance to the performance of an AI model for various radiology tasks on an evaluative set of data. Radiologists may be paid for performing radiological assessments for each individual case they undertake and perform. In some embodiments, the real-time radiology system features dynamic pricing of radiological work based on the AI-determined difficulty, urgency, and value of the radiological work (e.g., radiological assessment, interpretation, or review).
[0237]
[0266] In some embodiments, the real-time radiology system The AI algorithm may be configured to organize, prioritize, or stratify a plurality of medical image cases into subgroups of medical image cases for evaluation, interpretation, or review. Stratification of medical image cases may be performed by an AI algorithm based on image characteristics of the individual medical image cases to improve human efficiency in evaluating individual cases. For example, the algorithm may group visually similar or diagnostically similar cases together for human review, such as grouping identified cases with similar lesion types located in similar regions of the anatomy.
[0238]
[0267] Figure 20 shows an AI-assisted radiology evaluation workflow in a remote imaging setting. 1 illustrates a schematic example of a workflow. Using the systems and methods (e.g., including AI algorithms) of the present disclosure, a subject's images are retrieved from a PACS database and analyzed using AI algorithms to prioritize and filter out cases for radiological evaluation (e.g., based on the subject's breast density and / or breast cancer risk). The AI-assisted radiological evaluation workflow can optimize the routing of cases for radiological evaluation based on the skill level of the radiologist. For example, a first radiologist may have an average reading time of 45 seconds, expert-level expertise, and skill for evaluating extremely dense breasts. As another example, a second radiologist may have an average reading time of 401 seconds and beginner-level expertise. As another example, a third radiologist may have an average reading time of 323 seconds and beginner-level expertise. As another example, a fourth radiologist may have an average reading time of 145 seconds and beginner-level expertise. For example, a fifth radiologist may have an average reading time of 60 seconds, expert-level expertise, and good skills. The AI-assisted radiology review workflow may route a given subject case to a radiologist selected from the first, second, third, fourth, or fifth radiologist based on the radiologists' average reading time, expertise level, and / or skill level for the given subject case.
[0239]
[0268] In some embodiments, the AI-assisted radiology evaluation workflow comprises: (i) imaging (ii) sending the image or a derivative thereof to a first radiologist of a first set of radiologists for radiological evaluation and producing a screening result based at least in part on whether the image is classified as suspicious; (ii) sending the image or a derivative thereof to a second radiologist of a second set of radiologists for radiological evaluation and producing a screening result based at least in part on whether the image is classified as equivocal; or (iii) sending the image or a derivative thereof to a third radiologist of a third set of radiologists for radiological evaluation and producing a screening result based at least in part on whether the image is classified as normal.
[0240]
[0269] In some embodiments, the AI-assisted radiology evaluation workflow comprises at least If an image is classified as suspicious, the AI-assisted radiology evaluation workflow includes sending the image or a derivative thereof to a first radiologist of the first set of radiologists for radiological evaluation to produce a screening result. In some embodiments, if an image is classified as equivocal, the AI-assisted radiology evaluation workflow includes sending the image or a derivative thereof to a second radiologist of the second set of radiologists for radiological evaluation to produce a screening result. In some embodiments, if an image is classified as normal, the AI-assisted radiology evaluation workflow includes sending the image or a derivative thereof to a third radiologist of the third set of radiologists for radiological evaluation to produce a screening result.
[0241]
[0270] In some embodiments, the subject screening results are displayed as an image or a derivative thereof. In some embodiments, the first set of radiologists are located at an on-site clinic (e.g., the clinic where the images or their derivatives were acquired).
[0242]
[0271] In some embodiments, the second set of radiologists are radiologists (e.g., In some embodiments, the third set of radiologists includes a radiologist who is trained to classify images, or derivatives thereof, as normal or suspicious with greater accuracy than the trained algorithm. In some embodiments, the third set of radiologists is located remotely from the on-site clinic (e.g., the clinic where the images were acquired). In some embodiments, a third radiologist in the third set of radiologists performs radiological evaluation of images, or derivatives thereof, of a batch comprising multiple images (e.g., where the batch is selected to improve efficiency of the radiological evaluation).
[0243]
[0272] In some embodiments, the AI-assisted radiology evaluation workflow includes: The method further comprises performing a diagnostic procedure based at least in part on the screening results to produce a diagnostic result for the subject. In some embodiments, the diagnostic result for the subject is produced in the same clinic visit as the step of acquiring the images. In some embodiments, the diagnostic result for the subject is produced within about one hour of the step of acquiring the images.
[0244]
[0273] In some embodiments, the image or a derivative thereof may include additional images of parts of the subject's body. In some embodiments, the additional characteristics are sent to a first radiologist, a second radiologist, or a third radiologist based at least in part on the anatomical characteristics. , tissue characteristics (e.g., tissue density or physical properties), presence of foreign bodies (e.g., implants), type of finding, pathology (e.g., predicted by an algorithm such as a machine learning algorithm), or a combination thereof.
[0245]
[0274] In some embodiments, the image or a derivative thereof is provided by a first radiologist, a second radiologist, to the first radiologist, the second radiologist, or the third radiologist based at least in part on additional characteristics of the radiologist, or the third radiologist (e.g., the first radiologist, the second radiologist, or the third radiologist's personal ability to perform a radiological assessment of the at least one image or derivative thereof).
[0246]
[0275] In some embodiments, the AI-assisted radiology evaluation workflow comprises: and generating an alert based at least in part on sending the image or a derivative thereof to a first radiologist, or sending the image or a derivative thereof to a second radiologist. In some embodiments, the AI-assisted radiology evaluation workflow includes sending the alert to the subject or the subject's clinical caregiver. In some embodiments, the AI-assisted radiology evaluation workflow includes sending the alert to the subject through a patient mobile application. In some embodiments, the alert is generated in real time with (b) or near real time with (b).
[0247]
[0276] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that It will be apparent that such embodiments are provided by way of example only. The present invention is not intended to be limited by the specific examples given herein. While the present invention has been described with reference to the above-referenced specifications, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the present invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed in practicing the present invention. Accordingly, the present invention is also intended to encompass any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the present invention, and that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
1. A method for processing at least one image or derivative thereof of a subject's body part, (a) A step of obtaining at least one image or its derivative of the body of the subject, wherein the at least one image or its derivative includes a medical image. (b) Using a trained algorithm, classifying the at least one image or its derivative as normal, ambiguous, or suspicious as indicating cancer, wherein the classifying step includes applying an image processing algorithm to the at least one image or its derivative; (c)(b) When the at least one image or its derivative is classified, (i) If the at least one image or its derivative is classified as suspected to indicate the cancer, the step of sending the at least one image or its derivative to a first radiologist for radiological evaluation in order to produce a screening result or a diagnostic result, (ii) If the at least one image or its derivative is classified as ambiguous as indicating the cancer, the step of sending the at least one image or its derivative to a second radiologist, different from the first radiologist, for radiological evaluation to produce a screening result or a diagnostic result, (iii) If the at least one image or its derivative is classified as normal, the step of sending the at least one image or its derivative to a third radiologist for radiological evaluation in order to produce a screening result or a diagnostic result. The steps to perform, (d) A step of receiving a radiological evaluation of the subject from the first radiologist, the second radiologist, or the third radiologist, wherein the radiological evaluation of the subject is generated by the first radiologist, the second radiologist, or the third radiologist based at least in part on a radiological analysis of the at least one image or its derivative; A method that includes this.
2. A method according to claim 1, wherein the trained algorithm is configured to classify the at least one image or its derivative as normal, ambiguous, or suspicious with a sensitivity of at least about 80%.
3. A method according to claim 1, wherein the trained algorithm is configured to classify the at least one image or its derivative as normal, ambiguous, or suspicious with at least about 80% specificity.
4. A method according to claim 1, wherein the trained algorithm is configured to classify the at least one image or its derivative as normal, ambiguous, or suspicious with a positive prediction of at least about 80%.
5. A method according to claim 1, wherein the trained algorithm is configured to classify the at least one image or its derivative as normal, ambiguous, or questionable with at least about 80% negative prediction.
6. A method according to claim 1, wherein the trained algorithm is configured to identify at least one region of the at least one image or its derivative that contains or is suspected of containing abnormal tissue.
7. A method according to claim 1, wherein the cancer is breast cancer.
8. A method according to claim 1, wherein the trained algorithm is trained using at least about 100 independent training samples, each containing an image showing or suspected of showing the cancer.
9. A method according to claim 1, wherein the trained algorithm is trained using a first group of independent training samples comprising positive images showing or suspected of showing the cancer, and a second group of independent training samples comprising negative images not showing or suspected of showing the cancer.
10. A method according to claim 1, wherein the trained algorithm includes a supervised machine learning algorithm, the supervised machine learning algorithm includes a deep learning algorithm, a support vector machine (SVM), a neural network, or a random forest.
11. A method according to claim 1, further comprising the step of monitoring the subject, the monitoring step comprising evaluating images of the subject's body parts at a plurality of time points, the evaluating step being at least in part based on the classification of the at least one image or its derivative at each of the plurality of time points as normal, ambiguous or suspicious as indicating cancer.
12. A method according to claim 11, wherein the difference in the evaluation of the images of the subject's body at a plurality of time points indicates one or more clinical indicators selected from the group including (i) the diagnosis of the subject, (ii) the prognosis of the subject, and (iii) the effectiveness or ineffectiveness of a series of treatments for the subject.
13. A method according to any one of claims 1 to 12, wherein the second radiologist is a radiologist trained to classify the at least one image or its derivative as normal or suspicious with greater accuracy than the trained algorithm.
14. A method according to any one of claims 1 to 12, wherein the third radiologist is located remotely from an on-site clinic, and the at least one image is acquired at the on-site clinic.
15. A method according to any one of claims 1 to 12, wherein the third radiologist performs the radiological evaluation of at least one image or its derivative from a batch comprising a plurality of images, the batch being selected to improve the efficiency of the radiological evaluation.
16. A method according to any one of claims 1 to 15, further comprising the step of performing a diagnostic procedure for the subject on at least partly the screening results to produce a diagnostic result for the subject.
17. A method according to claim 16, wherein the diagnostic result of the subject is produced on the same day as the step of acquiring the at least one image.
18. A method according to any one of claims 1 to 17, wherein the at least one image or its derivative is sent to the first radiologist, the second radiologist, or the third radiologist, at least in part on additional characteristics of the body of the subject.
19. A method according to claim 18, wherein the additional characteristics include anatomical structure, tissue characteristics, presence of a foreign body, type of finding, pathological condition, or a combination thereof.
20. A method according to any one of claims 1 to 19, wherein the at least one image or its derivative is sent to the first radiologist, the second radiologist, or the third radiologist, at least in part on additional characteristics of the first radiologist, the second radiologist, or the third radiologist.
21. A method according to any one of claims 1 to 20, the method further comprising the step of sending the at least one image or its derivative to the first radiologist, or the step of sending the at least one image or its derivative to the second radiologist, or the step of sending the at least one image or its derivative to the third radiologist, wherein the method further comprises the step of generating an alert based at least in part on the step of sending the at least one image or its derivative to the third radiologist.
22. A method according to claim 21, further comprising the step of transmitting the alert to the subject or the subject's clinical healthcare provider.
23. A method according to claim 22, further comprising the step of transmitting the alert to the subject through the patient's mobile application.
24. A method according to claim 21, wherein the alert is generated in real time or near real time with respect to (b).
25. A method according to claim 1, wherein the step of applying the image processing algorithm includes the steps of identifying a region of interest within the at least one image or its derivative, and labeling the region of interest to produce at least one labeled image.
26. A method according to any one of claims 1 to 25, further comprising the step of storing one or more of the at least one image or its derivatives and the classification in a database.
27. A method according to claim 1, wherein the at least one image comprises a plurality of images obtained from the subject, the plurality of images being obtained using different modalities or at different points in time.
28. A method according to claim 1, the method comprising the steps of sending the at least one image or its derivative to a first radiologist for a first radiological evaluation if the at least one image is classified by the machine learning classifier into a specific category of the plurality of categories, and then sending the at least one image or its derivative to a second radiologist for a second radiological evaluation.
29. A method according to claim 28, wherein the specific category is a different category from the first radiological evaluation.