System and method for generating medical images
By dividing 3D medical images into 2D slices and using a 2D classifier trained with non-local labels to generate synthetic 2D interpretive images, the problems of resource-intensive and low-accuracy processing of 3D images in existing technologies are solved, achieving efficient and accurate 2D image generation, applicable to CT, MRI and 3D mammograms.
Patent Information
- Application Number
- CN202180089909.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-01-11
- Filing Date
- 2021-12-06
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2041-12-06
AI Technical Summary
Existing technologies struggle to efficiently generate 2D medical images to facilitate radiologists' viewing of 3D medical images, especially since processing 3D images requires significant computational resources and time, and the creation of training datasets relies on manual labeling, resulting in low classifier accuracy.
By dividing 3D medical images into 2D slices and using a 2D classifier trained on a non-localized labeled training dataset, interpretable weights of the interpretable map are calculated and aggregated to generate synthetic 2D interpretable images, providing non-localized indications for visual discovery while reducing computational resources and time consumption.
It improves the efficiency and accuracy of generating 2D interpretive images, reduces computational resources and processing time, and enhances the visualization of 3D medical images, especially the processing of CT, MRI, and 3D mammograms.
Smart Images

Figure CN116710956B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] In some embodiments of the invention, the invention relates to medical image processing, and more specifically, but not exclusively, to systems and methods for generating a 2D medical image from a 3D medical image. BACKGROUND
[0002] A 2D medical image can be created from a 3D medical image to assist a radiologist in navigating the 3D medical image. The radiologist can use the 2D medical image in order to determine which portions of the 3D medical image to focus on. For example, in a 2D image of a CT scan showing a lung nodule in a particular lobe of a particular lung, the radiologist can observe a slice of the CT scan corresponding to the particular lobe to obtain a better observation of the lung nodule. SUMMARY
[0003] According to a first aspect, a computer-implemented method for generating a synthetic 2D explanation image from a 3D medical image, comprises: inputting each 2D medical image of a plurality of 2D medical images created by partitioning the 3D medical image to a 2D classifier, the 2D classifier being trained on a training dataset of 2D medical images labeled with an indication of a visual finding depicted therein; for each respective 2D medical image of the plurality of 2D medical images, computing a respective explanation map of a plurality of explanation maps, the respective explanation map comprising a plurality of regions corresponding to a plurality of corresponding regions in the respective 2D image, each respective region in the respective explanation map being associated with a computed explainable weight, the computed explainable weight being indicative of an influence of the respective corresponding region in the respective 2D medical image on a result of the 2D classifier provided for the respective 2D medical image; computing a synthetic 2D explanation image, the synthetic 2D explanation image comprising a respective aggregate weight for each respective region thereof, each respective aggregate weight being computed by aggregating a plurality of the computed explainable weights for respective regions of the plurality of explanation maps corresponding to the respective region in the synthetic 2D explanation image; and providing the synthetic 2D explanation image for presentation on a display.
[0004] According to a second aspect, a method of generating a 2D classifier for analyzing 2D images of 3D medical images, comprising: accessing a plurality of training 3D medical images, for each respective 3D medical image of the plurality of 3D medical images: dividing the respective 3D medical image into a plurality of 2D medical images, inputting each 2D medical image of the plurality of 2D medical images into a 2D classifier, the 2D classifier trained on a training dataset of 2D medical images labeled with an indication of a visual finding depicted therein; for each respective 2D medical image of the plurality of 2D medical images, computing a respective explanation map of a plurality of explanation maps, the respective explanation map comprising a plurality of regions corresponding to a plurality of corresponding regions in the respective 2D image, each respective region in the respective explanation map being associated with a computed explainable weight, the computed explainable weight indicating an influence of the respective corresponding region in the respective 2D medical image on a result of the 2D classifier provided for the respective 2D medical image; computing a synthetic 2D explanation image, the synthetic 2D explanation image comprising a respective aggregated weight for each respective region thereof, each respective aggregated weight computed by aggregating a plurality of the computed explainable weights for respective regions of the plurality of explanation maps corresponding to the respective region in the synthetic 2D explanation image, assigning a label indicating a presence of the visual finding depicted therein to the synthetic 2D explanation image; generating an updated training dataset comprising a plurality of the synthetic 2D explanation images and corresponding labels; and generating an updated 2D classifier by updating the training of the 2D classifier using the updated training dataset.
[0005] According to a third aspect, a computer-implemented method for generating a synthetic 2D interpretation image from sequentially acquired video 2D medical images, comprises: receiving a plurality of 2D medical images as a sequence captured as a video over a time interval, wherein the plurality of 2D medical images are spaced in time, inputting each 2D medical image of the plurality of 2D medical images to a 2D classifier trained on a training dataset of 2D medical images labeled with an indication of visual findings depicted therein; for each respective 2D medical image of the plurality of 2D medical images, computing a respective interpretation map of a plurality of interpretation maps, the respective interpretation map comprising a plurality of regions corresponding to a plurality of corresponding regions in the respective 2D image, each respective region in the respective interpretation map being associated with a computed explainable weight, the computed explainable weight being indicative of an influence of the respective corresponding region in the respective 2D medical image on a result of the 2D classifier provided for the respective 2D medical image, and computing a synthetic 2D interpretation image, the synthetic 2D interpretation image comprising a respective aggregated weight for each respective region thereof, each respective aggregated weight being computed by aggregating a plurality of the computed explainable weights for respective regions of the plurality of interpretation maps corresponding to the respective region in the synthetic 2D interpretation image.
[0006] According to a fourth aspect, a computer-implemented method of generating a synthetic 2D interpretation image from a 3D medical image, comprises: inputting each 2D medical image of a plurality of 2D medical images obtained by at least one of: partitioning a 3D medical image, and being captured as a video over a time interval, to a 2D classifier trained on a training dataset of 2D medical images labeled with an indication of visual findings depicted therein; for each respective 2D medical image of the plurality of 2D medical images, computing a respective interpretation map of a plurality of interpretation maps, the respective interpretation map comprising a plurality of regions corresponding to a plurality of corresponding regions in the respective 2D image, each respective region in the respective interpretation map being associated with a computed explainable weight, the computed explainable weight being indicative of an influence of the respective corresponding region in the respective 2D medical image on a result of the 2D classifier provided for the respective 2D medical image; computing a synthetic 2D interpretation image, the synthetic 2D interpretation image comprising a respective aggregated weight for each respective region thereof, each respective aggregated weight being computed by aggregating a plurality of the computed explainable weights for respective regions of the plurality of interpretation maps corresponding to the respective region in the synthetic 2D interpretation image; and providing the synthetic 2D interpretation image for presentation on a display.
[0007] According to a fifth aspect, a device for generating a synthetic 2D explanation image from a 3D medical image, comprising: at least one hardware processor executing code for: inputting each 2D medical image of a plurality of 2D medical images obtained by at least one of: partitioning a 3D medical image, and being captured as a video over a time interval, to a 2D classifier trained on a training dataset of 2D medical images labeled with an indication of a visual finding depicted therein, for each respective 2D medical image of the plurality of 2D medical images, computing a respective explanation map of a plurality of explanation maps, the respective explanation map comprising a plurality of regions corresponding to a plurality of corresponding regions in the respective 2D image, each respective region in the respective explanation map being associated with a computed explainable weight, the computed explainable weight indicating an influence of the respective corresponding region in the respective 2D medical image on a result of the 2D classifier provided for the respective 2D medical image, computing a synthetic 2D explanation image, the synthetic 2D explanation image comprising a respective aggregated weight for each respective region thereof, each respective aggregated weight computed by aggregating a plurality of the computed explainable weights for respective regions of the plurality of explanation maps corresponding to the respective region in the synthetic 2D explanation image, and providing the synthetic 2D explanation image for presentation on a display.
[0008] According to a sixth aspect, a computer program product for generating a synthetic 2D explanation image from a 3D medical image, comprising a non-transitory medium storing a computer program which, when executed by at least one hardware processor, causes the at least one hardware processor to perform: inputting each 2D medical image of a plurality of 2D medical images obtained by at least one of: partitioning a 3D medical image, and being captured as a video over a time interval, to a 2D classifier trained on a training dataset of 2D medical images labeled with an indication of a visual finding depicted therein, for each respective 2D medical image of the plurality of 2D medical images, computing a respective explanation map of a plurality of explanation maps, the respective explanation map comprising a plurality of regions corresponding to a plurality of corresponding regions in the respective 2D image, each respective region in the respective explanation map being associated with a computed explainable weight, the computed explainable weight indicating an influence of the respective corresponding region in the respective 2D medical image on a result of the 2D classifier provided for the respective 2D medical image, computing a synthetic 2D explanation image, the synthetic 2D explanation image comprising a respective aggregated weight for each respective region thereof, each respective aggregated weight computed by aggregating a plurality of the computed explainable weights for respective regions of the plurality of explanation maps corresponding to the respective region in the synthetic 2D explanation image, and providing the synthetic 2D explanation image for presentation on a display.
[0009] In further implementations forms of the first, second, third, fourth, fifth, and sixth aspects, each respective aggregate weight represents a respective likelihood of presence of a visual finding at a corresponding respective region of the computed composite 2D interpretation image.
[0010] In further implementations forms of the first, second, fourth, fifth, and sixth aspects, the plurality of 2D medical images is computed by dividing the 3D medical image into a plurality of sequential 2D slices along a z-axis, wherein a respective aggregate weight is computed for each respective region of the plurality of sequential 2D slices having common x, y coordinates along x- and y-axes and a variable z coordinate along the z-axis.
[0011] In further implementations forms of the first, second, fourth, fifth, and sixth aspects, an orientation of the z-axis defining an axis into which the 3D medical image is sliced into the plurality of sequential 2D slices is obtained as a function of a view axis selected by a user viewing the 3D medical image rendered on a display, wherein the computed composite 2D interpretation image based on the z-axis corresponding to the view axis is rendered on the display together with the 3D medical image, and the method further comprises dynamically detecting a change in the view axis of the 3D medical image rendered on the display in at least one iteration, dynamically computing an updated composite 2D interpretation image based on the change in the view axis, and dynamically updating the display to render the updated composite 2D interpretation image.
[0012] In further implementations forms of the first, second, fourth, fifth, and sixth aspects, further comprising computing a particular orientation of the z-axis defining an axis into which the 3D medical image is sliced into the plurality of sequential 2D slices, the particular orientation of the z-axis generating an optimal composite 2D interpretation image having a maximum aggregate weight representing a minimum occlusion of the visual finding, automatically adjusting rendering of the 3D medical image on the display to the particular orientation of the z-axis, and rendering the optimal composite 2D interpretation image on the display.
[0013] In further implementations forms of the first, second, fourth, fifth, and sixth aspects, each 2D medical image of the plurality of 2D medical images comprises pixels corresponding to voxels of the 3D medical image, a respective interpretable weight is assigned to each pixel of each 2D medical image of the plurality of 2D medical images, and for each pixel of the composite 2D interpretation image having a particular (x, y) coordinate, a respective aggregate weight is computed by aggregating the interpretable weights of the pixels of the plurality of 2D medical images having corresponding (x, y) coordinates for a variable z coordinate.
[0014] In further implementations forms of the first, second, third, fourth, fifth, and sixth aspects, the indication of the visual finding of the training dataset is non-localized for the respective 2D image whole, and wherein the 2D classifier utilizes non-localized data to generate a result indicating the visual finding for an input 2D image whole.
[0015] In further implementations forms of the first, second, third, fourth, fifth, and sixth aspects, each respective explanation map has a respective explanation weight that represents a relative influence of the respective corresponding region on the result of the 2D classifier.
[0016] In further implementations forms of the first, second, third, fourth, fifth, and sixth aspects, each respective aggregate weight of the synthetic 2D explanation image is computed as a weighted average of the explainable weights computed for the respective region of the plurality of explanation maps corresponding to the respective region in the synthetic 2D explanation image.
[0017] In further implementations forms of the first, second, third, fourth, fifth, and sixth aspects, each respective explanation map comprises a plurality of pixels corresponding to pixels of the respective 2D medical image with pixel intensity values adjusted by the corresponding respective explainable weight, wherein the synthetic 2D explanation image comprises a plurality of pixels with pixel intensity values computed by aggregating the pixel intensity values adjusted by the corresponding respective explainable weights in the plurality of explanation maps.
[0018] In further implementations forms of the first, second, fourth, fifth, and sixth aspects, the 3D medical image is selected from a group comprising: CT, MRI, breast tomography, digital breast tomosynthesis (DBT), 3D ultrasound, 3D nuclear imaging, and PET.
[0019] In further implementations forms of the first, second, third, fourth, fifth, and sixth aspects, the visual finding represents cancer.
[0020] In further implementations forms of the first, second, third, fourth, fifth, and sixth aspects, further comprising: selecting a subset of the plurality of explanation maps, wherein each selected explanation map comprises at least one cluster of at least one region having a required explanation weight that is higher than explanation weights of other regions excluded from the cluster, wherein the synthetic 2D image is computed from the selected subset.
[0021] In further implementation forms of the fourth, fifth, and sixth aspects, further comprising: generating an updated 2D classifier of the 2D classifier for analyzing 2D images of the 3D medical images by: accessing a plurality of training 3D medical images, for each respective 3D medical image of the plurality of 3D medical images: partitioning the respective 3D medical image into a plurality of 2D medical images, inputting each 2D medical image of the plurality of 2D medical images into a 2D classifier trained on a training dataset of 2D medical images labeled with an indication of a visual finding depicted therein, for each respective 2D medical image of the plurality of 2D medical images, computing a respective explanation map of a plurality of explanation maps, the respective explanation map comprising a plurality of regions corresponding to a plurality of corresponding regions in the respective 2D image, each respective region in the respective explanation map being associated with a computed explainable weight indicating an influence of the respective corresponding region in the respective 2D medical image on the result of the 2D classifier of the respective 2D medical image, computing a synthetic 2D explanation image, the synthetic 2D explanation image comprising a respective aggregated weight for each respective region of the synthetic 2D explanation image, each respective aggregated weight being computed by aggregating a plurality of the computed explainable weights for respective regions of the plurality of explanation maps corresponding to the respective region in the synthetic 2D explanation image, assigning a label indicating a presence of the visual finding depicted therein to the synthetic 2D explanation image, generating an updated training dataset comprising a plurality of the synthetic 2D explanation images and corresponding labels, and generating the updated 2D classifier by updating the training of the 2D classifier using the updated training dataset.
[0022] In further implementation forms of the second, third, fourth, fifth, and sixth aspects, further comprising: after accessing the plurality of training 3D medical images, partitioning each 3D medical image of the plurality of 3D medical images into a plurality of 2D medical images, labeling each respective 2D medical image with a label indicating a presence of a visual finding depicted with the respective 2D medical image, wherein the label is non-localized and is assigned to the respective 2D medical image as a whole, creating the training dataset of 2D medical images comprising the plurality of 2D medical images and associated non-localized labels, and training the 2D classifier using the training dataset.
[0023] In further implementation forms of the third, fourth, fifth, and sixth aspects, the plurality of 2D medical images are captured by an imaging device selected from a group consisting of: a colonoscope, an endoscope, a bronchoscope, and a 2D ultrasound.
[0024] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the application pertains. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the application, exemplary methods and / or materials are described below. In case of conflict, the patent specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and are not intended to be limiting. BRIEF DESCRIPTION OF DRAWINGS
[0025] Some embodiments of the application are herein described, by way of example only, with reference to the accompanying drawings. With specific reference now to the drawings in detail, it is stressed that the particulars shown are by way of example and for purposes of illustrative discussion of embodiments of the application. In this regard, the description taken with the drawings makes apparent to those skilled in the art how embodiments of the application can be practiced.
[0026] In the drawings:
[0027] Figure 1 is a flowchart of a method of generating a synthetic 2D explanatory image from a 3D medical image, in accordance with some embodiments of the application;
[0028] Figure 2 is a block diagram of components of a system for generating a synthetic 2D explanatory image from a 3D medical image and / or for training a 2D classifier using a synthetic 2D explanatory image generated from a 3D medical image, in accordance with some embodiments of the application;
[0029] Figure 3 is a flowchart of a method of training a 2D classifier using a synthetic 2D explanatory image generated from a 3D medical image, in accordance with some embodiments of the application;
[0030] Figure 4 is a schematic depicting respective synthetic 2D explanatory images in comparison to other standard methods for computing 2D images from 3D images, in accordance with some embodiments of the application; and
[0031] Figure 5 is a schematic depicting automatic computation of a specific orientation of a z-axis defining an axis by which a 3D medical image is sliced into 2D slices, for generating an optimal synthetic 2D explanatory image having a maximum aggregate weight of presence representing a maximum likelihood of a visual finding, in accordance with some embodiments of the application. DETAILED DESCRIPTION
[0032] In some embodiments of the application, the present application relates to medical image processing, and more particularly, but not exclusively, to systems and methods for generating 2D medical images from 3D medical images.
[0033] One aspect of some embodiments of the present application relates to a system, method, device and / or code instructions (e.g., stored on a memory, executable by one or more hardware processors) for generating a synthetic 2D interpretation image from a 3D medical image, the synthetic 2D interpretation image comprising an indication of clinically and / or diagnostically most important regions aggregated from a plurality of 2D medical images created by partitioning the 3D medical image. The 2D medical images created by partitioning the 3D medical image are fed into a 2D classifier. The 2D classifier is trained on a training dataset of 2D medical images labeled with an indication of visual findings depicted therein, optionally, for the 2D image as a whole, i.e., non-localized data. The 2D classifier can generate a non-localized indication of the presence of a visual finding within the input 2D image as a whole, without having to provide an indication of the location of the visual finding within the 2D image, e.g., the 2D classifier is a binary classifier that outputs YES / NO for the presence of the visual finding and / or the likelihood of the presence of the visual finding within the 2D image as a whole, and does not have to generate a region-specific (e.g., per-pixel) output indicating which pixel(s) correspond to the visual finding. A respective interpretation map is computed for the corresponding respective 2D medical image fed into the 2D classifier. The respective interpretation map comprises regions (e.g., individual pixels, groups of pixels) corresponding to regions of the respective 2D image (e.g., pixel-to-pixel correspondence, groups of pixels correspond to individual pixels). Each respective region of the respective interpretation map is associated with a computed explainable weight indicating the influence of the respective corresponding region of the respective 2D medical image on the result of the 2D classifier of the respective 2D medical image. For example, pixels with relatively higher weights indicate that those pixels play a more important role in the 2D classifier determining the result of the presence of a visual finding in the 2D medical image. Pixels with higher weights indicate that the region depicted by the higher weight pixels can be clinically and / or diagnostically significant. The explainable weights can be computed, for example, using an artificial intelligence explainability (XAI) process. The synthetic 2D interpretation image is computed by projecting the 3D volume onto the synthetic 2D interpretation image using the weights. The respective aggregated weights are for each corresponding respective region in the plurality of interpretation maps. Each respective aggregated weight is computed by aggregating the explainable weights computed for the respective region in the interpretation map corresponding to the respective region of the synthetic 2D interpretation image, e.g., for each region in the x-y plane when the 2D images are along the x-y plane, the aggregated weight is computed along the z-axis of the plurality of interpretation maps. The synthetic 2D interpretation image can be provided for presentation on a display, e.g., together with a presentation of the 3D image. The 2D interpretation image can help an observer (e.g., a radiologist) decide which regions of the 3D image to focus on, e.g., according to the regions of the 3D image corresponding to the regions of the synthetic 2D interpretation image with the highest aggregated weights.
[0034] At least some implementations of the systems, methods, apparatuses, and / or code instructions described herein are directed to the technical problem of reducing the computational resources to process 3D medical images, for example, captured by CT, MRI, PET, and 3D mammograms. At least some implementations of the systems, methods, apparatuses, and / or code instructions described herein improve computer processing of 3D medical images by reducing the computational resources required to process 3D medical images within a reasonable time, and / or by reducing the time to process 3D medical images using existing resources. Processing 3D medical images requires a large amount of processing resources and / or memory resources due to the large amount of data stored in the 3D medical images. Processing such 3D medical images requires a large amount of time, making it impractical to process a large number of 3D images. For example, neural networks that apply 3D convolutions take a large amount of time and / or use a large amount of computational resources to process 3D images. Computing the locations of identified visual findings in 3D images consumes a particularly large amount of computational resources and / or a large amount of processing time. Some existing methods divide the 3D images into a plurality of 2D slices and feed each slice into a 2D classifier designed to identify the locations of visual findings within the respective 2D slice. However, this approach also consumes a large amount of computational resources and / or a large amount of processing time to compute the locations of each visual finding in each 2D image. Furthermore, generating training datasets of 2D and / or 3D images with the locations of visual findings marked to train 2D and / or 3D classifiers requires intensive resources, as in this case the labels are added manually by a trained user who manually reviews each 2D and / or 3D image in order to locate the visual findings and add the labels. Since creating such training datasets involves a large amount of work, they are scarce and have a small number of images. Classifiers trained using such small training datasets can have low accuracy.
[0035] At least some implementations of the systems, methods, apparatuses, and / or code instructions described herein provide a solution to the above-described technical problem and / or improve computer processing of 3D images by dividing a 3D medical image into 2D image slices. Each 2D slice is fed into a 2D classifier that is trained to output an indication of whether a visual finding is located within the 2D image as a whole, without determining a location of the visual finding within the 2D image. The 2D classifier can be trained on a training dataset of 2D images labeled with non-localizing labels for the image as a whole. For example, such labeling can be performed automatically based on natural language processing methods that analyze radiology reports to determine visual findings depicted in the images and generate non-localizing labels accordingly. The use of non-localizing labels enables an automated method and / or a method that consumes less manual and / or computational resources compared to the use of location labels. The 2D classifier that outputs non-localizing results consumes significantly less computational resources and / or processing time compared to 3D classifiers and / or 2D classifiers that output locations of visual findings. As described herein, the indication of a location of a visual finding within a 3D image is computed by aggregating the explanation maps with the weights to compute a composite 2D explanation image, which consumes significantly less computational resources and / or processing time compared to 3D classifiers and / or 2D classifiers that output locations of visual findings.
[0036] At least some implementations of the systems, methods, apparatuses, and / or code instructions described herein can be used with existing trained 2D classifiers without having to require retraining of the 2D classifiers. For example, the 2D composite images can be used with existing 2D classifiers that automatically analyze 2D slices of 3D CT images to detect lung nodules without having to require significant adjustments to the 2D classifiers.
[0037] At least some implementations of the systems, methods, apparatuses, and / or code instructions described herein can train the 2D classifiers using automated tools for creating training datasets, e.g., automated tools that analyze radiology reports and generate indications of which visual findings radiologists find in the images without having to label the images with labels that indicate the visual findings are located in the images.
[0038] At least some implementations of the systems, methods, apparatuses, and / or code instructions described herein are directed to the technical problem of improving visibility of visual features captured in 3D imaging data captured, e.g., by CT, MRI, PET, and 3D mammography. At least some implementations of the systems, methods, apparatuses, and / or code instructions described herein are directed to the technical problem of improving visibility of visual features captured as videos of 2D imaging data captured, e.g., by colonoscopy, endoscopy, bronchoscopy, and / or 2D ultrasound.
[0039] At least some implementations of the methods, systems, devices, and / or code instructions described herein address the technical problem of generating 2D reference images of 3D images and / or 2D videos and / or improve the technical field of generating 3D images and / or 2D videos. The 2D reference images can be used by a user in order to help navigate the 3D images and / or 2D videos. For example, a user reviews the 2D reference images in order to help determine which anatomical regions appear suspicious to include a visual finding (e.g., cancer) in order to spend more time reviewing the corresponding anatomical regions in the frames of the 3D images and / or 2D videos.
[0040] At least some implementations of the methods, systems, devices, and / or code instructions described herein address the technical problem of generating 2D reference images of 3D images and / or sequential 2D images by feeding 2D slices of the 3D images and / or frames of the 2D videos into a 2D classifier that generates non-localized results and / or improve the technical field. The 2D classifier is trained on a training dataset of 2D images with non-localized labels, i.e., the labels are for the 2D images as a whole without an indication of the location of a visual finding in the 2D images. An explanation map is computed for each fed 2D slice and / or frame. The explanation map includes weights that indicate the influence of the respective regions of the fed 2D slice and / or frame on the results of the 2D classifier. A 2D composite image is created by aggregating the weights of the explanation maps. Pixels in the 2D composite image that represent a visual finding are depicted with a higher relative weight relative to other pixels in the 2D composite image that do not depict the visual finding, e.g., they appear brighter.
[0041] At least some implementations of the methods, systems, apparatuses, and / or code instructions described herein differ from other standard methods for creating 2D reference images from 3D images. For example, in some methods, a 3D image is projected onto a 2D plane to create a 2D reference image that does not provide any context awareness, e.g., a standard CINE VIEW. In such images, important visual findings can be obscured by other non-important anatomical features and / or artifacts. In another example, a 3D image is projected onto a 2D plane to create a 2D reference image using context awareness, e.g., using maximum intensity projection (MIP). MIP is performed based on localization information provided by a 2D classifier. In yet another method, a 3D image is divided into 2D slices, where each slice is input into a 2D classifier that generates a result indicating the location of a visual finding in the corresponding image. Such 2D classifiers are trained on a training dataset of 2D images that are labeled with the location of a visual finding depicted therein. Such 2D classifiers that generate the location of a visual finding are difficult and / or resource intensive to create because it is difficult to create a training dataset with localization data because they require manual labeling and thus can not be available, or a limited number of images can be available. In contrast, at least some implementations of the methods, systems, apparatuses, and / or code instructions described herein use a 2D classifier that generates a non-localization indication of a visual finding. The 2D classifier can be trained on a training dataset with non-localization labels that can be created automatically from radiology reports using NLP methods to automatically extract the labels. By aggregating the weights of the computed explanation maps for each 2D slice of a 3D image and / or for frames of a 2D video, the location data of the generated composite 2D explanation image is obtained.
[0042] At least some implementations of the methods, systems, apparatuses, and / or code instructions described herein address the technical problem of increasing the accuracy of a 2D classifier that generates a non-localization indication of a visual finding in a 2D image, e.g., a slice of a 3D image and / or a frame of a 2D video, and / or improve the art. In addition to or instead of training on 2D slices of 3D images and / or frames of videos, the accuracy of the classifier is improved by computing a corresponding composite 2D explanation image of a 3D image and / or 2D video of a training dataset (as described herein) and training the 2D classifier on the composite 2D explanation image.
[0043] At least some implementations of the methods, systems, devices, and / or code instructions described herein address the technical problem of and / or improve the technical field of improving the ability to identify salient visual findings in 3D images. Viewing a 3D image in a non-optimal orientation can hinder important visual findings. For example, a small tumor located in the liver can be obscured by other anatomical features and / or artifacts at certain viewing orientations. At least some implementations of the methods, systems, devices, and / or code instructions described herein provide a technical solution to the technical problem and / or improve the technical field by computing an optimal viewing orientation of a 3D medical image in order to minimize the obstruction of visual findings by other anatomical features and / or artifacts. The optimal viewing orientation is computed as a respective axis along which the 3D medical image is sliced in order to generate a respective synthetic 2D explanatory image for which an aggregated weight of the explanatory map is maximized, e.g., in a cluster. The maximization of the aggregated weight (e.g., in a cluster) represents the best view of the visual finding. The 3D image can be presented to a user in the optimal viewing orientation.
[0044] Before one or more embodiments of the application are explained in detail, it is to be understood that the application is not limited in its application to the details of construction and the
[0045] The application can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present application.
[0046] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0047] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions to storage media within the respective computing / processing device for execution by a processor.
[0048] Computer readable program instructions for carrying out operations of the present application can be assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present application.
[0049] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0050] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including
[0051] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable apparatus or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0052] The flow diagrams and block diagrams in the drawings are illustrative of possible architectures, functions, and operations for systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flow diagrams and block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical functions (‘instructions’). In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0053] The flow diagrams and block diagrams in the drawings are illustrative of possible architectures, functions, and operations for systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flow diagrams and block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical functions (‘instructions’). In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0054] Reference is now made to Figure 1 which is a flowchart of a method of generating a synthetic 2D explanatory image from a 3D medical image, in accordance with some embodiments of the present application. Reference is also made to Figure 2 which is a block diagram of components of a system 200 for generating a synthetic 2D explanatory image from a 3D medical image and / or for training a 2D classifier using a synthetic 2D explanatory image generated from a 3D medical image, in accordance with some embodiments of the present application. Reference is also made to Figure 3 which is a flowchart of a method of training a 2D classifier using a synthetic 2D explanatory image generated from a 3D medical image, in accordance with some embodiments of the present application. The system 200 can be implemented by one or more hardware processors 202 of a computing device 204 executing code instructions stored in a memory (also referred to as program storage) 206 to implement the reference Figure 1 and / or Figure 3 features of the described methods.
[0055] The computing device 204 can be implemented, for example, as a client terminal, a server, a radiology workstation, a virtual machine, a virtual server, a computing cloud, a mobile device, a desktop computer, a thin client, a smart phone, a tablet computer, a laptop computer, a wearable computer, a glasses computer, and a watch computer.
[0056] The computing 204 can include an advanced visualization workstation, which is sometimes attached to a radiology workstation and / or other devices.
[0057] The computing device 204 and / or the client terminal 208 and / or the server 218 can be implemented, for example, as a radiology workstation, an image viewing station, a picture archiving and communication system (PACS) server, and an electronic medical record (EMR) server.
[0058] Multiple architectures of the system 200 based on the computing device 204 can be implemented. In an exemplary implementation, the computing device 204 storing the code 206A can be implemented as one or more servers (e.g., a web server, a network server, a computing cloud, a virtual server) providing services to one or more servers 218 and / or client terminals 208 over a network 210 (e.g., the reference Figure 1one or more of the described actions, e.g., providing software as a service (SaaS) to the server 218 and / or client terminal 208, providing software services accessible using a software interface (e.g., an application programming interface (API), a software development kit (SDK)), providing applications for local download to the server 218 and / or client terminal 208, and / or providing functionality to the server 218 and / or client terminal 208 using a remote access session, such as through a web browser and / or a viewing application. For example, a user uses the client terminal 208 to access the computing device 204 that acts as a PACS server or other medical image storage server. The computing device 204 computes a composite image from 3D medical images provided by the client terminal 208 and / or obtained from another data source (e.g., a PACS server). The composite image can be provided to the client terminal 208 for rendering on a display of the client terminal 208 (e.g., alongside a rendering of the 3D medical images) and / or provided for further processing and / or storage. Alternatively or additionally, the composite image is used to update training of the 2D classifier 220C as described herein. For example, the updated 2D classifier 220C can be used as described herein. Other features can be performed centrally by the computing device 204 and / or locally at the client terminal 208. In another implementation, the computing device 204 can include locally stored software (e.g., code 206A) that performs the functions described with reference to the server 218 and / or the client terminal 208. Figure 1 and / or Figure 3 one or more of the described actions, e.g., as a stand-alone client terminal and / or server. The composite image can be computed locally from the 3D images and / or 2D frames and rendered on a display of the computing device 204. In yet another implementation, the server 218 is implemented as a medical image storage server. A user uses the client terminal 208 to access the composite image from the server 218. The composite image(s) can be computed locally by the server 218 and / or by the computing device 204 using 3D images and / or 2D frames that can be stored on the server 218 and / or at another location. The composite image is rendered on a display of the client terminal 208. The computing device 204 can provide enhanced features to the image server 218 by computing composite images from 3D images and / or 2D frames stored by the image server 218. For example, the PACS server 218 communicates with the computing device 204 using an API to transfer 3D images and / or composite images to the computing device 204 and / or receive computed composite images.
[0059] The computing device 204 receives 3D medical images and / or 2D images (e.g., as a video acquisition) captured by the medical imaging device(s) 212. The medical imaging device 212 can capture 3D images, e.g., CT, MRI, breast tomosynthesis, 3D ultrasound, and / or nuclear images such as PET. Alternatively or additionally, the medical imaging device 212 can capture a video of 2D images, e.g., colonoscopy, bronchoscopy, endoscopy, and 2D ultrasound.
[0060] The medical images captured by the medical imaging device 212 can be stored in an anatomical image repository 214, e.g., a storage server, a computing cloud, a virtual memory, and a hard drive. As described herein, the 2D images 220D created by partitioning 3D images and / or 2D slices and / or 2D frames captured as a video can be stored in the medical image repository 214, and / or in other locations such as the data storage device 220 of the computing device 204, and / or on another server 218. As Figure 2 The storage of 2D images 220D by the data storage device 220 is one non-limiting example.
[0061] The computing device 204 can receive 3D images and / or 2D frames and / or sequential 2D medical images via one or more imaging interfaces 226 (e.g., wired connections (e.g., physical ports), wireless connections (e.g., antennas), network interface cards, other physical interface implementations, and / or virtual interfaces (e.g., software interfaces, application programming interfaces (APIs), software development kits (SDKs), virtual network connections)).
[0062] The memory 206 stores code instructions executable by the hardware processor 202. Exemplary memory 206 includes random access memory (RAM), read only memory (ROM), storage devices, non-volatile memory, magnetic media, semiconductor memory devices, hard drives, removable storage, and optical media (e.g., DVD, CD-ROM). For example, the memory 206 can be code 206A that performs one or more acts of the methods described with reference to FIGS. 1-3. Figure 1 and / or 3.
[0063] The computing device 204 can include a data storage device 220 for storing data, e.g., GUI code 220A (which can present a synthetic image, such as alongside a 3D image), XAI code 220B for computing an explanation graph, a 2D classifier that receives 2D images as input, and / or 2D images 220D obtained by dividing a 3D medical image and / or as video frames, as described herein. The data storage device 220 can be implemented, e.g., as a memory, a local hard drive, a removable storage unit, an optical disc, a storage device, a virtual memory, and / or a remote server 218 and / or a computing cloud (e.g., accessed over the network 210). Note that the GUI 220A and / or the XAI code 220B and / or the 2D classifier 220C and / or the 2D images 220D can be stored, e.g., in the data storage device 220, with the executing portions loaded into the memory 206 for execution by the processor(s) 202.
[0064] The computing device 204 can include a data interface 222 (optionally a network interface), e.g., a network interface card, a wireless interface for connecting to a wireless network, a physical interface for connecting to a cable for network connectivity, a software-implemented virtual interface, network communication software providing higher layer network connectivity, and / or one or more of other implementations, for connecting to the network 210.
[0065] The computing device 204 can use the network 210 (or another communication channel, such as over a direct link (e.g., cable, wireless) and / or an indirect link (e.g., via an intermediary computing unit such as a server, and / or via a storage device) to connect with one or more of:
[0066] * a client terminal 208, e.g., a user uses the client terminal 208 to access the computing device 204 to view a synthetic image computed based on 3D images (and / or sequential 2D images) stored on a server (e.g., the computing device 204 acts as a PACS server). The synthetic image can be computed by the computing device 204 and / or by the client terminal 208.
[0067] * a server 218, e.g., when the server 218 is implemented as a PACS server, where a user uses the client terminal 208 to access the PACS server. The computing device 204 provides enhanced features to the PACS server, receives 3D images and / or 2D video frames from the PACS server, and provides synthetic images to the PACS server, where the client terminal accesses the synthetic images from the PACS server.
[0068] * a medical image repository 214 that stores the captured 3D images and / or 2D video frames and / or synthetic image(s). The medical image repository 214 can store 2D images created by dividing 3D images.
[0069] The computing device 204 and / or the client terminal 208 and / or the server 218 includes and / or is in communication with one or more physical user interfaces 224 that include a display for presenting synthetic images and / or 3D images and / or video frames, and / or mechanisms for interacting with synthetic images and / or 3D images, such as rotating an axis of view of a 3D image, zooming in and / or out of a synthetic image and / or 3D image, and / or marking findings on a synthetic image. An exemplary user interface 208 includes one or more of, for example, a touchscreen, a display, a keyboard, a mouse, and voice-activated software using a speaker and a microphone.
[0070] Now referring back to Figure 1 At 102, a 3D medical image is obtained. Alternatively, sequential 2D images are obtained. The sequential 2D medical images can be captured as a video over a time interval. The 2D medical images can be spaced in time, for example, 2D frames per second, or other value. The sequential 2D images can be obtained at different locations along a body region, effectively mapping a 3D volume within a body, for example, along a colon, along an esophagus, along an airway, and / or along different 2D slices of a body (e.g., 2D ultrasound slices of a liver).
[0071] Examples of 3D medical images include: CT, MRI, mammotomography, digital breast tomosynthesis (DBT), 3D ultrasound, 3D nuclear imaging, and PET.
[0072] Examples of 2D medical images include: colonoscopy, endoscopy, bronchoscopy, and 2D ultrasound.
[0073] At 104, the 3D medical image can be divided into 2D images, optionally 2D slices. The 2D slices can be parallel to each other and sliced along a common slice plane. The 3D medical image can be automatically divided into 2D slices along a predetermined axis (e.g., by a PACS server, by a CT machine, by DICOM-based code, by viewing software). Alternatively, the slice axis is selected by a user and / or automatically selected by code, for example, as described herein.
[0074] The sequential 2D images are already considered to be divided. Alternatively, to obtain another slice axis, a 3D image can be reconstructed from the sequential 2D images, and then the 3D reconstructed image is sliced along the selected axis.
[0075] At 106, the respective 2D medical images are input into a 2D classifier.
[0076] Optionally, the 2D classifier has been pre-trained on a training dataset of 2D medical images labeled with an indication of a visual finding depicted therein.
[0077] Optionally, the indication of the visual finding of the training dataset is non-localized. The training 2D medical images can be associated with labels indicating the presence or absence of a visual finding in the respective training 2D medical image as a whole, without indicating where in the image the visual finding is located. The 2D classifier generates a result indicating the visual finding for the input 2D image as a whole, with non-localized data, e.g., outputting a yes / no value indicating the presence or absence of the visual finding for the image as a whole.
[0078] Optionally, the visual finding is an indication of cancer, e.g., breast cancer, lung nodule, colon cancer, brain cancer, bladder cancer, kidney cancer, and metastatic disease. Alternatively, other visual findings can be defined. The cancer can be treated by applying a treatment suitable for treating the cancer, optionally a treatment for the type of cancer, e.g., surgery, chemotherapy, radiation therapy, immunotherapy, and combinations thereof.
[0079] At 108, a respective explanation map is computed for the respective 2D medical image. The explanation map comprises a plurality of regions corresponding to regions in the respective 2D image, e.g., each respective pixel of the explanation map corresponds to each respective pixel of the 2D image, a group of respective pixels of the explanation map (e.g., 2x2, 3x3, or other region) corresponds to a single pixel of the 2D image, a single respective pixel of the explanation map corresponds to a group of pixels of the 2D image, and / or a group of respective pixels of the explanation map corresponds to a group of pixels of the 2D image. Each respective region of the respective explanation map is associated with a computed explainable weight. The explainable weight indicates the influence of the respective pair of corresponding regions of the respective 2D medical image on the result of the 2D classifier of the respective 2D medical image. Each explanation weight of each respective explanation map can represent a relative influence of the respective pair of corresponding regions on the result of the 2D classifier, e.g., a first pixel has a weight of 50 and a second pixel has a weight of 10, indicating that the first pixel is 5 times higher than the second pixel in the decision of the 2D classifier of the visual finding. Optionally, the sum of the explainable weights of the plurality of regions of each respective explanation map can total to 100% (or 1), or the sum of the weights does not have to be a fixed value.
[0080] The explanation map can be implemented as a heat map.
[0081] An explanation map can be computed using XAI code that computes weights of regions (e.g., each pixel or each group of pixels) that most influence decisions of the 2D classifier, optionally generating a heat map. Exemplary XAI code is described, for example, in R.C. Fong and A. Vedaldi, 2017. Interpretable explanations of blackboxes by meaningful perturbation. arXiv preprint arXiv: 1704.03296, and / or Shoshan, Y. and Ratner, V., 2018. Regularized adversarial examples for model interpretability. arXiv preprint arXiv: 1811.07311.
[0082] The explanation map can comprise pixels corresponding to pixels of the respective 2D medical image with pixel intensity values adjusted by the corresponding respective interpretable weights. For example, for a particular pixel of a 2D image, the pixel can have a pixel intensity value of 75, the corresponding interpretable weight is computed to be 0.6, obtaining a pixel intensity value of 45 for the explanation map.
[0083] In terms of mathematical representation, the 3D image can be represented as V, the explanation map (e.g., heat map) can be represented as H, where voxels of the 3D image are represented as V(x, y, z), and the corresponding interpretable weights are represented as H(x, y, z), which indicate influence on decisions of the 2D classifier when the number of feed slices of the volume V is represented as z.
[0084] At 110, the features described with reference to 106-108 can be iterated sequentially and / or in parallel for the 2D medical images. Optionally, 106-108 are iterated for each 2D medical image. Alternatively, a subset of the 2D medical images is selected, for example, by uniform sampling. The sampled subset of 2D medical images can be processed as described with reference to 106-108.
[0085] At 112, a subset of the explanation maps can be selected to create a composite 2D image. The explanation maps can be selected according to one or more clusters comprising one or more regions with explanation weights that meet a requirement, e.g., the cluster has a total explanation weight value that is higher than the explanation weight of other regions excluded from the cluster and / or a mean explanation weight value that is higher than a threshold and / or range. For example, the weight of pixels in the cluster is at least 0.5 higher than the weight of pixels not included in the cluster. In another example, the cluster of explanation weights above a threshold has at least a minimum dimension. For example, a cluster is defined as a group of pixels with an explanation weight greater than 0.6 and / or with a dimension of at least 5x5 pixels or 10x10 pixels and / or a mean explanation weight greater than 0.7 or other values.
[0086] At 114, a composite 2D explanation image is computed from the plurality of explanation maps. The composite 2D explanation image includes a respective aggregated weight for each respective region thereof, e.g., each pixel or each group of pixels (e.g., 2x2, 3x3, or other dimension). Each respective aggregated weight can represent a respective likelihood of presence of a visual finding at a corresponding respective region of the computed composite 2D explanation image.
[0087] The composite 2D explanation image can be a projection of the 3D image to 2D image via the weights of the explanation maps.
[0088] Each respective aggregated weight is computed by aggregating the explainable weights computed for the respective region of an explanation map corresponding to the respective region of the composite 2D explanation image. Optionally, each 2D medical image includes pixels corresponding to voxels of the 3D medical image. Each respective explainable weight is assigned to each pixel of each 2D medical image. For each pixel in the composite 2D explanation image having specific (x, y) coordinates, the respective aggregated weight can be computed by aggregating the explainable weights of the pixels in the 2D medical images having corresponding (x, y) coordinates for variable z coordinates. For example, for 2D medical images obtained by dividing the 3D medical image into 2D slices along the z axis, a respective aggregated weight is computed for each respective region (e.g., pixel or group of pixels) of a 2D slice having common x, y coordinates along the x and y axes and variable z coordinates along the z axis, which can indicate the slice number. For example, the explainable weights of the 5 explanation maps at (x, y, z) coordinates (10, 15, 1), (10, 15, 2), (10, 15, 3), (10, 15, 4), and (10, 15, 5) are aggregated into a single aggregated weight and assigned to the pixel of the composite 2D image at (x, y) coordinates (10, 15). The aggregated weight at coordinates (10, 15) of the composite 2D image corresponds to a voxel at (x, y, z) coordinates (10, 15, z) of the 3D image, where z is variable within the z value range of the 3D image.
[0089] Each respective aggregate weight of the synthetic 2D interpretation image can be computed as, for example, a weighted average, a median, a sum, a maximum, or a mode of the explainable weights computed for the respective region of the interpretation map corresponding to the respective region of the synthetic 2D interpretation image.
[0090] Optionally, when the interpretation map comprises pixels corresponding to pixels of the respective 2D medical image with pixel intensity values adjusted by the corresponding respective explainable weights, the synthetic 2D interpretation image comprises pixels with pixel intensity values computed by aggregating the pixel intensity values adjusted by the corresponding respective explainable weights of the interpretation map.
[0091] In mathematical notation, the 2D synthetic image is denoted as C, where each pixel denoted as (x, y) is an aggregation of the slices (e.g., all slices) denoted as V(z, y, :), weighted by the respective interpretation map (e.g., heat map) weights denoted as (H(x, y, :)), where the following example equation holds:
[0092] C(x, y) = sum_over_z(H(x, y, z) * V(x, y, z)) / sum_over_z(H(x, y, z))
[0093] At 116, optionally, a best viewing angle of the synthetic 2D image is computed. The 3D image can be presented at the determined best viewing angle. The best viewing angle represents a minimum occlusion of the visual findings within the 3D image.
[0094] The best viewing angle can represent a best angle via which the 3D image is projected to the synthetic 2D image by the weights of the interpretation map.
[0095] The best viewing angle can correspond to a slice angle used for creating the 2D slices from the 3D image, i.e., a particular orientation of the z-axis defining an axis by which the 3D medical image is sliced into 2D slices. The 2D slices sliced at the best viewing angle are used to generate the best synthetic 2D interpretation image with the maximum aggregate weight representing the minimum occlusion of the visual findings. For example, the best viewing angle can be selected by a sequential and / or parallel trial-and-error method, by evaluating a plurality of synthetic 2D images (e.g., randomly selected and / or sequentially selected starting from predetermined values) computed for different best viewing angles, wherein the best viewing angle is selected according to the best synthetic 2D image, wherein the occlusion of the visual findings within the 3D image is minimal when the 3D image is presented at the best viewing angle. Alternatively or additionally, the best viewing angle can be computed, e.g., based on code analyzing the 3D image and / or the synthetic 2D images to select the best orientation.
[0096] Note that the optimal viewing angle can be determined at one or more features of the process for computing and / or presenting a synthetic 2D image, e.g., prior to initially dividing a 3D image into 2D images (e.g., as described with reference to 104), and / or by iterating 104-114 in a trial-and-error process and / or in other appropriate parts of the process.
[0097] At 118, the 2D synthetic image is provided, e.g., presented on a display, stored in a memory and / or data storage device (e.g., a PACS server), forwarded to another device (e.g., from a PACS server to a client terminal for viewing thereon), and / or provided to another process, e.g., fed into another classifier, fed into a 2D classifier, and / or used to update training of a 2D classifier.
[0098] Optionally, the 2D synthetic image is presented concurrently with the 3D image, e.g., side-by-side. The 2D synthetic image can replace a standard summary image created for the 3D image, e.g., using CVIEW.
[0099] Optionally, when computing the optimal 2D synthetic image according to the optimal angle of the determined viewing axis, the presentation of the 3D medical image on the display is automatically adjusted to an orientation (e.g., an orientation of the z-axis) corresponding to the optimal viewing angle.
[0100] At 120, one or more features described with reference to 104-118 are iterated, optionally, e.g., according to real-time user navigation, the 2D synthetic image is dynamically updated to correspond to a real-time viewing axis of the 3D image.
[0101] A user can adjust the angle of the viewing axis of a 3D image presented on a display. A real-time value of the viewing axis of the 3D image can be tracked. The orientation of the z-axis defining an axis along which the 3D medical image is sliced into 2D slices (e.g., as described with reference to 104) can be set according to the real-time value of the viewing axis selected by a user viewing the 3D medical image presented on the display. A current synthetic 2D interpretation image is computed based on the z-axis corresponding to the value of the viewing axis (e.g., as described with reference to 106-114). The current synthetic 2D is presented on the display side-by-side with the 3D medical image (e.g., as described with reference to 118). Changes in the value of the viewing axis of the 3D medical image presented on the display are dynamically detected (e.g., as described with reference to 120). An updated synthetic 2D interpretation image is dynamically computed based on the changes in the value of the viewing axis (e.g., as described with reference to 106-114). The display is dynamically updated by presenting the updated synthetic 2D interpretation image (e.g., as described with reference to 118).
[0102] Now returning to reference to Figure 3At 302, a plurality of training 3D medical images of sample subjects is accessed. The training 3D medical images are optionally all of the same type of imaging modality, depicting the same body location, for finding the same type of visual finding, e.g., all chest CT scans for locating lung nodules, and / or all 3D mammograms for locating breast cancer.
[0103] At 304, the respective 3D medical images are divided into a plurality of 2D medical images, e.g., as described with reference to 104 of Figure 1 .
[0104] At 306, the respective 2D medical images (e.g., each) are input into a 2D classifier trained on a training dataset of 2D medical images labeled with an indication of a visual finding depicted therein, e.g., as described with reference to 106 of Figure 1 .
[0105] At 308, a respective explanation map of the respective 2D medical image is computed. The respective explanation map comprises regions corresponding to regions of the respective 2D image. Each respective region of the respective explanation map is associated with a computed explainable weight indicating an influence of the respective corresponding region of the respective 2D medical image on a result of the 2D classifier provided with the respective 2D medical image. For example, as described with reference to 108 of Figure 1 .
[0106] At 310, a respective synthetic 2D explanation image is computed. The synthetic 2D explanation image comprises a respective aggregate weight for each respective region thereof. Each respective aggregate weight is computed by aggregating the explainable weights computed for respective regions of explanation maps corresponding to respective regions of the synthetic 2D explanation image. For example, as described with reference to 110 of Figure 1 .
[0107] At 312, a respective label is assigned to the synthetic 2D explanation image, indicating a presence of a visual finding depicted therein. For example, the label can be created manually by a user based on a manual visual inspection of the 3D image and / or the synthetic 3D explanation image, and / or automatically, e.g., by natural language processing (NLP) code analyzing a radiology report created for the 3D image to extract the visual finding.
[0108] At 314, the respective synthetic 2D explanation image and the corresponding label can be added to an updated training dataset.
[0109] At 316, one or more features described with reference to 304-314 are iterated for the plurality of 3D training medical images, optionally for each 3D medical image.
[0110] In step 318, an updated 2D classifier can be created by updating its training using an updated training dataset. This updated 2D classifier can then be used to create new synthetic 2D images, for example, in reference... Figure 1 106 and / or Figure 3 The 306 description is used in the process.
[0111] Optionally, after accessing the training 3D medical images in 320, as in 302, and after segmenting the 3D images in 304, a 2D classifier can be created and / or updated. Each corresponding 2D medical image can be associated with a label (e.g., manually and / or automatically created, as described herein) indicating the presence of visual findings depicted with the corresponding 2D medical image. The label can be non-local, i.e., assigned to the corresponding 2D medical image as a whole. A training dataset of 2D medical images can be created by including the 2D medical images and non-local related labels created from the 3D medical images. The training dataset can be used to create and / or update the 2D classifier.
[0112] Now for reference Figure 4 This is a schematic diagram depicting corresponding synthetic 2D interpretation images 400A, 400B, and 400C, compared to other standard methods for calculating 2D images from 3D images, according to some embodiments of the present invention.
[0113] For a 2D image sliced along the z-axis of a 3D image, for a specific fixed y-value on the y-axis, and for a set of x-values along the x-axis (e.g., horizontal lines of pixels), composite 2D interpretive images 400A, 400B, and 400C are calculated. Composite 2D interpretive images 400A, 400B, and 400C represent horizontal lines of pixels, also known as composite interpretive lines. For clarity and simplicity of interpretation, horizontal lines of pixels are depicted. It should be understood that a complete composite 2D interpretive image comprises multiple parallel horizontal lines of pixels along the y-axis.
[0114] Each of the synthesized 2D interpretative images 400A, 400B, and 400C is based on a common 3D signal represented as F(x, y, z), for which a 3D image 402 is shown, i.e., a single horizontal line depicting the pixels at the same y-value for each 2D slice. Within the 3D image 402, a first circle 404 and a second circle 406 represent clinically meaningful visual findings, while rectangles 408 and ellipses 410 represent meaningless anatomical features and / or artifacts.
[0115] The sum of the lines in the 3D image 402 is calculated and represented as P(x, y) = The synthetic interpretation line 400A is computed using the standard method. Note that the first circle 404 is partially occluded by the rectangle 408, and the second circle 406 is fully occluded by the rectangle 408, making it difficult to discretize the presence of the first circle 404 and the second circle 406 within the synthetic interpretation line 400A.
[0116] The synthetic interpretation line 400B is computed using a state-of-the-art (SOTA) method, in which a heat map indicating the location of the visual finding is generated by a 2D classifier for each 2D image of the 3D image 402. The 2D classifier is trained on a training dataset of 2D images, in which labels are assigned to the location of the visual finding within the training images. The heat map output of the 2D classifier for each 2D slice is represented as D z (x, y). The formula The synthetic interpretation line 400B is computed. Note that although the first circle 404 is partially occluded by the rectangle 408, and the second circle 406 is fully occluded by the rectangle 408, the presence of the first circle 404 and the second circle 406 is discernible within the synthetic interpretation line 400B, as the higher heat map is higher at locations corresponding to the first circle 404 and the second circle 406, and lower at locations corresponding to the rectangle 408 and the ellipse 410.
[0117] The synthetic interpretation line 400C is computed using a 2D classifier that generates non-localization results, as described herein, in which the 2D classifier is trained on a training dataset with non-localization labels, and the synthetic interpretation image is computed by aggregating the interpretation weights of the interpretation maps. The interpretation map computed for each 2D slice is represented as D z (x, y). The formula The synthetic interpretation line 400C is computed. Note that the presence of the first circle 404 and the second circle 406 is discernible within the synthetic interpretation line 400C, at least as much as the synthetic interpretation line 400B computed using the state-of-the-art method, with the added advantage that the 2D classifier is trained using non-localization labels, which provides for the automatic training of the 2D classifier using labels such as automatically extracted from radiology reports using NLP methods.
[0118] Reference is now made to Figure 5 which is a schematic diagram illustrating the automatic computation of a particular orientation of a z-axis that defines an axis by which a 3D medical image is sliced into 2D slices, for generating an optimal synthetic 2D interpretation image having a maximum aggregated weight representing the maximum likelihood of the presence of a visual finding, according to some embodiments of the present application. The presentation of the 3D medical image on a display can be adjusted according to the particular orientation of the z-axis, and / or the optimal synthetic 2D interpretation image computed based on the particular orientation of the x-axis can be presented on the display.
[0119] Diagram 502 is for the case of the standard z-axis 504. The synthetic 2D interpretation image 506 is computed for 2D images that are slices along the z-axis 504 of the 3D image 508. For clarity and simplicity of interpretation, the synthetic 2D interpretation image 506 represents a horizontal line of pixels, also referred to as a synthetic interpretation line, which is computed from 2D images that are slices along the z-axis 504 of the 3D image for a particular fixed y value of the y-axis for a set of x values along the x-axis, e.g., a horizontal line of pixels. It will be appreciated that a complete synthetic 2D interpretation image includes multiple parallel horizontal lines of pixels along the y-axis.
[0120] Within the 3D image 508, the circle 510 represents a clinically meaningful visual finding, while the rectangle 512 and the ellipse 514 represent non-meaningful anatomical features and / or artifacts.
[0121] For the synthetic 2D interpretation image 506 created using the standard z-axis 504, the circle 510 and the rectangle 512 are along the same line parallel to the standard z-axis 512. As a result, the weight of the circle 510 is aggregated with the weight of the rectangle 512, which can make it more difficult to distinguish the circle 510. In the case where the rectangle 512 represents a clinically meaningful visual finding, the combination of the weight of the circle 510 with the weight of the rectangle 512 can make it more difficult to distinguish that there are two spatially separated visual findings in the 3D image 508.
[0122] In contrast, diagram 516 is for the case of the selected z-axis 518 used to generate the optimal synthetic 2D interpretation image 520 with the greatest aggregated weight representing the greatest likelihood of existence of a visual finding. For the synthetic 2D interpretation image 520 created using the selected z-axis 518, each of the circle 510, the rectangle 512, and the ellipse 514 are along different lines parallel to the selected z-axis 518. As a result, the weight of the circle 510 is not aggregated with the weight of the rectangle 512, and is not aggregated with the weight of the ellipse 514, which makes it possible to better distinguish the weights of the different visual findings on the optimal synthetic 2D interpretation image 520.
[0123] The description of various embodiments of the present application has been presented for purposes of illustration but is not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein was chosen to best explain the principles of the embodiments, the practical application or technical improvement over technology found in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0124] It is expected that many related classifiers and / or XIA processes will be developed during the life of a patent maturing from this application, and the scope of the terms classifiers and / or XIA processes are intended to include all such new technologies a priori.
[0125] As used herein, the term "about" means ±10%.
[0126] The terms "comprise," "comprising," "include," "including," "have," "having," and "contain," "containing," and variations thereof, mean "including but not limited to."
[0127] The phrase "consisting essentially of" means that the composition or method can include additional ingredients and / or steps, but only if the additional ingredients and / or steps do not materially alter the basic and novel characteristics of the claimed composition or method.
[0128] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a compound" or "at least one compound" can include a plurality of compounds, including mixtures thereof.
[0129] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.
[0130] The word "optionally" is used herein to mean "may or can not be present." Any particular embodiment of the application can include a plurality of "optional" features, unless such features conflict.
[0131] In this application, various embodiments of the application can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and is to be interpreted -in the context of the specification as a whole. Therefore, the description of a range it should be considered to include all possible subranges and individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to include the individual numbers 1, 2, 3, 4, 5, and 6, and subranges such as from 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, etc. This same logic should be applied to ranges reciting fractions or decimals.
[0132] Whenever a numerical range is indicated, it is meant to include any cited numeral (fractional or integral) within the indicated range. The phrases "ranging / range between" a first indicate number and a second indicate number and "ranging / range from" a first indicate number "to" a second indicate number are used herein interchangeably and are meant to include the first and second indicated numbers and all the fractional and integral numerals therebetween.
[0133] It is to be understood that certain features of the application described in the context of separate embodiments can also be provided in combination in a single embodiment. Conversely, various features of the application described in the context of a single embodiment can also be provided separately or in any appropriate
[0134] Although the application has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations as fall within the spirit and broad scope of the appended claims.
[0135] It is the intent of the applicant(s) that all publications, patents and patent applications referred to in this specification be incorporated in their entirety by reference into the specification for purposes of the disclosure and description of the application being in the spirit and scope of the application. In the case of inconsistencies between the disclosure of the specification and the disclosure of any publication, patent or patent application incorporated by reference, the disclosure of the specification shall prevail. Further, the citation or identification of any reference in this application shall not be construed as an admission by the applicant(s) that such reference is available as prior art to the present application. To the extent that section headings are used, they should not be construed as necessarily limiting. In addition, any priority document(s) of this application is / are hereby incorporated by reference in its / their entirety.
Claims
1. A computer-implemented method of generating a synthetic 2D explanatory image from a 3D medical image, comprising: inputting each 2D medical image of a plurality of 2D medical images obtained by at least one of: partitioning a 3D medical image, and being captured as a video over a time interval, into a 2D classifier trained on a training dataset of 2D medical images labeled with indications of visual findings depicted therein; for each respective 2D medical image of the plurality of 2D medical images, computing a respective explanatory map of a plurality of explanatory maps, the respective explanatory map comprising a plurality of regions corresponding to a plurality of corresponding regions in the respective 2D image, each respective region in the respective explanatory map being associated with a computed explainable weight indicative of an influence of the respective corresponding region in the respective 2D medical image on a result of the 2D classifier provided for the respective 2D medical image; computing a synthetic 2D explanatory image comprising a respective aggregated weight for each respective region thereof, each respective aggregated weight being computed by aggregating a plurality of the computed explainable weights for respective regions of the plurality of explanatory maps corresponding to the respective region in the synthetic 2D explanatory image; and providing the synthetic 2D explanatory image for presentation on a display.
2. The method of claim 1, wherein, Each respective aggregated weight represents a respective likelihood of presence of a visual finding at a corresponding respective region of the computed synthetic 2D explanatory image.
3. The method of claim 1, wherein, The plurality of 2D medical images is computed by partitioning the 3D medical image into a plurality of sequential 2D slices along a z-axis, wherein a respective aggregated weight is computed for each respective region of the plurality of sequential 2D slices having common x, y coordinates along x and y axes and variable z coordinates along the z-axis.
4. The method of claim 3, wherein, An orientation of the z-axis defining an axis by which the 3D medical image is sliced into the plurality of sequential 2D slices is obtained in accordance with an eye- axis selected by a user viewing the 3D medical image presented on a display, wherein the synthetic 2D explanatory image computed based on the z-axis corresponding to the eye-axis is presented on the display together with the 3D medical image, and the method further comprises, in at least one iteration, dynamically detecting a change in the eye-axis of the 3D medical image presented on the display; based on the change in the eye-axis, dynamically computing an updated synthetic 2D explanatory image; and dynamically updating the display to present the updated synthetic 2D explanatory image.
5. The method of claim 3, further comprising: computing a particular orientation of the z-axis defining an axis by which the 3D medical image is sliced into the plurality of sequential 2D slices, the particular orientation of the z-axis generating an optimal synthetic 2D explanatory image having a maximum aggregated weight representing a minimum occlusion of the visual findings; automatically adjusting a presentation of the 3D medical image on the display to the particular orientation of the z-axis; and presenting the optimal synthetic 2D explanatory image on the display. 6. The method of claim 3, wherein, Each 2D medical image of the plurality of 2D medical images includes pixels corresponding to voxels of the 3D medical image, respective interpretable weights are assigned to each pixel of each 2D medical image of the plurality of 2D medical images, and for each pixel in the synthetic 2D explanation image having specific (x, y) coordinates, a respective aggregated weight is computed by aggregating the interpretable weights of the pixels in the plurality of 2D medical images having corresponding (x, y) coordinates for a variable z coordinate.
7. The method of claim 1, wherein, The indication of the visual finding of the training dataset is non-localized for respective 2D image wholes, and wherein the 2D classifier generates results indicating the visual finding for input 2D image wholes utilizing non-localized data.
8. The method of claim 1, wherein, Each explanation weight of each respective explanation map represents a relative influence of the respective corresponding region on the results of the 2D classifier.
9. The method of claim 1, wherein, Each respective aggregated weight of the synthetic 2D explanation image is computed as a weighted average of the interpretable weights computed for respective regions in the plurality of explanation maps corresponding to respective regions in the synthetic 2D explanation image.
10. The method of claim 1, wherein, Each respective explanation map includes a plurality of pixels corresponding to pixels of a respective 2D medical image having pixel intensity values adjusted by corresponding respective interpretable weights, wherein the synthetic 2D explanation image includes a plurality of pixels having pixel intensity values computed by aggregating the pixel intensity values adjusted by corresponding respective interpretable weights in the plurality of explanation maps.
11. The method of claim 1, wherein, The 3D medical images are selected from the group consisting of: CT, MRI, mammotomography, digital breast tomosynthesis (DBT), 3D ultrasound, 3D nuclear imaging, and PET.
12. The method of claim 1, wherein, The visual finding represents cancer.
13. The method of claim 1, further comprising: A subset of the plurality of explanation maps is selected, wherein each selected explanation map includes at least one cluster of at least one region having explanation weights higher than required explanation weights of other regions excluded from the cluster, wherein the synthetic 2D image is computed from the selected subset.
14. The method of claim 1, further comprising: An updated 2D classifier of the 2D classifier is generated for analyzing 2D images of the 3D medical images by: accessing a plurality of training 3D medical images; for each respective 3D medical image in the plurality of 3D medical images: dividing the respective 3D medical image into a plurality of 2D medical images; inputting each 2D medical image in the plurality of 2D medical images into a 2D classifier trained on a training dataset of 2D medical images labeled with an indication of a visual finding depicted therein; for each respective 2D medical image of the plurality of 2D medical images, computing a respective explanation map of a plurality of explanation maps, the respective explanation map comprising a plurality of regions corresponding to a plurality of corresponding regions in the respective 2D image, each respective region in the respective explanation map being associated with a computed explainable weight, the computed explainable weight indicating an influence of the respective pair of corresponding regions in the respective 2D medical image on the result of the 2D classifier being provided with the respective 2D medical image; computing a synthetic 2D explanation image, the synthetic 2D explanation image comprising a respective aggregated weight for each respective region of the synthetic 2D explanation image, each respective aggregated weight being computed by aggregating a plurality of the computed explainable weights for respective regions of the plurality of explanation maps corresponding to the respective region in the synthetic 2D explanation image; assigning a label indicating a presence of the visual finding depicted therein to the synthetic 2D explanation image; generating an updated training dataset comprising a plurality of the synthetic 2D explanation images and corresponding labels; and generating the updated 2D classifier by updating the training of the 2D classifier using the updated training dataset.
15. The method of claim 14, further comprising: after accessing the plurality of training 3D medical images, partitioning each 3D medical image of the plurality of 3D medical images into a plurality of 2D medical images; labeling each respective 2D medical image with a label indicating a presence of a visual finding depicted with the respective 2D medical image, wherein the label is non-localized and is assigned to the respective 2D medical image as a whole; creating the training dataset of 2D medical images comprising the plurality of 2D medical images and associated non-localized labels; and training the 2D classifier using the training dataset.
16. The method of claim 1, wherein, The plurality of 2D medical images captured as a video are captured by an imaging device selected from the group comprising: a colonoscope, an endoscope, a bronchoscope, and a 2D ultrasound.
17. A device for generating a synthetic 2D explanation image from a 3D medical image, comprising: at least one hardware processor executing code to: input each 2D medical image of a plurality of 2D medical images obtained by at least one of: partitioning a 3D medical image, and being captured as a video over a time interval, to a 2D classifier trained on a training dataset of 2D medical images labeled with an indication of a visual finding depicted therein; for each respective 2D medical image of the plurality of 2D medical images, compute a respective explanation map of a plurality of explanation maps, the respective explanation map comprising a plurality of regions corresponding to a plurality of corresponding regions in the respective 2D image, each respective region in the respective explanation map being associated with a computed explainable weight, the computed explainable weight indicating an influence of the respective pair of corresponding regions in the respective 2D medical image on the result of the 2D classifier being provided with the respective 2D medical image; computing a synthetic 2D interpretation image, the synthetic 2D interpretation image comprising a respective aggregate weight for each respective region thereof, each respective aggregate weight computed by aggregating a plurality of the interpretable weights computed for a respective region of the plurality of interpretation images corresponding to the respective region in the synthetic 2D interpretation image; and providing the synthetic 2D interpretation image for presentation on a display.
18. A computer program product for generating a synthetic 2D interpretation image from a 3D medical image, comprising a non-transitory medium storing a computer program which, when executed by at least one hardware processor, causes the at least one hardware processor to perform: inputting each 2D medical image of a plurality of 2D medical images obtained by at least one of: partitioning a 3D medical image, and being captured as a video over a time interval, into a 2D classifier trained on a training dataset of 2D medical images labeled with an indication of a visual finding depicted therein; for each respective 2D medical image of the plurality of 2D medical images, computing a respective interpretation image of a plurality of interpretation images, the respective interpretation image comprising a plurality of regions corresponding to a plurality of corresponding regions in the respective 2D image, each respective region in the respective interpretation image associated with a computed interpretable weight, the computed interpretable weight indicating an influence of the respective corresponding region of the respective 2D medical image on a result of the 2D classifier provided for the respective 2D medical image; computing a synthetic 2D interpretation image, the synthetic 2D interpretation image comprising a respective aggregate weight for each respective region thereof, each respective aggregate weight computed by aggregating a plurality of the interpretable weights computed for a respective region of the plurality of interpretation images corresponding to the respective region in the synthetic 2D interpretation image; and providing the synthetic 2D interpretation image for presentation on a display.
Citation Information
Patent Citations
Adaptive weight aggregation stereo matching algorithm based on parallax information
CN107564044A
Three-dimensional image visual saliency detection method based on local comparison and global guidance
CN110555434A