Saliency Maps for Medical Imaging
Patent Information
- Application Number
- JP2024509022
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-08-20
- Filing Date
- 2022-08-11
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-08-11
AI Technical Summary
Existing medical imaging reconstruction techniques often treat all regions of an image equally important, failing to account for the varying levels of relevance to human observers, particularly in medical contexts where certain anatomical structures require closer attention for accurate diagnosis and surgical decisions.
A machine learning module trained to predict a saliency map that identifies regions of high user attention on medical images, allowing for personalized and relevance-weighted image reconstruction and analysis by incorporating user attention patterns.
Enhances the quality of medical images by focusing reconstruction efforts on clinically relevant areas, improving diagnostic accuracy and surgical decision-making by aligning image processing with human observer needs.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to medical imaging, and more particularly to saliency maps for medical imaging. [Background technology]
[0002] Image reconstruction is one of the fundamental components of medical imaging. Its main objective is to provide high-quality clinical medical images. Historically, this goal has been achieved using various kinds of reconstruction algorithms. Most of these algorithms, including for example SENSE and compressed SENSE in magnetic resonance imaging (MRI), are based on expert knowledge or hypotheses about the characteristics of the image to be reconstructed. Recently, data-driven techniques based on machine learning have been added to improve the quality of the reconstruction. These data-driven techniques based on machine learning have gradually reduced the dependency of image reconstruction on expert knowledge or hypotheses about the characteristics of the image. Data-driven techniques may even push image reconstruction to the extreme, where most of the models used for image reconstruction will be based on machine learning and the reliance on expert-set parameters may be minimal. Summary of the Invention
[0003] The invention provides a medical system, a computer program and a method as set forth in the independent claims. Embodiments are set forth in the dependent claims.
[0004] It is proposed to provide a machine learning module trained to provide a saliency map predicting the distribution of user attention on a medical image. The saliency map can identify regions of high interest in the medical image, i.e. regions predicted to receive high user attention. This distribution of user attention can be used to weight the relevance of various regions in the medical image. Thus, regions of interest that are more relevant than other regions can be identified and / or selected. The higher the predicted user attention for a certain region, the higher the relevance of this region. Depending on the predicted relevance, different regions can be considered in different ways, e.g. for image reconstruction and / or image analysis. For example, an image reconstruction method can be selected based on the predicted user attention. The selected method can be, for example, the method most suitable for providing a high-quality image reconstruction for the anatomical structures contained in the regions predicted to receive high user attention.
[0005] For example, a quality assessment that evaluates the image quality of the reconstructed medical image, i.e., the quality of the image reconstruction, may be weighted using the distribution of user attention. Thus, reconstruction errors may be identified and / or evaluated based on relevance. A reconstruction error in an area of high relevance may be considered more relevant than a similar reconstruction error in an area of low relevance. For example, a saliency map may be used to weight an out-of-distribution (OOD) map and / or to weight the resulting out-of-distribution (OOD) score that defines the uncertainty of the medical image.
[0006] For example, a saliency map predicting user attention may be used to weight the output of a machine learning module during training of the module to reconstruct medical images, such that training of the machine learning module may be more focused on regions of the reconstructed images that are more relevant based on the predicted user attention.
[0007] A training saliency map for training a machine learning module to predict a saliency map of a given medical image can be generated using eye tracking of a user reading the training medical image. The machine learning module can be trained to predict a personalized saliency map, for example, having a distribution of user attention personalized for an individual user. Thus, personalized weighting can be implemented using the personalized saliency map that predicts the distribution of personalized user attention.
[0008] In one aspect, the present invention provides a medical system. The medical system includes a memory storing machine executable instructions. The memory further stores a first machine learning module trained to output a saliency map as an output in response to receiving a medical image as an input. The saliency map predicts a distribution of a user's attention on the medical image. The medical system further includes a computing system. Execution of the machine executable instructions causes the computing system to receive the medical image. Execution of the machine executable instructions causes the computing system to further provide the medical image as an input to the trained first machine learning module. In response to providing the medical image, execution of the machine executable instructions causes the computing system to further receive a saliency map of the medical image as an output from the trained first machine learning module. The saliency map predicts a distribution of the user's attention on the medical image. Execution of the machine executable instructions causes the computing system to further provide the saliency map of the medical image.
[0009] The saliency map predicting the distribution of user attention may be generated by a trained machine learning module, e.g., a machine learning module trained using deep learning. The machine learning module may include, e.g., a neural network. The neural network may be, e.g., a U-Net neural network.
[0010] Using the trained first machine learning module to output a saliency map as an output in response to receiving a medical image as an input may be part of a testing phase for testing the trained first machine learning module. Using the trained first machine learning module to output a saliency map as an output in response to receiving a medical image as an input may be part of a prediction phase for predicting a saliency map for a received medical image using the trained first machine learning module.
[0011] Machine learning based reconstruction methods are becoming more and more important in medical imaging. In general, known image reconstruction techniques tend to treat different regions of the reconstructed image as equally important. Such an assumption may be perfectly reasonable if images in general are considered. Each region of the image may contribute more or less equally to the overall perception of the image. However, in the case of medical images, experts tend to look more closely at certain regions that show certain anatomical structures of interest than others. They may carefully select important regions, i.e. regions of high interest, and look for certain details shown there. Such details may be necessary, even essential, for making a correct diagnosis and / or surgical decision.
[0012] Using a saliency map to predict the distribution of user attention may have the advantage of incorporating such important information. The predicted distribution of user attention may be used to determine the relevance of different regions of a medical image. This may fill a large gap to a reliable and intelligent system that considers different levels of relevance of different regions of a medical image. Using such a saliency map may allow even data-driven reconstruction techniques to take into account the recipients of the reconstructed medical images and their needs.
[0013] It is proposed to provide a machine learning module, i.e., an algorithm based on machine learning (ML) (e.g., deep learning (DL)) configured to identify regions of interest in 2D and / or 3D medical images (i.e., 2D and / or 3D medical still images) and provide a saliency map predicting the distribution of a user's attention on the 2D and / or 3D medical images. As input to the machine learning module, a 2D or 3D medical image may be provided. For example, additional context for the input medical image may be provided. Such a context may identify, for example, the task for which the image is utilized, such as measuring a particular organ, detecting a tumor or lesion, and / or the context may identify, for example, the imaging method used to acquire the imaging data, e.g., magnetic resonance imaging (MRI), computed tomography (CT), or molecular imaging (AMI) (e.g., positron emission tomography (PET) or single photon emission computed tomography (SPECT)). Such a trained machine learning module, i.e., a trained saliency map estimation module, may be applied in various ways.
[0014] A trainable machine learning module may be provided that incorporates information about the attention of the expert to each region of the observed 2D and / or 3D medical image. Using this module, a saliency map may be provided that predicts the distribution of the user's attention on the medical image, which may help to provide an improved medical image. For example, by selecting a more appropriate reconstruction method, an improved medical image that is more in line with the needs of the expert may be provided. Thus, a method for personalized selection of the reconstruction method or personalization of the machine learning module for image reconstruction may be provided.
[0015] A saliency map is considered herein as a map representing the distribution of a user's attention on a medical image. The level of the user's attention can provide a weighting factor for differently weighting different regions of the medical image, which can capture how clinically relevant the different regions of the medical image are. Differently weighted image regions can be processed differently in image reconstruction and / or image analysis.
[0016] A training saliency map for training a machine learning module to predict a saliency map for a given medical image may be obtained by capturing the actions of a user, e.g., a radiologist or a scanning technician, while observing the medical image. For example, eye-tracking technology may be used to obtain a model of gaze direction and gaze patterns to identify where the user is looking within the displayed medical image. Alternatively and / or additionally, other users' interactions with the displayed medical image may be taken into account in determining the distribution of the user's attention level on the medical image. For example, selection of a particular area of the medical image, zooming in on a particular area of the medical image, processing of a particular area of the medical image, and / or the position and movement of a cursor controlled by the user within the medical image may be taken into account. The machine training module may be trained to predict a saliency map as an output for a medical image received as an input.
[0017] Known image reconstruction models may treat all regions of a reconstructed image (particularly a medical image) in the same way. A saliency map predicting the distribution of user attention may be incorporated during training so that the loss function obtained for a particular region in a medical image reconstructed by a machine learning module is weighted according to the importance of each region to a human observer. The importance of each region to a human observer may be predicted by the saliency map. This approach may allow the development of a reconstruction module that reduces the image quality of less relevant regions in the background or less relevant regions of anatomical features in the reconstructed medical image that are less important for the current clinical task. In exchange, the image quality of more relevant regions and anatomical features in the reconstructed medical image may be improved. Thus, an improvement in the image quality of a particular region based on the information provided by the saliency map may be achieved in exchange for a reduction in the image quality of other less important regions. Such a quality exchange may be particularly useful in high-speed MRI, since the information, i.e. image data, obtained is of comparable small amount.
[0018] Thus, the distribution of user attention predicted by the saliency map may represent a distribution of interest levels: the higher the predicted user attention to a particular region of the medical image, the higher the interest level for that image region. Regions with high user attention can be considered regions of high interest and therefore high relevance.
[0019] The medical image may be a tomographic image, such as a magnetic resonance (MR) image or a computed tomography (CT) image. The medical image may be generated using advanced molecular imaging (AMI) methods, such as positron emission tomography (PET) or single photon emission computed tomography (SPECT).
[0020] The trained first machine learning module may be trained by the same medical system as the medical system that uses the saliency map output by the trained first machine learning module. Additionally or alternatively, the trained first machine learning module may be trained by another medical system that is different from the medical system that uses the trained first machine learning module to output the saliency map.
[0021] For example, execution of the machine-executable instructions causes the computing system to further provide a trained first machine learning module. Providing the trained first machine learning module includes providing a first machine learning module. The first training data includes a first pair of training medical images and a training saliency map. The training saliency map represents a distribution of a user's attention on the training medical images. The first machine learning module is trained using the first training data. The resulting trained first machine learning module is trained to output the training saliency map of the first pair in response to receiving the training medical images of the first pair.
[0022] To train the machine learning module to predict the saliency map, a dataset providing training data may be collected. The training data may include pairs of training medical images and training saliency maps representing the distribution of user attention for each training medical image. To collect such training saliency maps, a set of training medical images may be provided. The training medical images may be displayed on a display device. An eye-tracking device such as a camera may be placed in front of a user, e.g., a radiologist, working with the training medical images displayed on a display device, e.g., a computer screen. The eye-tracking device may be placed, for example, above, below, or near the display device. The training medical images may be medical images from a particular field of interest, e.g., MRI or CT. The training medical images may be clinical images used by a user, e.g., a radiologist. Recordings of eye position and / or movement from an eye-tracking device such as a camera may be used to track eye movements and match them to positions in the displayed medical images observed during recording. Tracking and matching of eye position and movement may be performed, for example, using an attention determination module. The attention determination module may be configured to use the eye tracking device to determine a distribution of user attention on the displayed training medical image and determine points of attention within the displayed training medical image for a user of the medical system viewing the displayed training medical image. The longer and / or more frequently a user looks at a particular point on the displayed training medical image, the higher the user's attention level assigned to this point may be.
[0023] For example, the medical system further includes a display device. The provision of the first training data includes, for each of the training medical images of the first training data, Displaying a training medical image using a display device; Measuring a distribution of a user's attention on the displayed training medical images; and generating a training saliency map of a first pair of training data including the displayed training medical image using the measured distribution of user attention on the training medical image.
[0024] For example, the medical system further includes an eye tracking device configured to measure eye positions and movements of a user of the medical system. The memory further stores an attention determination module configured to use the eye tracking device to determine a distribution of user attention on the displayed training medical images and to determine points of attention within the displayed training medical images for a user of the medical system viewing the displayed training medical images.
[0025] For example, from camera tracking of a user's eyes, e.g., a radiologist's eye, a training saliency map may be provided providing a distribution of training user attention for multiple different types of medical images to be used as training medical images.
[0026] For example, the trained first machine learning module is trained to output a user-specific saliency map that predicts a user-specific distribution of the user's attention on the input medical image in response to receiving a medical image as an input. In some examples, the user-specific saliency map may be provided. For example, a trained saliency map generated by determining the user attention level of only a particular user may be used to obtain a machine learning module trained for this particular user. For example, multiple machine learning modules may be provided, each trained for and assigned to a different user of the multiple users. Depending on the user using the medical system, for example, depending on the user engaged, a machine learning module of the multiple machine learning modules may be selected to predict the saliency map. For example, a machine learning module assigned to the user using the medical system may be selected.
[0027] The saliency map may be provided for use within a medical image reconstruction chain, i.e. for the selection of an image reconstruction method, for improvement of an image reconstruction method, and / or for evaluation of an image reconstruction method, in particular an image reconstruction method performed using a machine learning module to perform the image reconstruction.
[0028] For example, the medical system may be further configured to use the saliency map to select a reconstruction method for reconstructing a medical image from a plurality of predefined reconstruction methods. The medical images are test medical images of a predefined type of anatomical structure for which the reconstructed medical image is to be directed. A plurality of test maps are provided. Each test map is assigned to a different reconstruction method. Each test map identifies a section of the test image that includes an anatomical substructure of the predefined type of anatomical structure that, when using the assigned reconstruction method, results in the highest quality of image reconstruction compared to other anatomical substructures of the predefined type of anatomical structure. Execution of the machine-executable instructions further causes the computing system to: Providing a test map; Comparing the test map with the saliency map; determining, among the test maps, a test map that has the highest structural similarity to the saliency map; selecting a reconstruction method assigned to the determined test map; and Reconstructing the medical image to be reconstructed using the selected reconstruction method; and
[0029] For example, the trained machine learning module, i.e., the saliency map estimation module, may be used as a tool for selecting a reconstruction method. For example, a reconstruction method that is optimal for a particular purpose and / or a particular user working with medical images may be selected. The various reconstruction methods may refer, for example, to various methods for reconstructing a medical image from a given set of acquired medical imaging data. The various reconstruction methods may also refer, for example, to various methods for acquiring medical imaging data used to reconstruct the medical image, for example, using various sampling patterns. The various methods for acquiring medical imaging data may refer, in some cases, to using various medical imaging systems for data acquisition, such as, for example, an MRI system, a CT imaging system, a PET imaging system, or a SPECT imaging system.
[0030] The trained machine learning module can be used to select the best reconstruction method for the user. By providing a saliency map that predicts the distribution of the user's attention, the trained machine learning provides a method that takes into account a measure of relevance relevant to the user. In this environment, multiple reconstruction methods with comparable properties can be deployed. For example, a certain anatomical structure should be depicted. Comparable properties may refer to the fact that all reconstruction methods can provide reconstructed medical images that represent the respective anatomical structure. Based on the prediction of the user's attention provided by the saliency map, the reconstruction method that is most suitable for the user can be selected. The saliency map can, for example, predict the distribution of the user's attention for an individual user. Thus, different saliency maps can be predicted for different users depending on the user's experience, references, and / or way of working. Different radiologists may view images in different ways. Thus, each radiologist may benefit from a reconstruction that best suits his or her needs.
[0031] Each reconstruction method may be designed to handle the specific characteristics of the input data, i.e., the specific characteristics of the acquired medical data used to reconstruct the medical image. For example, in MRI, some reconstruction methods produce medical images with high contrast between white and gray matter in the brain, while other reconstruction methods may provide better signal-to-noise ratios in areas closer to the skull. For each reconstruction method, a test map may be provided that identifies the section of the test image where the reconstruction method has the highest image reconstruction quality. For example, if a reconstruction method provides high image quality for white matter in the brain, a test map may be provided that highlights the white matter contained in the test image. For example, if a reconstruction method provides high image quality for gray matter in the brain, a test map may be provided that highlights the gray matter contained in the test image. For example, if a reconstruction method provides high image quality for cerebrospinal fluid in the brain, a test map may be provided that highlights the cerebrospinal fluid contained in the test image. The test map may have the appearance of a saliency map. The test map may be compared to a saliency map provided, for example, by machine learning configured to predict the distribution of attention of a particular radiologist on the test image. The most suitable reconstruction method is selected, i.e. the reconstruction method whose test map shows the highest similarity with the saliency map. The similarity may be determined by estimating the distance between the saliency map acquired for the radiologist and multiple test maps provided for different reconstruction methods.
[0032] A saliency map may be provided that predicts the distribution of user attention (e.g., for a particular user) on a test image. The test image may represent a particular type of image to be reconstructed. Information on the predicted distribution of user attention provided by the saliency map may be used to select the most appropriate of the available reconstruction models. Various reconstruction models may act differently on different regions of the medical image to be reconstructed. A saliency map that predicts user attention may help to select a model that better reconstructs regions where user attention tends to be concentrated, i.e., regions that are more valuable to a particular expert in a given context, than other regions of the medical image to be reconstructed.
[0033] To train a machine learning model to predict user attention, for example, a camera can be placed in front of an expert to track eye movements and / or position and identify areas in a template medical image that are attended to by the expert.
[0034] For example, the memory further stores an out-of-distribution estimation module configured to output an out-of-distribution map in response to receiving a medical image as an input. The out-of-distribution map represents a level of compliance of the input medical image with respect to a reference distribution defined by a set of reference medical images. Execution of the machine-executable instructions further causes the computing system to provide the medical image as an input to the out-of-distribution estimation module. In response to providing the medical image, execution of the machine-executable instructions further causes the computing system to receive an out-of-distribution map of the medical image as an output from the out-of-distribution estimation module. The out-of-distribution map represents a level of compliance of the medical image with respect to a predetermined distribution. Execution of the machine-executable instructions further causes the computing system to provide a weighted out-of-distribution map. Providing the weighted out-of-distribution map includes weighting the level of compliance represented by the out-of-distribution map using a distribution of a user's attention on the medical image predicted by the saliency map.
[0035] To more reliably detect regions of a medical image that are highly likely to contain reconstruction artifacts, a saliency map predicting a user attention distribution on the medical image may be combined with an uncertainty estimation map, such as an out-of-distribution (OOD) map. The saliency map may allow weighting of regions of the OOD map. The weighted OOD map may be less accurate because it gives low weights or zero to regions where false or erased anatomical structures may occur. However, the weighted OOD map allows to focus only on regions that are important to the user, so this information may be more valuable to the end user. According to an example, a set of reference medical images may also be provided to the out-of-distribution estimation module to calculate the ODD map. In the case of a trained out-of-distribution estimation module, the training medical images and the assigned training OOD map may be used to train the out-of-distribution estimation module to provide an OOD map.
[0036] For example, when a machine learning module is used to reconstruct a medical image from medical imaging data, if the medical imaging data is too similar to the training medical imaging data used to train the machine learning module, there may be no guarantee that an accurate result, i.e., an accurately reconstructed medical image, is provided. Thus, if the data input to the trained machine learning module is outside the training data distribution, the reconstructed medical image generated using the trained machine learning module may be inaccurate. The resulting reconstructed medical image may look like a correct medical image, but it is incorrect. A reconstructed medical image that is too similar to a set of reference medical images (e.g., the training medical images used to train the machine learning module that provides the reconstructed medical image) may be considered to be "out of distribution" with respect to the reference distribution defined by the set of reference medical images. The similarity, i.e., the level of conformance, may be determined, for example, on a pixel or voxel basis.
[0037] The out-of-distribution estimation module may be configured or trained to output an out-of-distribution map. As used herein, the out-of-distribution estimation module encompasses a software module that may be used to detect, for example, whether a reconstructed medical image is within the distribution of training medical images. The compliance level may represent the probability that the reconstructed medical image or a region of the reconstructed medical image is within the distribution of training medical images.
[0038] The out-of-distribution estimation module provided in the form of a trained machine learning module may include, for example, an out-of-distribution estimation neural network or a set of neural networks. The out-of-distribution estimation neural network is a neural network, e.g. a classifier network, configured to receive a medical image and provide as output a classification map of this medical image in the form of an out-of-distribution map. The out-of-distribution map represents the compliance level of the input medical image with respect to a reference distribution defined by a set of reference medical images. The set of reference medical images may, for example, be a set of training medical images used to train a further machine learning module to reconstruct medical images. The medical image for which the out-of-distribution map was generated may be reconstructed by this further machine learning module. The out-of-distribution map may indicate the distribution on the medical image of the probability that sections of the medical image are within a reference distribution defined by a set of reference medical images, e.g. a set of training medical images, for those sections.
[0039] Out-of-distribution (OOD) estimation methods can be used to estimate the uncertainty, and thus the reliability, of image reconstruction results, i.e., the reconstructed image. OOD estimation methods can be useful for identifying regions of the reconstructed image that are highly likely to contain reconstruction artifacts. In such methods, it is usually assumed that all regions of the reconstructed image have equal relevance. However, in practice, especially in clinical practice, this may not generally be the case. Anatomical structures shown by medical images may include anatomical substructures and / or features that are essential for diagnosis. However, at the same time, such anatomical structures may include anatomical substructures and / or features that are less important or negligible for diagnosis.
[0040] Incorrect predictions can occur because the medical image contains a mixture of relevant and non-relevant regions. For example, a false negative prediction can occur if the more relevant regions, i.e., regions that are more important to the expert, have a low OOD score compared to other less relevant regions, but the uncertainty indicated by the OOD score is still high enough to make a mistake. A false positive prediction can occur if a high OOD score, indicating high uncertainty, indicates that there may be problems in regions that are completely irrelevant to the expert performing the examination. Both cases can be avoided by building a significant and reliable saliency map that is used to weight the OOD results. The weighted uncertainty map may be presented to the user, for example as a warning.
[0041] For example, a saliency map predicting a user attention distribution on a medical image may be combined with an uncertainty map, e.g., an OOD map, providing a measure of the uncertainty of the reconstruction of the reconstructed medical image. For example, the uncertainty of an image reconstructed using a reconstruction neural network may not be equally important in each region of the image observed. For example, a medical image with low uncertainty in highly relevant regions and high uncertainty in less relevant or irrelevant regions may be more reliable for a user than a medical image with medium uncertainty in highly relevant regions and low uncertainty in less relevant or irrelevant regions.
[0042] Using the saliency map to weight the uncertainty levels of different regions may provide a more reliable estimation of the uncertainty levels. Such estimation may take into account different relevance levels of different regions of the medical image based on a prediction of the user's attention provided by the saliency map. If it is predicted that the user will pay more attention to a particular region, the uncertainty level of this region may be weighted higher than the corresponding level of a region that attracts less attention. In one example, a weighted uncertainty map, e.g., an OOD map, may be displayed for the user on a display device of the medical system. Such a weighted uncertainty map may show a distribution of uncertainty levels on the reconstructed medical image with the uncertainty levels weighted based on the relevance of the regions to which a reactivity level is assigned. In a further example, the weighted uncertainty map may be reduced to a scalar value, e.g., by averaging. This scalar value may be used to evaluate the overall reliability of the reconstructed image. Based on this evaluation, a decision is made, e.g., to reacquire the imaging data used to reconstruct the medical image or to directly warn the user of potential problems, e.g., in the provided image. For example, if the scalar value exceeds a predetermined threshold, a signal may be issued recommending reacquisition of imaging data and / or reacquisition of imaging data may be initiated. For example, if the scalar value exceeds a predetermined threshold, a signal may be issued to alert a user that the reliability of the reconstructed image may be insufficient.
[0043] For example, providing the weighted out-of-distribution map further includes calculating an out-of-distribution score using the weighted compliance levels provided by the weighted out-of-distribution map, the out-of-distribution score representing a probability that the medical image as a whole is within the reference distribution.
[0044] The aggregated OOD score of the weighted OOD map may be presented to a user, for example for warning purposes. Additionally or alternatively, the aggregated OOD score may be presented to a user or used by another automated system, for example to decide whether to discard / rescan the object. The aggregated OOD score may be calculated, for example, by averaging the compliance levels, i.e., local OOD scores, contained in the weighted OOD map.
[0045] For example, the memory further stores an image quality assessment module configured to output an image quality map in response to receiving the medical image and the saliency map as input. The image quality map represents a distribution of image quality levels of the input medical image weighted using a distribution of user attention on the input medical image predicted by the input saliency map. Execution of the machine-executable instructions further causes the computing system to provide the medical image and the saliency map as input to the image quality assessment module. In response to providing the medical image and the saliency map, execution of the machine-executable instructions further causes the computing system to receive an image quality map as output from the image quality assessment module. The image quality map represents a distribution of image quality levels of the medical image weighted using a distribution of user attention on the medical image predicted by the saliency map. Execution of the machine-executable instructions further causes the computing system to provide the received image quality map.
[0046] A saliency map that predicts the distribution of a user's attention on a medical image can be used to enhance image quality assessment metrics that can be applied, for example, during training and / or evaluation of a medical image reconstruction module (e.g., a deep learning-based reconstruction module). The saliency map can significantly enhance the informational value of currently used image quality assessment metrics, such as mean square error (MSE), peak signal-to-noise ratio (PSNR), or Structural Similarity Index Measure (SSIM). Such improvements can potentially enable training and deployment of higher quality reconstruction modules to reconstruct medical images.
[0047] Using the personalized user-specific saliency maps, personalized reconstruction modules can be trained and deployed to reconstruct medical images that are optimized for individual users.
[0048] In addition to the medical image and the saliency map, an expected medical image may be provided to the image quality assessment module. The image quality level may provide a measure of how well the medical image matches the expected medical image. The image quality level may quantify the similarity between the assessed medical image and the expected medical image. The higher the similarity, the higher the image quality. If the image quality assessment module is used to train a medical image reconstruction module, the reference medical image may be, for example, a training medical image that the medical reconstruction module should be trained to predict.
[0049] Researchers spend a lot of time and resources building new state-of-the-art algorithms for medical image reconstruction. The process of designing such algorithms requires that each algorithm is evaluated multiple times. For example, when training a machine learning module to reconstruct medical images, it may be necessary to evaluate the quality of the medical images reconstructed by the machine learning module during training. To evaluate the predictions of the machine learning module, a loss function is used. Such a loss function is a measure of how well the machine learning module predicts the expected result, i.e., the ground truth. In the case of image reconstruction, the ground truth may be provided, for example, in the form of training images to be reconstructed. The loss function compares the actual output of the machine learning module, e.g., the reconstructed medical image, with the expected or target output, e.g., the target medical image to be reconstructed. The result of the loss function is called the loss and is a measure of how well the actual output of the machine learning module matches the expected output. A high value of the loss indicates poor performance of the machine learning module, while a low value indicates good performance.
[0050] Loss functions for evaluating image quality may use image quality (IQ) evaluation metrics, such as mean square error (MSE), peak signal-to-noise ratio (PSNR), Structural Similarity Index Measure (SSIM), Blind / Referenceless Image Spatial Quality Evaluator (BRISQUE), Gradient Magnitude Similarity Deviation (GMSD), or Feature Similarity Index Measure (FSIM). However, none of these metrics are correlated with concepts of quality that may be specific to a task, modality, and / or individual. Therefore, to evaluate image quality more reliably, application experts should be involved. Using a saliency map that helps weight the IQ metrics according to the perception of importance of experts with deep knowledge could have the advantage of significantly improving the quality of the IQ metrics and simplifying the experimentation process. For example, no experts are needed to evaluate image quality. Thus, the human-in-the-loop and the resulting complexity of the experimentation process can be avoided.
[0051] For example, a saliency map predicting user attention can be used to improve the IQ evaluation metric used to train and validate machine learning modules used for image reconstruction. Different regions of a reconstructed medical image may play very different roles and therefore be of very different relevance to the user. The saliency map can be used to improve the performance of the IQ metric by non-uniformly weighting the reconstruction errors according to the relevance of the region of the reconstructed image in which each reconstruction error occurs. Thus, the network being trained may be penalized more heavily for errors occurring in more important regions of the image being reconstructed and less heavily for errors occurring in less important regions.
[0052] For example, the image quality assessment module is used to train a second machine learning module to output a medical image as an output in response to receiving medical imaging data as an input. The image quality estimated by the image quality assessment module may represent a loss of the output medical image of the second machine learning module relative to one or more reference medical images. Upon execution of the machine executable instructions, the computing system further provides a second machine learning module. Upon execution of the machine executable instructions, the computing system further provides second training data for training the second machine learning module. The second training data includes a second pair of training medical imaging data and a training medical image reconstructed using the training medical imaging data.
[0053] Execution of the machine-executable instructions further causes the computing system to train a second machine learning module. The second machine learning module is trained to output training medical images of the second pairs in response to receiving training medical imaging data of the second pairs. The training includes, for each second pair, providing the training medical imaging data as input to the second machine learning module and receiving a preliminary medical image as output. The received preliminary medical image is a medical image.
[0054] The distribution of image quality represented by the image quality map received for the medical image from the image quality assessment module is used as a distribution of loss of the medical image for a training medical image of the second pair provided as a reference medical image to the image quality assessment module to determine the received image quality map. The parameters of the second machine learning module are adjusted during training until the loss of the medical image meets a predetermined criterion.
[0055] To improve the performance of the index as an evaluation method for training machine learning modules for medical image reconstruction, a method for modifying the image quality assessment index may be provided by incorporating a saliency map that predicts the distribution of user attention. This criterion may, for example, require that the loss is less than a certain threshold.
[0056] The medical system may be configured for image reconstruction. For this purpose, the medical system may provide a second machine learning module. The second machine learning module may be trained to reconstruct medical images, i.e., trained as an image reconstruction module. The medical images may be tomographic images, such as magnetic resonance images or computed tomography images. The medical images may be generated using advanced molecular imaging methods, such as positron emission tomography or single photon emission tomography.
[0057] Based on the saliency map, the image reconstruction can be adapted. For example, when using a machine learning module to reconstruct a medical image, the weights can be adapted by the saliency map. The machine learning module for reconstructing a medical image can include a neural network, and during training, the weights in the neural network can be adapted by the saliency map. For example, the saliency map with a prediction of the user attention distribution can be used to adapt the image reconstruction, for example, when employing a deep learning reconstruction technique.
[0058] The saliency map predicts the distribution of the user's attention on the medical image. The distribution of the user's attention may represent the distribution of interest levels on the medical image. Based on the saliency map, the parameters of the image reconstruction module may be adapted. Thus, the saliency map may be used to operate the image reconstruction module. For example, the image reconstruction module may include a neural network trained using deep learning reconstruction techniques to adapt weights in the neural network by the saliency map. The saliency itself may also be generated using deep learning, for example, using eye tracking to determine the user's attention on medical images of different types of anatomical structures. Weighting the image reconstruction using the distribution of the user's attention predicted by the saliency map may have the beneficial effect of directing the reconstruction effort more accurately to clinically interesting parts or aspects of the reconstructed medical image. The clinical interest may be determined by the user attention of the user of the medical system.
[0059] For example, providing the received image quality map may include calculating an image quality score using the received image quality map, the image quality score representing an average image quality of the medical image.
[0060] For example, the medical system is configured to acquire medical imaging data for reconstructing a medical image, the medical imaging data being collected using any of a number of data acquisition methods, such as magnetic resonance imaging, computed tomography, positron emission tomography, single photon emission computed tomography, etc.
[0061] In another aspect, the present invention provides a medical system including a memory storing machine executable instructions and a computing system. Execution of the machine executable instructions causes the computing system to provide a trained machine learning module to output a saliency map as an output in response to receiving a medical image as an input. The saliency map predicts a distribution of a user's attention on the medical image. Providing the trained machine learning module includes providing the machine learning module. Further, training data is provided, the training data including pairs of training medical images and training saliency maps. The training saliency maps represent a distribution of the user's attention on the training medical images. The machine learning module is trained using the training data. The resulting trained machine learning module is trained to output the training saliency map of the pair in response to receiving the training medical images of the pair.
[0062] The medical system may provide the trained machine learning module for outputting a saliency map as an output in response to receiving a medical image as an input. The saliency map predicts the distribution of a user's attention on the medical image. The medical system may use the trained machine learning module on the medical system alone. Additionally or alternatively, the medical system may provide the trained machine learning module to another medical system for use. For example, the medical system may transmit the trained machine learning module to the other medical system.
[0063] Using the trained machine learning module may include providing a medical image as an input to the trained machine learning module. In response to providing the medical image, a saliency map of the medical image is received as an output from the trained machine learning module. The saliency map predicts a distribution of a user's attention on the medical image. The received saliency map of the medical image may be provided for further use.
[0064] In another aspect, the invention provides a computer program comprising machine executable instructions executed by a computing system for controlling a medical system. The computer program further comprises a machine learning module trained to, in response to receiving a medical image as an input, output a saliency map. The saliency map predicts a distribution of a user's attention on the medical image. Execution of the machine executable instructions causes the computing system to receive a medical image. The medical image is provided as an input to the trained machine learning module. In response to providing the medical image, a saliency map of the medical image is received as an output from the trained machine learning module. The saliency map predicts a distribution of a user's attention on the medical image. The saliency map of the medical image is provided.
[0065] In another aspect, the present invention provides a computer program comprising machine executable instructions executed by a computing system for controlling a medical system. Execution of the machine executable instructions causes the computing system to provide a trained machine learning module to output a saliency map as an output in response to receiving a medical image as an input. The saliency map predicts a distribution of a user's attention on the medical image. Providing the trained machine learning module comprises providing the machine learning module. Further, training data is provided comprising pairs of training medical images and training saliency maps. The training saliency maps represent a distribution of a user's attention on the training medical images. The machine learning module is trained using the training data. The resulting trained machine learning module is trained to output the training saliency map of a first pair in response to receiving a training medical image of the first pair.
[0066] The trained machine learning module may be used by the same medical system that trained the trained machine learning module. Additionally or alternatively, the trained machine learning module may be provided for use by another medical system. For example, the trained machine learning module may be transmitted to the other medical system.
[0067] Using the trained machine learning module may include providing a medical image as an input to the trained machine learning module. In response to providing the medical image, a saliency map of the medical image is received as an output from the trained machine learning module. The saliency map predicts a distribution of a user's attention on the medical image. The received saliency map of the medical image may be provided for further use.
[0068] In another aspect of the invention, there is provided a medical imaging method using a trained machine learning module to output a saliency map as an output in response to receiving a medical image as an input. The saliency map predicts a distribution of a user's attention on the medical image. The method includes receiving a medical image. The medical image is provided as an input to the trained machine learning module. In response to providing the medical image, a saliency map of the medical image is received as an output from the trained machine learning module. The saliency map predicts a distribution of the user's attention on the medical image. A saliency map of the medical image is provided.
[0069] In another aspect of the invention, a method is provided for providing a trained machine learning module to output a saliency map as an output in response to receiving a medical image as an input. The saliency map predicts a distribution of a user's attention on the medical image. Providing the trained machine learning module includes providing a machine learning module. Further, training data is provided that includes pairs of training medical images and a training saliency map. The training saliency map represents a distribution of the user's attention on the training medical images. The machine learning module is trained using the training data. The resulting trained machine learning module is trained to output the training saliency map of a first pair in response to receiving a training medical image of the first pair.
[0070] The trained machine learning module may be used by the same medical system that trained the trained machine learning module. Additionally or alternatively, the trained machine learning module may be provided for use by another medical system. For example, the trained machine learning module may be transmitted to the other medical system.
[0071] Using the trained machine learning module may include providing a medical image as an input to the trained machine learning module. In response to providing the medical image, a saliency map of the medical image is received as an output from the trained machine learning module. The saliency map predicts a distribution of a user's attention on the medical image. The received saliency map of the medical image may be provided for further use.
[0072] One or more of the above embodiments may be combined, unless the embodiments being combined are contradictory.
[0073] As will be appreciated by those skilled in the art, aspects of the present invention may be embodied as an apparatus, a method, or a computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects (all of which may be referred to generally herein as "circuits," "modules," or "systems"). Additionally, aspects of the present invention may take the form of a computer program product embodied by one or more computer-readable medium(s) having computer-executable code embodied thereon.
[0074] Any combination of one or more computer readable media may be utilized. The computer readable media may be computer readable signal media or computer readable storage media. As used herein, "computer readable storage media" may encompass any tangible storage media capable of storing instructions executable by a processor or computing system of a computing device. The computer readable storage media may also be referred to as computer readable non-transitory storage media. The computer readable storage media may also be referred to as tangible computer readable media. In some embodiments, the computer readable storage media may be capable of storing data accessible by the computing system of a computing device. Examples of computer readable storage media include, but are not limited to, floppy disks, magnetic hard disk drives, solid state hard disks, flash memory, USB thumb drives, Random Access Memory (RAM), Read Only Memory (ROM), optical disks, magneto-optical disks, and register files of a computing system. Examples of optical disks include Compact Disks (CDs) and Digital Versatile Disks (DVDs), such as CD-ROM, CD-RW, CD-R, DVD-ROM, DVD-RW, or DVD-R disks, etc. The term computer-readable storage medium also refers to various types of recording media that a computing device can access via a network or communication link. For example, data may be retrieved via a modem, the Internet, or a local area network. Computer executable code embodied in a computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic cable, RF, etc., or any suitable combination thereof.
[0075] A computer-readable signal medium may include a propagated data signal that includes computer-executable code (e.g., in baseband or as part of a carrier wave). Such a propagated signal may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium is not a computer-readable storage medium, but may be any computer-readable medium capable of communicating, propagating, or transmitting a program for use by or in connection with an instruction execution system, apparatus, or device.
[0076] "Computer memory" or "memory" is one example of a computer-readable storage medium. Computer memory is any memory directly accessible by a computing system. "Computer storage" or "storage" is another example of a computer-readable storage medium. Computer storage is any non-volatile computer-readable storage medium. In some implementations, computer storage may also be computer memory and vice versa.
[0077] As used herein, a "computing system" encompasses electronic components capable of executing programs, machine-executable instructions, or computer-executable code. References to a computing system, including examples of "computing system," should be interpreted as potentially including multiple computing systems or processing cores. A computing system may be, for example, a multi-core processor. A computing system may also refer to a collection of multiple processors aggregated in a single computing system or distributed among multiple computing systems. The term computing system should also be interpreted as meaning a collection or network of multiple computing devices, each of which includes one or more processors or computing systems. Machine-executable code or instructions may be executed by multiple computing systems or processors aggregated in the same computing device or distributed across multiple computing devices.
[0078] Machine-executable instructions or computer-executable code may include instructions or programs that cause a processor or other computing system to perform aspects of the invention. Computer-executable code for carrying out operations of aspects of the invention may be written in any combination of one or more of object-oriented programming languages such as Java, Smalltalk, C++, and traditional procedural programming languages such as C, or similar programming languages, and compiled into machine-executable instructions. In some cases, the computer-executable code may be in the form of a high-level language or in pre-compiled form, and may be used with an interpreter that generates machine-executable instructions on the fly. In other examples, the machine-executable instructions or computer-executable code may be in the form of programming for a programmable logic gate array.
[0079] The computer executable code may run entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or a connection may be established to an external computer (e.g., via the Internet using an Internet Service Provider).
[0080] Aspects of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It will be understood that each block or group of blocks of the flowcharts, illustrations, and / or block diagrams may be implemented by computer program instructions in the form of computer executable code, where applicable. It will also be understood that blocks in different flowcharts, illustrations, and / or block diagrams may be combined, where not mutually inconsistent. These computer program instructions may be provided to a computing system of a general purpose computer, special purpose computer, or other programmable data processing device to form a machine. The instructions executed via the computing system of the computer or other programmable data processing device create means for implementing the functions / acts identified in one or more blocks of the flowcharts and / or block diagrams.
[0081] These machine-executable instructions or computer program instructions may be stored on a computer-readable medium that can cause a computer, other programmable data processing apparatus, or other device to function in a particular manner. The instructions stored on the computer-readable medium create an article of manufacture that includes instructions that implement the functions / acts identified in one or more blocks of the flowcharts and / or block diagrams.
[0082] Machine-executable instructions or computer program instructions may be loaded onto a computer, other programmable data processing device, or other device, and a computer-implemented process may be generated by causing the computer, other programmable device, or other device to execute a series of operational steps. The instructions executed on the computer or other programmable device provide a process for performing the functions / operations identified in one or more blocks of the flowcharts and / or block diagrams. As used herein, a "user interface" is an interface that allows a user or operator to interact with a computer or computer system. A "user interface" may also be referred to as a "human interface device." A user interface may provide information or data to an operator and / or receive information or data from an operator. A user interface may allow a computer to receive input from an operator and may provide output from the computer to a user. In other words, a user interface may allow an operator to control or operate a computer, and an interface may allow a computer to display the effects of the operator's control or operation. Displaying data or information on a display or graphical user interface is an example of providing information to an operator. A keyboard, a mouse, a trackball, a touchpad, a pointing stick, a graphics tablet, a joystick, a gamepad, a webcam, a headset, pedals, wired gloves, a remote control, and receiving data via an accelerometer are all examples of user interface elements that allow for the receipt of information or data from an operator.
[0083] As used herein, a "hardware interface" encompasses an interface that allows a computing system of a computer system to interact with and / or control external computing devices and / or devices. A hardware interface may allow a computing system to send control signals or instructions to external computing devices and / or devices. A hardware interface may also allow a computing system to exchange data with external computing devices and / or devices. Examples of hardware interfaces include, but are not limited to, a universal serial bus, an IEEE 1394 port, a parallel port, an IEEE 1284 port, a serial port, an RS-232 port, an IEEE-488 port, a Bluetooth connection, a wireless local area network connection, a TCP / IP connection, an Ethernet connection, a control voltage interface, a MIDI interface, an analog input interface, and a digital input interface.
[0084] As used herein, a "display" or "display device" encompasses an output device or user interface adapted to display images or data. A display may output visual, auditory, or tactile data. Examples of displays include, but are not limited to, computer monitors, television screens, touch screens, tactile electronic displays, Braille screens, cathode ray tubes (CRTs), storage tubes, bi-stable displays, electronic paper, vector displays, flat panel displays, fluorescent display tubes (VFs), light emitting diode (LED) displays, electroluminescent displays (ELDs), plasma display panels (PDPs), liquid crystal displays (LCDs), organic light emitting diode displays (OLEDs), projectors, and head mounted displays.
[0085] The term "machine learning" (ML) refers to computer algorithms used to extract useful information from training data by automatically constructing probabilistic frameworks called machine learning modules. Machine learning can be performed using one or more learning algorithms such as linear regression, K-means, classification algorithms, reinforcement algorithms, etc. A "machine learning module" can be, for example, a set of equations or rules that allow predicting an unmeasured value from other known values.
[0086] As used herein, a "neural network" encompasses a computational system configured to learn (i.e., gradually improve its ability) to perform a task by studying examples, generally without the need for task-specific programming. A neural network includes a number of units called neurons, which are communicatively connected by connections for transmitting signals between the connected neurons. The connections between the neurons are called synapses. A neuron receives a signal as an input and changes its internal state, or activation function, in response to the input. In response to the input, learned weights and biases, an activation function is generated as an output and sent through one or more synapses to one or more connected neurons. The network forms a weighted directed graph, where the neurons are the nodes and the connections between the neurons are weighted directed edges. The weights and biases can be modified by a process called learning, which is governed by learning rules. A learning rule is an algorithm that modifies the parameters of the neural network so that a given input to the network produces a preferred output. As a result of this learning process, the weights and biases of the network can be modified.
[0087] Neurons may be organized in layers. Different layers may perform different types of transformations on the input. A signal applied to a neuronal network passes from the first layer, the input layer, to the last layer, the output layer, through intermediate (hidden) layers located between the input and output layers.
[0088] As used herein, "network parameters" include neuronal weights and biases that can change as learning progresses and can increase or decrease the strength of signals sent downstream by neurons through synapses.
[0089] As used herein, a "medical system" encompasses any system including a memory storing machine-executable instructions and a computing system configured to execute the machine-executable instructions, where execution of the machine-executable instructions causes the computing system to process medical images. The medical system may be configured to use a trained machine learning module to process the medical images and / or to train the machine learning module to process the medical images. The medical system may be further configured to generate medical images using medical imaging data and / or to acquire medical imaging data to generate medical images.
[0090] Medical imaging data is defined herein as a record of measurements made by a tomographic medical imaging system representing a subject. The medical imaging data may be reconstructed into a medical image. A medical image is defined herein as a reconstructed two- or three-dimensional visualization of the anatomical data contained within the medical imaging data. This visualization may be performed using a computer.
[0091] As used herein, a magnetic resonance imaging (MRI) image or MR image is defined as a reconstructed two- or three-dimensional visualization of anatomical data contained within magnetic resonance imaging data, which visualization may be performed using a computer.
[0092] A computed tomography (CT) image is defined herein as a two- or three-dimensional cross-sectional image reconstructed by a computer using data from X-ray measurements taken from different angles.
[0093] A positron emission tomography (PET) image is defined herein as a two- or three-dimensional visualization of the reconstructed distribution of a radiopharmaceutical injected into the body as a radiotracer. The gamma radiation caused by the radiotracer can be detected, for example, using a gamma camera. This imaging data can be used to reconstruct a two- or three-dimensional image. A PET scanner may, for example, be integrated into a CT scanner. A PET image may, for example, be reconstructed using a CT scan performed using one scanner during the same session.
[0094] Single photon emission computed tomography (SPECT) is defined herein as a two- or three-dimensional visualization of the reconstructed distribution of a gamma-ray-emitting radiopharmaceutical injected into the body as a radiotracer. SPECT imaging is performed by acquiring imaging data from multiple angles using a gamma camera. A computer is used to reconstruct a two- or three-dimensional tomographic image using the acquired imaging data. SPECT is similar to PET in its use of radiotracers and detection of gamma rays. In contrast to PET, the radiotracers used in SPECT emit gamma rays, which are measured directly. On the other hand, PET radiotracers emit positrons, which provide gamma rays when the positrons annihilate with the electrons. [Brief description of the drawings]
[0095] Preferred embodiments of the present invention will now be described, by way of example only, with reference to the following drawings, in which:
[0096] [Figure 1] FIG. 1 shows an example of a medical system. [Diagram 2] FIG. 2 shows a flow chart illustrating an exemplary method for predicting a saliency map. [Diagram 3] FIG. 2 shows a flow chart illustrating an exemplary method for predicting a saliency map. [Figure 4] FIG. 4 shows a further example of a medical system. [Diagram 5] FIG. 5 shows a further example of a medical system. [Figure 6] FIG. 6 shows a flow chart illustrating an exemplary method for predicting a saliency map. [Figure 7] FIG. 7 shows a flowchart illustrating an example of a method for training a machine learning module to predict a saliency map. [Figure 8] FIG. 9 illustrates an example of a medical system configured to train a machine learning module. [Figure 9] FIG. 9 illustrates an example of a medical system configured to train a machine learning module. [Figure 10] FIG. 10 shows an example of a machine learning module having a neural network trained to predict a saliency map. [Figure 11] FIG. 11 shows a flow chart illustrating an example of a method for generating a training saliency map. [Figure 12] FIG. 12 shows a flow chart illustrating an example of a method for generating a training saliency map. [Figure 13] FIG. 13 shows a flow chart illustrating an example of a method for selecting a reconstruction method using a saliency map. [Figure 14] FIG. 14 shows an example of how the saliency map can be used to select a reconstruction method. [Figure 15] FIG. 15 shows a flow chart illustrating an example of a method for using a saliency map to provide a weighted OOD map. [Figure 16] FIG. 16 shows a flow chart illustrating an example of a method for using a saliency map to provide a weighted OOD score. [Figure 17] FIG. 17 shows a flow chart illustrating an example of a method for using a saliency map to provide a weighted OOD map. [Figure 18] FIG. 18 shows a flow chart illustrating an example of a method for using a saliency map to provide an image quality map. [Figure 19] FIG. 19 shows a flow chart illustrating an example of a method for using a saliency map to provide an image quality score. [Figure 20] FIG. 20 shows a flow chart illustrating an example of a method for using a saliency map to provide an image quality map. [Figure 21] FIG. 21 shows a flowchart illustrating an example of a method for training a machine learning module to reconstruct medical images using saliency maps. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0097] Elements with like numbers in the figures are equivalent elements or perform the same function. An element described previously is not necessarily described in a later figure if the function is equivalent.
[0098] FIG. 1 illustrates an example of a medical system 100. The medical system is illustrated as including a computer 102. The computer 102 is intended to represent one or more computing devices. The computer 102 may be, for example, a computer that receives a medical image 124 to predict a saliency map 126 and / or receives acquired medical imaging data 123 to reconstruct the medical image 124. The computer 102 may be incorporated into a magnetic resonance imaging system, for example, as part of a control system of the magnetic resonance imaging system configured to acquire the medical imaging data. The medical imaging system may be, for example, a magnetic resonance imaging system, a computed tomography system, or an advanced molecular imaging system, for example, a positron emission tomography system or a single photon emission tomography system. In another example, the computer 102 may be a remote computer system used to reconstruct images remotely. For example, the computer 102 may be a server in a radiology department or a virtual computer system in a cloud computing system.
[0099] Computer 102 is further shown as including a computing system 104. Computing system 104 is intended to represent one or more processors or processing cores, or other computing systems, located in one or more locations. Computing system 104 is shown as being connected to an optional hardware interface 106. Optional hardware interface 106 may, for example, enable computing system 104 to control other components, such as a magnetic resonance imaging system, a computed tomography system, a positron emission tomography system, or a single photon emission computed tomography system.
[0100] The computing system 104 is further shown as being connected to an optional user interface 108, which may, for example, allow an operator to control and operate the medical system 100. The optional user interface 108 may include, for example, output and / or input devices that allow a user to interact with the medical system. The output device may include, for example, a display device configured to display medical images. The input device may include, for example, a keyboard and / or a mouse that allows a user to insert control commands to control the medical system 100. The optional user interface 108 may include, for example, an eye tracking device, such as a camera configured to track the eye position and / or movement of a user using the medical system 100. The computing system 104 is further shown as being connected to a memory 110. The memory 110 is intended to represent various types of memory that may be connected to the computing system 104.
[0101] The memory is shown as including machine-executable instructions 120. The machine-executable instructions 120 enable the computing system 104 to perform tasks, such as controlling other components and performing various data and image processing tasks. The machine-executable instructions 120 may, for example, enable the computing system 104 to control other components, such as a magnetic resonance imaging system, a computed tomography system, a positron emission tomography system, or a single photon emission computed tomography system.
[0102] The memory 110 is further shown as including a trained machine learning module 122. The trained machine learning module 122 is configured to receive a medical image 124 and, in response, provide a saliency map 126. The saliency map 126 predicts a distribution of a user's attention on the medical image 124. The medical image 124 may be, for example, an MRI image, a CT image, or an AMI image, such as a PET image or a SPECT image.
[0103] The memory 110 is further shown as including medical imaging data 123. The medical system 100 may be configured, for example, to reconstruct a medical image 124 using the medical imaging data 123. The medical imaging data 123 may be, for example, MRI data, CT image data, or AMI data such as PET image data or SPECT image data. For example, the memory 110 may further include a machine learning module 121 configured to receive the medical imaging data 123 and provide a reconstructed medical image 124 in response.
[0104] The memory 110 is further shown as including a set of test maps 128. The medical system 100 may be configured to perform different reconstruction methods for reconstructing a medical image, such as the medical image 124, using, for example, the medical imaging data 123. Each test map 128 may be assigned to a different one of the reconstruction methods and identify a section of the test image that has the highest quality of image reconstruction using the assigned reconstruction method. The test image, for example the medical image 124, may display a predetermined type of anatomical structure that is the subject of the medical image to be reconstructed using one of the reconstruction methods. If a medical image of the brain is to be reconstructed, the test image may show the structure of the brain. The section identified by the test map may include an anatomical substructure of the predetermined type of anatomical structure that has the highest quality of image reconstruction using the assigned reconstruction method, compared to other anatomical substructures of the predetermined type of anatomical structure. For example, if one of the reconstruction methods provides high image quality with respect to white matter in the brain, the test map assigned to this reconstruction method may indicate, for example, highlighting, the white matter included in the test image. For example, if one of the reconstruction methods provides high image quality with respect to gray matter in the brain, the test map assigned to the reconstruction method may point to, e.g., highlight, the gray matter contained in the test image. For example, if one of the reconstruction methods provides high image quality with respect to cerebrospinal fluid in the brain, the test map assigned to the reconstruction method may point to, e.g., highlight, the cerebrospinal fluid contained in the test image. For example, the medical image 124 may be a test image, and the saliency map 126 may be used to select one of the reconstruction methods for reconstructing one or more medical images. The saliency map 126 may be used to determine which of the test maps 128 has the highest structural similarity to the saliency map 126. The reconstruction method assigned to the determined test map may be selected, and one or more medical images may be reconstructed using the selected reconstruction method. For example, the medical imaging data 123 may be used to reconstruct a medical image.
[0105] The memory 110 is further shown as including an OOD estimation module 130. The OOD estimation module 130 is configured to receive the medical image 124 and, in response, provide an OOD map 132 representing a compliance level of the input medical image 124 with respect to a reference distribution defined by a set of reference medical images. The memory 110 may further include a weighted OOD map 134 generated by weighting the compliance level represented by the OOD map 132 using a distribution of the user's attention on the medical image 124 predicted by the saliency map 126. The weighted compliance level provided by the weighted OOD map 134 may be used, for example, to calculate an OOD score representing the probability that the medical image 124 as a whole is within the reference distribution.
[0106] The memory 110 is further shown as including an image quality assessment module 136. The image quality assessment module 136 is configured to receive the medical image 124 and, in response, provide an image quality map 138 representing a distribution of image quality levels of the input medical image 124. If the saliency map 126 is received along with the medical image 124 as input, the image quality levels may be weighted using a distribution of user attention on the input medical image 124 as predicted by the input saliency map 126. The weighted levels of image quality provided by the image quality map 138 may be used to calculate an image quality score, for example providing an average image quality of the medical image 124.
[0107] The image quality assessment module 136 may be used by the medical system 100 to train the machine learning module 121, for example. The image quality estimated by the image quality assessment module 136 may represent, for example, the loss of the medical image 124 output by the machine learning module 121 relative to one or more reference medical images. For example, training data for training the machine learning module 121 may be provided. Each training data may be included in the memory 110, for example. The training data for training the machine learning module 121 may include pairs of training medical imaging data and training medical images reconstructed using the training medical imaging data. The machine learning module 121 may be trained to output the training medical image upon receiving the training medical image data. The training may include, for each pair, providing the training medical imaging data to the machine learning module 121 as an input and receiving a preliminary medical image as an output. The received preliminary medical image may be, for example, the medical image 124. The distribution of image quality represented by the image quality map 138 for the medical image 124 received from the image quality assessment module 136 may be used as the distribution of loss of the medical image 124 relative to the training medical image of the pair. To determine the received image quality map 138, the training medical image may be provided to the image quality assessment module 136 as a reference medical image. Training the machine learning module 121 may include adjusting parameters of the machine learning module until the loss of the medical image 124 meets a predefined criterion. This criterion may, for example, require that the loss is less than a predefined threshold.
[0108] FIG. 2 shows a flow chart illustrating an exemplary method of predicting a saliency map using a trained machine learning module. The method of FIG. 2 may be executed, for example, by the medical system 100 of FIG. 1. A medical image is received in block 200. The medical image may be reconstructed using acquired medical imaging data or may be provided in a reconstructed form. In block 202, the medical image is provided as an input to the trained machine learning module. In block 204, in response to providing the medical image as an input, a saliency map of the medical image is received as an output from the trained machine learning module. The received saliency map predicts a distribution of a user's attention on the provided medical image. In block 204, the saliency map of the medical image is provided for further use. For example, the saliency map may be used to weight an OOD map and / or weight an image quality map to select an appropriate image reconstruction method. Such weighting of the image quality map may be used, for example, to train a further machine learning module to reconstruct a medical image from the medical imaging data.
[0109] 3 shows a flow chart illustrating an exemplary method for predicting a saliency map. As shown in FIG. 3, a medical image 124, for example of a brain, is provided to a trained machine learning module 122. The trained machine learning module 122 may be configured, i.e., trained, to provide a saliency map 126 upon receiving the medical image 124. The saliency map 126 predicts a distribution of a user's attention on the medical image 124. The brighter a pixel is in the saliency map 126 of FIG. 3, the higher the predicted level of attention of the user to that pixel in the medical image 124 may be.
[0110] Figure 4 illustrates a further example of a medical system 100 including a magnetic resonance imaging system 302. The medical system 100 illustrated in Figure 4 is similar to the medical system 100 of Figure 1, except that it additionally includes a magnetic resonance imaging system 302 controlled by a computing system 104.
[0111] The magnetic resonance imaging system 302 has a magnet 304. The magnet 304 is a superconducting cylindrical magnet with a bore 306 passing through it. Different types of magnets can be used. For example, it is possible to use both split cylindrical magnets and so-called open magnets. Split cylindrical magnets are similar to standard cylindrical magnets, but differ in that the cryostat is split into two parts to allow access to the iso-plane of the magnet; such magnets can be used, for example, with charged particle beam therapy. Open magnets have two magnet parts, one located above the other to allow enough space to accommodate a subject between them, the arrangement of the two parts being similar to that of a Helmholtz coil. Open magnets are popular because the subject is less occluded. Inside the cryostat of the cylindrical magnet is a collection of superconducting coils.
[0112] Within the bore 306 of the cylindrical magnet 304 is an imaging zone 308 in which there is a magnetic field strong enough and uniform to perform magnetic resonance imaging. A region of interest 309 is shown within the imaging zone 308. Magnetic resonance data is typically acquired about the region of interest. A subject 318 is shown supported by a subject support 320 such that at least a portion of the subject 318 is within the imaging zone 308 and the region of interest 309.
[0113] Also within the magnet bore 306 are a set of magnetic field gradient coils 310 used for preliminary magnetic resonance data acquisition to spatially encode magnetic spins within the imaging zone 308 of the magnet 304. The magnetic field gradient coils 310 are connected to a magnetic field gradient coil power supply 312. It should be understood that the magnetic field gradient coils 310 are representative. Typically, the magnetic field gradient coils 310 include three separate coil sets for spatial encoding in three orthogonal spatial directions. The magnetic field gradient power supply supplies current to the magnetic field gradient coils 310. The current supplied to the magnetic field gradient coils 310 is controlled as a function of time and may be ramped or pulsed.
[0114] Next to the imaging zone 308 is a radio frequency coil 314 for manipulating the orientation of the magnetic spins in the imaging zone 308 and for receiving radio signals from the spins in the imaging zone 308. The radio frequency antenna may include multiple coil elements. The radio frequency antenna may also be referred to as a channel or antenna. The radio frequency coil 314 is connected to a radio frequency transceiver 316. The radio frequency coil 314 and the radio frequency transceiver 316 may be replaced by separate transmit and receive coils and separate transmitters and receivers. It should be understood that the radio frequency coil 314 and the radio frequency transceiver 316 are representative. The radio frequency coil 314 is also intended to represent a dedicated transmit antenna and a dedicated receive antenna. Similarly, the transceiver 316 may also represent a separate transmitter and receiver. The radio frequency coil 314 may also have multiple receive / transmit elements, and the radio frequency transceiver 316 may have multiple receive / transmit channels. For example, if a parallel imaging technique such as SENSE is implemented, the radio frequency coil 314 has multiple coil elements.
[0115] The transceiver 316 and the gradient controller 312 are shown as being connected to the hardware interface 106 of the computer system 102 .
[0116] The memory 110 is further shown as including a plurality of pulse sequence commands 330. The pulse sequence commands 330 may be commands configured to control the magnetic resonance imaging system 302 to acquire medical image data 123 from the region of interest 309, or data convertible to such commands. In the case of the medical system 100 according to Fig. 4, the medical image data 123 is magnetic resonance imaging data, and the medical image 124 reconstructed using the medical image data 123 is an MRI image.
[0117] Fig. 5 shows a further example of the medical system 100. The medical system 100 shown in Fig. 4 is similar to the medical system 100 of Fig. 1, except that it additionally includes a computed tomography system 332 controlled by the computing system 104. The CT system 332 is shown as being controlled by the computer 102. The hardware interface 106 allows the computing system 104, e.g. a processor, to exchange data with and control the CT system 332. The memory 110 of the computer system 102 is shown as including additional CT system control commands 350 for controlling the CT system 332. The CT system control commands 350 can be used by the computing system 104 to control the CT system 332 to obtain medical imaging data 123 in the form of CT imaging data. The CT imaging data 123 can be used to reconstruct a medical image 124 in the form of a CT image.
[0118] The CT system 332 may include a rotating gantry 336. The gantry 336 may rotate about an axis of rotation 340. A subject 318 is shown on a subject support 320. Within the gantry 336 is an x-ray tube 342, which may be within an x-ray tube high voltage isolation tank, for example. A voltage stabilization circuit 338 may also include the x-ray tube 342, may be within the x-ray power supply 334, or may be external to both. The x-ray power supply 334 provides power to the x-ray tube 342.
[0119] The x-ray tube 334 produces x-rays 346 that pass through the subject 318 and are received by the detector 344. Within the box 309 is a region of interest within the imaging zone 308 from which a CT or computed tomography image 124 of the subject 318 may be created.
[0120] For positron emission tomography (PET) or single photon emission computed tomography (SPECT), a similar system may be used with a detector 344 including a gamma camera. For positron emission tomography or single photon emission computed tomography, no external radiation source is required. For example, the detector 344 of the CT system 332 may include a gamma camera and may be used for PET-CT imaging or SPECT-CT imaging, i.e., a combination of PET and CT or a combination of SPECT and CT, respectively.
[0121] FIG. 6 shows a flow chart illustrating an example of a method for predicting a saliency map, for example using the medical system of FIG. 4 or FIG. 5. Alternatively or additionally, the medical system may be configured for PET imaging or SPECT imaging. In block 210, medical imaging data is acquired using a medical imaging system included in the medical system. In the case of the medical system of FIG. 4, the medical imaging system is an MRI system and the medical imaging data is MRI data. In the case of the medical system of FIG. 5, the medical imaging system is a CT imaging system and the medical imaging data is CT data. In block 212, a medical image is reconstructed using the medical imaging data acquired in block 210. In the case of the medical system of FIG. 4, the medical image is an MRI image. In the case of the medical system of FIG. 5, the medical image is a CT image. Further blocks 214, 216, and 218 of FIG. 6 are equivalent to blocks 202, 204, and 206 of FIG. 2. In the case of PET and SPECT, a similar method can be used to predict the saliency map.
[0122] FIG. 7 shows a flow chart illustrating an example of a method for training a machine learning module to predict a saliency map. In block 220, a machine learning module to be trained is provided. The provided machine learning module may be, for example, an untrained machine learning module or a pre-trained machine learning module that is further trained. In block 222, training data for training the machine learning module is provided. The training data includes pairs of training medical images and training saliency maps. The training saliency maps represent the distribution of a user's attention on the training medical images. In block 224, the machine learning module is trained using the provided training data. The resulting trained machine learning module is trained to output a training saliency map of the pair in response to receiving a training medical image of the pair.
[0123] FIG. 8 illustrates an example of a medical system 101 configured to train a machine learning module 160 to output a saliency map as an output in response to receiving a medical image as an input. The resulting saliency map predicts the distribution of a user's attention on the medical image provided as an input. The machine learning module 160 may be, for example, an untrained machine learning module or a partially pre-trained machine learning module. Training the machine learning module 160 results in a trained machine learning module 122. The medical system 101 is illustrated as including a computer 102. The computer 102 is intended to represent one or more computing devices. For example, the computer 102 may be a remote computer system used to remotely train the machine learning module 160. For example, the computer 102 may be a server in a radiology department or a virtual computer system in a cloud computing system.
[0124] Computer 102 is further shown as including a computing system 104. Computing system 104 is intended to represent one or more processors or processing cores, or other computing systems, located in one or more locations. Computing system 104 is shown as being connected to an optional hardware interface 106. Optional hardware interface 106 may, for example, allow computing system 104 to control other components.
[0125] The computing system 104 is further shown as being connected to an optional user interface 108, which may, for example, allow an operator to control and operate the medical system 101. The optional user interface 108 may include, for example, output and / or input devices that allow a user to interact with the medical system. The output devices may include, for example, a display device configured to display medical images and saliency maps. The input devices may include, for example, a keyboard and / or a mouse that allow a user to insert control commands to control the medical system 101. The optional user interface 108 may include, for example, an eye tracking device, such as a camera configured to track the eye position and / or movement of a user using the medical system 101. The computing system 104 is further shown as being connected to a memory 110. The memory 110 is intended to represent various types of memory that may be connected to the computing system 104.
[0126] The memory is shown as including machine-executable instructions 120. The machine-executable instructions 120 enable the computing system 104 to perform tasks, such as controlling other components and performing various data and image processing tasks. The machine-executable instructions 120 may, for example, enable the computing system 104 to train a machine learning module 160 and provide a trained machine learning module 122 as a result of the training.
[0127] For training the machine learning module 160, training data 162 may be provided. The training data may include pairs of training medical images and training saliency maps. The training saliency maps represent the distribution of a user's attention on the training medical images. The training saliency maps may be generated using an eye tracking device, such as a camera configured to track the eye position and / or movement of a user using the medical system 101. Based on the tracking data provided by the eye tracking device, the distribution of the user's attention on the training medical images may be determined, and the saliency map may be obtained as a result. The machine learning module 160 is trained using the provided training data 162. The machine learning module 160 is trained to output the training saliency map of the pair in response to receiving the training medical images of the pair. In this manner, the machine learning module 122 may be generated.
[0128] 9 illustrates another example of the medical system 100 configured to train the machine learning module 160 to output a saliency map as an output in response to receiving a medical image as an input. The resulting saliency map predicts the distribution of a user's attention on the medical image provided as an input. The machine learning module 160 may be, for example, an untrained machine learning module or a partially pre-trained machine learning module. Training the machine learning module 160 results in a trained machine learning module 122.
[0129] The medical system 100 shown in Figure 9 corresponds to the medical system 100 shown in Figure 1. The medical system 100 shown in Figure 9 further includes a machine learning module 160 that is trained using training data 162 to provide a trained machine learning module 122. The machine-executable instructions 120 may, for example, enable the computing system 104 to train the machine learning module 160 and provide the trained machine learning module 122 as a result of the training.
[0130] For training the machine learning module 160, training data 162 may be provided. The training data may include pairs of training medical images and training saliency maps. The training saliency maps represent the distribution of a user's attention on the training medical images. The training saliency maps may be generated using an eye tracking device, such as a camera configured to track the eye position and / or movement of a user using the medical system 100. Based on the tracking data provided by the eye tracking device, the distribution of the user's attention on the training medical images may be determined, and the saliency map may be obtained as a result. The machine learning module 160 is trained using the provided training data 162. The machine learning module 160 is trained to output the training saliency map of the pair in response to receiving the training medical images of the pair. In this manner, the machine learning module 122 may be generated.
[0131] FIG. 10 illustrates an exemplary trained machine learning module 122 having a neural network trained to predict a saliency map. An exemplary architecture of the provided neural network may be a U-Net architecture. U-Net is a convolutional neural network. A convolutional neural network is a type of deep neural network, and is applied, for example, to the analysis of visual images. U-Net includes a contracting path and an expansive path, and thus has a U-shaped architecture. The contracting path is a typical convolutional network consisting of repeated application of convolutions, each of which may be followed by a ReLU (rectified linear unit) and max pooling operation. During contraction, spatial information may decrease while feature information may increase. The expansive path may combine feature information and spatial information through a sequence of up-convolution and concatenation with high-resolution features from the contracting path.
[0132] Such a machine learning module 122 (e.g., having a U-Net architecture) may be trained to mimic the behavior of a user, e.g., a radiologist, in relation to a displayed medical image. For example, the user attention detected in FIG. 12 below may be mimicked. For this purpose, training data may be provided that includes pairs of training medical images (e.g., medical image 406 displayed in FIG. 12) and training saliency maps (e.g., saliency map 408 determined in FIG. 12 below for the displayed medical image 406). The machine learning module 122 may be trained to receive such training medical images 406 as input and, in response, generate as output the training saliency map 408 assigned to the input training medical image 406. Thus, the machine learning module 122 may be enabled to predict saliency maps for other medical images as well.
[0133] For example, the machine learning module 122 for predicting the saliency map may be trained based on a training pair of acquired medical images 406 and corresponding saliency maps 408 as shown in FIG. 12 below, using a loss function designed for image-image transformation. For example, loss functions such as MSE, MAE, SSIM, or combinations thereof may be used. The trained machine learning module 122 may be configured to receive medical images as input and predict a saliency map. For example, such saliency may predict the distribution of a user's attention. Such a saliency map may be used, for example, as a user guide to indicate noteworthy regions of a given medical image. Alternatively or additionally, such a saliency map may be used in connection with image reconstruction and / or medical image analysis.
[0134] FIG. 11 illustrates an example of a method for providing a training saliency map to a machine learning module. In block 230, a training medical image is displayed using a display device of the medical system used to provide the training saliency map. In block 232, a distribution of a user's attention on the displayed training medical image is measured. The user's attention may be measured, for example, based on a user's interaction with the training medical image. For example, the position and movement of a cursor controlled by a user within the displayed training medical image may be measured. For example, the user's attention may be measured using an eye tracking device configured to measure the eye position and movement of a user of the medical system. An attention determination module configured to determine the distribution of the user's attention on the displayed training medical image using the eye tracking device may be used to determine points of attention within the displayed training medical image for a user of the medical system viewing the displayed training medical image. In block 234, a training saliency map is generated using the measured distribution of the user's attention on the training medical image.
[0135] FIG. 12 illustrates an example of a method for generating training data for training a machine learning module to generate a saliency map that predicts the distribution of a user's attention on a medical image. The training saliency data may be obtained, for example, in the form of user attention data determined using an eye-tracking device 144 to track areas of a medical image 406 that are the focus of a user. The obtained training saliency data is used to generate a saliency map 408. The saliency map 408 may be used as training data 407 in combination with a medical image 406 to train a machine learning module to receive a given medical image as input and provide a saliency map as output. The trained machine learning model may be used to generate a saliency map that, at run time, predicts the distribution of a user's attention on a given medical image.
[0136] FIG. 12 shows an example of a pipeline for constructing a training saliency map 408 using eye tracking of a user 500, e.g., a radiologist, observing medical images such as MR images. Training data including a set of medical images 406, i.e., the training medical images, are presented to the radiologist 500 and displayed, e.g., on a display device 140 of the medical equipment 100. The position and movement of the eyes 502 of the radiologist 500 are recorded using an eye tracking device 144, such as a camera. The collected eye tracking data is converted into a saliency map 408 representing the distribution of the user attention of the radiologist 500 on the displayed medical images 406. For this purpose, an attention determination module 150 may be provided. The attention determination module 150 may be configured to determine the distribution of the user attention on the displayed training medical images 406 using the eye tracking device 144. This may be used, for example, to determine points of attention within the displayed training medical images 406 for the user 500 of the medical system 100 viewing the displayed training medical images 406. At block 234, a training saliency map is generated using the distribution of the user's attention on the measured training medical images. The displayed medical images 406 can be paired with the generated saliency map 408 to provide pairs of training data 407 for training the machine learning module to predict the saliency map for a given medical image. Each pair of training data 407 includes a training medical image 406 and a training saliency map 408 representing the distribution of the user's attention on the respective training medical image 406 as determined using the eye tracking device 144.
[0137] FIG. 13 shows a flow chart illustrating an example of a method for selecting a reconstruction method using a saliency map. A medical system performing this method may be configured to use the saliency map to select a reconstruction method for reconstructing a medical image from a plurality of predefined reconstruction methods. In block 240, a plurality of test maps are provided for a test medical image. The test medical image is a medical image of a predefined type of anatomical structure for which the reconstructed medical image is to be targeted. Each test map is assigned to a different reconstruction method. Each test map identifies a section of the test image that includes an anatomical substructure of the predefined type of anatomical structure that has the highest image reconstruction quality when using the assigned reconstruction method, compared to other anatomical substructures of the predefined type of anatomical structure. In block 242, the test map is compared to the saliency map. The saliency map is a predicted saliency map for the test medical image. That is, the saliency map predicts the distribution of the user's attention on the test medical image. In block 244, the test map that has the highest structural similarity to the saliency map is determined. In block 246, the reconstruction method assigned to the determined test map is selected for reconstructing the medical image to be reconstructed. In block 246, the medical image to be reconstructed is reconstructed using the selected reconstruction method.
[0138] In this way, the trained machine learning module can be used to select a reconstruction method that is more convenient for the user. By providing a saliency map that predicts the distribution of the user's attention, the trained machine learning provides a method that takes into account a measure of relevance related to the user. In this environment, multiple reconstruction methods with comparable characteristics can be deployed. For example, a certain anatomical structure should be depicted. Comparable characteristics may refer to the fact that all reconstruction methods can provide reconstructed medical images that represent the respective anatomical structure. Based on the prediction of the user's attention provided by the saliency map, the reconstruction method that is most suitable for the user can be selected. The saliency map can, for example, predict the distribution of the user's attention for an individual user. Thus, different saliency maps can be predicted for different users depending on the user's experience, references, and / or way of working. Different radiologists may view images in different ways. Thus, each radiologist may benefit from a reconstruction that better meets their needs. Each reconstruction method can be designed to handle the specific characteristics of the input data, i.e., the specific characteristics of the acquired medical data used to reconstruct the medical image. For example, in MRI, some reconstruction methods produce medical images with high contrast between white and gray matter in the brain, while other reconstruction methods may provide better signal-to-noise ratios in regions closer to the skull. For each reconstruction method, a test map may be provided that identifies a section of the test image where the reconstruction method has the highest image reconstruction quality. For example, if a reconstruction method provides high image quality for white matter in the brain, a test map may be provided that highlights the white matter in the test image. For example, if a reconstruction method provides high image quality for gray matter in the brain, a test map may be provided that highlights the gray matter in the test image. For example, if a reconstruction method provides high image quality for cerebrospinal fluid in the brain, a test map may be provided that highlights the cerebrospinal fluid in the test image. The test map may have the appearance of a saliency map.The test map may be compared with a saliency map provided, for example, by a machine learning configured to predict the distribution of attention of a particular radiologist on test images. The most suitable reconstruction method is selected, i.e. the reconstruction method for which the test map shows the highest similarity with the saliency map. The similarity may be determined by estimating the distance between the saliency map obtained for the radiologist and multiple test maps provided for different reconstruction methods.
[0139] FIG. 14 illustrates an example of how a saliency map may be used to select a reconstruction method. For example, suppose a medical image of a brain is to be reconstructed. A medical image 124 of the brain may be provided as a test image, and the test image may be provided with multiple test maps 128 assigned to multiple different reconstruction methods. In this example, there are three available reconstruction methods, each with its own strengths in reconstructing different parts of the brain. For example, a first reconstruction method may have strengths in reconstructing white matter in the brain. A second reconstruction method may have strengths in reconstructing gray matter in the brain, and a third reconstruction method may have strengths in reconstructing cerebrospinal fluid. For each reconstruction method, a test map may be provided that highlights the area of the test image where that reconstruction method has the highest reconstruction quality. For example, for the first reconstruction method, a test map 422 may be provided that highlights the white matter contained in the test image. For example, for the second reconstruction method, a test map 424 may be provided that highlights the gray matter contained in the test image, and for the third reconstruction method, a test map 426 may be provided that highlights the cerebrospinal fluid. Furthermore, a saliency map 126 predicting the distribution of the user's attention on the test medical image 124 is generated using a trained machine learning module, for example a user-specific trained machine learning module. The trained machine learning module may be, for example, a deep learning neural network. In this case, the second reconstruction method may be selected, for example, because the test map 424 highlighting the gray matter contained in the test image 124 best matches the distribution of the user's attention predicted by the saliency map 126 for the medical test image 124 of the brain. Thus, it is predicted that the user will mainly pay attention to the gray matter. Thus, the model that is most compatible with the gray matter may be selected as the model that best suits the user's needs.
[0140] FIG. 15 shows a flow chart illustrating an example of a method for providing a weighted OOD map using a saliency map. A medical system performing this method may include an OOD estimation module configured to output an OOD map in response to receiving a medical image as an input. The OOD map represents a level of compliance of the input medical image with respect to a reference distribution defined by a set of reference medical images. In block 250, the medical image is provided as an input to the OOD estimation module. In block 252, in response to providing the medical image, an OOD map of the medical image is received as an output from the OOD estimation module. The OOD map represents a level of compliance of the medical image with respect to a predetermined distribution. In block 254, a weighted OOD map is provided. Providing the weighted OOD map may include weighting the level of compliance represented by the OOD map received in block 252 using a distribution of the user's attention on the medical image predicted by the saliency map.
[0141] Figure 16 shows a flow chart illustrating an example of a method for using a saliency map to provide a weighted OOD score. Blocks 260-264 of Figure 16 correspond to blocks 250-254 of Figure 15. In block 266, an OOD score is calculated using the weighted compliance levels of the weighted OOD map provided in block 264. The OOD score represents the probability that the medical image as a whole falls within the reference distribution.
[0142] FIG. 17 shows a flow chart illustrating an example of a method for providing a weighted OOD map 134 using the saliency map 126. The saliency map 126 is used to scale the OOD map 132, resulting in a weighted OOD map 134. The ODD map 132 is generated for the same medical image for which the saliency map 126 is provided. The OOD map 132 provides a distribution of OOD values on the medical image, i.e. a distribution of uncertainty levels. The saliency map 126 predicts the distribution of the user's attention on the medical image, and thus the distribution of relevance. The scaled OOD map 134 looks rather blank, because according to the user's attention as indicated by the saliency map 126, the areas with the highest OOD scores as indicated by the OOD map 132 seem to be the least important to the user. Thus, such medical images may still be reliable, even though they may contain uncertainty, since the uncertainty is limited to the unimportant areas of the medical image.
[0143] FIG. 18 shows a flow chart illustrating an example of a method for providing an image quality map using a saliency map. A medical system used to perform this method may include an image quality assessment module configured to output an image quality map in response to receiving a medical image and a saliency map as input. The image quality map represents a distribution of image quality levels of the input medical image weighted using a distribution of user attention on the input medical image predicted by the input saliency map. In block 270, the medical image and the saliency map are provided as input to the image quality assessment module. In block 272, in response to providing the medical image and the saliency map, an image quality map is received as an output from the image quality assessment module. The image quality map represents a distribution of image quality levels of the medical image weighted using a distribution of user attention on the medical image predicted by the saliency map. In block 274, the computing system provides the received image quality map. The image quality map may be used, for example, as a weighted loss function for training a machine learning module to reconstruct a medical image.
[0144] Figure 19 shows a flow chart illustrating an example of a method for providing an image quality score using a saliency map. Blocks 280-284 of Figure 19 correspond to blocks 270-274 of Figure 18. In block 286, an image quality score is calculated using the received image quality map. The image quality score may represent an average image quality of the medical image.
[0145] FIG. 20 shows a flow chart illustrating an example of a method for providing an image quality map using a saliency map, for example a weighted reconstruction error for an MSE metric. The image quality map 137, i.e. an error map of the distribution of measures of reconstruction error on medical images, is scaled using the saliency map 126. As a measure of reconstruction error, for example the MSE metric can be used. The reconstruction error values for the MSE metric are weighted using the distribution of user attention predicted by the saliency map 126. The scaled image quality map 138 shows that the neural network under training may need to be penalized in a very different way to generate images that meet the needs of the user determined based on the predicted user attention. Without scaling, the neural network under training would be penalized for errors occurring in the brighter areas of the image quality map 137, which are very different from the areas highlighted in the weighted image quality map 138. The areas highlighted in the image quality map 138 are areas that are important to the user. Thus, using image quality map 137 to adjust the parameters of a neural network during training may improve the overall image quality of the image being reconstructed by the neural network. However, the improvement may be negligible for areas that are important to the user. To improve image quality in a way that the user can actually benefit, it may be necessary to use a weighted image quality map such as map 138.
[0146] FIG. 21 shows a flowchart illustrating an example of a method for training a machine learning module to reconstruct a medical image using a saliency map. An image quality assessment module that provides an image quality map weighted using a saliency map, such as that shown in FIG. 18 and FIG. 20, can be used to train the machine learning module to output a medical image as an output in response to receiving medical image data as an input. The image quality estimated by the image quality assessment module may represent a loss of the output medical image of the machine learning module relative to one or more reference medical images. In block 290, a machine learning module to be trained may be provided. In block 292, training data for training the machine learning module is provided. The training data includes pairs of training medical imaging data and training medical images reconstructed using the training medical imaging data. In block 294, the machine learning module is trained using the training data and the image quality map. The machine learning module is trained to output a training medical image of the pair in response to receiving the training medical imaging data of the pair. The training includes, for each pair, providing the training medical imaging data to the machine learning module as an input and receiving a preliminary medical image as an output. For the preliminary medical image, an image quality map is generated that is weighted by the predicted saliency map for the preliminary medical image. The distribution of image quality represented by the image quality map received for the preliminary medical image from the image quality assessment module is used as a distribution of losses of the preliminary medical image relative to the training medical image of the pair that is provided as a reference medical image to the image quality assessment module to determine the received image quality map. The parameters of the machine learning module are adjusted during training until the losses of the medical image meet a predefined criterion. The criterion may, for example, require that the losses are below a threshold.
[0147] While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive. The invention is not limited to the disclosed embodiments.
[0148] From the drawings, the disclosure, and the appended claims, other variations of the disclosed embodiments can be understood and realized by those skilled in the art in practicing the claimed invention. In the claims, the terms "comprise" and "include" do not exclude other elements or steps, and singular elements do not exclude a plurality. A single processor or other unit may fulfill the functions of several items recited in the claims. The mere fact that several means are recited in mutually different dependent claims does not indicate that a combination of these means cannot be used advantageously. The computer program may be stored and / or distributed on a suitable medium, such as an optical storage medium or a solid-state medium, supplied together with or as part of other hardware, or distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems. Any reference signs in the claims should not be interpreted as limiting the scope thereof. [Explanation of symbols]
[0149] 100 Medical Systems 101 Medical Systems 102 Computer 104 Computing Systems 106 Optional Hardware Interface 108 Optional User Interface 110 Memory 120 Machine Executable Instructions 121 Machine Learning Module 122 Trained Machine Learning Modules 123 Medical Imaging Systems 124 Medical Imaging 126 Saliency Map Set of 128 test maps 130 OOD Estimation Module 132 OOD Map 134 Weighted OOD Map 136 Image Quality Assessment Module 138 Quality Map 140 Display Devices 144 Eye Tracking Devices 150 Attention Decision Module 160 Machine Learning Module 162 Training Data 302 Magnetic Resonance Imaging System 304 Magnet 306 Magnet Bore 308 Imaging Zone 309 Areas of Interest 310 Magnetic Gradient Coil 312 Magnetic field gradient coil power supply 314 Radio Frequency Coil 318 Transmitter / Receiver 318 Subjects 320 Subject support platform 330 Pulse Sequence Commands 332 CT System 334 X-ray power supply 336 Gantry 338 Voltage Stabilizer Circuit 340 Rotational Axis 342 X-ray tube 344 Detector 346 X-ray 350 CT Control Command 406 Medical Training Images 407 Training Data 408 Training Saliency Map 422 Test Map 424 Test Map 426 Test Map 500 users 502 Eye
Claims
**Claim 1**: A medical system comprising a computing system and a memory, wherein when machine-executable instructions are executed on the computing system, the memory stores the machine-executable instructions including a trained first machine learning module configured to output a saliency map as an output in response to receiving a medical image as an input, and the saliency map predicts the distribution of a user's attention on the medical image, and the memory A medical system for reconstructing a medical image, comprising Execution of the machine-executable instructions further causes the computing system to Provide a test map for a test medical image of a predetermined type of anatomical structure for which the medical image is to be reconstructed, each test map being assigned to a different one of the reconstruction methods, and each test map identifying a section of the test image including an anatomical sub-structure of the predetermined type of anatomical structure having the highest image reconstruction quality compared to other anatomical sub-structures of the predetermined type of anatomical structure when the assigned reconstruction method is used, the step of Comparing the test map with the saliency map Determining the test map having the highest structural similarity to the saliency map among the test maps Selecting the reconstruction method assigned to the determined test map Reconstructing the medical image using the selected reconstruction method Configured to execute Medical system. **Claim 2** By execution of the machine-executable instructions, the computing system further provides the trained first machine learning module, and the provision of the trained first machine learning module includes Providing the first machine learning module Providing first training data including a first pair of a training medical image and a training saliency map, the training saliency map representing the distribution of a user's attention on the training medical image, providing the first training data Training the first machine learning module using the first training data, wherein the resulting trained first machine learning module is trained to output the training saliency map of the first pair in response to receiving the training medical image of the first pair, the method for training the first machine learning module, the medical system according to claim 1.
3. The medical system further includes a display device, and the providing of the first training data includes, for each of the training medical images of the first training data, displaying the training medical image using the display device; measuring the distribution of the user's attention on the displayed training medical image; generating the training saliency map of the first pair of training data including the displayed training medical image using the measured distribution of the user's attention on the training medical image, the medical system according to claim 2.
4. The medical system further includes a gaze tracking device that measures the position and movement of the eyes of the user of the medical system, and the memory further stores an attention determination module, and the attention determination module uses the gaze tracking device to determine the distribution of the user's attention on the displayed training medical image, thereby determining the attention point in the displayed training medical image for the user of the medical system looking at the displayed training medical image, the medical system according to claim 3.
5. The trained first machine learning module is trained to output a user-specific saliency map that predicts the user-specific distribution of the user's attention on the input medical image in response to receiving the medical image as an input, the medical system according to any one of claims 1 to 4.
6. The memory further stores an out-of-distribution estimation module, By executing the machine-executable instructions, the computing system further provides the medical image as an input to the out-of-distribution estimation module; receives an out-of-distribution map of the medical image as an output from the out-of-distribution estimation module in response to the providing of the medical image, wherein the out-of-distribution map represents the compliance level of the medical image with respect to a predetermined distribution, receiving the out-of-distribution map The medical system according to claim 1, comprising providing a weighted out-of-distribution map that includes weighting the compliance level represented by the out-of-distribution map using the distribution of the user's attention on the medical image predicted by the saliency map.
7. The provision of the weighted out-of-distribution map further includes calculating an out-of-distribution score using the weighted compliance level provided by the weighted out-of-distribution map, wherein the out-of-distribution score represents the probability that the medical image is within the reference distribution as a whole. The medical system according to claim 6.
8. The memory further stores an image quality evaluation module. Upon execution of the machine-executable instructions, the computing system further provides the medical image and the saliency map as inputs to the image quality evaluation module. Upon providing the medical image and the saliency map, receiving, as an output from the image quality evaluation module, an image quality map, wherein the image quality map represents a distribution of the image quality levels of the medical image weighted using the distribution of the user's attention on the medical image predicted by the saliency map. Receiving the image quality map. The medical system according to claim 1, comprising providing the received image quality map.
9. The image quality evaluation module is used to train a second machine learning module to output a medical image as an output in response to receiving medical imaging data as an input, and the image quality estimated by the image quality evaluation module represents the loss of the output medical image of the second machine learning module with respect to one or more reference medical images. Upon execution of the machine-executable instructions, the computing system further provides the second machine learning module. providing second training data for training the second machine learning module, wherein the second training data includes a second pair of training medical imaging data and a training medical image reconstructed using the training medical image data. Providing the second training data. Training the second machine learning module, wherein the second machine learning module is trained to output the training medical image of the second pair in response to receiving the training medical image data of the second pair, and the training includes, for each second pair, providing the respective training medical imaging data as an input to the second machine learning module and receiving a preliminary medical image as an output, wherein the received preliminary medical image is the medical image, and executing training the second machine learning module. The distribution of the image quality represented by the image quality map received for the medical image from the image quality evaluation module is used as the distribution of the loss of the medical image with respect to the training medical image of each of the second pairs provided as a reference medical image to the image quality evaluation module for determining the received image quality map. The parameters of the second machine learning module are adjusted during the training until the loss of the medical image meets a predetermined criterion. The medical system according to claim 8.
10. The providing of the received image quality map includes calculating an image quality score using the received image quality map, and the image quality score represents the average image quality of the medical image. The medical system according to claim 8.
11. The medical system acquires medical imaging data for reconstructing the medical image, and the medical imaging data is acquired using any one of the data acquisition methods of magnetic resonance imaging, computed tomography, positron emission tomography, and single photon emission computed tomography. The medical system according to claim 1.
12. A method for reconstructing a medical image using a trained machine learning module, the method comprising: Receiving a medical image; Providing the medical image as an input to the trained machine learning module; Outputting a saliency map of the medical image from the trained machine learning module in response to the providing of the medical image, wherein the saliency map predicts the distribution of the user's attention on the medical image. Providing a test map for a test medical image of a predetermined type of anatomical structure from which the medical image is reconstructed, each test map being assigned to a different one of the reconstruction methods, each test map identifying a section of the test image that includes a sub-anatomical structure of the predetermined type of anatomical structure having the highest image reconstruction quality compared to other sub-anatomical structures of the predetermined type of anatomical structure when the assigned reconstruction method is used; Comparing the test map with the saliency map; Determining the test map among the test maps that has the highest structural similarity to the saliency map; Selecting the reconstruction method assigned to the determined test map; Reconstructing the medical image using the selected reconstruction method A method comprising. A computer program comprising machine-executable instructions, the execution of which is configured to cause a computing system to perform the method according to claim 12.