Bias mitigation of machine learning models
Patent Information
- Application Number
- PCT/US2024/037839
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-07-14
- Filing Date
- 2024-07-12
- Publication Date
- 2025-06-05
AI Technical Summary
Existing machine learning models used for image analysis, particularly in computer vision tasks, often suffer from biases that can lead to inaccurate classifications and unfair outcomes.
The proposed method involves generating a second training data set that includes composite images created by segmenting images into foreground and background portions and applying transformations such as blurring, rotation, or adversarial transformations. These composite images are used to retrain the initial machine learning model, thereby mitigating biases.
By using the composite images and transformations, the method effectively reduces bias in the machine learning model, improving its accuracy and fairness in image classification tasks.
Smart Images

Figure US2024037839_05062025_PF_FP_ABST
Abstract
Description
BIAS MITIGATION OF MACHINE LEARNING MODELSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Patent Application No. 63 / 513,841, filed July 14, 2023, and titled “Bias Mitigation of Machine Learning Models,” which is incorporated by reference herein in its entirety.FIELD
[0002] The present disclosure relates generally to techniques for reducing bias in trained machine learning models including those used for image analysis.BACKGROUND
[0003] Interpretation of digital images by computers, generally termed computer vision, includes the extraction of high-dimensional data from digital images. These digital images can take many forms, such as being drawn from video sequences, views from multiple cameras, or images from medical scanning devices or other medical processes. To interpret digital images, computer vision tasks may include tasks such as object detection, object recognition, pose estimation, motion estimation, shape recognition, and / or image classification tasks. Some computer vision techniques involve using machine learning models to perform computer vision tasks.SUMMARY
[0004] In some embodiments, the techniques described herein relate to a method of performing image classification using a trained machine learning model with mitigated biases, the method including: obtaining a test data set including first images; assigning, using the trained machine learning model, one or more images of the test data set to one or more classifications, wherein the trained machine learning model is trained using a training data set including second images, the training including: training, using the second images, a machine learning model to obtain an initial trained machine learning model; generating a second training data set including composite images by: segmenting images of the second images to obtain first image portions and second image portions; and generating the composite images using the first image portions and the second image portions; obtaining initial predictions using the initial trained machine learning model and the second training data set; computing importance scores using the obtained initial predictions; determining whether an average value of the importancescores is less than or equal to a threshold value; and generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, the trained machine learning model using the second training data set; and outputting the assigned one or more classifications obtained using the trained machine learning model.
[0005] In some embodiments, generating the composite images using the first image portions and the second image portions includes using foreground image portions and background image portions.
[0006] In some embodiments, generating the second training data set further includes applying a transformation to one or more of the first image portions and / or the second image portions before generating the composite images.
[0007] In some embodiments, applying the transformation includes applying one or more of a blurring transformation, rotation, adversarial transformation, textural transformation, and / or color transformation to one or more of the first image portions and / or the second image portions.
[0008] In some embodiments, obtaining initial predictions includes providing the composite images and the second image portions as input to the initial trained machine learning model to obtain initial predictions for the composite images and initial predictions for the second image portions.
[0009] In some embodiments, computing a first importance score of the importance scores includes: obtaining, for one first image portion of the first image portions, difference values by subtracting initial predictions for each second image portion from an average of the initial predictions for composite images including the one first image portion; and computing the first importance score by averaging the obtained difference values.
[0010] In some embodiments, generating the trained machine learning model includes generating a trained deep neural network or a trained convolutional neural network.
[0011] In some embodiments, generating the trained machine learning model using the second training data set includes either (i) retraining the initial trained machine learning model using the second training data set or (ii) training the machine learning model using the second training data set.
[0012] In some embodiments, the techniques described herein relate to at least one non- transitory computer-readable medium storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method of performing image classification using a trained machine learning model with mitigated biases, the method including: obtaining a test data set including first images; assigning, using thetrained machine learning model, one or more images of the test data set to one or more classifications, wherein the trained machine learning model is trained using a training data set including second images, the training including: training, using the second images, a machine learning model to obtain an initial trained machine learning model; generating a second training data set including composite images by: segmenting images of the second images to obtain first image portions and second image portions; and generating the composite images using the first image portions and the second image portions; obtaining initial predictions using the initial trained machine learning model and the second training data set; computing importance scores using the obtained initial predictions; determining whether an average value of the importance scores is less than or equal to a threshold value; and generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, the trained machine learning model using the second training data set; and outputting the assigned one or more classifications obtained using the trained machine learning model.
[0013] In some embodiments, generating the composite images using the first image portions and the second image portions includes using foreground image portions and background image portions.
[0014] In some embodiments, generating the second training data set further includes applying a transformation to one or more of the first image portions and / or the second image portions before generating the composite images.
[0015] In some embodiments, applying the transformation includes applying one or more of a blurring transformation, rotation, adversarial transformation, textural transformation, and / or color transformation to one or more of the first image portions and / or the second image portions.
[0016] In some embodiments, obtaining initial predictions includes providing the composite images and the second image portions as input to the initial trained machine learning model to obtain initial predictions for the composite images and initial predictions for the second image portions.
[0017] In some embodiments, computing a first importance score of the importance scores includes: obtaining, for one first image portion of the first image portions, difference values by subtracting initial predictions for each second image portion from an average of the initial predictions for composite images including the one first image portion; and computing the first importance score by averaging the obtained difference values.
[0018] In some embodiments, generating the trained machine learning model includes generating a trained deep neural network or a trained convolutional neural network.
[0019] In some embodiments, generating the trained machine learning model using the second training data set includes either (i) retraining the initial trained machine learning model using the second training data set or (ii) training the machine learning model using the second training data set.
[0020] In some embodiments, the techniques described herein relate to an image classification system, including: at least one processor; and at least one non-transitory computer-readable medium storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method of performing image classification using a trained machine learning model with mitigated biases, the method including: obtaining a test data set including first images; assigning, using the trained machine learning model, one or more images of the test data set to one or more classifications, wherein the trained machine learning model is trained using a training data set including second images, the training including: training, using the second images, a machine learning model to obtain an initial trained machine learning model; generating a second training data set including composite images by: segmenting images of the second images to obtain first image portions and second image portions; and generating the composite images using the first image portions and the second image portions; obtaining initial predictions using the initial trained machine learning model and the second training data set; computing importance scores using the obtained initial predictions; determining whether an average value of the importance scores is less than or equal to a threshold value; and generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, the trained machine learning model using the second training data set; and outputting the assigned one or more classifications obtained using the trained machine learning model.
[0021] In some embodiments, generating the composite images using the first image portions and the second image portions includes using foreground image portions and background image portions.
[0022] In some embodiments, generating the second training data set further includes applying a transformation to one or more of the first image portions and / or the second image portions before generating the composite images.
[0023] In some embodiments, applying the transformation includes applying one or more of a blurring transformation, rotation, adversarial transformation, textural transformation, and / or color transformation to one or more of the first image portions and / or the second image portions.
[0024] In some embodiments, obtaining initial predictions includes providing the composite images and the second image portions as input to the initial trained machine learning model to obtain initial predictions for the composite images and initial predictions for the second image portions.
[0025] In some embodiments, computing a first importance score of the importance scores includes: obtaining, for one first image portion of the first image portions, difference values by subtracting initial predictions for each second image portion from an average of the initial predictions for composite images including the one first image portion; and computing the first importance score by averaging the obtained difference values.
[0026] In some embodiments, generating the trained machine learning model includes generating a trained deep neural network or a trained convolutional neural network.
[0027] In some embodiments, generating the trained machine learning model using the second training data set includes either (i) retraining the initial trained machine learning model using the second training data set or (ii) training the machine learning model using the second training data set.
[0028] In some embodiments, the trained machine learning model was trained using a first data set including first images, and the method includes: obtaining initial predictions using the trained machine learning model and a second data set including second images; computing importance scores using the obtained initial predictions; determining, based on the computed importance scores, a degree of bias in the trained machine learning model; and outputting an indication of the degree of bias in the trained machine learning model.
[0029] In some embodiments, generating the second data set includes generating the second images by: segmenting images of the first images to obtain first image portions and second image portions; and generating the second images using the first image portions and the second image portions.
[0030] In some embodiments, generating the second images using the first image portions and the second image portions includes using foreground image portions and background image portions.
[0031] In some embodiments, generating the second data set further including applying a transformation to one or more of the first image portions and / or the second image portions before generating the second images.
[0032] In some embodiments, applying the transformation includes applying one or more of a blurring transformation, rotation, adversarial transformation, textural transformation,and / or color transformation to one or more of the first image portions and / or the second image portions.
[0033] In some embodiments, obtaining the initial predictions includes providing the second images and the second image portions as input to the trained machine learning model to obtain initial predictions for the second images and initial predictions for the second image portions.
[0034] In some embodiments, computing a first importance score of the importance scores includes: obtaining, for one first image portion of the first image portions, difference values by subtracting initial predictions for each second image portion from an average of the initial predictions for second images including the one first image portion; and computing the first importance score by averaging the obtained difference values.
[0035] In some embodiments, determining the degree of bias includes determining whether an average value of the computed importance scores is less than or equal to a threshold value.
[0036] In some embodiments, outputting the indication of the degree of bias includes generating a graphical output for display to a user indicating whether the average value of the computed importance scores is less than or equal to the threshold value.
[0037] In some embodiments, the method further includes generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, a retrained machine learning model using the second data set.
[0038] In some embodiments, the techniques described herein relate to at least one non- transitory computer-readable medium storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method of determining a degree of bias in a trained machine learning model, wherein the trained machine learning model was trained using a first data set including first images, the method including: obtaining initial predictions using the trained machine learning model and a second data set including second images; computing importance scores using the obtained initial predictions; determining, based on the computed importance scores, a degree of bias in the trained machine learning model; and outputting an indication of the degree of bias in the trained machine learning model.
[0039] In some embodiments, generating the second data set includes generating the second images by: segmenting images of the first images to obtain first image portions and second image portions; and generating the second images using the first image portions and the second image portions.
[0040] In some embodiments, generating the second images using the first image portions and the second image portions includes using foreground image portions and background image portions.
[0041] In some embodiments, generating the second data set further including applying a transformation to one or more of the first image portions and / or the second image portions before generating the second images.
[0042] In some embodiments, applying the transformation includes applying one or more of a blurring transformation, rotation, adversarial transformation, textural transformation, and / or color transformation to one or more of the first image portions and / or the second image portions.
[0043] In some embodiments, obtaining the initial predictions includes providing the second images and the second image portions as input to the trained machine learning model to obtain initial predictions for the second images and initial predictions for the second image portions.
[0044] In some embodiments, computing a first importance score of the importance scores includes: obtaining, for one first image portion of the first image portions, difference values by subtracting initial predictions for each second image portion from an average of the initial predictions for second images including the one first image portion; and computing the first importance score by averaging the obtained difference values.
[0045] In some embodiments, determining the degree of bias includes determining whether an average value of the computed importance scores is less than or equal to a threshold value.
[0046] In some embodiments, outputting the indication of the degree of bias includes generating a graphical output for display to a user indicating whether the average value of the computed importance scores is less than or equal to the threshold value.
[0047] In some embodiments, the method further includes generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, a retrained machine learning model using the second data set.
[0048] In some embodiments, the techniques described herein relate to a system, including: at least one processor; and at least one non-transitory computer-readable medium storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method of determining a degree of bias in a trained machine learning model, wherein the trained machine learning model was trained using a first data set including first images, the method including: obtaining initial predictions using the trained machine learning model and a second data set including second images; computing importancescores using the obtained initial predictions; determining, based on the computed importance scores, a degree of bias in the trained machine learning model; and outputting an indication of the degree of bias in the trained machine learning model.
[0049] In some embodiments, generating the second data set includes generating the second images by: segmenting images of the first images to obtain first image portions and second image portions; and generating the second images using the first image portions and the second image portions.
[0050] In some embodiments, generating the second images using the first image portions and the second image portions includes using foreground image portions and background image portions.
[0051] In some embodiments, generating the second data set further including applying a transformation to one or more of the first image portions and / or the second image portions before generating the second images.
[0052] In some embodiments, applying the transformation includes applying one or more of a blurring transformation, rotation, adversarial transformation, textural transformation, and / or color transformation to one or more of the first image portions and / or the second image portions.
[0053] In some embodiments, obtaining the initial predictions includes providing the second images and the second image portions as input to the trained machine learning model to obtain initial predictions for the second images and initial predictions for the second image portions.
[0054] In some embodiments, computing a first importance score of the importance scores includes: obtaining, for one first image portion of the first image portions, difference values by subtracting initial predictions for each second image portion from an average of the initial predictions for second images including the one first image portion; and computing the first importance score by averaging the obtained difference values.
[0055] In some embodiments, determining the degree of bias includes determining whether an average value of the computed importance scores is less than or equal to a threshold value.
[0056] In some embodiments, outputting the indication of the degree of bias includes generating a graphical output for display to a user indicating whether the average value of the computed importance scores is less than or equal to the threshold value.
[0057] In some embodiments, the techniques further include generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, a retrained machine learning model using the second data set.
[0058] In some embodiments, the techniques described herein relate to a method of performing analysis of one or more medical images using a trained machine learning model with mitigated biases, the method including: obtaining the one or more medical images; assigning, using the trained machine learning model, images of the one or more medical images to one or more classifications, each of the one or more classifications representing a prediction related to a human health status; and outputting the assigned one or more classifications obtained using the trained machine learning model, wherein: the trained machine learning model is trained using a training data set including second medical images including images of diseased tissue and images of healthy tissue, the training including: training, using the second medical images, a machine learning model to obtain an initial trained machine learning model; generating a second training data set including altered medical images by: altering the images of diseased tissue by applying one or more transformations; and generating the altered medical images by combining a transformed image of diseased tissue with an image of healthy tissue of the images of healthy tissue; obtaining initial predictions using the initial trained machine learning model and the second training data set; computing importance scores using the obtained initial predictions; determining whether an average value of the importance scores is less than or equal to a threshold value; and generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, the trained machine learning model using the second training data set.
[0059] In some embodiments, altering the images of diseased tissue includes applying a transformation to reduce a size of the image portions.
[0060] In some embodiments, altering the image portions of diseased tissue includes applying a transformation to alter an appearance of the diseased tissue such that the diseased tissue appears to be at an earlier stage in disease progression.
[0061] In some embodiments, the second medical images include radiographic images, magnetic resonance (MR) images, computed tomography (CT) images, computed axial tomography (CAT) images, ultrasound images, and / or positron emission tomography (PET) images.
[0062] In some embodiments, the second medical images include whole slide images (WSIs). In some embodiments, the second medical images include image tiles obtained from one or more WSIs.
[0063] In some embodiments, the images of diseased tissue include segmented image portions of diseased tissue.
[0064] In some embodiments, the techniques described herein relate to at least one non- transitory computer-readable medium storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method of performing analysis of one or more medical images using a trained machine learning model with mitigated biases, the method including: obtaining the one or more medical images; assigning, using the trained machine learning model, images of the one or more medical images to one or more classifications, each of the one or more classifications representing a prediction related to a human health status; and outputting the assigned one or more classifications obtained using the trained machine learning model, wherein: the trained machine learning model is trained using a training data set including second medical images including images of diseased tissue and images of healthy tissue, the training including: training, using the second medical images, a machine learning model to obtain an initial trained machine learning model; generating a second training data set including altered medical images by: altering the images of diseased tissue by applying one or more transformations; and generating the altered medical images by combining a transformed image of diseased tissue with an image of healthy tissue of the images of healthy tissue; obtaining initial predictions using the initial trained machine learning model and the second training data set; computing importance scores using the obtained initial predictions; determining whether an average value of the importance scores is less than or equal to a threshold value; and generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, the trained machine learning model using the second training data set.
[0065] In some embodiments, altering the images of diseased tissue includes applying a transformation to reduce a size of the image portions.
[0066] In some embodiments, altering the image portions of diseased tissue includes applying a transformation to alter an appearance of the diseased tissue such that the diseased tissue appears to be at an earlier stage in disease progression.
[0067] In some embodiments, the second medical images include radiographic images, magnetic resonance (MR) images, computed tomography (CT) images, computed axial tomography (CAT) images, ultrasound images, and / or positron emission tomography (PET) images.
[0068] In some embodiments, the second medical images include whole slide images (WSIs). In some embodiments, the second medical images include image tiles obtained from one or more WSIs.
[0069] In some embodiments, the images of diseased tissue include segmented image portions of diseased tissue.
[0070] In some embodiments, the techniques described herein relate to a medical image analysis system, including: at least one processor; and at least one non-transitory computer- readable medium storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method of performing analysis of one or more medical images using a trained machine learning model with mitigated biases, the method including: obtaining the one or more medical images; assigning, using the trained machine learning model, images of the one or more medical images to one or more classifications, each of the one or more classifications representing a prediction related to a human health status; and outputting the assigned one or more classifications obtained using the trained machine learning model, wherein: the trained machine learning model is trained using a training data set including second medical images including images of diseased tissue and images of healthy tissue, the training including: training, using the second medical images, a machine learning model to obtain an initial trained machine learning model; generating a second training data set including altered medical images by: altering the images of diseased tissue by applying one or more transformations; and generating the altered medical images by combining a transformed image of diseased tissue with an image of healthy tissue of the images of healthy tissue; obtaining initial predictions using the initial trained machine learning model and the second training data set; computing importance scores using the obtained initial predictions; determining whether an average value of the importance scores is less than or equal to a threshold value; and generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, the trained machine learning model using the second training data set.
[0071] In some embodiments, altering the images of diseased tissue includes applying a transformation to reduce a size of the image portions.
[0072] In some embodiments, altering the image portions of diseased tissue includes applying a transformation to alter an appearance of the diseased tissue such that the diseased tissue appears to be at an earlier stage in disease progression.
[0073] In some embodiments, the second medical images include radiographic images, magnetic resonance (MR) images, computed tomography (CT) images, computed axial tomography (CAT) images, ultrasound images, and / or positron emission tomography (PET) images.
[0074] In some embodiments, the second medical images include whole slide images (WSIs). In some embodiments, the second medical images include image tiles obtained from one or more WSIs.
[0075] In some embodiments, the images of diseased tissue include segmented image portions of diseased tissue.BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Various embodiments and embodiments of the disclosed technology will be described with reference to the following figures. It should be appreciated that the figures are not necessarily drawn to scale.
[0077] FIG. 1 is a diagram illustrating a process of training a machine learning model to perform image analysis with mitigated biases, in accordance with some embodiments of the technology described herein.
[0078] FIG. 2 is a diagram illustrating a process of constructing a training data set of composite images, in accordance with some embodiments of the technology described herein.
[0079] FIG. 3 is a diagram illustrating example composite images having different background features, in accordance with some embodiments of the technology described herein.
[0080] FIG. 4 is a diagram illustrating calculated average importance scores of foreground features as analyzed by different machine learning models, in accordance with some embodiments of the technology described herein.
[0081] FIG. 5 is a diagram illustrating a heatmap of average importance scores for different machine learning models and different image subjects, in accordance with some embodiments of the technology described herein.
[0082] FIG. 6 is a diagram illustrating average importance scores for different machine learning models, image perturbations, and different image subjects, in accordance with some embodiments of the technology described herein.
[0083] FIG. 7A is a diagram illustrating average importance scores for different foreground images of an airplane when evaluated by a ViT-L network, in accordance with some embodiments of the technology described herein.
[0084] FIG. 7B is a diagram illustrating average importance scores for different foreground images of a horse when evaluated by a ViT-L network, in accordance with some embodiments of the technology described herein.
[0085] FIG. 8A is a diagram illustrating average importance scores for different background images paired with a foreground image of an airplane when evaluated by a ViT-L network, in accordance with some embodiments of the technology described herein.
[0086] FIG. 8B is a diagram illustrating average importance scores for different background images paired with a foreground image of a horse when evaluated by a ViT-L network, in accordance with some embodiments of the technology described herein.
[0087] FIG. 8C is a diagram illustrating average importance scores for a given background across all foregrounds of an airplane with different transformations applied to the background, in accordance with some embodiments of the technology described herein.
[0088] FIG. 8D is a diagram illustrating average importance scores for a given background across all foregrounds of a horse with different transformations applied to the background, in accordance with some embodiments of the technology described herein.
[0089] FIG. 9 is a flowchart of an illustrative process 900 for performing image classification using a trained machine learning model, in accordance with some embodiments of the technology described herein.
[0090] FIG. 10 is a flowchart of an illustrative process 1000 for training a machine learning model with mitigated biases, in accordance with some embodiments of the technology described herein.
[0091] FIG. 11A is a diagram illustrating a process of training a machine learning model to perform analysis of medical images, in accordance with some embodiments of the technology described herein.
[0092] FIG. 1 IB is a diagram illustrating a process of analyzing biases in a trained machine learning model trained to perform analysis of medical images, in accordance with some embodiments of the technology described herein.
[0093] FIG. 11C is a diagram illustrating a process of training a machine learning model using a training data set configured to mitigate biases in the trained machine learning model, in accordance with some embodiments of the technology described herein.
[0094] FIG. 12A is an illustrative whole slide image (WSI).
[0095] FIG. 12B is an illustrative tiling of a WSI.
[0096] FIG. 13 is a diagram illustrating a process of training a machine learning model to analyze tiled WSIs and mitigating biases in the trained machine learning model, in accordance with some embodiments of the technology described herein.
[0097] FIG. 14 is a diagram of an illustrative computer system on which embodiments described herein may be implemented.DETAILED DESCRIPTION
[0098] Described herein are techniques for mitigating bias in trained machine learning models configured to perform image analysis tasks. For example, described herein are techniques for mitigating bias and improving accuracy in trained machine learning models configured to analyze medical and / or pathology images to assist in diagnostic and predictive tasks. The techniques described herein include generating a set of composite or transformed images for use in quantitatively assessing a trained machine learning model’s bias. Using the trained machine learning model and the generated set of composite or transformed images, initial predictions are obtained and used to calculate importance scores which quantify the weight the trained machine learning model places on features in making predictions. Based on an average value of the calculated importance scores (e.g., whether the average of the calculated importance scores is greater than or less than a threshold value), the degree of bias of the trained machine learning model may be assessed. To mitigate bias, the base machine learning model (e.g., an untrained machine learning model) or the already-trained machine learning model may be trained using the generated composite or transformed images. This process may be iterated, with new composite and / or transformed images being generated and used to assess the trained machine learning model’s level of bias.
[0099] Machine learning models include a variety of algorithms, including, as non-limiting examples, support-vector machines (SVMs), Bayesian networks, and neural networks (e.g., deep neural networks (DNNs), convolutional neural networks (CNNs), artificial neural networks (ANNs), etc.). A machine learning model may be trained to accomplish a specific task, including image analysis or computer vision tasks. The training of machine learning models may take a number of forms, including supervised and unsupervised learning. In the case of supervised learning, training data consisting of training examples is provided to the machine learning model. For example, to train a machine learning model to complete an image classification task, the machine learning model may be presented with training data consisting of pairs of images and classification labels. The machine learning model may then compare its classifications of the images with the provided labels to provide feedback to the model parameters during training.
[0100] In the case of a machine learning model trained using supervised learning, the training data has a significant impact on the performance of the trained machine learning model. However, once trained, a machine learning model functions much like a black box, and it is not well understood how trained machine learning models make certain predictions or whata trained machine learning model may consider in making a prediction. While machine learning models are powerful, they have a tendency to learn low-level features and biases present in training data. The “black box” nature of machine learning models further makes it difficult to assess whether the trained machine learning model has learned biases that affect its performance.
[0101] Biases in a trained machine learning model are a concern not only because they affect the performance of a model but also because their biases may result in biased outcomes affecting real people. For example, machine learning models have been previously trained to perform histopathology and identify certain diseases in whole slide images (WSIs). However, some of these trained models have been found to weight human-made markings on slides heavily in making a diagnosis because, in the training data set, the WSIs of diseased tissue included similar markings while the WSIs of healthy tissue did not include such markings. Such a bias could result in a misdiagnosis based on whether the pathologist had marked the slide in a certain way rather than based on actual cell characteristics. As another example, in the field of facial recognition, there is a known bias in facial recognition algorithms resulting in a reduced identification accuracy when identifying faces having darker skin tones and / or non-Westem features, potentially resulting in improper identification of an individual, which may have serious ramifications including false criminal accusations and / or arrests.
[0102] Currently, interpretability tools can identify biases present in trained machine learning models, but these tools fail to quantify the size of the bias or to disentangle multiple contributions to bias from different features in the same image type. The inventors have recognized and appreciated that biases caused by individual features can be quantified using global importance analysis (GIA) tools and the generation of targeted test images configured to assess the effects of particular image features on model biases. GIA tools, as described herein, allow for systematic comparison of how a trained machine learning model learns particular features of an image and the generation, based on the quantification of the trained machine learning model’s biases, of a trained machine learning model with reduced or mitigated biases. For example, GIA tools may provide insight into how the trained machine learning model learns foreground features such that the effects of foreground and background biases can be quantitatively measured and mitigated in a final trained machine learning model that may be deployed as a software tool.
[0103] Accordingly, the inventors have developed systems and techniques for determining a degree of bias in a trained machine learning model trained using an initial data set including images. In some embodiments, the techniques include providing a second data set to the trainedmachine learning model to obtain initial predictions (e.g., outputs) from the trained machine learning model. The images of the second data set may be crafted to provide insight regarding the trained machine learning model’s biases towards certain image features.
[0104] In some embodiments, to generate the images of the second data set, the images of the initial data set may be segmented to obtain two or more image portions. The images of the second data set may thereafter be generated by combining image portions obtained from the segmenting process. In some embodiments, a transformation may be applied to one or more of the segmented image portions before generating the second images. For example, one or more of a blurring, rotation, adversarial, textural, and / or color transformation to one or more of the segmented image portions.
[0105] As one non-limiting example, the images of the initial data set may be segmented into foregrounds and backgrounds, and images of the second data set may be generated by combining different foregrounds with different backgrounds. It should be appreciated that segmenting the images of the initial data set may be performed based on other features, alternatively or in addition to the foregrounds and backgrounds of the images, including segmenting based on the presence of certain colors, shapes, image sharpness, contrast, edges, or any other suitable image features.
[0106] In some embodiments, the initial predictions may be obtained by providing the generated images of the second data set and a set of segmented image portions as input to the initial trained machine learning model. The set of segmented image portions may correspond to those image portions exhibiting a feature of interest (e.g., background portions, portions having a certain shape, color, sharpness, etc.).
[0107] In some embodiments, the obtained initial predictions may be used to compute importance scores (e.g., using GIA). In some embodiments, an importance score for a first image portion (e.g., a foreground portion) may be computed by subtracting an initial prediction for a second image portion (e.g., a background portion) from an average of the initial predictions for second images including the first image portion. These importance scores may be averaged across all backgrounds to obtain an average importance score for that first image portion.
[0108] In some embodiments, the computed importance scores may then be used to determine a degree of bias in the trained machine learning model. The computed importance scores may be averaged to obtain an average value, and the degree of bias may be determined by comparing the average value with a threshold value. If the average value is less than or equalto the threshold value, the trained machine learning model may be deemed to be less affected by biases in its training.
[0109] In some embodiments, the techniques may further include outputting an indication of the degree of bias in the trained machine learning model based on the determination of the degree of bias in the trained machine learning model. For example, a graphical indication may be generated and output to a display to a user (e.g., on a monitor or other screen). In some embodiments, the graphical indication may be determined based on whether the average value of the computed importance scores is determined to be less than or equal to a threshold value.
[0110] The inventors have further recognized and appreciated that biases in a machine learning model may be mitigated by training or retraining the machine learning model. The retraining may be implemented using the composite images that were generated for the purposes of quantifying the machine learning model’s biases. In some embodiments, mitigating biases may include (i) retraining the trained machine learning model using the data set including composite images or (ii) training the base machine learning model from scratch using the data set including composite images. The processes described herein of evaluating a machine learning model for biases and retraining the model may be iterated until biases in the model are satisfactorily mitigated.
[0111] Following below are more detailed descriptions of various concepts related to, and embodiments of, methods and apparatus for denoising medical images. It should be appreciated that although the techniques described herein may be described in connection with denoising medical images obtained using medical devices, the techniques developed by the inventors and described herein are not so limited and may be applied to other types of images acquired using non-medical imaging devices. It should be appreciated that various embodiments described herein may be implemented in any of numerous ways. Examples of specific implementations are provided herein for illustrative purposes only. In addition, the various embodiments described in the embodiments below may be used alone or in any combination and are not limited to the combinations explicitly described herein.I. Mitigating Biases in Trained Machine Learning Models
[0112] FIG. 1 is a diagram illustrating a process 100 of training a machine learning model to perform image analysis with mitigated biases, in accordance with some embodiments of the technology described herein. The process 100 may be executed using any suitable computing device. For example, in some embodiments, the process 100 may be performed by a single computing device. As another example, in some embodiments, the process 100 may beperformed by one or more processors (e.g., co-located in a same facility. Alternatively, in some embodiments, the process 100 may be performed by one or more processors located remotely from one another (e.g., as part of a cloud computing environment).
[0113] In some embodiments, the process 100 may take as input a training data set 102 and a machine learning model 103. The training data set 102 may be, for example, a series of images paired with annotations. As a simple example, the goal may be to obtain a trained machine learning model arranged to classify images as containing a horse or not containing a horse. In such an example, the training data set 102 may include a series of images, some of which contain a horse and some of which do not contain a horse, and a series of annotations indicating which images of the series of images contain or do not contain a horse.
[0114] In some embodiments, the machine learning model 103 may be any suitable machine learning model for image analysis tasks. For example, the machine learning model 103 may be a support vector machine, a Bayesian machine learning model, and / or a clustering machine learning model. In some embodiments, the machine learning model 103 may be a neural network, including but not limited to a deep neural network (DNN), a convolutional neural network (CNN), a deep convolutional neural network (DCN), a vision transformer, a perceptron, and / or a multi-layer perceptron (MLP). In some embodiments, the machine learning model 103 may be a neural network configured to perform image classification tasks. For example, the machine learning model 103 may be one of the non-limiting selection of ResNet50, Robust ResNet50, EfficientNet-B7, Swin-L, BEiT-L, and / or ViT-L. It should be appreciated that the machine learning model 103 may be any suitable type of machine learning model, as aspects of the technology described herein are not limited in this respect.
[0115] In some embodiments, the process 100 includes training 104 of the machine learning model 103 using the training data set 102. The training 104 generates the initial trained machine learning model 106 which may be evaluated for biases in later portions of the process 100. The initial trained machine learning model 106 may include a number of parameters (e.g., thousands, millions, billions, or trillions of parameters), the values of which have been determined by the training 104 based on the training data set 102.
[0116] In some embodiments, the process 100 includes generating 108 a second training data set 110 using the training data set 102. The generating 108 may include segmenting 108a features of interest from the images of the training data set 102. Segmenting 108a may generate first image portions and second image portions that have been segmented out from the images of the training data set 102. For example, segmenting 108a may include segmenting out foreground image portions and background image portions from the images of the training dataset 102. As another example, segmenting 108a may include segmenting out regions bounded by specified shapes, regions including certain colors, regions including certain pixel patterns, and / or regions including any particular feature of interest. In some embodiments, the segmenting 108a may be optional, as the training data set 102 may include pre-segmented images or portions of images that have been extracted from whole images. An example of such a training data set is described in connection with Example 1 herein.
[0117] In some embodiments, generating 108 may also include generating 108b composite images using the segmented features obtained by segmenting 108a. In some embodiments, generating 108b may include combining one or more of the first image portions and the second image portions to obtain a composite image. In the example where segmenting 108a produces foreground and background image portions, generating 108b may include combining a foreground image portion from a first image of the training data set 102 and a background image portion from a second image of the training data set 102 to obtain a new composite image.
[0118] FIG. 2 is a diagram illustrating a process 200 of constructing a second training data set (e.g., second training data set 110 of FIG. 1) of composite images, in accordance with some embodiments of the technology described herein. From segmenting 108a, a database 202 of foreground image portions and a database 204 of background image portions may be obtained. To generate composite image 206, a foreground image portion 203 may be selected from the database 202 and a background image portion 205 may be selected from the database 204. The foreground image portion 203 may then be overlaid on top of the background image portion 205 to generate the composite image 206. By generating this second training data set including images with unique combinations of foregrounds and backgrounds, the effects of the combinations on how the initial trained machine learning model 106 makes predictions may be quantified.
[0119] Returning to FIG. 1, generating 108b of the composite images may further include applying one or more transformations to one or both of the first image portions and / or the second image portions obtained from segmenting 108a. The transformation(s) may be applied to the first and / or second image portions prior to generating the composite images. In some embodiments, applying the transformation(s) to the first and / or second image portions includes applying one or more of a blurring transformation, a sharpening transformation, a rotation, a skewing transformation, an adversarial transformation, a textural transformation, and / or a color transformation to one or more of the first image portions and / or the second image portions. As an example, and in the context of the example of FIG. 2, a transformation may be applied toone or both of the foreground image portion 203 and / or the background image portion 205 prior to generating the composite image 206.
[0120] In generating 108b the composite images, selection of the second image portions may affect the accuracy of the calculated importance scores. The high dimensionality of image data makes it computationally intractable to sample all possible backgrounds that could be observed in practice (e.g., image 302). Moreover, many backgrounds are not that meaningful, with little probability density of being observed from the data distribution. Here, “natural” background images 304 may be used in analysis with an assumption that each background image is equiprobable irrespective of the image class. This gross assumption precludes the exact value of the effect size as a meaningful causal quantity. Nevertheless, the employment of such arbitrary background images is still useful to test how well the model responds to interventions that are out-of-distribution (e.g., “unnatural” images 306). In addition, a systematic GIA analysis can enable quantitative comparisons across different models’ responses to specific interventional hypotheses.
[0121] In some embodiments, the initial trained machine learning model 106 and the second training data set 110 may be used in acquiring 112 initial predictions. Acquiring 112 the initial predictions may include providing the second training data set 110 as input to the initial trained machine learning model 106. The initial predictions may be obtained as output from the initial trained machine learning model 106. to obtain initial predictions for the composite images and initial predictions for the second image portions.
[0122] In some embodiments, some segmented features of interest 111, obtained by segmenting 108a, may also be provided as input to the initial trained machine learning model 106. The initial predictions may further include the output from the initial trained machine learning model 106 generated based on the segmented features of interest 111. In some embodiments, the segmented features of interest 111 may include only the first image portions or only the second image portions. As an example, and returning to the example of FIG. 2, the segmented features of interest 111 provided to the initial trained machine learning model 106 may include only the background image portions 205 such that the initial predictions include output generated by the initial trained machine learning model 106 using the second training data set 110 and the segmented features of interest 111.
[0123] In some embodiments, after acquiring 112 the initial predictions, the process 100 may include calculating 114 importance scores using the acquired initial predictions. To calculate an importance score for a composite image including a first segmented feature of interest, an initial prediction for the segmented feature of interest may be subtracted from aninitial prediction for the composite image including the segmented feature of interest. For example, to calculate an importance score for a composite image including a foreground image portion and a background image portion, an initial prediction made for a background image portion may be subtracted from an initial prediction for the composite image including the background image portion.
[0124] In some embodiments, an importance score may be calculated for each composite image in the second training data set 110. Thereafter, an average importance score for a first image portion may be calculated by averaging the calculated importance scores for all composite images including the first image portion. For example, to calculate an average importance score for a foreground image portion, the importance scores for each composite image including the foreground image portion and different background image portions may be averaged. It should be appreciated that average importance scores may alternatively or additionally be calculated for each second image portion by averaging the calculated importance scores associated with all composite images including a particular second image portion.
[0125] In some embodiments, calculating the average importance scores can, in mathematical notation, be described as below. Given machine learning model that takes as input an image x (G IRLxL, where L is the image size) and learns a function y(x) (G IR), the average treatment effect (ATE), which serves to control for confounding relationships between features in any given individual is given by:ATE = E[ydo( )3,(x)] - E[y(x)] where yo(^,)ydenotes the machine learning model’s output when simultaneously setting the set of pixels in x to a value , y(x) represents the machine learning model’s output for an unperturbed data sample, and E is an expectation taken over the data distribution.
[0126] The average importance (e.g., the ATE) can further be estimated by directly averaging the effect size of individual treatment effects, according to:where N is the number of samples to approximate the expectation and p(xn) is the probability of the nthimage.
[0127] In some embodiments, after calculating 114 of the importance scores, a determination 116 may be made as to whether the initial trained machine learning model 106has biases within sufficient limits or whether the initial trained machine learning model 106 has unacceptable biases that should be mitigated. To make determination 116, the calculated average importance scores may be compared to a threshold value. For example, the calculated average importance scores may be in a range from 0.0 to 1.0, with scores closer to 1.0 indicating that the initial trained machine learning model 106 strongly used features of the first image portions (e.g., the foreground image portions, in some embodiments) to make predictions. The threshold value of interest may be any suitable value within the range from 0.0 to 1.0 (e.g., 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, etc.).
[0128] In some embodiments, the determination 116 may be made by comparing a mean value or a median value of the average importance scores to the threshold value to determine a degree of bias in the initial trained machine learning model 106. In some embodiments, after determination 116, the computing system implementing the process 100 may generate an output indicating one or both of an indication of the degree of bias (e.g., an indication of the mean or median values of the average importance scores, an indication summarizing all the average importance scores as in the example of FIGs. 4 or 5) and / or an indication that the initial trained machine learning model 106 has either suitably low bias or unsuitably high bias. The output may be, for example, a graphical output to a display for viewing by a user or an output that is transmitted to another computing device and / or stored in computer memory.
[0129] In some embodiments, if the determination 116 determines that the mean value or the median value is greater than or equal to the threshold value (e.g., the model more strongly weights the first image portions in making predictions), then the determination 116 may be that the initial trained machine learning model 106 is suitably trained with acceptable biases. In such a determination, the process 100 may proceed to the end 118 and the initial trained machine learning model 106 may be deployed for use in image analysis.
[0130] In some embodiments, the determination 116 may result in a finding that the initial trained machine learning model 106 has an unacceptable level of biases. For example, the determination 116 may determine that the mean value or the median value of the average importance scores is less than the threshold value (e.g., the model more strongly weights the second image portions in making predictions). In such a determination, the process 100 may proceed to generating 120 a trained machine learning model 122 with mitigated biases.
[0131] In some embodiments, generating 120 the trained machine learning model 122 may be implemented by retraining the initial trained machine learning model 106 using the second training data set 110 as the training input. Alternatively or additionally, generating 120 thetrained machine learning model 122 may be implemented by training, from scratch, the machine learning model 103 using the second training data set 110.
[0132] In some embodiments, the process 100 may be performed iteratively. For example, after generating 120 the trained machine learning model 122, another process of generating 108 may be performed to generate a third training data set and corresponding segmented features of interest. Initial predictions may then be obtained using the third training data set, the segmented features of interest, and the trained machine learning model 122, and the initial predictions may be used to calculate importance scores for the trained machine learning model 122. If the trained machine learning model 122 is determined to be unacceptably biased, a new trained machine learning model may be generated using the third training data set. In this manner, a trained machine learning model may be generated with mitigated biases.
[0133] FIG. 4 is an example diagram illustrating calculated importance scores of foreground features for different machine learning models, in accordance with some embodiments of the technology described herein. Each box-violin describes the calculated importance scores of each foreground object in the training data set across all image categories. The horizontal line represents the median value, and the white dot represents the mean value. FIG. 5 shows a heatmap of calculated importance scores for each machine learning model and for each image category.
[0134] To generate FIGs. 4 and 5, a database of foreground and background images was used. The database included 3,436 foreground images that were generated using dense foreground segmentations. The original images were of different dimensions, so extracted foregrounds were resized to fit into a 224 x 224 box, unless otherwise stated. The dataset of background images contained 4,050 224 x224 pixel images which had the foregrounds from the original images replaced with other parts of the background. Images were processed using the Python Imaging Library.
[0135] Each foreground image was embedded so that it was centered on each background image. The importance score was computing by subtracting the initial prediction generated using the background image from the initial prediction generated using the background embedded with a foreground. An average importance score for each foreground was calculated by averaging the individual importance scores across all backgrounds for a given foreground. Average importance scores closer to 1.0 indicate that the model strongly used features of the foreground in making its predictions.
[0136] Machine learning models that accurately model true causal effects are more robust to distribution shifts and improve generalizability. To evaluate how well previously trainedmodels learn robust representations of foreground objects, which would be a causal feature for image classification, a systematic set of experiments was performed using previously trained CNNs and vision transformers. Specifically, the average importance score was calculated for each foreground object across 4,050 out-of-distribution (OOD) background images (e.g., background images that are “unnatural”). FIG. 4 shows a comparison of the average importance scores for each foreground object across all image categories for different models including ResNet50, Robust ResNet50, EfficientNet-B7, Swin-L, BEiT-L, and ViT-L. All models exhibited a wide distribution of average importance scores; some foregrounds were classified robustly across all backgrounds, while others were not.
[0137] As seen in FIG. 4, ResNet-50 and Robust ResNet-50 (trained adversarially) performed similarly, which was surprising considering that adversarially trained models are hypothesized to learn better feature representations, beyond low-level noise. EfficientNet-B7, a much larger model that is near state-of-the-art (84.3% top-1 accuracy), yielded a bimodal distribution, with many average importance scores above 0.6 or below 0.2. The vision transformers also exhibited a bimodal distribution, but more foregrounds had higher average importance scores. These results suggest that (1) vision transformers learn foregrounds more robustly than neural networks; (2) models trained on larger datasets (ImageNet-22K for Swin- L, BEiT-L, and ViT-L) are more robust than models trained on smaller datasets (ImageNet- 1K); and (3) all of the tested models are non-robust to some foregrounds.
[0138] Additional analysis was performed to better understand the robustness of foregrounds according to their image category. Specifically, for each model, the global importance scores for the foreground images were separated into semantic categories (e.g., “airplane,” “horse,” “hot air balloons,” etc.) and averaged within each category. It appears that some categories are universally robust across all models while others are not, as shown in FIG. 5. The majority of categories fall in a more variable regime where some models perform better than others. ViT-L consistently performs better with the exception of a few categories, which include “sign,” “Taj Mahal,” and “bike.” Vision transformers are especially more robust than the neural networks in the following classes: “Gymnastic,” “Airplane,” “WomanSoccer,” “Cow,” and “Christ.” Evidently, the models are not making confident predictions based on the true causal features of the foregrounds, suggesting that machine learning models may depend on features outside of the foreground to make predictions. GIA enables a direct quantification of the degree to which models use particular image features in making their predictions.
[0139] FIG. 6 is a diagram illustrating average importance scores for different machine learning models, image perturbations, and different image subjects, in accordance with someembodiments of the technology described herein. To generate FIG. 6, foreground images were perturbed in various ways, and an average importance score was measured for each perturbed foreground. Each dot represents the average importance score of all foregrounds in a semantic category, and error bars represent ± 95% confidence interval with 1,000 bootstraps. Different image categories are plotted along rows, with “Airplane” being in the top row, “Horse” being in the middle row, and “Hot Air Balloons” being in the bottom row. Different perturbations are plotted along columns. From left to right are blur, size, pixelation, and miscellaneous perturbations. Results are presented for each machine learning model. Examples of perturbations are shown along the top row of FIG. 6 for an “Airplane” example foreground. The left-most image shows foreground without any perturbations. From left-to-right, examples show blur values of 1.0 and 2.0; foreground size of 200 pixels and 100 pixels; pixelation of 192 and 64; contrast of 0.5; grayscale; and color inversion.
[0140] As seen in FIG. 6, machine learning models are sensitive to mild textural perturbations. It is known that some networks leverage texture information over shapes. To evaluate the extent that networks leverage fine texture information, varying levels of Gaussian blur were applied to the foreground images but not the background images and calculated average importance scores for comparison. Visually, the level of blur with a2= 2 yields an image where the foreground is still identifiable to humans. Nevertheless, neural networks exhibited a dose-dependent effect of blur on the average importance scores among airplanes and a slight effect in horses, but none in the highly robust hot air balloons.
[0141] Interestingly, the rank-order of models as determined in the experiments of FIGs. 4 and 5 was largely maintained, though the models seem to be converging to an average importance score of 0 as the level of blur increases. Vision transformers reached similar performance at a blur of <J2= 2 as the neural networks had without blur. Noticeably, Swin-L had the lowest robustness in the “Hot Air Balloon” class.
[0142] Another source of textural perturbation was tested by pixelating foreground images. Pixelation replaces a patch of pixels with a single color and in this way removes textural information, while introducing a small effect on the shape of the foreground (i.e., stronger levels of pixelation will appear “blockier”). It was found that pixelation decreased average importance scores with similar trends to that of blurring. While vision transformers are more robust to textural perturbations than neural networks, these results suggest that vision transformers’ predictive abilities still have a strong dependency on local texture. GIA techniques, as described herein, allow for the measurement of effect sizes of targetedperturbations to a feature of interest and the decoupling of background features from image analysis processes.
[0143] It was also found that larger foregrounds are classified more robustly. A major bias in ImageNet datasets is that the primary object of interest is often curated to be centered and magnified relative to other objects in the background. Thus, models trained on ImageNet are likely biased by the size of classified objects. To test the effect of foreground size on model predictions, the foreground images were systematically resized, and the average importance scores were calculated. Interestingly, smaller foregrounds seemed to have a stronger effect on vision transformers, particularly on airplane images, albeit their average importance scores remained higher than neural network-based models. Again, Swin-L yielded variable results for hot air balloon images, suggesting that vision transformers are not universally more robust than neural networks.
[0144] Finally, color perturbations have varying effects across foregrounds. Color-based perturbations that were tested include: contrast (factor of 0.5), conversion to gray scale, and color inversion. The machine learning models’ performance was variable across these perturbations and depended heavily on the image category.
[0145] Together, this demonstrates that the GIA techniques described herein are a powerful method for quantitatively comparing the effects of perturbations on a model’s predictions. GIA enables a direct quantification of the biases on model predictions. It should be appreciated that, the use of GIA to quantify biases is not limited to the transformations depicted in FIG. 6, but could include other biases such as translations, rotations, skews, etc., as aspects of the technology described herein are not limited in this respect.
[0146] As labels are curated by human eye, a selection bias of what is included in the dataset or a sampling bias, where a particular style and / or angle of a foreground image may be introduced to the dataset. To explore the extent of these bias further, average importance scores of foreground images for a given image category analyzed using ViT-L were ranked, as shown in the examples of FIGs. 7A and 7B. FIG. 7A illustrate ranked average importance scores for different foreground images of an airplane when evaluated by a ViT-L network, and FIG. 7B illustrates average importance scores for different foreground images of a horse when evaluated by a ViT-L network.
[0147] The airplane foreground images of FIG. 7A that were robust across all out-ofdistribution backgrounds tended to be large airliners, as illustrated across the top of FIG. 7 A. In contrast, the airplane foreground images that were least robust across out-of-distributionbackgrounds were military aircraft, smaller jets, and propeller planes as illustrated across the bottom of FIG. 7 A.
[0148] By contrast, the average importance scores of horse foregrounds decayed linearly. The properties of a horse that was robust across all out-of-distribution backgrounds was a horse that was large, brown, photographed from the side, and with a grassy field or farm in the background, as illustrated across the top of FIG. 7B. On the other hand, the least robust horses were smaller in size, photographed from different angles, and placed in different backgrounds, like beaches or tunnels. Viewpoint and lighting also seemed important in horse robustness. Many non-robust horses were photographed from the front and in low-light conditions or included glare from the sun, as illustrated across the bottom of FIG. 7B. Together, GIA enables an investigation into properties of foregrounds that are robust and non-robust to distribution shifts in the predictions made by a machine learning model.
[0149] It has generally been established that background context may be used by machine learning models when making image classification, it remains unclear to what extent models rely on background context. Using GIA, the effect size of background context can be measured for a given image category. This may be accomplished by embedding different foreground objects from the same image category across the same background and calculating the average importance score for each background. This technique marginalizes over the foreground objects instead of the background context.
[0150] FIGs. 8 A and 8B are plots ranking the average importance scores for different backgrounds paired with a same foreground image from the airplane and horse categories as evaluated by a ViT-L network, in accordance with some embodiments of the technology described herein. By ranking the calculated average importance scores for each background image for a given image category, it can be seen that monochrome backgrounds, illustrated across the tops of FIGs. 8A and 8B, seem to be a robust background across all foregrounds. In contrast, complex out-of-distribution images with complex color schemes and objects, as illustrated across the bottoms of FIGs. 8 A and 8B, serve as a very poor background context.
[0151] Next, it was asked whether an unfavorable background context can be transformed to become more like the positive background context. For airplanes, this transformation was accomplished by removing the complex features of poor backgrounds with a strong Gaussian blur and converting the color scheme to grayscale, as shown in FIG. 8C. It was found that each intervention alone significantly increased the model’s predictions, as indicated by the boxviolin plots summarizing the average importance scores when calculated for each background 1intervention. It was found that blur provides a more substantial effect size than color, and the combination of blur and monochrome provided a similar effect as blur alone.
[0152] For horses, the monochrome perturbation was to a green color shift, to represent grassy context, as illustrated in FIG. 8D. Strikingly, the intervention of green color-shift and blur were each significant and comparable on their own, but in combination their effect size increased further, providing the best context for horses as indicated by the box-violin plots summarizing the average importance scores when calculated for each background intervention.
[0153] These background manipulations show that despite yielding state-of-the-art prediction performance, the performance of ViT-L strongly depends on background features and that there is a strong class dependence on background features. Of course, learning such dependencies could be a beneficial strategy as statistical relationships may help with predictions. However, there is also an argument against such learning as it limits the ability of the model to generalize to out-of-distribution images. For instance, a dataset that contains images of a horse in a grassy background may learn this dependency and then misclassify an image containing a horse on a beach.
[0154] FIG. 9 is a flowchart of an illustrative process 900 for performing image classification using a trained machine learning model, in accordance with some embodiments of the technology described herein. The process 900 may be executed using any suitable computing device. For example, in some embodiments, the process 900 may be performed by a single computing device. As another example, in some embodiments, the process 900 may be performed by one or more processors (e.g., co-located in a same facility. Alternatively, in some embodiments, the process 900 may be performed by one or more processors located remotely from one another (e.g., as part of a cloud computing environment).
[0155] In some embodiments, the process 900 may begin with an act 902 in which a test data set including first images may be obtained. The test data set may be, for example, one or more images that a user would like the trained machine learning model to analyze (e.g., to make a prediction and / or a classification). In some embodiments, the test data set may be obtained by retrieving, using one or more computer processors, the one or more images from one or more computer-readable memories (e.g., remotely or locally located). In some embodiments, the test data set may be obtained by receiving, via transmission from another computing device, the one or more images. In some embodiments, the test data set may be obtained by generating the one or more images. For example, the one or more images may be captured by an imaging device (e.g., a digital or analog camera, a medical imaging device) or generated (e.g., by artistic creation and / or artificial generation using a computing device).
[0156] In some embodiments, after obtaining the test data set, the process 900 may proceed to act 904 in which one or more images of the test data set may be assigned to one or more classifications using a trained machine learning model. In some embodiments, the trained machine learning model may be a neural network, including but not limited to a deep neural network (DNN), a convolutional neural network (CNN), a deep convolutional neural network (DCN), a vision transformer, a perceptron, and / or a multi-layer perceptron (MLP). In some embodiments, the trained machine learning model may be a neural network configured to perform image classification tasks. For example, the trained machine learning model may be one of the non-limiting selection of ResNet50, Robust ResNet50, EfficientNet-B7, Swin-L, BEiT-L, and / or ViT-L. It should be appreciated that the trained machine learning model may be any suitable type of machine learning model, as aspects of the technology described herein are not limited in this respect.
[0157] In some embodiments, the trained machine learning model may be trained so as to mitigate biases in the trained machine learning model’s predictive abilities. For example, the trained machine learning model may be trained using global importance analysis techniques as described in connection with FIGs. 1, 10, 11A-C, and / or 13 described herein.
[0158] In some embodiments, the one or more images of the test data set may be provided as input to the trained machine learning model, and the trained machine learning model may be configured to generate an output based on each of the one or more images, the output being a classification and / or other type of prediction related to the contents of the one or more images. For example, the trained machine learning model may be configured to classify an image as containing or not containing an image subject and / or an anomaly in the image. Alternatively or additionally, the trained machine learning model may be configured to generate predictions based on the contents of the one or more images. As one non-limiting example, the trained machine learning model may be configured to generate diagnostic predictions related to a patient’s health. It should be appreciated that the applications of the trained machine learning model may be any suitable image analysis task, as aspects of the technology described herein are not limited in this respect.
[0159] In some embodiments, after assigning the one or more images to one or more classifications using the trained machine learning model, the process 900 may proceed to act 906 in which the assigned one or more classifications may be output. The assigned one or more classifications may be output using any suitable method. For example, the assigned one or more classifications may be output by being saved for subsequent access, transmitted to a recipientover a network, and / or displayed to a user of the computing system(s) implementing the process 900.
[0160] FIG. 10 is a flowchart of an illustrative process 1000 for training a machine learning model with mitigated biases, in accordance with some embodiments of the technology described herein. In some embodiments, the process 1000 may be used to train the trained machine learning model used in the process 900 described in connection with FIG. 9.
[0161] In some embodiments, the process 1000 may be executed using any suitable computing device. For example, in some embodiments, the process 1000 may be performed by a single computing device. As another example, in some embodiments, the process 1000 may be performed by one or more processors (e.g., co-located in a same facility. Alternatively, in some embodiments, the process 1000 may be performed by one or more processors located remotely from one another (e.g., as part of a cloud computing environment).
[0162] In some embodiments, the process 1000 may begin with act 1002, in which a machine learning model is trained using second images different than the images of the test data set to obtain an initial trained machine learning model. Training the machine learning model using the second images may include providing the second images to the machine learning model as input along with respective annotations and / or other information describing an ideal outcome or classification for each of the second images. For example, in training a machine learning model to identify whether a horse is present in an image, the images used for training may be paired with human-generated annotations indicating whether each of the images used for training depict a horse or do not depict a horse.
[0163] In some embodiments, after obtaining the initial trained machine learning model, the process 1000 may proceed to act 1004, in which a second training data set may be generated. Act 1004 may optionally include a first act 1004a, in which images of the second images are segmented to obtain first image portions and second image portions. For example, the first image portions may include foreground features that are segmented out from background features, which may constitute the second image portions. As another example, the first image portions may include regions bounded by specific shapes, regions including certain colors, textures, or patterns, and / or regions including any other suitable features of interest. In some embodiments, first act 1004a may be optional, as the second images may already be segmented prior to the initiation of process 1000.
[0164] In some embodiments, act 1004 may also include a second act 1004b, in which the composite images are generated using the first image portions and the second image portions. In some embodiments, the composite images may be generated by combining ones of the firstimage portions with ones of the second image portions. Prior to generating the composite image, portions of either the first image portions or the second image portions may be interpolated or otherwise filled to correct for any gaps caused by segmenting the images in act 1004a.
[0165] For example, where foreground image portions are segmented from background image portions, generating the composite images may include combining a foreground image portion with one or more background image portions. Prior to generating the composite images, the background image portions may have had regions corresponding to the foreground image portions filled in with interpolated textures, colors, etc. to generate a full background image from the background image portion.
[0166] In some embodiments, act 1004b may additionally include applying a transformation to one or more of the first image portions and / or the second image portions. The transformation(s) may be applied to the first and / or second image portions prior to generating the composite images. In some embodiments, applying the transformation(s) to the first and / or second image portions may include applying one or more of a blurring transformation, a sharpening transformation, a rotation, a skewing transformation, an adversarial transformation, a textural transformation, and / or a color transformation to one or more of the first image portions and / or the second image portions.
[0167] After generating the second training data set, in some embodiments process 1000 may proceed to act 1006 in which initial predictions are obtained using the initial trained machine learning model and the second training data set. In some embodiments, obtaining initial predictions may include providing the composite images of the second training data set and the second image portions as input to the initial trained machine learning model to obtain as output initial predictions for each of the composite images and initial predictions for each of the second image portions.
[0168] In some embodiments, after obtaining the initial predictions, the process 1000 may proceed to act 1008, in which importance scores are computed using the obtained initial predictions. In some embodiments, computing an importance score includes obtaining difference values for each first image portion. A difference value for a first image portion may be obtained by subtracting an initial prediction made using a second image portion from an initial prediction made using a composite image including the first image portion and the second image portion.
[0169] In some embodiments, an average importance score may further be computed by averaging the obtained difference values for the first image portion. That is, the averageimportance score may be obtained by averaging across all importance scores calculated for composite images including the first image portion. Alternatively, it should be appreciated that in some embodiments the average importance scores may be computed for each of the second image portions (e.g., averaging importance scores calculated for all composite images including the second image portions).
[0170] In some embodiments, after computing the average importance scores, the process 1000 may proceed to act 1010, in which it is determined whether a statistically relevant parameter related to the average importance scores is less than a threshold value. For example, a mean or median value of the average importance scores may be compared to the threshold value. If it is determined that the mean or median value of the average importance scores is less than or equal to the threshold value, it may be determined that the initial trained machine learning model has an unsuitable level of bias. If it is determined that the mean or median value of the average importance scores is greater than the threshold value, it may be determined that the initial trained machine learning model has a suitable level of bias for use in an image analysis task.
[0171] In some embodiments, the process 1000 may additionally include generating and outputting an indication of a determined degree of bias in the initial trained machine learning model. For example, the computing system implementing the process 1000 may output an indication of the statistical parameters describing the average importance scores. Alternatively or additionally, the computing system implementing the process 1000 may output for display a graphical indication of the initial trained machine learning model’s degree of bias (e.g., a graph, plot, shape, color, or other graphical indicia).
[0172] In some embodiments, in response to determining that the mean or median of the average importance scores is less than or equal to the threshold value, the process 1000 may proceed to act 1012. In act 1012, a new trained machine learning model may be generated using the second training data set. For example, the base machine learning model may be trained from scratch using the second training data set as input. Alternatively, the initial trained machine learning model may be retrained using the second training data set as additional training data input.
[0173] In some embodiments, the process 1000 may be performed iteratively. For example, after generating the trained machine learning model, another iteration of the process 1000 may be performed to generate a third training data set and corresponding importance scores for the trained machine learning model. If the trained machine learning model is determined to be unacceptably biased based on the computed importance scores, another new trained machinelearning model may be generated using the third training data set. In this manner, a trained machine learning model may be generated with mitigated biases in an iterative fashion.II. Example 1: Mitigating Biases in Predictions Based on Pathology Images
[0174] Pancreatic cancer occurs when cancerous cells develop in the pancreas. Most commonly, pancreatic cancer takes the form of pancreatic adenocarcinoma, which is a cancer that starts in the part of the pancreas that makes digestive enzymes. Pancreatic adenocarcinoma is often detected at a later stage, as physical symptoms do not usually appear at earlier stages of the disease. Unfortunately, late detection of pancreatic adenocarcinoma may reduce treatment efficacy or the ability to treat the disease, leading to drastically lower survival rates for patients whose cancer was detected in clinical stages III or IV relative to patients whose cancer was detected at stages I or II.
[0175] Pancreatic cancer is often first detected using noninvasive medical imaging techniques, including but not limited to computerized tomography (CT) scans, magnetic resonance imaging (MRI), and / or ultrasound. However, suspicious features indicating pancreatic cancer are often only detected after the tumor has grown large enough to be noticeable to a human eye. By the time tumors have reached such a size, the disease has usually progressed to a later stage. Additionally, patients with less skilled doctors or doctors who are unfamiliar with pancreatic cancer (e.g., in rural or less-developed regions) may be less likely to be diagnosed at an earlier stage of the disease.
[0176] The inventors have appreciated that the debiasing techniques described herein may be used to generate a trained machine learning model that could detect anomalies in medical images indicating an earlier stage of pancreatic cancer that human doctors might not notice. Accordingly, the inventors have developed techniques to (i) determine what image features a machine learning model is using to make a prediction of disease in a medical image and (ii) to train a machine learning model to detect earlier stages of disease than are typically detected. It should be appreciated that these techniques, while described in the context of pancreatic cancer, could be implemented for almost any disease or syndrome detectable by medical imaging and by any suitable medical imaging modality.
[0177] FIG. 11A is a diagram illustrating a process 1100a of training a machine learning model to perform analysis of medical images; FIG. 1 IB is a diagram illustrating a process 1100b of analyzing biases in a trained machine learning model trained to perform analysis of medical images; and FIG. 11C is a diagram illustrating a process 1100c of training a machine learning model using a training data set configured to mitigate biases in the trained machinelearning model, in accordance with some embodiments of the technology described herein. The processes 1100a, 1100b, and / or 1100c may be executed using any suitable computing device. For example, in some embodiments, the processes 1100a, 1100b, and / or 1100c may be performed by a single computing device. As another example, in some embodiments, the processes 1100a, 1100b, and / or 1100c may be performed by one or more processors (e.g., colocated in a same facility. Alternatively, in some embodiments, the processes 1100a, 1100b, and / or 1100c may be performed by one or more processors located remotely from one another (e.g., as part of a cloud computing environment).
[0178] In some embodiments, and as depicted in FIG. 11 A, the process 1100a may be implemented by training 1104 a machine learning model 1103 using initial training data 1102. The initial training data 1102 may be, for example, medical images 1102a and corresponding expert annotations and / or image segmentations 1102b. The medical images 1102a may be radiographic images, magnetic resonance (MR) images, computed tomography (CT) images, computed axial tomography (CAT) images, ultrasound images, and / or positron emission tomography (PET) images of a subject (e.g., a patient). The expert annotations and / or image segmentations 1102b may be human-made annotations describing the contents of each medical image and / or segmented portions of each medical image (e.g., isolating diseased or healthy portions of the medical image). For example, the medical images 1102a may be CT images obtained from the NIH Pancreas-CT dataset which include Digital Imaging and Communications in Medicine (DICOM) images with manual annotations describing types, severity, location, and other information related to pancreatic cancer present in the images. The medical images 1102a may include medical images obtained from healthy and diseased patients. Additionally or alternatively, the medical images 1102a may further include segmented portions of the CT images obtained from the NIH Pancreas-CT dataset, the segmented portions including portions corresponding to pancreatic tumors.
[0179] In some embodiments, the machine learning model 1103 may be any suitable machine learning model for image analysis tasks. For example, the machine learning model 1103 may be a support vector machine, a Bayesian machine learning model, and / or a clustering machine learning model. In some embodiments, the machine learning model 1103 may be a neural network, including but not limited to a deep neural network (DNN), a convolutional neural network (CNN), a deep convolutional neural network (DCN), a vision transformer, a perceptron, and / or a multi-layer perceptron (MLP). In some embodiments, the machine learning model 1103 may be a neural network configured to perform image classification tasks. For example, the machine learning model 1103 may be one of the non-limiting selection ofResNet50, Robust ResNet50, EfficientNet-B7, Swin-L, BEiT-L, and / or ViT-L. It should be appreciated that the machine learning model 1103 may be any suitable type of machine learning model, as aspects of the technology described herein are not limited in this respect.
[0180] In some embodiments, the process 1100 includes training 1104 of the machine learning model 1103 using the training data set 1102. The training 1104 generates the initial trained machine learning model 1106 which may be evaluated for biases in later portions of the processes 1100a, 1100b, and 1100c. The initial trained machine learning model 1106 may include a number of parameters (e.g., thousands, millions, billions, or trillions of parameters), the values of which have been determined by the training 1104 based on the training data set 1102.
[0181] In some embodiments, after obtaining the initial trained machine learning model 1106 using process 1100a, process 1100b may be initiated. As shown in FIG. 11B, process 1100b may begin by generating 1108 the training data set 1110 by transforming medical images 1102a of the initial training data 1102. The segmented images may be transformed using transformations mimicking “time reversal” effects so that the transformed segmented images mimic earlier stages of pancreatic cancer. For example, size transformations may be applied to the segmented images to make the depicted tumors smaller. Alternatively or additionally, the segmented images may be transformed in other ways (e.g., by applying blur transformations, skew transformations, rotational transformations, texture transformations, or any other suitable transformations). After transforming the segmented images of the medical images 1102a, the transformed images may be superimposed over the subset of the medical images 1102a depicting healthy pancreases to generate composite images making up the training data set 1110.
[0182] In some embodiments, after acquiring the training data set 1110, the computing device implementing process 1100b may proceed to acquiring 1112 of initial predictions. Acquiring 1112 the initial predictions may be performed by providing training data set 1110 as input to the initial trained machine learning model 1106. Additionally, the subset of the medical images 1102a depicting healthy pancreases may be provided as input to the initial trained machine learning model 1106 to acquire initial predictions.
[0183] In some embodiments, after acquiring 1112 the initial predictions, the computing device implementing process 1100b may proceed to calculating 1114 importance scores for each image of the training data set 1110. Calculating 1114 may proceed as described in connection with calculating 114 of FIG. 1 herein. In particular, an importance score for each image of the training data set 1110 may be calculated by subtracting an initial prediction for arespective image of the subset of the medical images 1102a depicting healthy pancreases from an initial prediction for the image of the training data set 1110. Thereafter, an average importance score may be calculated by averaging the importance scores associated with each image of the training data set 1110 containing a same transformed segmented image.
[0184] In some embodiments, after calculating 1114, the computing device implementing process 1100b may proceed to determining 1116 whether the average importance scores are greater than a threshold value. The value of the average importance scores relative to the threshold value may provide an indication as to whether the initial trained machine learning model 1106 is suitably unbiased for diagnostic use.
[0185] In some embodiments, determining 1116 may be implemented by comparing a mean value or a median value of the average importance scores to the threshold value to determine a degree of bias in the initial trained machine learning model 1106. In some embodiments, after determining 1116, the computing system implementing the process 1000 may generate an output indicating one or both of an indication of the degree of bias (e.g., an indication of the mean or median values of the average importance scores, an indication summarizing all the average importance scores) and / or an indication that the initial trained machine learning model 1106 has either suitably low bias or unsuitably high bias. The output may be, for example, a graphical output to a display for viewing by a user or an output that is transmitted to another computing device and / or stored in computer memory.
[0186] In some embodiments, if the outcome of determining 1116 is that the mean value or the median value is greater than or equal to the threshold value (e.g., the model more strongly weights the first image portions in making predictions), then it may be found that the initial trained machine learning model 1106 is suitably trained and has acceptable levels of bias. In such an event, the process 1100b may proceed to the end 1118 and the initial trained machine learning model 1106 may be deployed for diagnostic use.
[0187] In some embodiments, the outcome of determining 1116 may result in a finding that the initial trained machine learning model 1106 has an unacceptable level of bias. For example, the outcome of determining 1116 may be that the mean value or the median value of the average importance scores is less than the threshold value (e.g., the model more strongly weights the subset of the medical images 1102a depicting healthy pancreases in making predictions). In such a determination, the process 1100b may proceed to generating 1120 a trained machine learning model having mitigated biases. Generating 1120 may be performed using process 1100c of FIG. 11C.
[0188] In some embodiments, and as shown in FIG. 11C, training 1122 the machine learning model 1103 may be implemented by training the machine learning model 1103 from scratch using the training data set 1110. Alternatively or additionally, generating the trained machine learning model 1124 may be implemented by retraining the initial trained machine learning model 1106 using the training data set 1110.
[0189] In some embodiments, the processes 1100b and 1100c may be performed iteratively. For example, after generating 1120 the trained machine learning model 1124, another process of generating 1108 may be performed to generate a third training data set. Initial predictions may then be obtained using the third training data set and the trained machine learning model 1124, and the initial predictions may be used to calculate importance scores for the trained machine learning model 1124. If the trained machine learning model 1124 is determined to be unacceptably biased, a new trained machine learning model may be generated using the third training data set. In this manner, a trained machine learning model may be generated with mitigated biases in an iterative fashion.III. Example 2: Mitigating Biases in Predictions Made Using Whole Slide Images
[0190] Histopathology is the study of tissue at a microscopic scale to detect disease in a subject (e.g., a patient). Histopathology techniques often rely on the analysis of whole slide images (WSIs). An illustrative WSI of breast cancer tissue is shown in FIG. 12A. WSI analysis is a good candidate for the application of machine learning techniques, as WSI analysis largely relies on pattern recognition within cellular arrangements and intracellular structures.
[0191] WSIs may be analyzed using a machine learning model using so-called “multiple instance learning” techniques. In multiple instance learning, the areas of the WSI corresponding to tissue may be tiled into smaller regions (e.g., 128 x 128 pixels, 256 x 256 pixels, 512 x 512 pixels, etc.). An example of a tiled portion of the WSI of FIG. 12A is shown in FIG. 12B.
[0192] After tiling the WSI, the machine learning model may analyze each tile of the WSI individually and generate a separate prediction for each tile. The idea of multiple instance learning is that not all tiles have to be from the positive class to make an ultimate positive prediction. Rather, there just needs to be at least one tile in the positive class to make a positive prediction. Illustrative predictions that may be made based on WSI analysis includes identification of biomarkers for a particular cancer, basal cells, cancer with particular genetic mutations. Additional predictions that may be made based on WSI analysis further include arisk of relapse, likelihood of patient survival, the origin of tumor, tumor grade, and / or cancer subtype.
[0193] The inventors have appreciated that the features most relied on by machine learning models configured to perform WSI analysis using multiple instance learning may not be well understood or unbiased. Accordingly, the inventors have developed techniques for applying global importance analysis techniques as described herein to multiple instance learning WSI analysis techniques. It should be appreciated that the techniques described herein may be applied to any use of multiple instance learning, not only WSI analysis, as aspects of the technology described herein are not limited in this respect.
[0194] FIG. 13 is a diagram illustrating a process 1300 of training a machine learning model with mitigated biases to analyze tiled WSIs, in accordance with some embodiments of the technology described herein. The process 1300 may be executed using any suitable computing device. For example, in some embodiments, the process 1300 may be performed by a single computing device. As another example, in some embodiments, the process 1300 may be performed by one or more processors (e.g., co-located in a same facility. Alternatively, in some embodiments, the process 1300 may be performed by one or more processors located remotely from one another (e.g., as part of a cloud computing environment).
[0195] In some embodiments, the process 1300 may begin by generating 1302 a training data set 1304 by transforming one or more pathology images. For example, pathology images including healthy tiles 1302a (e.g., tiles with no histopathological indication of disease) and diseased tiles 1302b (e.g., tiles with a histopathological indication of disease). In some embodiments, features of interest of the diseased tiles 1302b may be segmented out from the diseased tiles 1302b and combined with tiles of the healthy tiles 1302a to generate composite tiles. Features of interest may include intracellular features (e.g., nuclei structures, membrane shapes) or intercellular features (e.g., larger arrangements of cells forming patterns).
[0196] In some embodiments, prior to generating the composite tiles to generate the training data set 1304, the segmented features of interest may be altered by applying a transformation to the segmented features. In some embodiments, applying the transformation to the segmented features includes applying one or more of a blurring transformation, a sharpening transformation, a rotation, a skewing transformation, an adversarial transformation, a textural transformation, and / or a color transformation to one or more of the segmented features.
[0197] In some embodiments, after generating 1302 the training data set 1304, the computing device implementing the process 1300 may proceed to acquiring 1308 initialpredictions using the initial trained machine learning model 1306 and the training data set 1304. The training data set 1304 may be provided to the initial trained machine learning model 1306 as input to generate initial predictions for each composite image in the training data set 1304. Additionally, healthy tiles 1302a may be provided to the initial trained machine learning model 1306 to generate initial predictions for each of the healthy tiles 1302a.
[0198] In some embodiments, after acquiring 1308 the initial predictions, the computing device implementing the process 1300 may proceed to calculating 1310 importance scores for each of the segmented features used to generate the training data set 1304. Calculating 1310 the importance scores may proceed in a similar manner as described in connection with the example of FIG. 1. For example, an importance score for each composite image of the training data set 1304 may be calculated by subtracting an initial prediction for a healthy tile 1302a from an initial prediction for a composite image of the training data set 1304 generated using the healthy tile 1302a. Thereafter, an average importance score for a segmented and / or transformed feature generated using diseased tiles 1302b may be calculated by averaging the importance scores associated with each composite image containing a same segmented and / or transformed feature.
[0199] In some embodiments, after calculating 1310, the computing device implementing process 1300 may proceed to determining 1312 whether the average importance scores are greater than a threshold value. The value of the average importance scores relative to the threshold value may provide an indication as to whether the initial trained machine learning model 1306 is suitably unbiased for diagnostic use.
[0200] Alternatively, in some embodiments, generating 1302, acquiring 1308, and calculating 1310 may proceed on a tile-by-tile basis. For example, rather than segmenting diseased tiles 1302b, generating 1302 may proceed by adding or subtracting diseased tiles 1302b to a set of healthy tiles 1302a. The training set 1304 may then include different sets of tiles comprises of healthy tiles 1302a and differing numbers of diseased tiles 1302b. These sets of tiles may then be used in acquiring 1308 to acquire initial predictions based on multiple instance learning. Thereafter, calculating 1310 may proceed by calculating importance scores for each set of tiles. For example, an importance score for each set of tiles of the training data set 1304 may be calculated by subtracting an initial prediction for a set of only healthy tiles 1302a from an initial prediction for a set of tiles including the healthy tiles 1302a and a number of diseased tiles 1302b. Thereafter, an average importance score for a set of healthy tiles 1302a may be calculated by averaging the importance scores associated with each set of tiles of the training set 1304 containing a same set of healthy tiles 1302a. It should be appreciated thattechniques of applying global importance analysis to tiled images may be used in applications other than the analysis of WSIs, as the technology described herein is not limited in that respect.
[0201] In some embodiments, determining 1312 may be implemented by comparing a mean value or a median value of the average importance scores to the threshold value to determine a degree of bias in the initial trained machine learning model 1306. In some embodiments, after determining 1312, the computing system implementing the process 1300 may generate an output indicating one or both of an indication of the degree of bias (e.g., an indication of the mean or median values of the average importance scores, an indication summarizing all the average importance scores) and / or an indication that the initial trained machine learning model 1306 has either suitably low bias or unsuitably high bias. The output may be, for example, a graphical output to a display for viewing by a user or an output that is transmitted to another computing device and / or stored in computer memory.
[0202] In some embodiments, if the outcome of determining 1312 is that the mean value or the median value is greater than or equal to the threshold value (e.g., the model more strongly weights the first image portions in making predictions), then it may be found that the initial trained machine learning model 1306 is suitably trained and has acceptable levels of bias. In such an event, the process 1300 may proceed to the end 1314 and the initial trained machine learning model 1306 may be deployed for diagnostic use.
[0203] In some embodiments, the outcome of determining 1312 may result in a finding that the initial trained machine learning model 1306 has an unacceptable level of bias. For example, the outcome of determining 1312 may be that the mean value or the median value of the average importance scores is less than the threshold value (e.g., the model more strongly weights the healthy tiles 1302a in making predictions). In such a determination, the process 1300 may proceed to generating 1316 a trained machine learning model having mitigated biases.
[0204] In some embodiments, generating 1316 a new trained machine learning model may be performed by training the base machine learning model from scratch using the training data set 1304. Alternatively or additionally, generating the trained machine learning model may be implemented by retraining the initial trained machine learning model 1306 using the training data set 1304.
[0205] In some embodiments, the process 1300 may be performed iteratively. For example, after generating 1316 a new trained machine learning model, another process of generating 1302 may be performed to generate a second training data set. Initial predictions may then be obtained using the second training data set and the trained machine learning model, and the initial predictions may be used to calculate importance scores for the trained machine learningmodel. If the trained machine learning model is determined to be unacceptably biased, a new trained machine learning model may be generated using the second training data set. In this manner, a trained machine learning model for WSI analysis may be generated with mitigated biases in an iterative fashion.IV. Conclusion
[0206] An illustrative implementation of a computer system 1400 that may be used in connection with any of the embodiments of the disclosure provided herein is shown in FIG. 14. In some embodiments, any one of the processes described herein may be implemented on and / or using the computer system 1400. The computer system 1400 may include one or more processors 1410 and one or more articles of manufacture that comprise tangible (e.g., non- transitory) computer-readable storage media (e.g., memory 1420 and one or more non-volatile storage media 1430). The processor 1410 may control writing data to and reading data from the memory 1420 and the non-volatile storage device 1430 in any suitable manner. To perform any of the functionality described herein, the processor 1410 may execute one or more processor-executable instructions stored in one or more non-transitory computer-readable storage media (e.g., the memory 1420), which may serve as non-transitory computer-readable storage media storing processor-executable instructions for execution by the processor 1410.
[0207] Having thus described several aspects and embodiments of the technology set forth in the disclosure, it is to be appreciated that various alterations, modifications, and improvements will readily occur to those skilled in the art. Such alterations, modifications, and improvements are intended to be within the spirit and scope of the technology described herein. For example, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the embodiments described herein. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, inventive embodiments may be practiced otherwise than as specifically described. In addition, any combination of two or more features, systems, articles, materials, kits, and / or methods described herein, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the scope of the present disclosure.
[0208] The above-described embodiments can be implemented in any of numerous ways. One or more aspects and embodiments of the present disclosure involving the performance of processes or methods may utilize program instructions executable by a device (e.g., a computer, a processor, or other device) to perform, or control performance of, the processes or methods. In this respect, various inventive concepts may be embodied as a computer readable storage medium (or multiple computer readable storage media) (e.g., a computer memory, one or more floppy discs, compact discs, optical discs, magnetic tapes, flash memories, circuit configurations in Field Programmable Gate Arrays or other semiconductor devices, quantum computing or information processing devices, or other tangible computer storage medium) encoded with one or more programs that, when executed on one or more computers or other processors, perform methods that implement one or more of the various embodiments described above. The computer readable medium or media can be transportable, such that the program or programs stored thereon can be loaded onto one or more different computers or other processors to implement various ones of the aspects described above. In some embodiments, computer readable media may be non-transitory media.
[0209] The terms “program” or “software” are used herein in a generic sense to refer to any type of computer code or set of computer-executable instructions that can be employed to program a computer or other processor to implement various aspects as described above. Additionally, it should be appreciated that according to one aspect, one or more computer programs that when executed perform methods of the present disclosure need not reside on a single computer or processor but may be distributed in a modular fashion among a number of different computers or processors to implement various aspects of the present disclosure.
[0210] Computer-executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically the functionality of the program modules may be combined or distributed as desired in various embodiments.
[0211] Also, data structures may be stored in computer-readable media in any suitable form. For simplicity of illustration, data structures may be shown to have fields that are related through location in the data structure. Such relationships may likewise be achieved by assigning storage for the fields with locations in a computer-readable medium that convey relationship between the fields. However, any suitable mechanism may be used to establish a relationship between information in fields of a data structure, including through the use of pointers, tags or other mechanisms that establish relationship between data elements.
[0212] When implemented in software, the software code can be executed on any suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers.
[0213] Further, it should be appreciated that a computer may be embodied in any of a number of forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer, as non-limiting examples. Additionally, a computer may be embedded in a device not generally regarded as a computer but with suitable processing capabilities, including a Personal Digital Assistant (PDA), a smartphone or any other suitable portable or fixed electronic device.
[0214] Also, a computer may have one or more input and output devices. These devices can be used, among other things, to present a user interface. Examples of output devices that can be used to provide a user interface include printers or display screens for visual presentation of output and speakers or other sound generating devices for audible presentation of output. Examples of input devices that can be used for a user interface include keyboards, and pointing devices, such as mice, touch pads, and digitizing tablets. As another example, a computer may receive input information through speech recognition or in other audible formats.
[0215] Such computers may be interconnected by one or more networks in any suitable form, including a local area network or a wide area network, such as an enterprise network, and intelligent network (IN) or the Internet. Such networks may be based on any suitable technology and may operate according to any suitable protocol and may include wireless networks, wired networks or fiber optic networks.
[0216] Also, as described, some aspects may be embodied as one or more methods. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.
[0217] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.
[0218] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”
[0219] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elementslisted with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[0220] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0221] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively.
Claims
CLAIMSWhat is claimed is:
1. A method of performing image classification using a trained machine learning model with mitigated biases, the method comprising: obtaining a test data set comprising first images; assigning, using the trained machine learning model, one or more images of the test data set to one or more classifications, wherein the trained machine learning model is trained using a training data set comprising second images, the training comprising: training, using the second images, a machine learning model to obtain an initial trained machine learning model; generating a second training data set comprising composite images by: segmenting images of the second images to obtain first image portions and second image portions; and generating the composite images using the first image portions and the second image portions; obtaining initial predictions using the initial trained machine learning model and the second training data set; computing importance scores using the obtained initial predictions; determining whether an average value of the importance scores is less than or equal to a threshold value; and generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, the trained machine learning model using the second training data set; and outputting the assigned one or more classifications obtained using the trained machine learning model.
2. The method of claim 1, wherein generating the composite images using the first image portions and the second image portions comprises using foreground image portions and background image portions.
3. The method of claim 1 or 2, wherein generating the second training data set further comprises applying a transformation to one or more of the first image portions and / or the second image portions before generating the composite images.
4. The method of claim 3, wherein applying the transformation comprises applying one or more of a blurring transformation, rotation, adversarial transformation, textural transformation, and / or color transformation to one or more of the first image portions and / or the second image portions.
5. The method of any one of claims 1-4, wherein obtaining initial predictions comprises providing the composite images and the second image portions as input to the initial trained machine learning model to obtain initial predictions for the composite images and initial predictions for the second image portions.
6. The method of claim 5, wherein computing a first importance score of the importance scores comprises: obtaining, for one first image portion of the first image portions, difference values by subtracting initial predictions for each second image portion from an average of the initial predictions for composite images including the one first image portion; and computing the first importance score by averaging the obtained difference values.
7. The method of any one of claims 1-6, wherein generating the trained machine learning model comprises generating a trained deep neural network or a trained convolutional neural network.
8. The method of any one of claims 1-7, wherein generating the trained machine learning model using the second training data set comprises either (i) retraining the initial trained machine learning model using the second training data set or (ii) training the machine learning model using the second training data set.
9. At least one non-transitory computer-readable medium storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method of performing image classification using a trained machine learning model with mitigated biases, the method comprising: obtaining a test data set comprising first images;assigning, using the trained machine learning model, one or more images of the test data set to one or more classifications, wherein the trained machine learning model is trained using a training data set comprising second images, the training comprising: training, using the second images, a machine learning model to obtain an initial trained machine learning model; generating a second training data set comprising composite images by: segmenting images of the second images to obtain first image portions and second image portions; and generating the composite images using the first image portions and the second image portions; obtaining initial predictions using the initial trained machine learning model and the second training data set; computing importance scores using the obtained initial predictions; determining whether an average value of the importance scores is less than or equal to a threshold value; and generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, the trained machine learning model using the second training data set; and outputting the assigned one or more classifications obtained using the trained machine learning model.
10. The at least one non-transitory computer-readable medium of claim 9, wherein generating the composite images using the first image portions and the second image portions comprises using foreground image portions and background image portions.
11. The at least one non-transitory computer-readable medium of claim 9 or 10, wherein generating the second training data set further comprises applying a transformation to one or more of the first image portions and / or the second image portions before generating the composite images.
12. The at least one non-transitory computer-readable medium of claim 11, wherein applying the transformation comprises applying one or more of a blurring transformation, rotation, adversarial transformation, textural transformation, and / or color transformation to one or more of the first image portions and / or the second image portions.
13. The at least one non-transitory computer-readable medium of any one of claims 9-12, wherein obtaining initial predictions comprises providing the composite images and the second image portions as input to the initial trained machine learning model to obtain initial predictions for the composite images and initial predictions for the second image portions.
14. The at least one non-transitory computer-readable medium of claim 13, wherein computing a first importance score of the importance scores comprises: obtaining, for one first image portion of the first image portions, difference values by subtracting initial predictions for each second image portion from an average of the initial predictions for composite images including the one first image portion; and computing the first importance score by averaging the obtained difference values.
15. The at least one non-transitory computer-readable medium of any one of claims 9-14, wherein generating the trained machine learning model comprises generating a trained deep neural network or a trained convolutional neural network.
16. The at least one non-transitory computer-readable medium of any one of claims 9-14, wherein generating the trained machine learning model using the second training data set comprises either (i) retraining the initial trained machine learning model using the second training data set or (ii) training the machine learning model using the second training data set.
17. An image classification system, comprising: at least one processor; and at least one non-transitory computer-readable medium storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method of performing image classification using a trained machine learning model with mitigated biases, the method comprising: obtaining a test data set comprising first images; assigning, using the trained machine learning model, one or more images of the test data set to one or more classifications, wherein the trained machine learning model is trained using a training data set comprising second images, the training comprising:training, using the second images, a machine learning model to obtain an initial trained machine learning model; generating a second training data set comprising composite images by: segmenting images of the second images to obtain first image portions and second image portions; and generating the composite images using the first image portions and the second image portions; obtaining initial predictions using the initial trained machine learning model and the second training data set; computing importance scores using the obtained initial predictions; determining whether an average value of the importance scores is less than or equal to a threshold value; and generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, the trained machine learning model using the second training data set; and outputting the assigned one or more classifications obtained using the trained machine learning model.
18. The image classification system of claim 17, wherein generating the composite images using the first image portions and the second image portions comprises using foreground image portions and background image portions.
19. The image classification system of claim 17 or 18, wherein generating the second training data set further comprises applying a transformation to one or more of the first image portions and / or the second image portions before generating the composite images.
20. The image classification system of claim 19, wherein applying the transformation comprises applying one or more of a blurring transformation, rotation, adversarial transformation, textural transformation, and / or color transformation to one or more of the first image portions and / or the second image portions.
21. The image classification system of any one of claims 17-20, wherein obtaining initial predictions comprises providing the composite images and the second image portions as inputto the initial trained machine learning model to obtain initial predictions for the composite images and initial predictions for the second image portions.
22. The image classification system of claim 21, wherein computing a first importance score of the importance scores comprises: obtaining, for one first image portion of the first image portions, difference values by subtracting initial predictions for each second image portion from an average of the initial predictions for composite images including the one first image portion; and computing the first importance score by averaging the obtained difference values.
23. The image classification system of any one of claims 17-22, wherein generating the trained machine learning model comprises generating a trained deep neural network or a trained convolutional neural network.
24. The image classification system of any one of claims 17-23, wherein generating the trained machine learning model using the second training data set comprises either (i) retraining the initial trained machine learning model using the second training data set or (ii) training the machine learning model using the second training data set.
25. A method of determining a degree of bias in a trained machine learning model, wherein the trained machine learning model was trained using a first data set comprising first images, the method comprising: obtaining initial predictions using the trained machine learning model and a second data set comprising second images; computing importance scores using the obtained initial predictions; determining, based on the computed importance scores, a degree of bias in the trained machine learning model; and outputting an indication of the degree of bias in the trained machine learning model.
26. The method of claim 25, wherein generating the second data set comprises generating the second images by: segmenting images of the first images to obtain first image portions and second image portions; andgenerating the second images using the first image portions and the second image portions.
27. The method of claim 26, wherein generating the second images using the first image portions and the second image portions comprises using foreground image portions and background image portions.
28. The method of claim 26 or 27, wherein generating the second data set further comprising applying a transformation to one or more of the first image portions and / or the second image portions before generating the second images.
29. The method of claim 28, wherein applying the transformation comprises applying one or more of a blurring transformation, rotation, adversarial transformation, textural transformation, and / or color transformation to one or more of the first image portions and / or the second image portions.
30. The method of any one of claims 26-29, wherein obtaining the initial predictions comprises providing the second images and the second image portions as input to the trained machine learning model to obtain initial predictions for the second images and initial predictions for the second image portions.
31. The method of claim 30, wherein computing a first importance score of the importance scores comprises: obtaining, for one first image portion of the first image portions, difference values by subtracting initial predictions for each second image portion from an average of the initial predictions for second images including the one first image portion; and computing the first importance score by averaging the obtained difference values.
32. The method of any one of claims 25-31, wherein determining the degree of bias comprises determining whether an average value of the computed importance scores is less than or equal to a threshold value.
33. The method of claim 32, wherein outputting the indication of the degree of bias comprises generating a graphical output for display to a user indicating whether the average value of the computed importance scores is less than or equal to the threshold value.
34. The method of claim 32 or 33, further comprising generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, a retrained machine learning model using the second data set.
35. At least one non-transitory computer-readable medium storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method of determining a degree of bias in a trained machine learning model, wherein the trained machine learning model was trained using a first data set comprising first images, the method comprising: obtaining initial predictions using the trained machine learning model and a second data set comprising second images; computing importance scores using the obtained initial predictions; determining, based on the computed importance scores, a degree of bias in the trained machine learning model; and outputting an indication of the degree of bias in the trained machine learning model.
36. The at least one non-transitory computer-readable medium of claim 35, wherein generating the second data set comprises generating the second images by: segmenting images of the first images to obtain first image portions and second image portions; and generating the second images using the first image portions and the second image portions.
37. The at least one non-transitory computer-readable medium of claim 36, wherein generating the second images using the first image portions and the second image portions comprises using foreground image portions and background image portions.
38. The at least one non-transitory computer-readable medium of claim 36 or 37, wherein generating the second data set further comprising applying a transformation to one or more ofthe first image portions and / or the second image portions before generating the second images.
39. The at least one non-transitory computer-readable medium of claim 38, wherein applying the transformation comprises applying one or more of a blurring transformation, rotation, adversarial transformation, textural transformation, and / or color transformation to one or more of the first image portions and / or the second image portions.
40. The at least one non-transitory computer-readable medium of any one of claims 36- 39, wherein obtaining the initial predictions comprises providing the second images and the second image portions as input to the trained machine learning model to obtain initial predictions for the second images and initial predictions for the second image portions.
41. The at least one non-transitory computer-readable medium of claim 40, wherein computing a first importance score of the importance scores comprises: obtaining, for one first image portion of the first image portions, difference values by subtracting initial predictions for each second image portion from an average of the initial predictions for second images including the one first image portion; and computing the first importance score by averaging the obtained difference values.
42. The at least one non-transitory computer-readable medium of any one of claims 35- 41, wherein determining the degree of bias comprises determining whether an average value of the computed importance scores is less than or equal to a threshold value.
43. The at least one non-transitory computer-readable medium of claim 42, wherein outputting the indication of the degree of bias comprises generating a graphical output for display to a user indicating whether the average value of the computed importance scores is less than or equal to the threshold value.
44. The at least one non-transitory computer-readable medium of claim 42 or 43, further comprising generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, a retrained machine learning model using the second data set.
45. A system, comprising: at least one processor; and at least one non-transitory computer-readable medium storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method of determining a degree of bias in a trained machine learning model, wherein the trained machine learning model was trained using a first data set comprising first images, the method comprising: obtaining initial predictions using the trained machine learning model and a second data set comprising second images; computing importance scores using the obtained initial predictions; determining, based on the computed importance scores, a degree of bias in the trained machine learning model; and outputting an indication of the degree of bias in the trained machine learning model.
46. The system of claim 45, wherein generating the second data set comprises generating the second images by: segmenting images of the first images to obtain first image portions and second image portions; and generating the second images using the first image portions and the second image portions.
47. The system of claim 46, wherein generating the second images using the first image portions and the second image portions comprises using foreground image portions and background image portions.
48. The system of claim 46 or 47, wherein generating the second data set further comprising applying a transformation to one or more of the first image portions and / or the second image portions before generating the second images.
49. The system of claim 48, wherein applying the transformation comprises applying one or more of a blurring transformation, rotation, adversarial transformation, textural transformation, and / or color transformation to one or more of the first image portions and / or the second image portions.
50. The system of any one of claims 46-49, wherein obtaining the initial predictions comprises providing the second images and the second image portions as input to the trained machine learning model to obtain initial predictions for the second images and initial predictions for the second image portions.
51. The system of claim 50, wherein computing a first importance score of the importance scores comprises: obtaining, for one first image portion of the first image portions, difference values by subtracting initial predictions for each second image portion from an average of the initial predictions for second images including the one first image portion; and computing the first importance score by averaging the obtained difference values.
52. The system of any one of claims 45-51, wherein determining the degree of bias comprises determining whether an average value of the computed importance scores is less than or equal to a threshold value.
53. The system of claim 52, wherein outputting the indication of the degree of bias comprises generating a graphical output for display to a user indicating whether the average value of the computed importance scores is less than or equal to the threshold value.
54. The system of claim 52 or 53, further comprising generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, a retrained machine learning model using the second data set.
55. A method of performing analysis of one or more medical images using a trained machine learning model with mitigated biases, the method comprising: obtaining the one or more medical images; assigning, using the trained machine learning model, images of the one or more medical images to one or more classifications, each of the one or more classifications representing a prediction related to a human health status; and outputting the assigned one or more classifications obtained using the trained machine learning model, wherein:the trained machine learning model is trained using a training data set comprising second medical images comprising images of diseased tissue and images of healthy tissue, the training comprising: training, using the second medical images, a machine learning model to obtain an initial trained machine learning model; generating a second training data set comprising altered medical images by: altering the images of diseased tissue by applying one or more transformations; and generating the altered medical images by combining a transformed image of diseased tissue with an image of healthy tissue of the images of healthy tissue; obtaining initial predictions using the initial trained machine learning model and the second training data set; computing importance scores using the obtained initial predictions; determining whether an average value of the importance scores is less than or equal to a threshold value; and generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, the trained machine learning model using the second training data set.
56. The method of claim 55, wherein altering the images of diseased tissue comprises applying a transformation to reduce a size of the images.
57. The method of claim 55 or 56, wherein altering the images of diseased tissue comprises applying a transformation to alter an appearance of the diseased tissue such that the diseased tissue appears to be at an earlier stage in disease progression.
58. The method of any one of claims 55-57, wherein the second medical images comprise radiographic images, magnetic resonance (MR) images, computed tomography (CT) images, computed axial tomography (CAT) images, ultrasound images, and / or positron emission tomography (PET) images.
59. The method of any one of claims 55-57, wherein the second medical images comprise whole slide images (WSIs).
60. The method of claim 59, wherein the second medical images comprise image tiles obtained from one or more WSIs.
61. The method of any one of claims 55-60, wherein the images of diseased tissue comprise segmented image portions of diseased tissue.
62. At least one non-transitory computer-readable medium storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method of performing analysis of one or more medical images using a trained machine learning model with mitigated biases, the method comprising: obtaining the one or more medical images; assigning, using the trained machine learning model, images of the one or more medical images to one or more classifications, each of the one or more classifications representing a prediction related to a human health status; and outputting the assigned one or more classifications obtained using the trained machine learning model, wherein: the trained machine learning model is trained using a training data set comprising second medical images comprising images of diseased tissue and images of healthy tissue, the training comprising: training, using the second medical images, a machine learning model to obtain an initial trained machine learning model; generating a second training data set comprising altered medical images by: altering the images of diseased tissue by applying one or more transformations; and generating the altered medical images by combining a transformed image of diseased tissue with an image of healthy tissue of the images of healthy tissue; obtaining initial predictions using the initial trained machine learning model and the second training data set; computing importance scores using the obtained initial predictions;determining whether an average value of the importance scores is less than or equal to a threshold value; and generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, the trained machine learning model using the second training data set.
63. The at least one non-transitory computer-readable medium of claim 62, wherein altering the images of diseased tissue comprises applying a transformation to reduce a size of the images.
64. The at least one non-transitory computer-readable medium of claim 62 or 63, wherein altering the images of diseased tissue comprises applying a transformation to alter an appearance of the diseased tissue such that the diseased tissue appears to be at an earlier stage in disease progression.
65. The at least one non-transitory computer-readable medium of any one of claims 62- 64, wherein the second medical images comprise radiographic images, magnetic resonance (MR) images, computed tomography (CT) images, computed axial tomography (CAT) images, ultrasound images, and / or positron emission tomography (PET) images.
66. The at least one non-transitory computer-readable medium of any one of claims 62- 64, wherein the second medical images comprise whole slide images (WSIs).
67. The at least one non-transitory computer-readable medium of claim 66, wherein the second medical images comprise image tiles obtained from one or more WSIs.
68. The at least one non-transitory computer-readable medium of any one of claims 62- 67, wherein the images of diseased tissue comprise segmented image portions of diseased tissue.
69. A medical image analysis system, comprising: at least one processor; and at least one non-transitory computer-readable medium storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor toperform a method of performing analysis of one or more medical images using a trained machine learning model with mitigated biases, the method comprising: obtaining the one or more medical images; assigning, using the trained machine learning model, images of the one or more medical images to one or more classifications, each of the one or more classifications representing a prediction related to a human health status; and outputting the assigned one or more classifications obtained using the trained machine learning model, wherein: the trained machine learning model is trained using a training data set comprising second medical images comprising images of diseased tissue and images of healthy tissue, the training comprising: training, using the second medical images, a machine learning model to obtain an initial trained machine learning model; generating a second training data set comprising altered medical images by: altering the images of diseased tissue by applying one or more transformations; and generating the altered medical images by combining a transformed image of diseased tissue with an image of healthy tissue of the images of healthy tissue; obtaining initial predictions using the initial trained machine learning model and the second training data set; computing importance scores using the obtained initial predictions; determining whether an average value of the importance scores is less than or equal to a threshold value; and generating, in response to determining that the average value of the importance scores is less than or equal to the threshold value, the trained machine learning model using the second training data set.
70. The system of claim 69, wherein altering the images of diseased tissue comprises applying a transformation to reduce a size of the images.
71. The system of claim 69 or 70, wherein altering the images of diseased tissue comprises applying a transformation to alter an appearance of the diseased tissue such that the diseased tissue appears to be at an earlier stage in disease progression.
72. The system of any one of claims 69-71, wherein the second medical images comprise radiographic images, magnetic resonance (MR) images, computed tomography (CT) images, computed axial tomography (CAT) images, ultrasound images, and / or positron emission tomography (PET) images.
73. The system of any one of claims 69-71, wherein the second medical images comprise whole slide images (WSIs).
74. The system of claim 73, wherein the second medical images comprise image tiles obtained from one or more WSIs.
75. The system of any one of claims 69-74, wherein the images of diseased tissue comprise segmented image portions of diseased tissue.
Citation Information
Patent Citations
Automated lesion detection, segmentation, and longitudinal identification
US20200085382A1
System and method for feature extraction and classification on ultrasound tomography images
US20210035296A1
Active learning for inspection tool
US20220262104A1