Systems and methods for machine learning-based diagnosis and prognosis of cancer
Machine learning systems using synthetic data from GANs provide an efficient and effective approach to diagnosing and prognosing prostate cancer, addressing issues of overdiagnosis and overtreatment in current methods.
Patent Information
- Application Number
- PCT/US2024/055687
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2024-11-13
- Publication Date
- 2025-05-22
AI Technical Summary
Current methods for diagnosing and prognosing prostate cancer are invasive, costly, and often lead to overdiagnosis and overtreatment, particularly for men under active surveillance.
The development of machine learning-based systems that utilize synthetic data, generated by generative adversarial networks (GANs), to train models for diagnosing and prognosing prostate cancer, potentially reducing the need for original patient data.
These systems demonstrate comparable performance to models trained with original data, offering a more efficient and potentially less invasive approach to prostate cancer diagnosis and prognosis, while minimizing overdiagnosis.
Smart Images

Figure US2024055687_22052025_PF_FP_ABST
Abstract
Description
[0001] DESCRIPTION
[0002] SYSTEMS AND METHODS FOR MACHINE LEARNING-BASED DIAGNOSIS AND PROGNOSIS OF CANCER
[0003] CROSS-REFERENCE TO RELATED APPLICATION
[0004] This application claims the benefit of U.S. Provisional Application Serial No. 63 / 598,207, filed November 13, 2023, the disclosure of which is hereby incorporated by reference in its entirety, including all figures, tables, and drawings.
[0005] BACKGROUND
[0006] Prostate cancer (PCa) is the third most-diagnosed cancer worldwide and the fifth-leading cause of cancer-specific death in males. The pipeline of PCa diagnosis starts with detecting levels of prostate-specific antigen (PSA); if PSA levels are above the normal range, the patient is tested for 4Kscore, PC A3, or phi. Positive results on these assays prompt imaging studies (e.g., magnetic resonance imaging (MRI)) to identify potential areas of PCa. Clinicians extract biopsies for inspection by pathologists and genomic testing. This entire pipeline is common for the prognosis of PCa. Multiple follow-ups are recommended, especially for men subjected to active surveillance. Over 77% of men with localized PCa are eligible for active surveillance, and this number has increased significantly in the last 10 years, which means, if enrolled, these men are subjected to repeat follow-ups every 6 months or 1 year. This overdiagnosis is a huge contributor to issues with physical, mental, and sexual health.
[0007] BRIEF SUMMARY
[0008] In view of the issues discussed in the Background section, there is a clear need for more effective methods of diagnosing prostate cancer (PCa) and assessing prognosis to improve patient outcomes. Embodiments of the subject invention provide novel and advantageous systems and methods for machine learning (ML)-based diagnosis and / or prognosis of PCa and / or other cancers. Synthetic data (e.g., digital pathology data) can be used to train at least one ML model, which can then be used for diagnosis and / or prognosis of PCa and / or other cancers. A generative adversarial network (GAN) (e.g., a customized GAN) can be used to generate the synthetic data. Either a small number (e.g., less than 100, such as in a range of from 1 to 25) or zero original (i.e., non-synthetic) data (e.g., images, such as radical prostatectomy images) from patients in different stages of cancer can be used in the training of the at least one ML model. The at least one ML model performs as well when trained with the synthetic data as when trained with only original data. Embodiments of the subject invention overcome the problems associated with using ML models trained with original data for diagnosis and / or prognosis of PCa and / or other cancers.
[0009] In an embodiment, a system for ML-based diagnosis and / or prognosis of cancer comprises: a processor; and a machine-readable medium in operable communication with the processor and having instructions stored thereon that, when executed by the processor, perform the following steps: a) utilizing a GAN to generate synthetic data; b) training an ML model using the synthetic data to give a trained ML model; and c) using the trained ML model to aid in diagnosis and / or prognosis of cancer. The GAN can be a deep convolutional GAN (dcGAN). The synthetic data can comprise at least 100 images (e.g., at least 200 images, at least 300 images, at least 400 images, at least 500 images, at least 600 images, at least 700 images, at least 800 images, at least 900 images, or at least 1,000 images). The training of the ML model can comprise using the synthetic data and a quantity of non-synthetic images from clinical trials. The quantity of non-synthetic images can be in a range of, for example, 1 to 100, 1 to 75, 1 to 50, or 1 to 25. The training of the ML model can comprise using only the synthetic data (i.e., no non-synthetic images from clinical trials). The GAN can comprise: a generator neural network comprising 1 layers; and a discriminator neural network comprising 12 layers. The generator neural network can comprise a transpose function with batch normalization and rectified linear unit (ReLU) functions. The discriminator neural network can comprise Conv2d, LeakyReLU functions with batch normalization. The system can further comprise a display in operable communication with the processor and / or the machine-readable medium. The instructions when executed can further perform the step of displaying, on the display, the synthetic images and / or results of the trained ML model. The cancer can be, for example, PCa, though embodiments are not limited thereto.
[0010] In another embodiment, a method for ML-based diagnosis and / or prognosis of cancer comprises: a) utilizing (e.g., by a processor) a GAN to generate synthetic data; b) training (e.g., by the processor) an ML model using the synthetic data to give a trained ML model; and c) using (e.g., by the processor) the trained ML model to aid in diagnosis and / or prognosis of cancer. The GAN can be a deep convolutional GAN (dcGAN). The synthetic data can comprise at least 100 images (e.g., at least 200 images, at least 300 images, at least 400 images, at least 500 images, at least 600 images, at least 700 images, at least 800 images, at least 900 images, or at least 1,000 images). The training of the ML model can comprise using the synthetic data and a quantity of non-synthetic images from clinical trials. The quantity of non-synthetic images can be in a range of, for example, 1 to 100, 1 to 75, 1 to 50, or 1 to 25. The training of the ML model can comprise using only the synthetic data (i.e., no non-synthetic images from clinical trials). The GAN can comprise: a generator neural network comprising 13 layers; and a discriminator neural network comprising 12 layers. The generator neural network can comprise a transpose function with batch normalization and rectified linear unit (ReLU) functions. The discriminator neural network can comprise Conv2d, LeakyReLU functions with batch normalization. The method can further comprise displaying (e.g., on a display in operable communication with the processor) the synthetic images and / or results of the trained ML model and / or the machine- readable medium. The cancer can be, for example, PCa, though embodiments are not limited thereto.
[0011] BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1A shows images of output from AlexNet, ResNet50, and EfficientNet models for a given tissue image. Regions in an individual image with different Gleason scores are distinguished using different colors (gradations of colors from green to red as the Gleason score increases)
[0013] Figure IB shows a table of artificial intelligence (Al)-model-assigned Gleason scores. Figure 2 A shows a workflow for images from the cancer genome atlas (TCGA) used in developing a training database to be used in a generative adversarial network (GAN), according to an embodiment of the subject invention.
[0014] Figure 2B shows original (top six images) and synthetic (bottom six images) histology images generated for a normal prostate.
[0015] Figure 3A shows a workflow for a GAN that can be used with embodiments of the subject invention. Images can be normalized, run through the GAN, and then optionally put through quality control (QC).
[0016] Figure 3B shows original (left 18 images) and synthetic (right 18 images) images that were generated for each primary Gleason Score of 3 (top 12 images), 4 (middle 12 images), and 5 (bottom 12 images). Figure 4 A shows a workflow for how an image can be chopped and then divided by red-green-blue (RGB) color matrices.
[0017] Figure 4B shows results of a principal component analysis (PCA) for Gleason scores of 3 (left plot), 4 (middle plot), and 5 (right plot).
[0018] Figure 4C shows color variance results for Gleason scores of 3 (left two plots), 4 (middle two plots), and 5 (right two plots). The results for the real images are given in the top three plots, and the results for the synthetic images are given in the bottom three plots.
[0019] Figure 5 shows a plot of Frechet Inception Distance score (FID) versus number of epochs, showing a loss function that stabilizes over time in order to discern an optimal epoch cut-off. The (blue) curve with the higher FID at epoch 7000 is for discriminator loss, and the (red) curve with he lower FID at epoch 7000 is for generator loss.
[0020] Figure 6 shows a table with results of testing different convolutional neural networks.
[0021] Figure 7 shows a table with results for a GAN-only model and original models.
[0022] Figure 8 shows a table of similarity index values for different groups.
[0023] Figure 9A shows an illustration of a pipeline used in generating synthetic images from prostate cancer digital histology, according to an embodiment of the subject invention. Images were pre-processed by PyHist and HistoQC. Those that passed quality control (QC) were then given to a pathologist for scoring, and then cut into small patches for modeling.
[0024] Figure 9B shows original and synthetic images that were generated for each primary Gleason pattern 3, 4, and 5, respectively.
[0025] Figure 10A shows a workflow for needle biopsy images that were used in developing the training database to be used in a GAN, according to an embodiment of the subject invention. Images were normalized, then fed into the GAN, and then assessed for quality.
[0026] Figure 10B shows example original and synthetic histology images generated for prostate cancer needle biopsies.
[0027] Figure 11 A shows the distributions of spatial recurrence properties 961 (in the first 10 principal components (PCs), which contain 95% of data variability) underlying different Gleason patterns for both real and synthetic patches on radical prostatectomy (RP). The dark (purple) lines indicate the mean values of each feature, and the gray area shows the 95% confidence interval. The results indicate that while the distributions of spatial properties are closely aligned between real and synthetic images under the same Gleason pattern, they markedly differ when comparing different Gleason patterns. Figure 11B shows a comparison of spatial recurrence properties between real and synthetic on the first four PCs (contain 90% of data variability). The distributions of the four PCs are similar between real and synthetic.
[0028] Figure 12 A shows the distributions of spatial recurrence properties (in the first eight PCs, which contain 90% of data variability) underlying different Gleason patterns for both real and synthetic patches on needle biopsy (NB). The dark (purple) lines indicate the mean values of each feature, and the gray area shows the 95% confidence interval. The results indicate that while the distributions of spatial properties are closely aligned between real and synthetic images under the same Gleason pattern, they markedly differ when comparing different Gleason patterns.
[0029] Figure 12B shows a comparison of spatial recurrence properties between real and synthetic on the first four PCs (contain 84% of data variability). The distributions of the four PCs are similar between real and synthetic.
[0030] Figure 13A shows a distribution of granular features associated with Gleason pattern
[0031] 3, as identified by the SHRQA quantification and verified by pathologists.
[0032] Figure 13B shows a distribution of granular features associated with Gleason pattern
[0033] 4, as identified by the SHRQA quantification and verified by pathologists.
[0034] Figure 13C shows a distribution of granular features associated with Gleason pattern
[0035] 5, as identified by the SHRQA quantification and verified by pathologists.
[0036] Figure 14 shows plots of true positive rate versus false positive rate for RP sections and NB sections, showing cumulative improvement in accuracy through ROC curves between synthetic + original against the original dataset. In both cases, p < 0.05.
[0037] Figure 15 A shows a table of similarity index values returned from a ten-fold cross validation run for RP images.
[0038] Figure 15B shows a table of similarity index values returned from a ten-fold cross validation run for NB images.
[0039] Figure 16A shows a table of Hotelling’s T-squared two-sample test results, showing a comparison of spatial recurrence properties between real and synthetic images under different Gleason patterns for RP. The test results indicated that there is no significant difference between real and synthetic images in RP.
[0040] Figure 16B shows a table of Hotelling’s T-squared two-sample test results, showing a comparison of spatial recurrence properties between real and synthetic images under different Gleason patterns for NB. The test results indicated that there is no significant difference between real and synthetic images in NB.
[0041] DETAILED DESCRIPTION
[0042] Embodiments of the subject invention provide novel and advantageous systems and methods for machine learning (ML)-based diagnosis and / or prognosis of PCa and / or other cancers. Synthetic data (e.g., digital pathology data) can be used to train at least one ML model, which can then be used for diagnosis and / or prognosis of PCa and / or other cancers. A generative adversarial network (GAN) (e.g., a customized GAN) can be used to generate the synthetic data. Either a small number (e.g., less than 100, such as in a range of from 1 to 75, 1 to 50, or 1 to 25) or zero original (i.e., non-synthetic) data (e.g., images, such as radical prostatectomy images) from patients in different stages of cancer can be used in the training of the at least one ML model. The at least one ML model performs as well when trained with the synthetic data as when trained with only original data. Embodiments of the subject invention overcome the problems associated with using ML models trained with original data for diagnosis and / or prognosis of PCa and / or other cancers.
[0043] The use of ML for automating the steps involved in PCa diagnosis and prognosis has had significant advancements. ML algorithms are able to learn from data and make predictions or decisions based on that learning. In PCa, these algorithms can be trained on, for example, data from tissue samples, medical images, and other clinical information to identify patterns and features associated with the disease. One area where ML has been particularly useful is in the automated analysis of medical images, such as magnetic resonance imaging (MRI) scans or biopsy slides. The algorithms can identify with a high degree of accuracy prostate areas suspicious for cancer and provide a more precise diagnosis than possible with manual analysis. For example, technologies such as PathAI, HTL Limited, PaigeAI, and Deciplex focus on digital imagery to aid in diagnosis by transforming the digital histology into meaningful data which could differentiate between cancer and non-cancer areas with a decent accuracy.
[0044] However, these technologies on their own are limited regarding proactive adaptation in clinics because the available tools do not effectively take into account cancer heterogeneity. One of the reasons for this is these tools are built using original (i.e., non-synthetic) training data, which is generally derived from clinical trials. There are several limitations to using clinical trial data in digital pathology for training ML models. One major limitation is the potential for bias in the data. Clinical trial data is typically collected from a specific group of patients, which may not be representative of the general population or a specific disease stage. This can result in ML models that perform well on the training data, but poorly on new or unseen data. Another limitation is the lack of diversity in the data. Clinical trial data is often collected from a homogeneous group of patients, which may not adequately capture the variability of the disease. This can result in ML models that are unable to generalize to different subtypes of the disease or to different populations. Additionally, there may be issues with the quality and reliability of the data. Digital pathology images can be affected by various factors, such as staining techniques, slide preparation, and the resolution of the microscope. These factors can introduce noise and variability into the data, which can impact the performance of ML models. Further, there may be ethical concerns related to the use of clinical trial data in digital pathology. The data may contain sensitive information about patients, such as their medical history and personal information. This information must be protected and used in accordance with ethical guidelines to ensure the privacy and dignity of the patients.
[0045] Embodiments of the subject invention overcome all of these limitations on the accuracy of ML models by significantly limiting (or outright eliminating) the use of clinical trial data for training the artificial intelligence (AI) / ML models and instead use synthetic data. One or more GAN models (e.g., one or more customized GAN models) can be used to generate synthetic data (e.g., digital pathology data) while requiring a limited number of (or zero) original images (e.g., radical prostatectomy images) from patients (e.g., PCa patients) in different stages of disease (see, e.g., Figures 3A and 3B). The synthetic data can then be used to train at least one ML model for diagnosis and / or prognosis purposes. The synthetic data can include a large quantity (e.g., at least 200, at least 300, at least 400, at least 500, or at least 1,000) images. The Al models trained with synthetically generated data perform as well as Al models trained with the original patient data. Radical prostatectomy (RP) sections from the cancer genome atlas, in-house active surveillance trial sections (RP sections), and needle biopsy data from Radbound University Medical Center (Karolinska Institute) were used for data validation and comparison (see the examples). The conceptual and technical innovations of embodiments of the subject invention overcome important challenges in developing efficient diagnostic and prognostic ML models in the domain of digital imagery.
[0046] In the context of digital imagery, image analysis algorithms can be used for grading, classification, and identification of metastases in multiple cancer types. Related art models have limitations. For example, models like PyTorch and Tensorflow are easy to implement, but slow to run. Another important factor in model accuracy is the architecture of the neural network themselves. The more layers the algorithm runs, the more accurate it is, but the slower it runs. Al models can be trained to differentiate between non-cancer and cancer areas with a significant reliability. However, related art Al tools consider only morphological architecture of a tumor to grade the cancer. This single consideration (tumor morphology) limits the application of these tools to appropriately characterize the tumor heterogeneity, especially when considering active surveillance, which requires characterization of proximal cancer areas that have minimal differences in the tissue architecture. Embodiments of the subject invention provide AI / ML models that can analyze the tissue architecture of PCa (and other cancers) at a granular level and integrate this information with the normalized genomic signatures that are specific to primary and secondary scores of low-grade cancer regions. This allows the AI / ML models to reliably identify and score the low grade PCa. The Al model was validated using data from RP sections from an in-house active surveillance trial and needle biopsy data, and the Al annotations were compared in consideration of the resulting diagnoses performed by three pathologists (see the examples). The Al model showed a concordance of 60% to 75% (kappa value 0.60-0.75) with respect to pathologist scoring, demonstrating its reliability. In general, the Al model is highly sensitive for detecting: 1) cancer versus non-cancer areas; and 2) low grade cancer areas, including the stratification into the different Gleason grade patterns, confirming the model’s excellent cancer detection performance.
[0047] The AI / ML models of embodiments of the subject invention are unique compared to other ML models in the digital pathology domain in several respects. For example, the present Al model originally showed a proportion of false positives (15% of the tissue) that was similar to other digital pathology models. In a majority of the cases, the false positive areas were mechanical distortions, artefacts, or folds that may be difficult to classify at a macro level annotation. However, tweaking the Al model with granular level annotation and genomics hyperparameter allowed a significant reduction in the false positive rate (down to 5% of the tissue). The granular level annotation was incorporated based on the rationale that, in pathology practice, even small suspicious areas can have a prognostic value. Moreover, the robustness of the models of embodiments of the subject invention were demonstrated by using inhouse trial RP sections and needle biopsy images from publicly available resources (see the examples). Further, detailed annotations were used for both training and testing images. This allowed for training and annotating the image requiring less computational power, making the models of embodiments of the subject invention more practical for routine clinical use.
[0048] Embodiments of the subject invention provide ML models that detect low grade PCa (and other cancers) by focusing on tumor architecture and genomic signatures. The ML models can assign a Gleason grade that is similar in accuracy (e.g., within 10% or within 5%) to that given by experienced pathologists.
[0049] Some related art Al models attempt to grade diseases via digital imagery, such as MRI or digital histology section analysis. However, they all have a key limitation in that the data on which the models are primarily trained on is derived from clinical trials. Embodiments of the subject invention can be applied to multiple cancer types and allow for the minimization or elimination of use of clinical trial training data, thereby improving the digital imagery classification across multiple platforms. Embodiments can be applied to multiple disease types and allow for the use of synthetically generated data to train the ML models, which can then aid in diagnosis and / or prognosis of PCa and / or other cancer types while minimizing or eliminating the requirement for cost-inefficient clinical trials and other invasive procedures.
[0050] Embodiments of the subject invention utilize GAN models (e.g., customized GAN models, such as a deep convolutional GAN model), with either a small quantity or no original training data, to generate a large cohort of synthetic data that can be used to train the ML model(s) to perform effective diagnosis and / or prognosis of PCa and / or other cancers (e.g., grading of PCa specimens). In application, this can replace the need for extensive clinical data for training any Al model in the domain of digital imagery, allowing for more cost-effective and quicker diagnosis and prognosis.
[0051] Embodiments of the subject invention provide a focused technical solution to the focused technical problem of how to increase the accuracy of ML-aided diagnosis and / or prognosis of PCa and / or other cancers. The solution is provided by using a GAN to generate synthetic data to train ML models for diagnosis and / or prognosis of PCa and / or other cancers. Embodiments of the subject invention can improve the computer system performing the ML- aided diagnosis and / or prognosis by increasing the efficiency and efficacy of the training of the ML model(s), for example by requiring a much lower quantity of original images for training (which can free up memory and / or processor usage).
[0052] The methods and processes described herein can be embodied as code and / or data. The software code and data described herein can be stored on one or more machine-readable media (e.g., computer-readable media), which may include any device or medium that can store code and / or data for use by a computer system. When a computer system and / or processor reads and executes the code and / or data stored on a computer-readable medium, the computer system and / or processor performs the methods and processes embodied as data structures and code stored within the computer-readable storage medium.
[0053] It should be appreciated by those skilled in the art that computer-readable media include removable and non-removable structures / devices that can be used for storage of information, such as computer-readable instructions, data structures, program modules, and other data used by a computing system / environment. A computer-readable medium includes, but is not limited to, volatile memory such as random access memories (RAM, DRAM, SRAM); and nonvolatile memory such as flash memory, various read-only-memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM), and magnetic and optical storage devices (hard drives, magnetic tape, CDs, DVDs); network devices; or other media now known or later developed that are capable of storing computer- readable information / data. Computer-readable media should not be construed or interpreted to include any propagating signals. A computer-readable medium of embodiments of the subject invention can be, for example, a compact disc (CD), digital video disc (DVD), flash memory device, volatile memory, or a hard disk drive (HDD), such as an external HDD or the HDD of a computing device, though embodiments are not limited thereto. A computing device can be, for example, a laptop computer, desktop computer, server, cell phone, or tablet, though embodiments are not limited thereto.
[0054] When ranges are used herein, combinations and subcombinations of ranges (including any value or subrange contained therein) are intended to be explicitly included. When the term “about” is used herein, in conjunction with a numerical value, it is understood that the value can be in a range of 95% of the value to 105% of the value, i.e. the value can be + / - 5% of the stated value. For example, “about 1 kg” means from 0.95 kg to 1.05 kg.
[0055] A greater understanding of the embodiments of the subject invention and of their many advantages may be had from the following examples, given by way of illustration. The following examples are illustrative of some of the methods, applications, embodiments, and variants of the present invention. They are, of course, not to be considered as limiting the invention. Numerous changes and modifications can be made with respect to embodiments of the invention.
[0056] MATERIALS AND METHODS
[0057] Images were deconstructed by first taking individual areas in patch-wise fashion using the software package PYHist. Briefly, the image was taken from raw .svs format and scanned for areas that are defined as tissue regions. This was done by color definition compared to whitespace background.
[0058] Each predefined block was taken at a 512x512 pixel area, which was identified as the lowest amount of space that the pathologist graded any given image. These patches were then used as the training database from which other images were annotated. The image annotation for a given test image then was a whole image from which a sliding window approach was taken that moved through each of the given pixel windows and scored against the training database.
[0059] Samples were taken from The Cancer Genome Atlas (TCGA). Histology images from 500 individuals were taken along with 499 RNAseq data points downloaded from the Broad Firehose repository. Thirty-eight local samples were taken from the University of Miami Medical Group clinical trial. The IRB protocol was approved by the University of Miami Miller School of Medicine, Miami, FL. A large number (10,318) needle biopsies were taken from a Kaggle challenge in 2020, which came from patients in the Karolinska institute.
[0060] Three initial algorithms were selected from which to test out the convolutional network. These were defined as the major stepwise breakthroughs in Al with the AlexNET, ResNet, and Xception models. Initially, training images were taken from a single pathologist defined areas and ten random images were used as test images. Granular level annotation was used to select the model with the highest accuracy. Post model selection, the hyperparameters were turned. Here a Tree-structured Parzen Estimator was implemented to complete a sequential model optimization. In addition, tree weights were investigated by a population-based training approach. From this, the highest performing hyperparameters were used along with the appropriate tree weights that were used to define the network.
[0061] Accuracy was defined in two different ways. First, in granular accuracy the areas that were selected were started with and annotated with the exact same Gleason grade with all 3 pathologists. Gold standard granular accuracy is when the convolutional network overlaps with the agreement between all pathologists. This is defined by the number of overlapping pixels in the defined area compared with the definition between the annotated areas of the pathologist grade. Second, patient level accuracy was defined when the Al pipeline has correctly identified a patient’s Gleason grade as defined by TOGA pathology confirmed grade.
[0062] The deep convolutional GAN (dcGAN) weights were initialized randomly from a normal distribution with a mean of 0 and a standard deviation of 0.02. The generator neural network was constructed using thirteen layers including transpose function with batch normalization and rectified linear unit (ReLU) functions. The walkthrough of the generator layers is as follows:
[0063] ConvTranspose2d -> BatchNorm2d -> ReLU -> ConvTranspose2d -> BatchNorm2d -> ReLU -> ConvTranspose2d -> BatchNorm2d -> ReLU -> ConvTranspose2d -> BatchNorm2d -> ReLU -> ConvTmaspose2d -> Tanh
[0064] The discriminator neural network was constructed using twelve layers of Conv2d, Leaky ReLU functions with batch normalization. The walkthrough of the layers is as follows:
[0065] Conv2d -> LeakyReLU -> Conv2d -> BatchNorm2d -> LeakyReLU -> Con v2d -> BatchNorm2d - > LeakyReLU -> Conv2d -> BatchNorm2d -> LeakyReLU -> Conv2d -> Sigmoid
[0066] Estimation of the number of parameters used for a single run ranged from 1.3 trillion to 1 .4 trillion calculations per run to generate 1 ,000 synthetic images that were run on standard graphics processing unit (GPU) chips. The time taken for a single run was dependent on the number of GPU processors available. For a multi-threaded GPU, the time taken ranged from 1.5 hours to 5 hours for a run while for a single GPU it ranged from 3 hours to 12 hours. Time was dependent on the level of resolution desired.
[0067] Finally, the binary cross entropy (BCE) loss function was used and the Adam optimizer was implemented for both the generator and discriminator. Iteration level statistics were generated at the end of each run and saved for further analysis in matplot.
[0068] The EfficientNET baseline model was setup with the B3 function within the TensorFlow python backend. Initial testing for the B6 function was found to be too restrictive and Bl was too simple to form complex patterns so B3 was selected as the model function. In order to form the model based on the images, a GlobalMaxPooling2D layer was added after the initial base as well as a Dropout layer which would help avoid overfitting. The dropout rate was set to 0.2 after initial estimations were too high. The number of classes for prediction layer was set to 4, which represented the normal tissue group as compared to the primary Gleason patterns GS3 (Gleason score of 3), GS4, and GS5. The EfficientNET model was set up using pre-trained weights from the “imageNet” to take advantage of transfer learning to reduce analysis time.
[0069] Image augmentation was completed using the Keras ImageDataGenerator function. Image rotation range was set to 45, width shift range and height shift range was set to 0.2, and the horizontal flip was set to true for flipping the image. Fill mode was defaulted to “nearest”. Validation data was not augmented, but the images were put through a rescale function to ensure that every test image was uniform before annotation.
[0070] The Frechet Inception Distance score (FID) was implemented in custom scripts developed in house. The FID model was pre-trained using Inception V3 weights for transfer learning. In-house code was centered around the FID model and inserted into the dcGAN to be run during each iteration. Stats were reported at intervals of 1000 and graphed with inhouse python scripts.
[0071] Principal component analysis (PCA) was performed by first transforming the images into numerical arrays. Images were separated into normal and synthetic batches and then distributed by primary Gleason score.
[0072] Intensity was calculated (using the R package imgpalr and magick) as the average of the color of the entire image while keeping the matrix framework (i.e., positional arguments were retained). PCA was conducted using the general prcomp function in R and plotted results were displayed in ggplot2.
[0073] Image Preprocessing: Many factors could induce variability in tissue biopsy images and thereby influence the ability of an Al model to make accurate predictions. These include variability in staining protocols, tissue quality, section thickness, tissue folding, and the amount of tissue on the slide. TCGA images were subjected to pre-processing. All 500 images were selected and inspected for color distribution. Color distribution, or staining intensity, was calculated based on the mean value of red-green-blue (RGB) colors and normalized. Images with an RGB mean intensity value that was two standard deviations away from the total mean value of all samples were considered to be an outlier. This resulted in 21 images being discarded out of the original 500 samples. Similarly, the image pre-processing was conducted for the in-house RP sections (of which there were 34; n=34). Sections were used to calculate the mean intensity, which was then compared to each of the images. One single sample had a color intensity value that was too far from the mean of the entire group, and thus was discarded.
[0074] Needle biopsy slides were considered from Radbound University Medical Center, Karolinska Institute (n=3949). The normalization procedure for needle biopsies was different than for the RP sections considering the limited amount of tissue that can be accommodated on a needle biopsy slide. For normalization, a more robust method was adopted first to identify areas of tissue in the image and then those areas were subjected to color normalization. Two hundred random images were selected and subjected to color normalization. Fifteen images were excluded due to having a mean color intensity greater than two standard deviations from the mean value of all 200 images.
[0075] Generating Synthetic Images from Normal Prostate: Digital pathology images for the prostate were downloaded from the GTEx Portal, a repository of 25,713 images. Downloaded prostate Images (n=599) were subjected to color normalization. Post-correction and normalization, 572 images were used to train the GAN model. Here, a dc-GAN network was applied to the training images. The dc-GAN contains two neural networks where one network (called a generator) creates random images, and the other network (called a discriminator) evaluates the image for realness based on a set of real training images. The time taken for a typical GAN run was benchmarked at an average of 2.5 hours per run, running on a single NVIDIA Al 000 GPU chip. Each GAN run yielded 1000 synthetic images that were chosen when the generator and discriminator loss functions were equal, which was between 25,544 to 62,426 epochs (see Figure 5). In order to ensure the correct patch size was taken, random synthetic prostate images generated from the 128x128 and 256x256 pixel patches were subjected to model based quality control (QC) assessment. A manual board certified pathologist inspection of these images was centered around the sharpness of the image and resolution (see also Figure 2). The response of the pathologists suggested a significant 80% approval for the data to be sufficiently good for both patch sizes, but noted that distinguishable features could not be evaluated because the images were too granular in nature. In order to further analyze the QC, a simple classification scalable vector machine model was created where synthetic images were classified into correct Gleason groups. Further, the synthetic images from the GAN were subjected to index calculations to evaluate the similarity between the images. This was computed via the relative inception score (RIS). Inception distance (similarity between images) was also calculated and was found to be 17.2 + / - 0.15, which is well above the threshold of 5 typically used in facial recognition GANs, therefore showing that the synthetic images were of sufficient quality. Together, the data suggested that high-quality synthetic images from prostate digital pathology sections were generated using the GAN.
[0076] Generating Synthetic Images from Prostate Cancer: Digital pathology slides were downloaded from 500 patients from the TCGA PRAD study and separated into pseudo-cohorts according to the primary and secondary Gleason scores. Additionally, 34 digital histology images were obtained from men subjected to active surveillance (MD select trial at the University of Miami). The tissue images were inspected using the HistoQC software to remove samples of low quality.
[0077] A random selection of 108 sections was then given to two board-certified pathologists for Gleason pattern annotation. In order to build the gold standard for training data, both pathologists were required to match each other and TCGA scoring as well as exact overlapping annotation of tumor region, which left 25 images. Image patches from these 25 individuals were then sliced using PyHIST software in dimensions of 96x96 and 256x256 pixel length. Areas of annotation obtained from the pathologists were overlaid with these sections. Image patches were required to contain at least 75% tissue image and no more than 25% whitespace. The smallest area of overlap between pathologists was 96 pixels, and this was considered to be the smallest size considered for image analysis. A total of 219 patches were generated from 25 images, which were subjected to enhancement by augmentation to give a total of 876 patches. The generator in the dcGAN networks was used to create random synthetic images from these 876 patches to create an effective training dataset. The resulting image patches (Gleason scores: normal (n=175), GS3 (n=726), GS4 (n=1029), and GS5 (n=152)) were subjected to an Adam optimization algorithm to determine the optimal iteration value. From this an optimal iteration of 14,000 was selected (see Figure 5). Upon completion of running the dcGAN, a board-certified pathologist was provided with a mixed pool of real and synthetically generated images.
[0078] The quality control approval rates were 80% for the synthetic images, which supports the usage and applications of synthetic images in training data. EXAMPLE 1 - Network Architecture and Model Selection
[0079] In order to identify the most suitable baseline model for the study, the AlexNet network architecture was first evaluated. This was the first large-scale convolutional neural network to be defined, featuring a series of convolutional pooling layers and an interconnected output layer. In order to test the accuracy of this model, 20 histology images were randomly selected from the Prostate Adenocarcinoma dataset (obtained from the Cancer Genome Atlas), which included images from 500 patients, with four images from each of the following scores: normal; GS6; GS7; GS8; and GS9. These images were graded using default parameters without pretraining the model and the results were compared against TCGA scoring and the assessments of two independent pathologists (considered the gold standard).
[0080] The first CNN model evaluated was AlexNet. AlexNet operates using deep convolutional layers and activation functions to parse and learn features from images for classification. Its efficiency stems from techniques like max-pooling, ReLU activation, and dropout for regularization, optimizing its learning process in neural network training. The AlexNet model demonstrated an accuracy rate of 55% in alignment with the gold standard.
[0081] Next, an improved model from AlexNet called the ResNet model was evaluated. ResNet incorporates residual networks based on skip connections within the network, which significantly improves upon AlexNet. ResNet was tested across the same 10 images and it was found that it, too, achieved an accuracy rate of approximately 55% in alignment with the gold standard.
[0082] The Xception model was then assessed. Based on the Inception model, this improves traditional network architectures such as AlexNet. Traditional models attempt to evaluate space and depth in their convolutional layers, whereas Xception separates the two into different dimensions. Xception achieved an accuracy rate of 60% in alignment with the gold standard, outperforming the other two models.
[0083] Last, the EfficientNet model was tested. EfficientNet utilizes a novel scaling method that employs a compound coefficient to scale up a CNN model in a more structured manner. This improves the model’s performance and efficiency while requiring minimal layers to train on existing images. EfficientNet achieved an accuracy rate of 65% in alignment with the gold standard, outperforming the other three models. Therefore, EfficientNet was selected as a base convolutional neural network (CNN) (see also the table in Figure 6). EXAMPLE 2 - Convolutional Network Classification
[0084] The applicability of the high quality synthetic images generated by the dcGAN were evaluated as an effective training dataset. The EfficientNet CNN model was trained with the generated synthetic data Gleason (normal (n=175), GS3 (n=726), GS4 (n=1029), and GS5 (n=152)) and used for assigning the Gleason grading to the test images in the TCGA (n=475). Additionally, a probability score (the highest probability that a given patch to be of a specific Gleason score) was assigned to the Gleason grades. In parallel to this, the EfficientNet CNN model was trained using the image patches derived from original digital pathology (normal (n==321), GS3 (n=2237), GS4 (n=3829), and GS5 (n=523). Probability scores were assigned to the Gleason grades similar to the previous step. Next, the benchmarking was performed using the probability scores between the EfficientNet model trained with the synthetic patches (GAN-enhanced) and with patches from the original images (non-GAN enhanced). These probability scores were compared to estimate the difference in the output (Gleason scoring).
[0085] Referring to the table in Figure 7, results showed that in normal tissue, there was no significant difference in scoring accuracy of the Al model when using the images either from GAN or from the original source. In the other Gleason scores, minimal non-significant differences were found between the models: GS3 (0.47 vs 0.53 p=0.115), GS4 (0.51 vs 0.55 p=0.221), GS5 (0.55 vs 0.57 p = 0.870). Together, the results highlighted that the quality of synthetically generated images are high enough to compete with the original images as training data.
[0086] EXAMPLE 3 - Enhancing Model Performance
[0087] Consideration was given to determine the optimal number of synthetic images that can be used to train the classification model to achieve the highest possible accuracy while ensuring that the dcGAN model was not overfitted by the similarity in the synthetic images. The output of EfficientNet model was optimized using synthetic data; it was ensured that the dcGAN model was not overfitted by the similarity in the synthetic images.
[0088] The synthetically generated images were subjected to similarity indexing (the closer to 1 the similarity index is, the better the alignment rate is). In this process, images were generated from primary Gleason grading in batches of 10,000, 50,000, and 100,000. Next, in order to confirm if through this optimization the maximum unique synthetic images that could be used to train the model without compromising accuracy can be found, each image batch was combined with the original images and run through the CNN. The output images were compared for model overfitting using the FID (a metric that calculates the distance between feature vectors calculated for real and generated images). Results suggested that with the 10,000 batch size, there was an increase in the accuracy of predicting the primary scores from 33% to 63% for GS3, 35% to 61% for GS4, and 41% to 68% for GS5. For the 50,000 batch size, there was an increase in the accuracy from 33% to 70% for GS3, 35% to 65% for GS4, and 41% to 71% for GS5. For the 100,000 batch size, there was an increase in the accuracy from 33% to 65% for GS3, 35% to 62% for GS4, and 41% to 67% for GS5. From this, it can be deduced that the sample complexity in the images topped out at about 50,000 and, at best, leveled off at that point. Thus, a maximum of 50,000 synthetic images were considered in the model.
[0089] Moreover, in order to avoid the possibility of overfitting the model due to little to no difference within the synthetically generated images within the same batch, a cross-validation approach to measuring inherent variability was applied. For this, all the image tiles (real and synthetic) were lined up randomly, and the data was divided into ten subgroups. These were expected to define the subgroups used for training and testing in a randomized way to allow for evaluation of mean square error, misclassification error rate, and confidence intervals. Here, 10-fold cross-validation was performed. Referring to the table in Figure 8, the results suggested that the similarity index values range from 0.8 to 1.0, while the accuracy remained consistent in GS3 (0.68-0.71), GS4 (0.63-0.70), and GS5 (0.66-0.73). This suggests that the model did not produce any significant outlier images that could confound accuracy in the EfficientNet CNN and reinforces the quality and variability of generated synthetic images.
[0090] Consideration was then give to the extent of technical variations that may exist in the synthetically regenerated images compared to the original images. This is important because each Gleason pattern has specific morphological characteristics that ideally is expected to be emulated by the synthetically generated images. For instance, Gleason pattern 3 includes well- formed glands of various sizes including branching glands (glands are discrete units that one could draw a circle around as long as they are not fused together). Gleason pattern 4 includes poorly-formed, fused, and cribriform glands (a confluent sheet of contiguous malignant epithelial cells with multiple glandular lumina that are easily visible at low power). Gleason pattern 5 includes sheets of tumor, individual, and cords of cells. Nests of cells with vague micro-acinar or only occasional gland space formation are also considered pattern 5. In order to delineate the technical variations between original and synthetically regenerated images, the original images were divided into primary Gleason patterns 3, 4, and 5. In order to look at the individual gland formation, the images were subjected to a PC A to study the color distribution between the GS3 / 4 / 5. The PCA allowed identification of features, and cellular morphology inherited within each of the Gleason patterns as discussed above. Each image was represented as a large multi-layered 3D matrix. The x and y coordinates in the matrix represented the location on the image, and the z coordinate was defined by the color intensity. For color intensity, three different numerical values were considered (each for red, green, and blue intensity) to represent one single point, or pixel, on the image; Figure 4 shows the illustration of color distribution. For Gleason pattern 3, the morphological appearance of the gland is well-formed and mainly circular in shape; the metrics in the patterns therefore is represented by a white color on the image and zeroes in each of the levels of matrices. The whitespace is surrounded by darker color staining where a high intensity value exists for the red and blue layers. Defining Gleason pattern 4 is a change in this shape that can be visually seen by a pathologist by the well-formed structure becoming elongated or fused gland. In the numerical matrix the whitespace (or 0) values in the array is represented in different locations. Thus, when considering a general PCA between Gleason pattern 3 and Gleason pattern 4, the differences in whitespace and surrounding color is in different locations in the matrix and represented in the PCA as two unique groups of images (see Figure 4). These principal considerations allowed for the use of PCA as a way to understand the technical variations that may exist between the different primary Gleason patterns in synthetically generated or original images.
[0091] EXAMPLE 4 - Technical Variations
[0092] In order to study technical variations, technical staining artifacts were identified that could confound the color variance. For this, a single image PCA color analysis showed 98.8% of the variance in red, 96.1% of the variance in blue, and only 92% of the variance in green, which is a cumulative representation of all primary Gleason patterns in the original images (see Figure 5). The least amount of variance was found in the red and blue colors, which are the dominant colors in the tissue images. The greatest amount of variance was found in the green color, which is the least dominant color in the tissue images. The results of low variance in red and blue suggest there are no hidden confounding staining artifacts in either the synthetic or the original images.
[0093] Further, in order to understand if the level of color distribution is similar between synthetic and original images, the PCA results from the three color (RGB) were combined from original and synthetic images. The results showed non-significant (p=0.78) technical variations when comparing the color distribution between the original (n=180) and synthetic (n=187) images.
[0094] The outcomes remained similar when comparing the primary scores as well: Gleason 3 original versus Gleason 3 synthetic, p=0.527; Gleason 4 original versus Gleason 4 synthetic, p=0.421; and Gleason 5 original versus Gleason 5 synthetic, p=0.802. Overall, the results suggested that the technical variability was the same within the entire image dataset and uniform in both synthetic and original images.
[0095] EXAMPLE 5 - Validation of CNN Performance
[0096] An experiment was performed to determine whether GAN-enhanced synthetic images could substitute the RP and needle biopsy images for training a CNN model and potentially improving its grading capabilities. For this purpose, the CNN model was trained with two sources of image patches. The first source was the patches derived from original RP sections from TCGA, which were classified according to Gleason pattern (normal (n=175), GS3 (n=726), GS4 (n=1029), and GS5 (n=152)).
[0097] The second source of patches was derived from original digital pathology combined with the synthetically generated images, classified according to Gleason pattern. In the second source, a total of 5000 image patches were used for each Gleason pattern. A probability score, which was the highest probability that a given patch is of a specific Gleason pattern, was assigned. The grading capabilities of this CNN model were then compared by allowing it to assign Gleason scoring to the RP images in the TCGA (n=475).
[0098] Interestingly, the results demonstrated that the CNN’s accuracy improved in GS3 from 0.53 to 0.67 (p=0.0010), in GS4 from 0.55 to 0.63 (p=0.0274), and in GS5 from 0.57 to 0.75 (p<0.0001) when trained with the combination of original and synthetic image patches compared to just being trained with original image patches. Moreover, the comparative analysis revealed a notable enhancement in accuracy and the receiver operating characteristic curve (ROC) for the combined (original and synthetic) dataset relative to the original (p= 0.0381).
[0099] The examples demonstrate that generation and use of synthetically generated digital pathology data from PCa can effectively train ML models, which later were able to perform Gleason grading as effective as when trained with original data from patients. Embodiments of the subject invention can minimize or avoid use of clinical data (and the several limitations that come with it) to make the ML models more effective. The ML models can improve reproducibility, reduce variability, and facilitate the diagnosis and prognosis of PCa and other cancers.
[0100] EXAMPLE 6 - Network Architecture and Model Selection
[0101] For Examples 6-12, the Materials and Methods were largely the same as for Examples 1-5 (and as discussed in the Materials and Methods section above). Though, 32 local samples were taken from the University of Miami pathology core. The IRB protocol was approved by the University of Miami Miller School of Medicine, Miami, FL, to ensure that the research adhered to ethical guidelines and principles. Further, the study included 3,949 needle biopsies sourced from Radboud University Medical Center and Karolinska Institute. Also, a preliminary conditional generative adversarial network (cGAN) was designed and implemented to assess the performance accuracy of various GAN architectures. The cGAN was developed utilizing Python 3.7.3 and the Tensorflow Keras 2.7.0 package. The generator component of the cGAN included three input layers and a single output layer. In parallel, the discriminator component was configured with analogous input, hidden, and output layers. The cGAN’s total parameter count was 19.2 million for each of the evaluated Gleason patterns. In addition, StyleGAN, a progressive generative adversarial network architecture, served as a baseline for comparison to the cGAN, featuring a distinct generator configuration. The architecture was adopted from the original StyleGAN with minimal alterations to the generator and discriminator networks. A notable modification involved substituting human face images in the StyleGAN with tissue images to create a tissue image GAN. The generator’s total parameter count amounted to 28.5 million, in contrast to 26.2 million in the original StyleGAN and 23.1 million in a conventional generator. Particular emphasis was placed on refining the GS4 and GS5 images to ensure adequate representation of tumor heterogeneity. In order to select the most appropriate CNN model, four different models were evaluated, starting with the AlexNet. AlexNet was the first CNN to be defined, and it includes a series of convolutional pooling layers and an interconnected output layer. To test the accuracy of this model, 20 histology images were randomly selected from the TCGA, which contained images from 500 patients, each with four images corresponding to scores of normal, GS6, GS7, GS8, and GS9. These images were graded using default parameters, without pretraining the model and the results were compared to the assessments of two independent pathologists (the gold standard) and the TCGA scoring. The AlexNet model achieved an accuracy rate of 55%, aligned with the gold standard. Next, the ResNet model was evaluated, which incorporates residual networks based on skip connections within the network, significantly improving AlexNet. ResNet achieved an accuracy rate of approximately 55% across the same 20 images, aligned with the gold standard. The next to be tested was the Xception model based on the Inception model and improving traditional network architectures such as AlexNet. Unlike traditional models, Xception separates space and depth into different dimensions in its convolutional layers. Xception achieved an accuracy rate of 60%, outperforming the other two models. Last, the EfficientNet model was tested, which utilizes a novel scaling method that employs a compound coefficient to scale up a CNN model in a more structured manner. This improves the model’s performance and efficiency, requiring minimal layers to train on existing images. EfficientNet achieved an accuracy rate of 65%, outperforming the other three models. Based on these results, the EfficientNet model was selected.
[0102] EXAMPLE 7 - Image Preprocessing
[0103] Several factors, such as staining protocols, tissue quality, section thickness, tissue folding, and the amount of tissue on the slide, could negatively impact the accuracy of Al models in making predictions from tissue biopsy images. To account for this, pre-processing normalization of the TCGA images was conducted. Specifically, all 500 images were selected and had their color distribution evaluated by calculating the mean value of RGB colors and normalizing this value. Images with an RGB mean intensity value two standard deviations away from the total mean value of all samples were identified as outliers and removed from the dataset. In total, 21 images were discarded due to being outliers. Similar pre-processing was also conducted for the radical prostatectomy (RP) section images obtained from the University of Miami Pathology core (n-32), where the mean intensity value was calculated and compared to each image. One sample was identified as an outlier and removed from the dataset. For the normalization of needle biopsy (NB) slides from Radbound University Medical Center and Karolinska Institute (n=3949), the images were first evaluated for areas of tissue before subjecting them to color normalization. Normalization led to the exclusion of 257 images from the dataset. Overall, the pre-processing steps helped to reduce the variability in tissue biopsy images and ensured a more consistent dataset for the Al models to make predictions.
[0104] EXAMPLE 8 - Quality-Controlled Annotation and Patch Generation
[0105] RP sections (from TCGA and University of Miami (UM)) were separated into pseudocohorts based on the primary and secondary Gleason pattern and were inspected using the HistoQC software to remove samples of low quality. A random selection of 143 sections was then provided to two pathologists for Gleason pattern annotation. To build gold standard training data, both pathologists were required to agree with each other, as well as TCGA scoring and precise overlapping annotation of the tumor region. This resulted in a total of 33 images. Image patches from these 25 individuals from TCGA and eight images from in-house selection were sliced using PyHIST software with dimensions of 96x96 and 256x256 pixel lengths. The areas of annotation obtained from the pathologists were overlaid with these sections, and image patches were required to contain at least 75% tissue image and no more than 25% whitespace. The smallest area of overlap between pathologists was 96 pixels, considered the smallest size for image analysis. In total, 219 patches were generated from 33 images, which were enhanced by augmentation to create 2082 patches (see Figure 9A).
[0106] EXAMPLE 9 - GAN Model Selection
[0107] In order to evaluate the performance of various GAN architectures and select the most appropriate one, 2082 RP image patches extracted from 33 individuals were divided into training cohorts according to their Gleason pattern (GS3, GS4, or GS5). Each training cohort was subjected to cGAN, StyleGAN, and dcGAN architectures. A total of 1000 synthetic images generated by each GAN were fed into a generic CNN for classification into their respective Gleason categories. The cGAN achieved an accuracy of 0.59, while the StyleGAN and dcGAN demonstrated accuracies of 0.65 and 0.64, respectively. Although the StyleGAN and dcGAN exhibited similar accuracies, their execution times differed significantly when utilizing a standard NVIDIA T100 GPU processor. The generation of 1,000 images in the StyleGAN required 2,372 minutes, whereas the dcGAN completed the task in 901 minutes. Based on these findings, the deep convolutional GAN (dcGAN) network generator was selected and used to create random synthetic images from 2082 patches to create an effective training dataset. These image patches were analyzed using the Adam optimization algorithm. This process helped to find the best iteration value for the model, which was 14,000 iterations. To verify the appropriateness of the patch sizes, random synthetic prostate images were generated from patches sized 128x128 and 256x256 pixels and evaluated through a quality control (QC) assessment. The manual assessment was done by board-certified pathologists, focusing on the images’ sharpness and resolution, which received an 80% approval rate for adequacy in both patch sizes.
[0108] Further, a training set was assembled from the PANDA challenge for NB analysis, including both tumor and benign sections. A pathologist annotated 300 needle biopsies, leading to the extraction of 1,712 cancer-specific patches and 539 benign tissue patches, each sized 256x256. Considering the differences in feature sizes between needle biopsies and RP sections, various patch sizes were examined from 512x512 to 32x32 pixels. The accuracy of cancergrade identification by CNNs dropped from 95.9% to 57.0% as the patch size decreased, with a significant decline noted below 64x64 pixels. Hence, 64x64 was selected as the initial patch size for needle biopsies, upscaled to 256x256 for enhanced model precision. These images underwent color normalization before being used to train the dc-GAN model. On a single NVIDIA A1000 GPU, the average GAN run took 2.5 hours, producing 1,000 synthetic images per run when the generator and discriminator loss functions balanced out, typically occurring at around 50000 epochs. For NB, a repository of 2,000 patches each were generated for tumor and normal samples (see Figures 10A and 10B). Similar to the RP sections, the manual assessment was done by board-certified pathologists, which received an 80% approval rate for adequacy in patch size.
[0109] EXAMPLE 10 - Benchmarking
[0110] In order to identify the optimal number of synthetic images necessary for training the EfficientNet classification model to maximize accuracy without inducing overfitting due to image similarity, the EfficientNet model was fine-tuned with synthetic data produced by the dcGAN model. This process included assessing the synthetic images’ similarity index to prevent or inhibit overfitting the dcGAN model. A similarity index closer to 1 indicates better alignment. Synthetic images were generated based on the primary Gleason pattern in batches of 10K, 50K, and 100K. To determine the maximum number of unique synthetic images for model training without affecting accuracy, each batch was combined with original images and processed through a CNN model. The resulting images were evaluated for potential model overfitting using the Frechet Inception Distance (FID) score, which measures the disparity between feature vectors of real and synthetic images. The FID for RP images for the 10K batch was 25.1, the 50K batch was 18.8, and the 100K batch was 36.2. The findings revealed that for the 10K batch, accuracy in predicting the primary pattern increased from 33% to 63% for GS3, 35% to 61% for GS4, and 41% to 68% for GS5. With the 50K batch, accuracy improved from 33% to 70% for GS3, 35% to 65% for GS4, and 41% to 71% for GS5. For the 100K batch, accuracy increased from 33% to 65% for GS3, 35% to 62% for GS4, and 41% to 67% for GS5. These results indicated that the diversity of images peaked at around 50K and stabilized beyond that point. Hence, a maximum of 50K synthetic images were integrated into the model. To further reduce the risk of overfitting due to limited variation within synthetically generated images from the same batch, a cross-validation method was employed to assess variability. All image tiles, both real and synthetic, were randomly organized, and the dataset was divided into ten subsets. These subsets were designed to randomly assign training and testing sets, enabling evaluation of the mean square error, the misclassification error rate, and confidence intervals. The similarity index values ranged between 0.8 to 1.0, and the accuracy consistently ranged from 0.68 to 0.71 for GS3, 0.63 to 0.70 for GS4, and 0.66 to 0.73 for GS5. These findings indicate that the model did not produce significant outlier images that could affect the accuracy of the EfficientNet CNN, affirming confidence in the quality and variability of the synthetic images.
[0111] Similar to RP sections, for NB the FID for the 10K batch was 21.2, the 50K batch was 20.2, and the 100K batch was 33.7. To prevent or inhibit overfitting in the context of NB, a ten-fold cross-validation method was applied to evaluate variability (see also the tables in Figures 15A and 15B). One thousand samples were randomly selected from the combined pool of real and synthetic images and divided into ten subsets. The similarity index values spanned from 0.92 to 1.0, while the accuracy remained stable for GS3 (0.67-0.71), GS4 (0.63-0.70), and GS5 (0.65-0.73). These results suggest that the model did not produce any significant outliers that could undermine the accuracy of the EfficientNet, ensuring the quality and variability of the synthetic needle biopsy images.
[0112] EXAMPLE 11 - Quality Evaluation of Synthetic Images
[0113] The extent of technical variations that exist between the synthetically generated images and the original images was investigated. This examination is important because each Gleason pattern possesses specific morphological characteristics that synthetically generated images should ideally replicate. For instance, Gleason pattern 3 is characterized by well-formed glands (discrete units that can be circled individually as long as they are not fused together) of varying sizes, including branching glands. These glands may be angulated or compressed, with the key feature of retaining at least a wisp of stroma between neighboring glands. If the latter is missing, then it is rated as pattern 4. The cribriform pattern is also included under pattern (I. The cribriform glands are a confluent sheet of contiguous malignant epithelial cells with multiple glandular lumina that are easily visible at low power. Gleason pattern 5 has two patterns, namely comedonecrosis and cords. In comedonecrosis, central necrosis with intraluminal necrotic cells is seen within papillary / cribriform spaces, while in singular form, cells forming cords without glandular lumens are visible.
[0114] To characterize technical and structural variations between synthetic and real images, spatial heterogeneous recurrence quantification analysis (SHRQA) was utilized, a robust technique capable of measuring complex microstructures based on spatial patterns (see also; Wang et al., Recurrence Network Analysis of Histopathological Images for the Detection of Invasive Ductal Carcinoma in Breast Cancer, IEEE / ACM Trans Comput Biol Bioinform., 20:3234-44, 2023; Chen et al., Recurrence network modeling and analysis of spatial data, Chaos, 28:085714, 2018; and Yang et al., Heterogeneous recurrence analysis of spatial data, Chaos, 30:013119; 2020; all three of which are hereby incorporated herein by reference in their entireties). The SHRQA process involves six key steps. It begins with the two-dimensional (2D) discrete wavelet transform (2D-DWT) using the Haar wavelet to reveal patterns not visible in the original image. Then, each image is transformed into an attribute vector via the space-filling curve (SFC), which importantly preserves the spatial proximity between pixels in the image within the vector. This step is important for analyzing the image’s geometric recurrence in vector form. A trajectory is formed in state space by projecting this attribute vector, highlighting the image’s geometric structure. Through Quadtree segmentation, the state space is divided into unique subregions to discern spatial transition patterns. An Iterated Function System projection is then applied, converting each attribute vector into a fractal plot that represents recurrence within the fractal topology. Finally, these fractal structures are quantified to illuminate the intricate geometric properties of the image, providing a detailed profile.
[0115] In a previous study, the SHRQA method was used to quantify synthetically generated image patches from eight different genitourinary organs, including the testis, kidney, prostate, bladder, vagina, cervix, ovary, and uterus (Van Booven et al., Synthetic Genitourinary Image Synthesis via Generative Adversarial Networks: Enhancing Artificial Intelligence Diagnostic Precision. J Pers Med., 14, 2024; which is hereby incorporated herein by reference in its entirety). This reflects the strength of the models, which have already been tested for quantification on histology image patches. Following a similar approach, the SHRQA method was applied to examine the spatial recurrence properties of real and synthetic image patches across three Gleason patterns (GS 3, 4, and 5) in the RP section. The sample set included an equal number of patches from real and synthetic sources, with a balanced representation of each Gleason pattern. Three thousand image patches were analyzed, each 256*256 pixels, evenly split between real and synthetic. A one-layer 2D-DWT with Haar wavelet decomposed each image into four sub-images, revealing fine details. SHRQA quantitatively outlined each patch’s microstructures. From an initial extraction of 1997 spatial recurrence features per patch, LASSO selected 1819 features as significant to the Gleason pattern. Hotelling’s T-squaredtest, a multivariate extension of the two-sample T-test, compared the spatial recurrence attributes of real versus synthetic patches. The resulting p- values of 0.8991 signified no significant differences in spatial recurrence properties between real and synthetic patches, as confirmed by the T-squared tests’ p- values for each Gleason pattern (see also the table in Figure 16A).
[0116] PCA was also employed on the spatial recurrence properties for RP, visualized using radar charts, revealing that the top ten principal components capture 90% of the variability. This allowed the distributions of spatial properties to be mapped for real and synthetic images across Gleason patterns, as depicted in Figures 11 A and 1 IB. Notably, while distributions for the same Gleason pattern aligned closely between real and synthetic images, significant differences were evident across different patterns. A similar analysis was conducted for the NB section (Figures 12A and 12B), examining 1200 patches with an equal split between real and synthetic images distributed evenly across GS 3, 4, and 5. SHRQA extracted 2585 initial features, with LASSO identifying 1578 as significant. Hotelling’s T-squared tests in the NB section corroborated the RP section’s results (see the table in Figure 16B), demonstrating that synthetic images reliably replicate the geometric nuances of real images for each Gleason pattern. These findings across both RP and NB sections validate the model’s efficiency in capturing the geometric intricacies consistent with real images.
[0117] In addition to SHQRA quantification, in order to delineate the technical variations between the original and synthetically generated images, the original images were categorized into primary Gleason pattern 3, 4, and 5. To examine individual gland formation, the images were subjected to a PCA to investigate the color distribution among GS3, GS4, and GS5. The PCA enabled identification of features and cellular morphology inherent to each of the Gleason pattern, as described previously. Each image was represented as a large multi-layered 3D matrix, with the x and y coordinates in the matrix indicating the location on the image and the z coordinate defined by the color intensity. Three distinct numerical values were considered for color intensity, one each for red, green, and blue intensity, representing a single point or pixel on the image (Figures 13A-13C illustrating color distribution). For Gleason pattern 3, the morphological appearance of the gland is well-formed and predominantly circular in shape. The matrices in these patterns are represented by a white color on the image and 0’s at each level of the matrices. The whitespace is encircled by darker color staining, where high-intensity values are present for the red and blue layers. Defining Gleason pattern 4 entails a change in shape that can be visually discerned by a pathologist as the well-formed structure becomes fused or shows a cribriform architecture. In the numerical matrix, the whitespace (or 0) values in the array were represented in different locations. Consequently, when examining a general PCA between Gleason pattern 3 and 4, the differences in whitespace and surrounding color appear in distinct locations within the matrix and represented in the PCA as two unique groups of images (Figures 13A-13c).
[0118] These principal considerations allowed PCA to be employed as a method for understanding the technical variations that may exist between different primary Gleason pattern in both synthetically generated and original images. In the continued investigation of technical variations, identification of technical staining artifacts that could confound color variance was sought. To achieve this, single image PCA color analysis revealed 98.8% of the variance in red, 96.1% of the variance in blue, and only 92% of the variance in green, which collectively represents all primary Gleason pattern in the original images. The least amount of variance was observed in the red and blue colors, which are the dominant colors in the tissue images. The greatest amount of variance was found in the green color, which is the least dominant color in the tissue images. The low variance results for red and blue suggest the absence of hidden confounding staining artifacts in both synthetic and original images. To determine if the level of color distribution was similar between synthetic and original images, the PCA results from the three colors (red, green, and blue) were combined for both original and synthetic images. The results revealed nonsignificant (p=0.78) technical variations when comparing the color distribution between the original (n=480) and synthetic (n=187) images. The outcomes remained consistent when comparing each primary pattern as well: Gleason 3 original versus Gleason 3 synthetic (p=0.527); Gleason 4 original versus Gleason 4 synthetic (p=0.421); and Gleason 5 original versus Gleason 5 synthetic (p=0.802), respectively. Overall, the findings suggest that the technical variability was uniform within the entire image dataset and consistent across both synthetic and original images.
[0119] Further, a randomized number of synthetic image patches allocated into Gleason 3, 4, and 5 via SHRQA quantification models were subjected to cross-validation for characterization by pathologists. Specific features associated with each phenotype (e.g., uniform glandular arrangement, moderate differentiation, fibrous stroma, mild nuclear atypia, clear lamina, and minimal desmoplasia associated with Gleason 3; moderate to severe nuclear atypia, poor differentiation, irregular and cribriform patterns, inflammatory infiltrate associated with Gleason 4; and sheet of cells, scant stromal tissue, undifferentiated tumor cells, sever nuclear atypia, complete loss of glandular structure associated with Gleason 5) were accurately identified by the SHRQA models. These findings across both RP and NB sections validate the model’s efficiency in capturing the geometric intricacies consistent with real images (see also Figures 13A-13C)
[0120] EXAMPLE 12 - Validation of CNN Performance Post-Training with Enhanced Synthetic Data Subsequently, determination was sought on whether synthetic images could substitute original images of the RP and NP for training a CNN model and potentially improving its grading capabilities. For this purpose, the CNN model was trained with two sources of image patches. The first source was the patches derived from original RP sections from TCGA and in-house images, which were classified according to the Gleason pattern (normal (n=l 75), GS3 (n=726), GS4 (n=1029), and GS5 (n=l 52)). The second source of patches was derived from a combination of original and synthetically generated images, classified according to the Gleason pattern. In the second source, a total of 5000 image patches were used for each Gleason pattern. The grading capabilities of this CNN model were then compared by allowing it to assign Gleason scoring to the RP images in the TCGA (n=475). Interestingly, the results demonstrated that the CNN’s accuracy significantly improved in GS3 from 0.53 to 0.67 (p=0.0010), in GS4 from 0.55 to 0.63 (p^0.0274), and in GS5 from 0.57 to 0.75 (pO.0001) when trained with the combination of original and synthetic image patches compared to just being trained with original image patches. Moreover, the comparative analysis revealed a notable enhancement in accuracy and the receiver operating characteristic curve (ROC) for the combined (original and synthetic) dataset relative to the original (p= 0.0381) (see also left side of Figure 14). Additionally, the in-house RP images (n=24) were subjected to grading using CNN models described herein. Results demonstrated a significant improvement of the CNN model to accurately assign the grade from GS6 from 0.53 to 0.67 (p=0.0010), in GS7 from 0.55 to 0.63 (p=0.0274), and in GS8 from 0.57 to 0.75 (pO.0001).
[0121] Further, the validation of the CNN model was extended using NB. Similar to the RP sections, the CNN model was trained with two sources of image patches. The first source was the patches derived from original NB sections, which were classified according to the Gleason pattern (normal (n=539), GS3 (n=610), GS4 (n=890), and GS5 (n=212)). The second source of patches was derived from original digital pathology combined with the synthetically generated images, classified according to the Gleason pattern. In the second source, a total of 2000 image patches were used for each Gleason pattern. The grading capabilities of this CNN model were then compared by allowing it to assign Gleason grading to the NB images (n=3649). The comparative analysis revealed a significant enhancement in accuracy and the ROC for the amalgamated dataset relative to the original (see also right side of Figure 14). Specifically, the enhancement was consistent across both benign and malignant samples, with an overall accuracy increase from 91% when using the original training images to an accuracy of 95% when using original and synthetic images combined (p = 0.0402). The original and synthetic combined training database yielded a sensitivity of 0.81, and a specificity of 0.92.
[0122] It should be understood that the examples and embodiments described herein are for illustrative purposes only and that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and purview of this application.
[0123] All patents, patent applications, provisional applications, and publications referred to or cited herein (including in the “References” section, if present) are incorporated by reference in their entirety, including all figures and tables, to the extent they are not inconsistent with the explicit teachings of this specification.
[0124] REFERENCES
[0125] 1 . Brawley OW. Prostate cancer epidemiology in the United States. World J Urol 2012; 30(2): 195-200.
[0126] 2. Badalament RA, Drago JR. Prostate cancer. Dis Mon 1991; 37(4): 199-268.
[0127] 3. Carthon B, Sibold HC, Blee S, R DP. Prostate Cancer: Community Education and Disparities in Diagnosis and Treatment. Oncologist 2021 ; 26(7): 537-48.
[0128] 4. Cook ED, Nelson AC. Prostate cancer screening. Curr Oncol Rep 201 1; 13(1): 57- 62.
[0129] 5. Litwin MS, Tan HJ. The Diagnosis and Treatment of Prostate Cancer: A Review. Jama 2017; 317(24): 2532-42.
[0130] 6. Moon TD. Prostate cancer. J Am Geriatr Soc 1992; 40(6): 622-7.
[0131] 7. Schatten H. Brief Overview of Prostate Cancer Statistics, Grading, Diagnosis and Treatment Strategies. Adv Exp Med Biol 2018; 1095: 1-14.
[0132] 8. Wozniak-Petrofsky J. The significance of prostatic specific antigen in men with prostate disease. An elevated PSA level may indicate prostate cancer— and it may not. Geriatr Nurs 1993; 14(3): 150-1.
[0133] 9. Aghdam AM, Amiri A, Salarinia R, Masoudifar A, Ghasemi F, Mirzaei H. MicroRNAs as Diagnostic, Prognostic, and Therapeutic Biomarkers in Prostate Cancer. Crit Rev Eukaryot Gene Expr 2019; 29(2): 127-39.
[0134] 10. Albertsen PC. PSA testing, cancer treatment, and prostate cancer mortality reduction: What is the mechanism? Urol Oncol 2023; 41(2): 78-81.
[0135] 11 . Barry MJ, Simmons LH. Prevention of Prostate Cancer Morbidity and Mortality: Primary Prevention and Early Detection. Med Clin North Am 2017; 101(4): 787-806.
[0136] 12. Borley N, Feneley MR. Prostate cancer: diagnosis and staging. Asian J Androl 2009; 11(1): 74-80.
[0137] 13. Jalloh M, Cooperberg MR. Implementation of PSA-based active surveillance in prostate cancer. Biomark Med 2014; 8(5): 747-53.
[0138] 14. Merriel SWD, Funston G, Hamilton W. Prostate Cancer in Primary Care. Adv Ther 2018; 35(9): 1285-94.
[0139] 15. Nguyen-Nielsen M, Borre M. Diagnostic and Therapeutic Strategies for Prostate Cancer. Semin Nucl Med 2016; 46(6): 484-90. 16. Pezaro C, Woo HH, Davis ID. Prostate cancer: measuring PSA. Intern Med J 2014; 44(5): 433-40.
[0140] 17. Sharma S, Zapatero-Rodriguez J, O'Kennedy R. Prostate cancer diagnostics: Clinical challenges and the ongoing need for disruptive and effective diagnostic tools. Biotechnol Adv 2017; 35(2): 135-49.
[0141] 18. Troyer DA, Mubiru J, Leach RJ, Naylor SL. Promise and challenge: Markers of prostate cancer detection, diagnosis and prognosis. Dis Markers 2004; 20(2): 117-28.
[0142] 19. Venderbos LD, Roobol MJ. PSA-based prostate cancer screening: the role of active surveillance and informed and shared decision making. Asian J Androl 2011; 13(2): 219-24.
[0143] 20. Parekh DJ, Punnen S, Sjoberg DD, et al. A multi-institutional prospective trial in the USA confirms that the 4Kscore accurately identifies men with high-grade prostate cancer. Eur Urol 2015; 68(3): 464-70.
[0144] 21 . Borque-Femando A, Rubio-Briones J, Esteban LM, et al. Role of the 4Kscore test as a predictor of reclassification in prostate cancer active surveillance. Prostate Cancer Prostatic Dis 2019; 22(1): 84-90.
[0145] 22. Konety B, Zappala SM, Parekh DJ, et al. The 4Kscore® Test Reduces Prostate Biopsy Rates in Community and Academic Urology Practices. Rev Urol 2015; 17(4): 231 -40.
[0146] 23. Mi C, Bai L, Yang Y, Duan J, Gao L. 4Kscore diagnostic value in patients with highgrade prostate cancer using cutoff values of 7.5% to 10%: A meta-analysis. Urol Oncol 2021; 39(6): 366.el-.el0.
[0147] 24. Punnen S, Nahar B, Prakash NS, Sjoberg DD, Zappala SM, Parekh DJ. The 4Kscore Predicts the Grade and Stage of Prostate Cancer in the Radical Prostatectomy Specimen: Results from a Multi-institutional Prospective Trial. Eur Urol Focus 2017; 3(1): 94-9.
[0148] 25. Scuderi S, Tin A, Gandaglia G, et al. Implementation of 4Kscore as a Secondary Test Before Prostate Biopsy: Impact on US Population Trends for Prostate Cancer. Eur Urol Open Sci 2023; 52: 1-3.
[0149] 26. Zappala SM, Dong Y, Linder V, et al. The 4Kscore blood test accurately identifies men with aggressive prostate cancer prior to prostate biopsy with or without DRE information. Int J Clin Pract 2017; 71(6).
[0150] 27. Hessels D, Klein Gunnewiek JM, van Oort I, et al. DD3(PCA3)-based molecular urine analysis for the diagnosis of prostate cancer. Eur Urol 2003; 44(1): 8-15; discussion -6.
[0151] 28. Day JR, Jost M, Reynolds MA, Groskopf J, Rittenhouse H. PC A3: from basic molecular science to the clinical lab. Cancer Lett 2011; 301(1): 1 -6. 29. de la Taille A. Progensa PCA3 test for prostate cancer detection. Expert Rev Mol Diagn 2007; 7(5): 491-7.
[0152] 30. Filella X, Foj L, Mila M, Auge JM, Molina R, Jimenez W. PCA3 in the detection and management of early prostate cancer. Tumour Biol 2013; 34(3): 1337-47.
[0153] 31. Gunelli R, Fragala E, Fiori M. PCA3 in Prostate Cancer. Methods Mol Biol 2021; 2292: 105-13.
[0154] 32. Hessels D, Schalken JA. The use of PCA3 in the diagnosis of prostate cancer. Nat Rev Urol 2009; 6(5): 255-61.
[0155] 33. Yang Z, Yu L, Wang Z. PCA3 and TMPRSS2-ERG gene fusions as diagnostic biomarkers for prostate cancer. Chin J Cancer Res 2016; 28(1 ): 65-71 .
[0156] 34. Kim L, Boxall N, George A, et al. Clinical utility and cost modelling of the phi test to triage referrals into image-based diagnostic services for suspected prostate cancer: the PRIM (Phi to Refine Mri) study. BMC Med 2020; 18(1): 95.
[0157] 35. Barisiene M, Bakavicius A, Stanciute D, et al. Prostate Health Index and Prostate Health Index Density as Diagnostic Tools for Improved Prostate Cancer Detection. Biomed Res Int 2020; 2020: 9872146.
[0158] 36. Tosoian JJ, Druskin SC, Andreas D, et al. Use of the Prostate Health Index for detection of prostate cancer: results from a large academic practice. Prostate Cancer Prostatic Dis 2017; 20(2): 228-33.
[0159] 37. Yan JQ, Huang D, Huang JY, et al. Prostate Health Index (phi) and its derivatives predict Gleason score upgrading after radical prostatectomy among patients with low-risk prostate cancer. Asian J Androl 2022; 24(4): 406-10.
[0160] 38. Ye C, Ho JN, Kim DH, et al. The Prostate Health Index and multi-parametric MRI improve diagnostic accuracy of detecting prostate cancer in Asian populations. Investig Clin Urol 2022; 63(6): 631-8.
[0161] 39. Zhang G, Li Y, Li C, Li N, Li Z, Zhou Q. Assessment on clinical value of prostate health index in the diagnosis of prostate cancer. Cancer Med 2019; 8(1 1): 5089-96.
[0162] 40. Kaplan I, Oldenburg NE, Meskell P, Blake M, Church P, Holupka EJ. Real time MRI- ultrasound image guided stereotactic prostate biopsy. Magn Reson Imaging 2002; 20(3): 295-9.
[0163] 41. Fernandes MC, Yildirim O, Woo S, Vargas HA, Hricak H. The role of MRI in prostate cancer: current and future directions. Magma 2022; 35(4): 503-21. 42. Mendhiratta N, Taneja SS, Rosenkrantz AB. The role of MRI in prostate cancer diagnosis and management. Future Oncol 2016; 12(21): 2431-43.
[0164] 43. Stabile A, Giganti F, Rosenkrantz AB, et al. Multiparametric MRI for prostate cancer diagnosis: current status and future directions. Nat Rev Urol 2020; 17(1 ): 41-61 .
[0165] 44. Stempel CV, Dickinson L, Pendse D. MRI in the Management of Prostate Cancer. Semin Ultrasound CT MR 2020; 41(4): 366-72.
[0166] 45. Wibmer AG, Vargas HA, Hricak H. Role of MRI in the diagnosis and management of prostate cancer. Future Oncol 2015; 11(20): 2757-66.
[0167] 46. Ikeda S, Elkin SK, Tomson BN, Carter JL, Kurzrock R. Next-generation sequencing of prostate cancer: genomic and pathway alterations, potential actionability patterns, and relative rate of use of clinical-grade testing. Cancer Biol Ther 2019; 20(2): 219-26.
[0168] 47. Abdulmajed MI, Hughes D, Shergill IS. The role of transperineal template biopsies of the prostate in the diagnosis of prostate cancer: a review. Expert Rev Med Devices 2015; 12(2): 175-82.
[0169] 48. Ahdoot M, Wilbur AR, Reese SE, et al. MRI-Targeted, Systematic, and Combined Biopsy for Prostate Cancer Diagnosis. N Engl J Med 2020; 382(10): 917-28.
[0170] 49. He Y, Shen Q, Fu W, Wang H, Song G. Optimized grade group for reporting prostate cancer grade in systematic and MRI-targeted biopsies. Prostate 2022; 82(1 1): 1125-32.
[0171] 50. Pinto F, Totaro A, Calarco A, et al. Imaging in prostate cancer diagnosis: present role and future perspectives. Urol Int 201 1; 86(4): 373-82.
[0172] 51. Bangma CH, Roemeling S, Schroder FH. Overdiagnosis and overtreatment of early detected prostate cancer. World J Urol 2007; 25(1): 3-9.
[0173] 52. Loeb S, Bjurlin MA, Nicholson J, et al. Overdiagnosis and overtreatment of prostate cancer. Eur Urol 2014; 65(6): 1046-55.
[0174] 53. Thompson IM. Overdiagnosis and overtreatment of prostate cancer. Am Soc Clin Oncol Educ Book 2012: e35-9.
[0175] 54. Resnick MJ, Lee DJ, Magerfleisch L, et al. Repeat prostate biopsy and the incremental risk of clinically insignificant prostate cancer. Urology 2011; 77(3): 548-52.
[0176] 55. Tataru OS, Vartolomei MD, Rassweiler JJ, et al. Artificial Intelligence and Machine Learning in Prostate Cancer Patient Management-Current Trends and Future Perspectives. Diagnostics (Basel) 2021 ; 1 1 (2). 56. Choi RY, Coyner AS, Kalpathy-Cramer J, Chiang MF, Campbell JP. Introduction to Machine Learning, Neural Networks, and Deep Learning. Transl Vis Sci Technol 2020; 9(2): 14.
[0177] 57. Deo RC. Machine Learning in Medicine. Circulation 2015; 132(20): 1920-30.
[0178] 58. Lo Vercio L, Amador K, Bannister JJ, et al. Supervised machine learning tools: a tutorial for clinicians. J Neural Eng 2020; 17(6).
[0179] 59. Villoutreix P. What machine learning can do for developmental biology. Development 2021 ; 148(1).
[0180] 60. Colling R, Pitman H, Oien K, et al. Artificial intelligence in digital pathology: a roadmap to routine use in clinical practice. J Pathol 2019; 249(2): 143-50.
[0181] 61 . Yousif M, van Diest PJ, Laurinavicius A, et al. Artificial intelligence applied to breast pathology. Virchows Arch 2022; 480(1): 191-209.
[0182] 62. Jovic S, Miljkovic M, Ivanovic M, Saranovic M, Arsic M. Prostate Cancer Probability Prediction By Machine Learning Technique. Cancer Invest 2017; 35(10): 647-51.
[0183] 63. Arvaniti E, Fricker K.S, Moret M, et al. Automated Gleason grading of prostate cancer tissue microarrays via deep learning. Sci Rep 2018; 8(1): 12054.
[0184] 64. Bhattacharya I, Lim DS, Aung HL, et al. Bridging the gap between prostate radiology and pathology through machine learning. Med Phys 2022; 49(8): 5160-81 .
[0185] 65. Bulten W, Pinckaers H, van Boven H, et al. Automated deep-learning system for Gleason grading of prostate cancer using biopsies: a diagnostic study. Lancet Oncol 2020; 21(2): 233- 41.
[0186] 66. Duenweg SR, Brehler M, Bobholz SA, et al. Comparison of a machine and deep learning model for automated tumor annotation on digitized whole slide prostate cancer histology. PLoS One 2023; 18(3): e0278084.
[0187] 67. Le MH, Chen J, Wang L, et al. Automated diagnosis of prostate cancer in multiparametric MRI based on multimodal convolutional neural networks. Phys Med Biol 2017; 62(16): 6497-514.
[0188] 68. Li H, Lee CH, Chia D, Lin Z, Huang W, Tan CH. Machine Learning in Prostate MRI for Prostate Cancer: Current Status and Future Opportunities. Diagnostics (Basel) 2022; 12(2).
[0189] 69. Mehralivand S, Yang D, Harmon SA, et al. Deep learning-based artificial intelligence for prostate cancer detection at biparametric MRI. Abdom Radiol (NY) 2022; 47(4): 1425-34.
[0190] 70. Xiang J, Wang X, Wang X, et al. Automatic diagnosis and grading of Prostate Cancer with weakly supervised learning on whole slide images. Comput Biol Med 2023; 152: 106340. 71. Bhargava HK, Leo P, Elliott R, et al. Computationally Derived Image Signature of Stromal Morphology Is Prognostic of Prostate Cancer Recurrence Following Prostatectomy in African American Patients. Clin Cancer Res 2020; 26(8): 1915-23.
[0191] 72. Safarpoor A, Kalra S, Tizhoosh HR. Generative models in pathology: synthesis of diagnostic quality pathology images(f). J Pathol 2021 ; 253(2): 131-2.
[0192] 73. Xu IRL, Van Booven DJ, Goberdhan S, et al. Generative Adversarial Networks Can Create High Quality Artificial Prostate Cancer Magnetic Resonance Images. J Pers Med 2023; 13(3).
[0193] 74. Alfano R, Bauman GS, Gomez JA, et al. Prostate cancer classification using radiomics and machine learning on mp-MRI validated using co-registered histology. Eur J Radiol 2022; 156: 110494.
[0194] 75. Bertelli E, Mercatelli L, Marzi C, et al. Machine and Deep Learning Prediction Of Prostate Cancer Aggressiveness Using Multiparametric MRL Front Oncol 2021; 11 : 802964.
[0195] 76. Bonekamp D, Schlemmer HP. [Machine learning and multiparametric MRI for early diagnosis of prostate cancer], Urologe A 2021 ; 60(5): 576-91 .
[0196] 77. Cuocolo R, Cipullo MB, Stanzione A, et al. Machine learning for the identification of clinically significant prostate cancer on MRI: a meta-analysis. Eur Radiol 2020; 30(12): 6877-87.
[0197] 78. Cuocolo R, Cipullo MB, Stanzione A, et al. Machine learning applications in prostate cancer magnetic resonance imaging. Eur Radiol Exp 2019; 3(1): 35.
[0198] 79. Hosseinzadeh M, Saha A, Brand P, Slootweg I, de Rooij M, Huisman H. Deep learning-assisted prostate cancer detection on bi-parametric MRI: minimum training data size requirements and effect of prior knowledge. Eur Radiol 2022; 32(4): 2224-34.
[0199] 80. Michaely HJ, Aringhieri G, Cioni D, Neri E. Current Value of Biparametric Prostate MRI with Machine-Learning or Deep-Learning in the Detection, Grading, and Characterization of Prostate Cancer: A Systematic Review. Diagnostics (Basel) 2022; 12(4).
[0200] 81. Vente C, Vos P, Hosseinzadeh M, Pluim J, Veta M. Deep Learning Regression for Prostate Cancer Detection and Grading in Bi-Parametric MRI. IEEE Trans Biomed Eng 2021; 68(2): 374-83.
[0201] 82. Madabhushi A, Lee G. Image analysis and machine learning in digital pathology: Challenges and opportunities. Med Image Anal 2016; 33: 170-5.
[0202] 83. Freeman K, Geppert J, Stinton C, et al. Use of artificial intelligence for image analysis in breast cancer screening programmes: systematic review of test accuracy. BMJ 2021; 374: nl872. 84. Twilt JJ, van Leeuwen KG, Huisman HJ, Futterer JJ, de Rooij M. Artificial Intelligence Based Algorithms for Prostate Cancer Classification and Detection on Magnetic Resonance Imaging: A Narrative Review. Diagnostics (Basel) 2021; 11(6).
[0203] 85. Van Booven DJ, Kuchakulla M, Pai R, et al. A Systematic Review of Artificial Intelligence in Prostate Cancer. Res Rep Urol 2021 ; 13: 31-9.
[0204] 86. Bulten W, Balkenhol M, Belinga JA, et al. Artificial intelligence assistance significantly improves Gleason grading of prostate biopsies by pathologists. Mod Pathol 2021; 34(3): 660-71.
[0205] 87. Bulten W, Kartasalo K, Chen PC, et al. Artificial intelligence for diagnosis and Gleason grading of prostate cancer: the PANDA challenge. Nat Med 2022; 28(1): 154-63.
[0206] 88. Shah P, Kendall F, Khozin S, et al. Artificial intelligence and machine learning in clinical development: a translational perspective. NPJ Digit Med 2019; 2; 69.
[0207] 89. Weissler EH, Naumann T, Andersson T, et al. The role of machine learning in clinical research: transforming the future of evidence generation. Trials 2021; 22(1): 537.
[0208] 90. Banerjee I, Li K, Seneviratne M, et al. Weakly supervised natural language processing for assessing patient-centered outcome following prostate cancer treatment. JAMIA Open 2019; 2(1): 150-9.
[0209] 91. Yagi Y, Riedlinger G, Xu X, et al. Development of a database system and image viewer to assist in the correlation of histopathologic features and digital image analysis with clinical and molecular genetic information. Pathol Int 2016; 66(2): 63-74.
[0210] 92. Collet JP. [Limitations of clinical trials]. Rev Prat 2000; 50(8): 833-7.
[0211] 93. Cheng JY, Abel JT, Balis UGJ, McClintock DS, Pantanowitz L. Challenges in the Development, Deployment, and Regulation of Artificial Intelligence in Anatomic Pathology. Am J Pathol 2021; 191(10): 1684-92.
[0212] 94. van der Laak J, Litjens G, Ciompi F. Deep learning in histopathology: the path to the clinic. Nat Med 2021; 27(5): 775-84.
[0213] 95. Krzyszczyk P, Acevedo A, Davidoff EJ, et al. The growing role of precision and personalized medicine for cancer treatment. TECHNOLOGY 2018; 06(03n04): 79-100.
[0214] 96. Zerbe N, Hufnagl P, Schliins K. Distributed computing in image analysis using open source frameworks and application to image sharpness assessment of histological whole slide images. Diagn Pathol 201 1 ; 6 Suppl 1 (Suppl 1): SI 6. 97. Aeffner F, Adissu HA, Boyle MC, et al. Digital Microscopy, Image Analysis, and Virtual Slide Repository. ILAR Journal 2018; 59(1): 66-79.
[0215] 98. McAlpine ED, Pantanowitz L, Michelow PM. Challenges Developing Deep Learning Algorithms in Cytology. Acta Cytol 2021; 65(4): 301-9.
[0216] 99. Niazi MKK, Keluo Y, Zynger DL, et al. Visually Meaningful Histopathological Features for Automatic Grading of Prostate Cancer. IEEE J Biomed Health Inform 2017; 21(4): 1027- 38.
[0217] 100. Erratum to "Visually Meaningful Histopathological Features for Automatic Grading of Prostate Cancer". IEEE J Biomed Health Inform 2017; 21(5): 1473-4.
[0218] 101. Hou L, Samaras D, Kurc TM, Gao Y, Davis JE, Saltz JH. Patch-based Convolutional Neural Network for Whole Slide Tissue Image Classification. Proc IEEE Comput Soc Conf Comput Vis Pattern Recognit 2016; 2016: 2424-33.
[0219] 102. Ren J, Sadimin ET, Wang D, Epstein JI, Foran DJ, Qi X. Computer aided analysis of prostate histopathology images Gleason grading especially for Gleason score 7. Annu Int Conf IEEE Eng Med Biol Soc 2015; 2015: 3013-6.
[0220] 103. Fauzi MF, Pennell M, Sahiner B, et al. Classification of follicular lymphoma: the effect of computer aid on pathologists grading. BMC Med Inform Decis Mak 2015; 15: 115.
[0221] 104. Kothari S, Phan JH, Young AN, Wang MD. Histological image classification using biologically interpretable shape-based features. BMC Med Imaging 2013; 13: 9.
[0222] 105. Kong J, Cooper LA, Wang F, et al. Machine-based morphologic analysis of glioblastoma using whole-slide pathology images uncovers clinically relevant molecular correlates. PLoS One 2013; 8(11): e81049.
[0223] 106. Dundar MM, Badve S, Bilgin G, et al. Computerized classification of intraductal breast lesions using histopathological images. IEEE Trans Biomed Eng 2011 ; 58(7): 1977-84.
[0224] 107. Sertel O, Kong J, Shimada H, Catalyurek UV, Saltz JH, Gurcan MN. Computer-aided Prognosis of Neuroblastoma on Whole-slide Images: Classification of Stromal Development. Pattern Recognit 2009; 42(6): 1093-103.
[0225] 108. Kong J, Sertel O, Boyer KL, Saltz JH, Gurcan MN, Shimada H. Computer-assisted grading of neuroblastic differentiation. Arch Pathol Lab Med 2008; 132(6): 903-4; author reply 4.
[0226] 109. Munjal R, Arif S, Wendler F, Kanoun O. Comparative Study of Machine-Learning Frameworks for the Elaboration of Feed-Forward Neural Networks by Varying the Complexity of Impedimetric Datasets Synthesized Using Eddy Current Sensors for the Characterization of Bi- Metallic Coins. Sensors (Basel) 2022; 22(4). 110. Hodas NO, Stinis P. Doing the Impossible: Why Neural Networks Can Be Trained at All. Front Psychol 2018; 9: 1185.
[0227] 111. Koh DM, Papanikolaou N, Bick U, et al. Artificial intelligence and machine learning in cancer imaging. Commun Med (Lond) 2022; 2: 133.
Claims
CLAIMSWhat is claimed is:
1. A system for machine learning (ML)-based diagnosis and / or prognosis of cancer, the system comprising: a processor; and a machine-readable medium in operable communication with the processor and having instructions stored thereon that, when executed by the processor, perform the following steps: a) utilizing a generative adversarial network (GAN) to generate synthetic data; b) training an ML model using the synthetic data to give a trained ML model; and c) using the trained ML model to aid in diagnosis and / or prognosis of cancer.
2. The system according to claim 1, wherein the GAN is a deep convolutional GAN (dcGAN).
3. The system according to any of claims 1 -2, wherein the synthetic data comprises at least 100 images.
4. The system according to any of claims 1-3, wherein the training of the ML model comprises using the synthetic data and a quantity of non-synthetic images from clinical trials, and wherein the quantity of non-synthetic images is in a range of from 1 to 75.
5. The system according to claim 4, wherein the quantity of non-synthetic images is in a range of from 1 to 25.
6. The system according to any of claims 1-3, wherein the training of the ML model comprises using only the synthetic data and no non-synthetic images from clinical trials.
7. The system according to any of claims 1-6, wherein the GAN comprises: agenerator neural network comprising 13 layers; and a discriminator neural network comprising 12 layers.
8. The system according to claim 7, wherein the generator neural network comprises a transpose function with batch normalization and rectified linear unit (ReLU) functions, and wherein the discriminator neural network comprises Conv2d, LeakyReLU functions with batch normalization.
9. The system according to any of claims 1-8, further comprising a display in operable communication with the processor and the machine-readable medium, wherein the instructions when executed further perform the step of displaying, on the display, results of the trained ML model.
10. The system according to any of claims 1-9, wherein the cancer is prostate cancer.
11. A method for machine learning (ML)-based diagnosis and / or prognosis of cancer, the method comprising: a) utilizing a generative adversarial network (GAN) to generate synthetic data; b) training an ML model using the synthetic data to give a trained ML model; and c) using the trained ML model to aid in diagnosis and / or prognosis of cancer.
12. The method according to claim 11, wherein the GAN is a deep convolutional GAN (dcGAN).
13. The method according to any of claims 11-12, wherein the synthetic data comprises at least 100 images.
14. The method according to any of claims 11-13, wherein the training of the ML model comprises using the synthetic data and a quantity of non-synthetic images from clinical trials, and wherein the quantity of non-synthetic images is in a range of from 1 to 75.
15. The method according to claim 14, wherein the quantity of non-synthetic images is in a range of from 1 to 25.
16. The method according to any of claims 11-13, wherein the training of the ML model comprises using only the synthetic data and no non-synthetic images from clinical trials.
17. The method according to any of claims 11-16, wherein the GAN comprises: a generator neural network comprising 13 layers; and a discriminator neural network comprising 12 layers.
18. The method according to claim 17, wherein the generator neural network comprises a transpose function with batch normalization and rectified linear unit (ReLU) functions, and wherein the discriminator neural network comprises Conv2d, LeakyReLU functions with batch normalization.
19. The method according to any of claims 11-18, further comprising displaying results of the trained ML model.
20. The method according to any of claims 11-19, wherein the cancer is prostate cancer.