Methods for reducing deep learning data requirements and systems for same
Autoencoders generate synthetic images to enhance training datasets, addressing data scarcity and bias, thereby improving deep learning model performance and equity in data-limited fields.
Patent Information
- Application Number
- PCT/US2025/023679
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-10
- Filing Date
- 2025-04-08
- Publication Date
- 2025-10-16
AI Technical Summary
The development of deep learning models, particularly in data-limited fields like healthcare, is hindered by the intensive data requirements and biases in training datasets, leading to sub-optimal performance and potential equity issues.
The use of autoencoders to generate high-fidelity synthetic images that enhance training datasets, addressing data scarcity and bias by combining real-world images with synthetic ones to improve model performance.
Enhanced training datasets improve model accuracy and equity by increasing data volume and diversity, reducing reliance on big data and mitigating biases.
Smart Images

Figure US2025023679_16102025_PF_FP_ABST
Abstract
Description
[0001] METHODS FOR REDUCING DEEP LEARNING DATA REQUIREMENTS AND SYSTEMS FOR SAME
[0002] CROSS-REFERENCE TO RELATED APPLICATION
[0003] Pursuant to 35 U.S.C. § 1 19(e), this application claims priority to the filing dates of United States Provisional Patent Application Serial No. 63 / 632,380 filed April 10, 2024, the disclosure of which application is herein incorporated by reference in its entirety.
[0004] INTRODUCTION
[0005] Training data for artificial intelligence (Al) models fuel and shape the development of Al models and tools. Intensive data requirements are a major bottleneck limiting the success of Al models and tools. This limitation is particularly acute in sectors with inherently scarce data. In healthcare, training data for Al models can be difficult to curate, triggering growing concerns that the current lack of access to healthcare by certain populations, such as, for example, under-privileged social groups, will translate into future bias in healthcare-related Al models and tools.
[0006] Developing accurate Al models, such as deep learning models, for image analyses requires vast quantities of high-quality training data. This poses a significant bottleneck for many fields of Al development, particularly in data- limited fields such as healthcare. See Rajpurkar, P., Chen, E., Banerjee, O. and Topol, E.J., 2022. Al in health and medicine. Nature medicine, 28(1 ), pp.31 -38. Traditional techniques for data cultivation, such as data gathering (e.g. vehicles equipped with various sensors and configured to drive around capturing training data for an autopilot Al), data mining (e.g. extraction of product purchasing preferences from existing sources such as social media) or collaborative data collection (e.g. pooling of pockets of data) are difficult to implement in medicine. Patient recruitment is challenging for all fields of clinical studies. There are also growing concerns that under-privileged social groups, which may have greater numbers of health concerns but less access to healthcare, are under- represented in training data sets, leading to biases in the resultant Ais, i.e., Ais trained on such limited data sets. See Luo, Y., Tian, Y., Shi, M., Elze, T. and Wang, M., 2023. Harvard Glaucoma Fairness: A Retinal Nerve Disease Dataset for Fairness Learning and Fair Identity Normalization. arXiv preprint arXiv:2306.09264. See also Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K. and Galstyan, A., 2021 . A survey on bias and fairness in machine learning.
[0007] ACM computing surveys (CSUR), 54(6), pp.1-35; Parikh, R.B., Teeple, S. and Navathe, A.S., 2019. Addressing bias in artificial intelligence in health care. JAMA, 322(24), pp.2377-2378. Dedicated equipment is needed for medical testing and highly specialized expertise is required for annotation of collected medical data. Consequently, only select institutions are endowed for medical data gathering and mining, and these efforts often result in sub-optimal datasets, frequently too small and / or too uniform for deep learning model training. See Yu, S., et al. “A re-balancing strategy for class-imbalanced classification based on instance difficulty.” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2022. Collaborations to pool pockets of data are actively being explored but face challenges, such as the lack of standard data formats and the need to protect sensitive and identifying patient information. See Bommakanti, N., et al. “Application of the sight outcomes research collaborative ophthalmology data repository for triaging patients with glaucoma and clinic appointments during pandemics such as COVID-19.” JAMA ophthalmology 138.9 (2020): 974-980. See also Deng, J., et al. “ImageNet: A large-scale hierarchical image database.” 2009 IEEE conference on computer vision and pattern recognition, IEEE, 2009. Furthermore, the narrow nature of deep learning models necessitates datasets that are tailored to the particular research question, thus traditional data cultivation approaches are often cost prohibitive and / or insufficient for healthcare Al model development.
[0008] SUMMARY
[0009] Thus, there is a need for improved and useful methods and systems for growing and enhancing inherently scarce datasets to alleviate dependence on big data. This invention provides such new and useful methods and systems, addressing the limitations mentioned above. Embodiments of the present invention relate to a novel approach for training data expansion and enhancement that can be applied towards any vision-based or image-based Al model development. Embodiments of the present invention can be used to generate synthetic images based on real-world subject images, e.g., based on real world patient imaging data, addressing the increasingly untenable data volume and quality requirements for Al model development. While embodiments of the present invention find use in healthcare-related context, they are not so limited and nonetheless have implications beyond healthcare, towards empowering Al adoption for all similarly data-challenged contexts.
[0010] As described, embodiments of the present invention find use in healthcare-related contexts, such as, for example, in the context of ophthalmological diseases, such as the non-limiting example of Al tools for use in detecting glaucoma. Glaucoma is a leading cause of irreversible blindness, affecting more than 75 million people worldwide in 2020 and is projected to increase to more than 111 million by 2040. See Tham, Y., et al. “Global prevalence of glaucoma and projections of glaucoma burden through 2040: a systematic review and meta-analysis.” Ophthalmology 121 .1 1 (2014): 2081 - 2090. Early detection and treatment are key to preserving vision. See Thompson, A.C., Jammal, A.A. and Medeiros, F.A., 2020. A review of deep learning for screening, diagnosis, and detection of glaucoma progression. Translational Vision Science & Technology, 9(2), pp.42-42. However, screening, or other preventative measures, for glaucoma is challenging due to the often- asymptomatic nature of early glaucoma - as much as 50% of glaucoma cases in developed countries remain undetected while in resource-poor, under-developed countries, up to 90% of individuals with glaucoma are undiagnosed. See Soh, Z., et al. “The global extent of undetected glaucoma in adults: a systematic review and meta-analysis.” Ophthalmology. 2021 . Furthermore, the increasing prevalence of glaucoma from aging populations is leading to a growing discrepancy between the supply and demand of glaucoma care; and as glaucoma progresses, the costs of care increases, further hindering the treatment of this blinding condition. Therefore, there is a strong need for more accessible, objective, and high-throughput detection or screening methods for glaucoma.
[0011] Deep learning, a branch of artificial intelligence (Al), has demonstrated effectiveness in detecting glaucomatous changes on clinical testing modalities such as optic disc photos, optical coherence tomography (OCT), and Humphrey visual fields (HVF). See Wu, J.H., Nishida, T., Weinreb, R.N. and Lin, J.W., 2022. Performances of machine learning in detecting glaucoma using fundus and retinal optical coherence tomography images: a meta-analysis. American Journal of Ophthalmology, 237, pp.1 -12. See also Christopher, M., et al. “Performance of deep learning architectures and transfer learning for detecting glaucomatous optic neuropathy in fundus photographs.” Scientific reports 8.1 (2018): 16685; Liao, W., et al. “Clinical interpretable deep learning model for glaucoma diagnosis.” IEEE journal of biomedical and health informatics 24.5 (2019): 1405-1412; Yu, S., Xiao, D., Frost, S. and Kanagasingam, Y., 2019. Robust optic disc and cup segmentation with deep learning for glaucoma detection. Computerized Medical Imaging and Graphics, 74, pp.61 -71 ; Christopher, M., et al. “Effects of study population, labeling and training on glaucoma detection using deep learning algorithms.” Translational Vision Science & Technology 9.2 (2020): 27-27. Computer vision models such as convolutional neural networks (CNNs), in particular, have the potential to offer objective, quantitative, and high-throughput glaucoma detection capabilities needed for population-based screening. See Diaz-Pinto, A., et al. “CNNs for automatic glaucoma assessment using fundus images: an extensive validation.” Biomedical engineering online 18 (2019): 1 -19. See also Fan, R., et al. “Detecting glaucoma from fundus photographs using deep learning without convolutions: Transformer for improved generalization.” Ophthalmology science 3.1 (2023): 100233; Xu, Y., et al. “Deep learning classifiers for automated detection of gonioscopic angle closure based on anterior segment OCT images.” American journal of ophthalmology 208 (2019): 273-280; Shan, J., et al. “Deep Learning Classification of Angle Closure based on Anterior Segment Optical Coherence Tomography.” Ophthalmology Glaucoma (2023). Recently, a novel deep learning architecture, the vision transformer (ViT), has demonstrated superior performance over CNNs in the detection of glaucoma. See Dosovitskiy, A., et al. “An image is worth 16x16 words: Transformers for image recognition at scale.” arXiv preprint arXiv:2010.11929 (2020). ViT models utilize an attention mechanism to capture feature dependencies and relationships within an image, and have achieved state-of-the-art performance in various computer vision tasks, such as image classification, object detection, and semantic segmentation, all of which are highly relevant to glaucoma detection. See Hugo T., et al. “Training data-efficient image transformers & distillation through attention.” International conference on machine learning. PMLR, 2021 . See also Carion, N., et al. “End- to-end object detection with transformers.” European conference on computer vision. Cham: Springer International Publishing, 2020; Zheng, S., et al. “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers.” Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, 2021 .
[0012] The development and performance of deep learning models, such as ViT models or CNN models, rely heavily on large-scale datasets, with increasingly untenable volume and labeling requirements. Image classification tasks, for instance, can require training CNNs or ViTs with over 10 million images. See Dosovitskiy, A., et al. “An image is worth 16x16 words: Transformers for image recognition at scale.” arXiv preprint arXiv:2010.1 1929 (2020). See also Hugo T., et al. “Training data-efficient image transformers & distillation through attention.” International conference on machine learning. PMLR, 2021. Similarly, achieving state-of-the-art performance in other computer vision tasks like object detection and semantic segmentation can require training backbone models with millions of images. See Carion, N., et al. “End-to-end object detection with transformers.” European conference on computer vision. Cham: Springer International Publishing, 2020. See also Zheng, S., et al. “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers.” Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, 2021 ; Wang, Y., et al. “Investigation of probability maps in deep-learning-based brain ventricle parcellation.” Medical Imaging 2023: Image Processing. Vol. 12464. SPIE, 2023.
[0013] Such requirements and constraints pose significant challenges for data- limited fields such as healthcare and medicine, where data acquisition and labeling is lengthy, resource-intensive and can require specialized expertise. As a result, public glaucoma datasets typically consist of only a few hundred to several thousand images. See Diaz-Pinto, A., et al. “CNNs for automatic glaucoma assessment using fundus images: an extensive validation.” Biomedical engineering online 18 (2019): 1 -19. See also Fang, H., et al. “REFUGE2 Challenge: A Treasure Trove for Multi-Dimension Analysis and Evaluation in Glaucoma Screening.” arXiv preprint arXiv:2202.08994 (2022); Sivaswamy, J., Krishnadas, S.R., Joshi, G.D., Jain, M. and Tabish, A. U.S., 2014, April. Drishti-gs: Retinal image dataset for optic nerve head (onh) segmentation. In 2014 IEEE 11 th international symposium on biomedical imaging (ISBI) (pp. 53- 56), IEEE; Budai, A., Bock, R., Maier, A., Hornegger, J. and Michelson, G., 2013. Robust vessel segmentation in fundus images. International journal of biomedical imaging, 2013; Fumero, F., Alayon, S., Sanchez, J.L., Sigut, J. and Gonzalez-Hernandez, M., 2011 , June. RIM-ONE: An open retinal image database for optic nerve evaluation. In 2011 24th international symposium on computer-based medical systems (CBMS) (pp. 1 -6), IEEE; sjchoi86: sjchoi86- HRF Database. GitHub. https: / / github.com / yiweichen04 / retina_dataset. In addition to limiting model performance, these constraints and limitations also raise concerns regarding the fairness and equity of Al models in this domain. In the field of deep learning, it is widely accepted that an increase in the amount of data generally results in improved model performance. Underrepresented aspects of data sets, such as, for example, underrepresented social groups, contribute significantly less healthcare data, thus potentially propagating their current lack of access to healthcare into future biases in healthcare Als. See Raju, M., Shanmugam, K.P. and Shyu, C.R., 2023. Application of Machine Learning Predictive Models for Early Detection of Glaucoma Using Real World Data. Applied Sciences, 13(4), p.2445.
[0014] The recent emergence of generative Al technologies has introduced the concept of synthetically enhanced datasets as a means to increase model performance while decreasing training data requirements. See Burlina, P.M., Joshi, N., Pacheco, K.D., Liu, T.A. and Bressler, N.M., 2019. Assessment of deep generative models for high-resolution synthetic retinal image generation of age-related macular degeneration. JAMA ophthalmology, 137(3), pp.258-264. See also Goodfellow, I., et al. “Generative adversarial nets.” Advances in neural information processing systems 27 (2014). While exciting, generative Als suffer from the problem of hallucination, a phenomenon in which the model outputs are so divorced from reality that the results are nonsensical. See Rajpurkar, P., Jia, R. and Liang, P., 2018. Know what you don't know: Unanswerable questions for SQuAD. arXiv preprint arXiv:1806.03822. Hallucination is believed to arise from a number of factors, including insufficient or biased training data, the very issues we are looking to address. Embodiments of the present invention circumvent this active problem by employing, for example, an autoencoder, to produce high- fidelity synthetic images that can be used to improve healthcare Al model performance and alleviate our debilitating dependence on big data.
[0015] Methods and systems for enhancing training data for deep learning models are provided. Aspects of the present invention include computer implemented methods of generating an enhanced image-based training data set for deep learning models comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images, selecting a first image from the image-based training data based at least in part on a first characteristic of the first image, generating a first synthetic image based on the first image by modifying a first aspect of the first image, and combining the image-based training data set and the first synthetic image to generate an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image. In other embodiments, methods comprise generating a synthetic image based on first and second images by combining the first image with the second image.
[0016] Aspects of the present invention further include computer implemented methods of training a deep learning model to detect a result using an enhanced training data set, comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images, applying an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on an image of the first subset of images of the image-based training data set, training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images, obtaining an experimental image, and using the trained deep learning model to predict whether a result is present in the experimental image.
[0017] Aspects of the present invention still further include computer implemented methods of training a deep learning model to detect a result using an enhanced training data set, comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images, selecting a first subset of images from the image-based training data based at least in part on a first characteristic of the first subset of images, wherein the first characteristic is associated with a data- collection bias of the image-based training data set, generating synthetic images based on the images of the first subset of images by modifying aspects of each image of the first subset of images such that the synthetic images comprises variations of the first characteristic, training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images, obtaining an experimental image from a population that is not subject to the data-collection bias, and using the trained deep learning model to predict whether a result is present in the experimental image.
[0018] Aspects of the present invention still further include computer implemented methods of generating an enhanced image-based training data set for deep learning models, comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images, and wherein the plurality of images exhibit a first characteristic at a first frequency, selecting a first subset of images from the image-based training data based at least in part on a presence of the first characteristic in each image of the first subset of images, generating a first number of synthetic images based on the first subset of images by modifying aspects of one or more images of the first subset of images, combining the image-based training data set and the first number of synthetic images to generate an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image such that the images of the enhanced image-based training data set exhibit the first characteristic at a second frequency.
[0019] Aspects of the present invention still further include computer implemented methods of generating an enhanced image-based training data set for deep learning models, comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; selecting first and second images from the image-based training data, wherein the first and second images each exhibit a common first characteristic; generating a first synthetic image based on the first and second images by combining the first image with the second image; and combining the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the common characteristic of the first and second images. Other embodiments comprise: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; applying an autoencoder to a first subset of images of the imagebased training data to generate a plurality of synthetic images, wherein each synthetic image is based on a combination of at least two images of the first subset of images of the image-based training data set; training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; obtaining an experimental image; and using the trained deep learning model to predict whether a result is present in the experimental image.
[0020] The methods and systems find use in a variety of different applications, and contexts, e.g., in a healthcare-related context, such as, for example, training Al models to screen subjects for diseases, such as training a ViT model to screen for glaucoma based on an optical disc image. The methods and systems may relate to images obtained using a variety of different imaging modalities. Embodiments of the present invention utilize images obtained using a variety of imaging modalities in connection with a variety of disease applications. For example, images obtained using photographic imaging modalities may be utilized in connection with the following non-limiting exemplary disease applications: glaucoma screening and monitoring, cancer surveillance, degenerative retinal diseases or acquired retinal diseases (e.g. diabetic retinopathy). For example, images obtained using MRI imaging modalities may be utilized in connection with the following non-limiting exemplary disease applications: aneurysm detection and monitoring, multiple sclerosis diagnosis and monitoring, spinal cord pathology, stroke, tumors, traumatic brain injury, joint injuries and diseases, heart and vascular diseases, cancer surveillance, liver disease, Alzheimer’s, epilepsy, peripheral nerve compression, renal disease or inflammatory bowel disease. For example, images obtained using X-ray imaging modalities may be utilized in connection with the following non-limiting exemplary disease applications: fractures, arthritis, lung diseases (TB, pneumonia, COPD, emphysema, pulmonary edema, pneumothorax), dental diseases, bone infections, scoliosis, osteoporosis, kidney stones or detection of foreign objects (e.g. ingested button batteries). For example, images obtained using CT imaging modalities may be utilized in connection with the following non-limiting exemplary disease applications: pulmonary embolism, trauma, stroke, cancer staging, abdominal disorders (appendicitis, diverticulitis, ovarian cysts, tumors), sinusitis, thyroid gland diseases or vascular diseases. For example, images obtained using ultrasound imaging modalities may be utilized in connection with the following non-limiting exemplary disease applications or health conditions: pregnancy monitoring, heart disease diagnosis and monitoring, gallstones, vascular diseases, kidney diseases, ascites, pelvic organ prolapse or testicular torsion.
[0021] BRIEF DESCRIPTION OF THE FIGURES
[0022] The invention may be best understood from the following detailed description when read in conjunction with the accompanying drawings. Included in the drawings are the following figures:
[0023] FIG. 1 illustrates a flow diagram for generating an enhanced image-based training data set for deep learning models according to an embodiment of the present invention.
[0024] FIG. 2 illustrates a flow diagram for generating an enhanced image-based training data set for deep learning models according to another embodiment of the present invention.
[0025] FIG. 3 illustrates a flow diagram for generating an enhanced image-based training data set for deep learning models according to another embodiment of the present invention.
[0026] FIG. 4 illustrates a flow diagram for training a deep learning model to detect a result using an enhanced training data set according to an embodiment of the present invention.
[0027] FIG. 5 illustrates a flow diagram for training a deep learning model to detect a result using an enhanced training data set according to another embodiment of the present invention.
[0028] FIG. 6 illustrates a flow diagram for generating an enhanced image-based training data set for deep learning models by combining first and second images according to another embodiment of the present invention.
[0029] FIG. 7 depicts aspects of, and operation of, an autoencoder according to an embodiment of the present invention. FIG. 8 depicts aspects of, and operation of, an autoencoder according to another embodiment of the present invention.
[0030] FIG. 9 depicts aspects of, and operation of, an autoencoder for generating a synthetic image based on a combination of two input images according to an embodiment of the present invention.
[0031] FIG. 10 panels (a) through (d) present illustrations of autoencoder building blocks according to embodiments of the present invention.
[0032] FIG. 11 depicts exemplary Level 1 encoder and decoder according to embodiments of the present invention.
[0033] FIG. 12 depicts exemplary Level 2 encoder and decoder according to embodiments of the present invention.
[0034] FIG. 13 depicts an exemplary autoencoder with a dual-level structure comprising combined Level 1 and Level 2 encoders and decoders according to embodiments of the present invention.
[0035] FIG. 14, panel (a) depicts an overview of a method of using an autoencoder for generating a synthetic image based on a subject image, according to an embodiment of the present invention. FIG. 14, panel (b) depicts different representations of various types of noise that may be added to an embedding of an input image in embodiments of the present invention.
[0036] FIG. 15, panels (a) and (b) depict exemplary autoencoder-generated synthetic images and corresponding subject images used to generate such synthetic images.
[0037] FIG. 16, panel (a) depicts a schematic diagram of a method of generating an enhanced image-based training data set and using such training data to train a deep learning model to predict a result according to an embodiment of the present invention. FIG. 16, panel (b) depicts aspects of a ViT architecture according to an embodiment of the present invention.
[0038] FIG. 17 depicts a functional block diagram for a computer system according to certain embodiments.
[0039] FIG. 18 depicts a general architecture of an example computing device according to certain embodiments. FIG. 19 depicts a summary of available image-based training data sets used in connection with experimental results obtained applying embodiments of the present invention.
[0040] FIG. 20 depicts a summary of aspects of the experimental results obtained applying embodiments of the present invention.
[0041] FIG. 21 presents experimental results for an image-based training data set showing the AUC vs. percentage of original image-based training data set being used in the training process.
[0042] FIG. 22 illustrates image pre-processing to extract a region around the optic nerve head from a fundus image of an image-based training data set in connection experimental results obtained applying embodiments of the present invention.
[0043] FIG. 23 presents certain training hyperparameters used in connection with obtaining experimental results.
[0044] FIG. 24 presents a training and validation schema according to embodiments used in connection with collecting experimental results.
[0045] FIGS. 25A-25C summarize subject information used in connection with collecting experimental results according to embodiments.
[0046] FIG. 26 presents model performance results under various conditions according to embodiments.
[0047] FIG. 27 presents results of synthetic image generation according to embodiments, where synthetic images were generated based on images of random patient pairs at different pre-defined ratios of scalar values for weighting the constituent images used to generate the synthetic images.
[0048] FIGS. 28A-28B present experimental results of model performance utilizing synthetic image augmentation according to embodiments.
[0049] In the figures, elements having the same or similar reference numerals have the same or similar features, unless explicitly stated otherwise. DETAILED DESCRIPTION
[0050] Aspects of the present invention include computer-implemented methods of generating an enhanced image-based training data set for deep learning models comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images, selecting a first image from the image-based training data based at least in part on a first characteristic of the first image, generating a first synthetic image based on the first image by modifying a first aspect of the first image, and combining the image-based training data set and the first synthetic image to generate an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image. Aspects of the present invention further include methods of training a deep learning model to detect a result using an enhanced training data set, comprising: obtaining an image-based training data set for a deep learning model, wherein the imagebased training data set comprises a plurality of images, applying an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on an image of the first subset of images of the image-based training data set, training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images, obtaining an experimental image, and using the trained deep learning model to predict whether a result is present in the experimental image. In other embodiments, methods comprise generating a synthetic image based on first and second images by combining the first image with the second image. Also provided are systems for performing the methods described herein as well as non-transitory computer readable storage media.
[0051] Before the present invention is described in greater detail, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.
[0052] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.
[0053] Certain ranges are presented herein with numerical values being preceded by the term “about.” The term “about” is used herein to provide literal support for the exact number that it precedes, as well as a number that is near to or approximately the number that the term precedes. In determining whether a number is near to or approximately a specifically recited number, the near or approximating unrecited number may be a number which, in the context in which it is presented, provides the substantial equivalent of the specifically recited number.
[0054] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, representative illustrative methods and materials are now described.
[0055] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates which may need to be independently confirmed.
[0056] It is noted that, as used herein and in the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements, or use of a “negative” limitation.
[0057] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present invention. Any recited method can be carried out in the order of events recited or in any other order which is logically possible.
[0058] While the system and method may be described for the sake of grammatical fluidity with functional explanations, it is to be expressly understood that the claims, unless expressly formulated under 35 U.S.C. § 112, are not to be construed as necessarily limited in any way by the construction of “means” or “steps” limitations, but are to be accorded the full scope of the meaning and equivalents of the definition provided by the claims under the judicial doctrine of equivalents, and in the case where the claims are expressly formulated under 35 U.S.C. § 112 are to be accorded full statutory equivalents under 35 U.S.C. § 1 12.
[0059] METHODS FOR REDUCING DEEP LEARNING DATA REQUIREMENTS
[0060] Aspects of the present invention include computer-implemented methods of generating an enhanced image-based training data set for deep learning models comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images, selecting a first image from the image-based training data based at least in part on a first characteristic of the first image, generating a first synthetic image based on the first image by modifying a first aspect of the first image, and combining the image-based training data set and the first synthetic image to generate an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image. In other embodiments, methods comprise generating a synthetic image based on first and second images by combining the first image with the second image.
[0061] FIG. 1 illustrates a flow diagram 100 for generating an enhanced imagebased training data set for deep learning models according to embodiments of the present invention. The embodiment of the present invention depicted relates to generating an enhanced image-based training data set based on existing images in an image-based training data set. In embodiments, the enhanced image-based training data set comprises additional synthetic images that are sufficiently related to the images of an existing image-based training data set. By sufficiently related, it is meant that such synthetic images can be used, in some cases, along with the original images of the existing image-based training data set, to train a deep learning model, for example, to train a deep learning model to identify image patterns of interest. Image patterns of interest may relate to image features or qualities, e.g., the presence of a structure in an image, and may relate to a variety of different contexts. For example, embodiments of the present invention may be used to generate synthetic images for an enhanced image-based training data set in the context of medicine and healthcare. For example, synthetic images may be generated that exhibit patterns related to disease states, e.g., disease structure. In other cases, synthetic images may be generated that exhibit patterns related to certain subject populations, such as patterns associated with patient populations, such as underrepresented or underserved patient populations.
[0062] Flow diagram 100 is an exemplary embodiment of the present invention provided for illustrative purposes. While certain explanations and examples of aspects of flow diagram 100 may be presented in certain contexts, such as the context of healthcare or medicine, such explanations and examples are provided for illustrative purposes, and it is to be understood that the present invention is not limited to methods solely related to healthcare or medicine, but rather may be employed in any context in which it is desired to augment or enhance imagebased training data sets such that one or more characteristics of the imagebased training data set are not underrepresented when such training data set is used to train a deep learning model.
[0063] Flow diagram 100 starts at step 101. At step 101 , an image-based training data set for a deep learning model is obtained. By image-based training data set, it is meant a collection or a plurality of images. Such images may relate to a common theme or purpose, such as, as described, images originating from a subject population, certain members of which exhibit a disease symptom or a disease structure or a disease state or other medical or health-related condition. In embodiments, the plurality of images of the image-based training data set comprises subject images. By subject images, it is meant images of subject that is generally a human subject and may be male or female of any age and with no specific medical history or history of disease or family history of disease. In some embodiments, the plurality of images of the image-based training data set corresponds to images of a plurality of subjects. In embodiments, when images of the image-based training data set are subject images, each image of the image-based training data set may correspond to a different subject, such that as many subjects are represented in the image-based training data set as images are present therein. In such cases, the plurality of subjects represented in the image-based training data set may exhibit certain common characteristics, such as, for example common demographic characteristics, such as an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location or other demographic or other characteristic. For example, in some cases, such subjects may be subjects that receive medical care at a common facility, such as a common treatment center or the like. In embodiments, the plurality of images of the image-based training data set comprises images of biologic tissues. Any biologic tissue capable of imaging may be applied and is not limited. An example of images of biologic tissues comprise retinal fundus images. Other exemplary images of biologic tissues are discussed throughout herein.
[0064] In embodiments, the image-based training data set comprises any number of images, such as one or more images, ten or more images, 100 or more images, 1 ,000 or more images, 10,000 or more images, 100,000 or more images, 1 ,000,000 or more images, 10,000,000 or more images or 100,000,000 or more images. In embodiments, a first subset of the images in the imagebased training data set are associated with a result, such as, for example, a disease state or other health-related condition, and a second subset are associated with the absence of a result, i.e., the absence of a disease or other health-related condition. Images belonging to such a first subset of images in the image-based training data set typically exhibit a pattern of interest, whereas images belonging to such a second subset of images in the image-based training data set typically do not exhibit the same pattern of interest. Such first subset of images in the image-based training data set may be referred to as positive samples, whereas such second subset of images in the image-based training data set may be referred to as negative or control samples. In embodiments, the image-based training data set comprises a first subset of images that exhibit a condition or a disease and a second subset of images that do not exhibit the condition of the disease. For example, in the context of an image-based training data set for training a deep learning model to predict the presence of glaucoma, a first subset of images of the training data set depict subjects (i.e., biologic tissue of a subject) that exhibits glaucoma or is positive for glaucoma, whereas a second subset of images of the training data set depict subjects (i.e., biologic tissue of subjects) that does not exhibit glaucoma or is negative for glaucoma.
[0065] In embodiments, image-based training data sets comprise labels associated with images indicating whether each image is a positive example or a negative or control example. For example, in the context of an image-based training data set for training a deep learning model to predict the presence of glaucoma, a first subset of images of the training data set may comprises labels indicating the images represent subjects or biologic tissue that are positive for glaucoma, whereas a second subset of images of the training data set may comprises labels indicating the images represent subjects or biologic tissue that are negative for glaucoma. In embodiments, the first and second subsets may comprise any convenient percentage of the image-based training data set. For example, the first subset of images (i.e., positive images or images exhibiting a condition or pattern of interest) may comprise 0.01 % or more, 0.1% or more, 1% or more, 5% or more, 10% or more, 50% or more, 75% or more of the imagebased training data set. For example, the second subset of images (i.e., negative or control images or images exhibiting the absence of a condition or pattern of interest) may comprise 0.01 % or more, 0.1% or more, 1 % or more, 5% or more, 10% or more, 50% or more, 75% or more of the image-based training data set. In embodiments, the image-based training data set may be suitable for training a deep learning model using supervised learning, semi-supervised learning or unsupervised learning or combinations thereof or related techniques such as hold-out or round-robin techniques.
[0066] As described, images of the image-based training data set may be generated using any convenient imaging modality. For example, the imagebased training data set may comprise images generated by one or more of the following modalities: photographic images or MRI images or X-ray images or CT images or ultrasound images. In certain embodiments, at least some of the plurality of images of the image-based training data set exhibit one or more of the following: glaucoma, cancer, a degenerative retinal diseases, an acquired retinal disease, diabetic retinopathy, aneurysm, multiple sclerosis, a spinal cord pathology, stroke, tumor, traumatic brain injury, joint injury or disease, heart disease, vascular disease, a liver disease, Alzheimer’s disease, epilepsy, peripheral nerve compression, renal disease, inflammatory bowel disease, a fracture, arthritis, a lung disease, tuberculosis, pneumonia, chronic obstructive pulmonary disease (COPD), emphysema, pulmonary edema, pneumothorax, a dental disease, a bone infection, scoliosis, osteoporosis, kidney disease, kidney stone, a presence of a foreign object, pulmonary embolism, trauma, stroke, an abdominal disorder, appendicitis, diverticulitis, an ovarian cyst, sinusitis, a thyroid gland disease, pregnancy, gallstone, ascites, pelvic organ prolapse or testicular torsion. In such embodiments, the image-based training data set may be configured for use training a deep learning model to recognize the presence of such conditions. For example, the training data set may be configured for use training a deep learning model to recognize the presence of glaucoma in an experimental image. In such example, a first subset of images exhibit the presence of glaucoma, and a second subset of images exhibit the absence of glaucoma.
[0067] As described, in embodiments of the present invention, the image-based training data set is a training data set for training a deep learning model. Deep learning models refer to any machine learning technique comprising an artificial neural network. Similarly, in embodiments of the present invention, the enhanced image-based training data set is a training data set for training a deep learning model. In embodiments, the image-based training data set is configured for training a deep learning model to perform one or more of: image classification, object detection or semantic segmentation. In some embodiments, the image-based training data set is configured for training a deep learning model to diagnose or screen for a disease or monitor a disease state or disease progression. In certain embodiments, the image-based training data set is a training data set for a convolutional neural network (CNN) or a vision transformer (ViT). In other embodiments, the image-based training data set is a training data set for image-based disease diagnosis or disease monitoring.
[0068] In embodiments, the image-based training data set comprises a bias. By bias, it is meant that the images of the image-based training data set are not reflective of a desired population (i.e., population of interest) or of certain aspects of a desired population (i.e., population of interest). For example, images of the image-based training data set may exhibit a bias such that such images relate to a first population of subjects that is not representative of a second population. For example, the images of the image-based training data set may be associated with a first population of subjects that exhibit a bias relative to the population of subjects on which the trained deep learning model may be applied. For example, the images of the image-based training data set may be associated with a first population of subjects in which there is a preponderance of a particular age or a particular gender or a particular race or a particular ethnicity or a particular culture or a particular socioeconomic characteristic or a particular geographic location or other characteristic, whereas the population of subjects on which the trained deep learning model may be applied does not exhibit such preponderance of such characteristic. For example, in a case where images of the image-based training data set are associated with an exclusively male subject population, such image-based training data set comprises a bias. Specifically, such image-based training data set comprises a bias relative to a second population that is 50% female and 50% male. For example, in a case where images of the image-based training data set are associated with a subject population that is exclusively over the age of 50, such image-based training data set comprises a bias. Specifically, such image-based training data set comprises a bias relative to a second population with 50% of subjects over age 50 and 50% of subjects under age 50.
[0069] In some cases, when the image-based training data set exhibits a bias and a deep learning model is trained using such biased training data, the deep learning model reflects such bias. In certain cases, such bias exhibited by the deep learning model degrades performance of the deep learning model when such deep learning model is applied to a second population without such bias. For example, in a case where images of the image-based training data set are associated exclusively with a first characteristic (e.g., an exclusively male population) and such image-based training data set is used to train a deep learning model, then subsequently applying the trained deep learning model to a second population that is not associated exclusively with the first characteristic (e.g., a 50% male and 50% female population) may cause the deep learning model to fail to perform as effectively or otherwise cause model performance to degrade. In the context of healthcare and medicine, such degraded performance may be associated with worse health-related outcomes when the deep learning model is used in connection with diagnosing, or screening for, disease or otherwise augmenting medical care. Biases of interest may comprise any potential bias that an image-based training data set may be subject to. For example, biases may comprise one or more of a bias related to an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location or the like.
[0070] Upon completing step 101 of flow diagram 100, the process next moves to step 102 of flow diagram 100.
[0071] At step 102, a first image from the image-based training data is selected based at least in part on a first characteristic of the first image. As described, the first characteristic may relate to any discernable characteristic of images of the image-based training data set. Characteristics of interest may relate to the image itself, such as a characteristic of an imaged structure or, when the image is a subject image, may relate to a characteristic of the subject, for example. For example, the first characteristic may relate to the presence of a structure, such as a tumor or cancer tissue or a fracture or other abnormal structure or growth. In some embodiments, in which the first image is a subject image, the first characteristic optionally comprises one or more of: a diagnosis or a disease state or an injury state or a condition of the subject of the first image. In other embodiments, in which the first image is a subject image, the first characteristic is a demographic characteristic of the subject optionally comprising one or more of: an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location of the subject. In some cases, the first characteristic is relatively rare in the image-based training data set, such as, for example, exhibited by 50% or less or 25% or less or 10% or less or 5% or less or 4% or less or 3% or less or 2% or less or 1 % or less or 0.1% or less or 0.01 % or less or 0.001 % or less of the images of the image-based training data set. In some cases, the first image is a subject image, and it is relatively rare for a subject in the image-based training data set to exhibit the first characteristic, such as, for example, exhibited by 50% or less or 25% or less or 10% or less or 5% or less or 4% or less or 3% or less or 2% or less or 1 % or less or 0.1 % or less or 0.01 % or less or 0.001 % or less of the subjects associated with images in the image-based training data set. In other cases, the first characteristic is underrepresented in the image-based training data set. For example, the first characteristic may be represented in the image-based training data set at a frequency that is lower than the frequency at which the first characteristic is observed in another population, such as a potential subject population or a potential patient population or a national population or another population at large. In still other embodiments, the first image is an image of a rare disease. By rare disease, it is meant, for example, a disease occurring with a frequency of less than 1 in 10,000 or less than 1 in 100,000 in a population. Diagnosing rare diseases can require clinical expertise, which can be in limited supply, and significant diagnostic delays can be common.
[0072] In embodiments, the first image corresponds to an image of an underrepresented group. In certain embodiments, the first image is a subject image, and the subject is a member of an underrepresented group. By underrepresented group, it is meant, for example, that a characteristic, such as a demographic characteristic, such as a socioeconomic aspect, of a subject is not represented in the image-based training data set at a level that is comparable with the level at which another population exhibits such demographic characteristic. For example, a majority of images of an imaged based training set may be of subjects exhibiting specific socioeconomic characteristics whereas another population, such as a population of potential patients at a healthcare center at which a deep learning model may be employed in connection with patient care, may not exhibit such socioeconomic characteristics. Employing embodiments of the present invention facilitates training deep learning models such that they are not subject to, or less subject to, certain biases that arise when exposed to, i.e., trained on, such limited training data sets.
[0073] In embodiments, the first image is selected to address a bias in the imagebased training data set. In certain embodiments, the first image is selected to enhance performance of the deep learning model. In other embodiments, the first image is selected to enhance performance of the deep learning model across a more diverse range of subjects or across a more diverse range of images.
[0074] Upon completing step 102 of flow diagram 100, the process next moves to step 103 of flow diagram 100.
[0075] At step 103, a first synthetic image is generated based on the first image by modifying a first aspect of the first image. In embodiments, generating a first synthetic image based on the first image comprises modifying a representation of a structure present in the first image. By structure present in the first image, it is meant any convenient aspect of the first image, such as, for example, a biological or anatomical structure present in, i.e. , depicted by, the image. Exemplary structures of interest include, without limitation, retinal fundus, optic nerve or optic vasculature. By representation of a structure present in the first image, it is meant, in embodiments, a digital representation, such as a representation based on, or reflected in or by, a computer data structure. By modifying a structure, it is meant adjusting one or more aspects of such representation, i.e., such data representation or data structure, in any convenient manner, such as, for example, adding noise or other stochastic patterns or other non-stochastic variations to aspects of the representation. In some embodiments, generating a first synthetic image based on the first image comprises introducing a variation in a structure present in the first image. In other embodiments, generating a first synthetic image based on the first image comprises introducing random variations into the first image. In still other embodiments, generating a first synthetic image based on the first image comprises introducing a variation in the first image other than mirroring, rotating, smoothing, or contrast reduction of the first image.
[0076] In embodiments, generating a first synthetic image based on the first image comprises applying a tool or an algorithm, e.g., a data processing technique or image processing technique, to the first image. For example, in some embodiments, generating a first synthetic image based on the first image comprises: applying an autoencoder or a diffusion-based model or a generative adversarial network (GAN) to the first image to generate the first synthetic image. Embodiments of the present invention leverage an autoencoder to generate synthetic images whereby controlled variations can be introduced to, for example, real-world medical images. See Kingma, D.P. and Welling, M., 2019. An introduction to variational autoencoders. Foundations and Trends® in Machine Learning, 12(4), pp.307-392. See also Hu, X., Chen, Y.J., Ho, T.Y. and Shi, Y., 2023. Conditional Diffusion Models for Weakly Supervised Medical Image Segmentation. arXiv preprint arXiv:2306.03878. Such autoencoder-based variations go beyond simple rotations, for example. In other embodiments, generating a first synthetic image based on the first image comprises applying an autoencoder to the first image to generate the first synthetic image, wherein the autoencoder is a trained autoencoder. In some cases, the autoencoder is trained on a second image-based training data set, wherein the second image-based training data set does not comprise any images of the image-based training data set. In embodiments, applying an autoencoder to the first image to generate the first synthetic image comprises: introducing a controlled variation of the first image into the first synthetic image, i.e., an intended variation or deliberate variation.
[0077] In embodiments, applying an autoencoder to the first image to generate the first synthetic image comprises: encoding the first image into an embedding, and decoding the embedding into the first synthetic image, wherein the embedding comprises a representation of the first image. By embedding it is meant any convenient representation of an image, such as, for example, an array, such as a one dimensional array or a two-dimensional array, of numerical values, such as floating-point values, representing the first image. In some embodiments, the embedding comprises a compressed representation of the first image. In certain embodiments, the embedding comprises a compressed stochastic representation of the first image. Other embodiments of the present invention further comprise introducing noise into the embedding. In some cases, the noise introduced into the embedding comprises one or more of uniform noise or step noise or random noise or patterned noise, wherein patterned noise is associated with a mathematical function, optionally comprising one or more of a sine or cosine function. As described, in some cases, the embedding comprises a linear representation of the first image. In certain cases, the noise introduced into the embedding comprises a linear representation of noise values.
[0078] In embodiments utilizing an autoencoder in connection with generating a first synthetic image, the autoencoder comprises an encoder and a decoder. In embodiments, the autoencoder comprises a dual-level structure. In other embodiments, the autoencoder comprises: an encoder and a decoder, wherein the encoder comprises a first level encoder configured to encode a subset of the first image, and a second level encoder configured to encode an output of the first level encoder, and the decoder comprises a first level decoder, configured to decode a subset of the first synthetic image, and a second level decoder configured to decode an output of the first level decoder. In embodiments, the encoder and decoder are modular components, meaning different encoders and decoders, or different hierarchical levels of encoders and decoders, can be utilized in different combinations to generate different results or different types of results. In some cases, the encoder and decoder are independently trained. In other cases, each level of the encoder and decoder are independently trained. In embodiments, the encoder and decoder comprise one or more of neural networks or residual networks. In some embodiments, each level of the encoder and decoder comprises a neural network.
[0079] Upon completing step 103 of flow diagram 100, the process next moves to step 104 of flow diagram 100.
[0080] At step 104, the image-based training data set and the first synthetic image are combined to generate, i.e., to form, an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image. Some embodiments of the present invention further comprise generating a plurality of synthetic images based on the first image by one or more modifications to one or more aspects of the first image, and combining the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models. Other embodiments of the present invention further comprise generating a plurality of synthetic images, wherein the synthetic images are generated based on more than one of the plurality of images, and combining the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models. In some embodiments, the imagebased training data does not reflect a demographic characteristic of a select population, and the enhanced image-based training data set is generated to reflect the demographic characteristic of the select population. In other embodiments, the enhanced image-based training data set comprises at least one synthetic image corresponding to each image of a plurality of images of the image-based training set. In still other embodiments, the enhanced image-based training data set comprises at least one synthetic image corresponding to each image of the image-based training set. In some cases, the enhanced imagebased training data set comprises three synthetic images corresponding to each image of the image-based training set. In other cases, the image-based training data set reflects a first population of subjects, and the method is a method of generating an enhanced image-based training data set for training a deep learning model for use with individuals that are not adequately represented in the image-based training data set.
[0081] As described, embodiments of the present invention comprise generating synthetic images and combining those with the image-based training data set to form an enhanced image-based training data set. In embodiments, synthetic images are generated from, or based on, a selection of one or more images within the image-based training set training set. That is, in embodiments, the synthetic images that are added to the enhanced image-based training data set are “selected” by selecting specific original images of the image-based training data set as the basis for synthesis of new images. This approach of selecting which of the images of the image-based training data set are used to generate synthetic images, is particularly useful for addressing unrepresentative samples within the training data, i.e., the image-based training data set. For instance, in training a model, such as a deep learning model, to identify a rare disease, the scarcity of images depicting the disease in the dataset, i.e., the image-based training data set, may necessitate generating additional synthetic images to bolster this category and subsequently bolstering the model’s effectiveness at identifying such rare disease.
[0082] Moreover, addressing biases in medical datasets is an important consideration in generating deep learning models for use in medical contexts. Disparities in healthcare access among different groups, such as different racial or cultural or socioeconomic groups, can lead to underrepresentation in the data. By synthesizing images that reflect aspects of these underrepresented groups, embodiments of the present invention can be used to enhance model performance across a more diverse range of individuals. That is, the enhanced imaged-based training data set, generated in embodiments of the present invention and resulting from step 104, can be used to address various types of issues associated with training datasets, such as issues related to available training data being small or limited or available training data being imbalanced in a way that one or more certain types of data is not representative, i.e., not adequately representative to train a model for use on a certain population.
[0083] Upon completing step 104 of flow diagram 100, the process ends.
[0084] FIG. 2 illustrates a flow diagram 200 for generating an enhanced imagebased training data set for deep learning models according to embodiments of the present invention. The embodiment of the present invention depicted relates to generating an enhanced image-based training data set based on existing images in an image-based training data set. Descriptions of aspects of the steps of flow diagram 200 that are elsewhere described are not duplicated in connection with flow diagram 200.
[0085] Flow diagram 200 is an exemplary embodiment of the present invention provided for illustrative purposes. While certain explanations and examples of aspects of flow diagram 200 may be presented in the context of healthcare or medicine, such explanations and examples are provided for illustrative purposes, and it is to be understood that the present invention is not limited to methods solely related to healthcare or medicine, but rather may be employed in any context in which it is desired to augment or enhance image-based training data sets such that one or more characteristics of the image-based training data set are not underrepresented when such training data set is used to train a deep learning model.
[0086] Flow diagram 200 starts at step 201 . At step 201 , an image-based training data set for a deep learning model is obtained, wherein the image-based training data set comprises a plurality of images.
[0087] Upon completing step 201 of flow diagram 200, the process next moves to step 202 of flow diagram 200.
[0088] At step 202, a first subset of images is selected from the image-based training data based at least in part on a first characteristic of the images of the first subset of images. Any convenient number of images may be selected for the first subset of images, such as one image, two images, three images, four images, five images, ten images, 20 images, 30 images, 40 images 50 images, 100 images, 200 images, 300 images, 400 images, 500 images, 1 ,000 images or more images. Images of the first subset of images may be selected based on the presence of a first characteristic by, for example, visual inspection by an individual, such as a trained professional, such as a medical doctor or another trained individual. Such first subset may be manually selected, i.e., selected by an individual, to identify images exhibiting a first characteristics, such as, for example, a bone fracture. In other cases, images of the first subset of images may be selected based on an automated selection process or based on labels or other information, e.g., metadata, associated with images of the image-based training set, such as label information identifying one or more demographic characteristics of a subject depicted in an image. For example, the first subset of images may be automatically selected to comprise images of male subjects, or the first subset of images may be automatically selected to comprise images of female subjects, or the first subset of images may be automatically selected to comprise images of subjects over age 50, or the first subset of images may be automatically selected to comprise images of subjects under age 50. Such first subset may be automatically selected by identifying and selecting based on one or more labels, e.g., labels indicating a gender or an age of a subject depicted in an image.
[0089] Upon completing step 202 of flow diagram 200, the process next moves to step 203 of flow diagram 200.
[0090] At step 203, synthetic images are generated based on images of the first subset of images by modifying one or more aspects of images of the first subset of images. In embodiments, any number of synthetic images may be generated based on one or more images of the first subset of images. In embodiments, any number of the images of the first subset of images may be used to generate synthetic images, such as one image of the first set of images, 50% of the images of the first set of images or 100% of the images of the first subset of images. In embodiments, one or more synthetic images may be generated by modifying one or more aspects of images of the first subset of images. Such modifications may vary as desired. In some cases, the same modification or the same type of modification of each image of the first subset of images may be employed (e.g., adding the same or similar value to an embedding representation of the images), and in other cases, different modifications or different types of modifications of each image of the first subset of images may be employed (e.g., adding different or different types of values to an embedding representation of the images).
[0091] Upon completing step 203 of flow diagram 200, the process next moves to step 204 of flow diagram 200.
[0092] At step 204, the image-based training data set and the synthetic images generated at step 203 are combined into an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image.
[0093] Upon completing step 204 of flow diagram 200, the process ends. FIG. 3 illustrates a flow diagram 300 for generating an enhanced imagebased training data set for deep learning models according to embodiments of the present invention. The embodiment of the present invention depicted relates to generating an enhanced image-based training data set based on existing images in an image-based training data set. Descriptions of aspects of the steps of flow diagram 300 that are described elsewhere are not duplicated in connection with flow diagram 300.
[0094] Flow diagram 300 is an exemplary embodiment of the present invention provided for illustrative purposes. While certain explanations and examples of aspects of flow diagram 300 may be presented in the context of healthcare or medicine, such explanations and examples are provided for illustrative purposes, and it is to be understood that the present invention is not limited to methods solely related to healthcare or medicine, but rather may be employed in any context in which it is desired to augment or enhance image-based training data sets such that one or more characteristics of the image-based training data set are not underrepresented when such training data set is used to train a deep learning model.
[0095] Flow diagram 300 starts at step 301 . At step 301 , an image-based training data set for a deep learning model is obtained, wherein the image-based training data set comprises a plurality of images, and wherein the plurality of images exhibit a first characteristic at a first frequency. The first frequency may comprise any convenient frequency, such as 100% or 90% or 80% or 70% or 60% or 50% or 40% or 30% or 20% or 10% or 9% or 8% or 7% or 6% or 5% or 4% or 3% or 2% or 1 % or 0.1 % or 0.01 % or 0.001% of images of the imagebased training data set exhibit the first characteristic. That is, in embodiments the first characteristic may be exhibited in, or associated with, 100% or 90% or 80% or 70% or 60% or 50% or 40% or 30% or 20% or 10% or 9% or 8% or 7% or 6% or 5% or 4% or 3% or 2% or 1 % or 0.1 % or 0.01 % or 0.001 % of images of the image-based training data set.
[0096] Upon completing step 301 of flow diagram 300, the process next moves to step 302 of flow diagram 300. At step 302, a first subset of images from the image-based training data are selected based at least in part on a presence of the first characteristic in each image of the first subset of images. In embodiments, the presence of the first characteristic need not imply the presence of a particular structure or pattern or aspect in an image or a subject thereof but may also reflect, for example, the absence of a particular structure or pattern or aspect in an image of a subject thereof. In embodiments, the first subset comprises any convenient number of images. In some cases, the first subset comprises one, two, three, four, five, six, seven, eight, nine, ten, 100, 1 ,000, 10,000, 100,000, 1 ,000,000 or more images. In other cases, the first subset comprises 100% or 50% or 40% or 30% or 20% or 10% or 5% or 4% or 3% or 2% or 1% or less than 1% of images of the imagebased training data set that exhibit the first characteristic.
[0097] Upon completing step 302 of flow diagram 300, the process next moves to step 303 of flow diagram 300.
[0098] At step 303, a first number of synthetic images based on the first subset of images are generated by modifying aspects of one or more images of the first subset of images. Any convenient number of synthetic images may be generated at step 303 and such may be determined based on, for example, a desired second frequency, as described below in connection with step 304.
[0099] Upon completing step 303 of flow diagram 300, the process next moves to step 304 of flow diagram 300.
[0100] At step 304, the image-based training data set and the first number of synthetic images are combined to generate an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image such that the images of the enhanced image-based training data set exhibit the first characteristic at a second frequency. In embodiments, the second frequency may be selected as desired and typically differs from the first frequency. For example, the second frequency may be selected such that the frequency at which the first characteristic appears in the enhanced imagebased training data set corresponds to a frequency at which the first characteristic appears in a subject population, such as a patient population. Typically, the second frequency is greater than the first frequency. The second frequency may be greater than the first frequency by any convenient amount as such depends in part on the number of synthetic images generated at step 303.
[0101] Upon completing step 304 of flow diagram 300, the process ends.
[0102] FIG. 4 illustrates flow diagram 400 for training a deep learning model to detect a result using an enhanced training data set according to embodiments of the present invention. The embodiment of the present invention depicted relates to training a deep learning model to detect a result using an enhanced training data set. Descriptions of aspects of the steps of flow diagram 400 that are described elsewhere are not duplicated in connection with flow diagram 400.
[0103] Flow diagram 400 is an exemplary embodiment of the present invention provided for illustrative purposes. While certain explanations and examples of aspects of flow diagram 400 may be presented in the context of healthcare or medicine, such explanations and examples are provided for illustrative purposes, and it is to be understood that the present invention is not limited to methods solely related to healthcare or medicine, but rather may be employed in any context in which it is desired to augment or enhance image-based training data sets to improve training a deep learning model to detect a result using an enhanced training data set.
[0104] Flow diagram 400 starts at step 401 . At step 401 , an image-based training data set for a deep learning model is obtained, wherein the image-based training data set comprises a plurality of images.
[0105] Upon completing step 401 of flow diagram 400, the process next moves to step 402 of flow diagram 400.
[0106] At step 402, a first subset of images from the image-based training data set is selected based at least in part on a first characteristic of the first subset of images, wherein the first characteristic is associated with a data-collection bias of the image-based training data set. By data collection bias, it is meant any potential bias that may affect collection of an image-based training data set. That is, in embodiments, a data collection bias may comprise a bias such that the image-based training data set is not a random sample with respect to any selected characteristic. For example, the data collection bias may reflect that subjects depicted in the image-based training data set are not randomly sampled. For example, the data collection bias may reflect that subjects depicted in the image-based training data set receive medical care in one or more healthcare settings, e.g., one or more academic healthcare settings, such that the subjects are not geographically or socioeconomically diverse (or randomly sampled) or otherwise are not representative of a larger population or a more diverse population. That is, in embodiments, the image-based training data set does not reflect the first characteristic in a way that is representative of how the first characteristic would be represented in images of a different population, e.g., a broader population or a more diverse population.
[0107] Upon completing step 402 of flow diagram 400, the process next moves to step 403 of flow diagram 400.
[0108] At step 403, synthetic images are generated based on the images of the first subset of images by modifying aspects of each image of the first subset of images such that the synthetic images comprises variations of the first characteristic. Therefore, upon completion of step 403, synthetic images are generated such that an additional number of images are available that reflect the first characteristic and / or additional images are generated which make available additional variations of how the first characteristic is reflected in images. That is, at step 403, both or either of the following are generated: (i) the number of images that reflect the first characteristic as well as (ii) additional variations of how the first characteristic is represented. Different embodiments may be configured to emphasize making a greater number of images depicting the first characteristic available or generating additional variations of how the first characteristic is depicted or combinations thereof.
[0109] Upon completing step 403 of flow diagram 400, the process next moves to step 404 of flow diagram 400.
[0110] At step 404, a deep learning model is trained to detect a result using images from the image-based training data set and the plurality of synthetic images. By detecting a result, it is meant estimating a likelihood that a result is present in an image that is presented as input to the deep learning model. In some embodiments, detecting a result refers to estimating a likelihood that a subject would be diagnosed with a disease using techniques available in the prevailing standard of care. For example, detecting a result may refer to estimating a likelihood that an image of subject indicates that the subject has glaucoma.
[0111] Any convenient deep learning model may be applied, such as, for example, a vision transformer (ViT) or a convolutional neural network (CNN). The deep learning model may be trained to detect any result of interest, for which it is possible to train a deep learning model. For example, in certain embodiments, the deep learning model is trained, using images from the imagebased training data set and the plurality of synthetic images, for image-based disease diagnosis or disease monitoring or the like using the enhanced imagebased training data set. Any convenient technique for training a deep learning model may be employed, such as, for example, supervised learning, semisupervised learning, unsupervised learning, in each case, employing any of a held-out approach or a round robin approach or the like. Images from the imagebased training data set and the plurality of synthetic images may be divided and utilized as training and validation data in any convenient manner and such may vary.
[0112] In embodiments, training a model refers to configuring, fitting or otherwise preparing a model to make predictions, such as, for example, to estimate a result, such as, for example, the presence of a disease such as glaucoma, and is distinguished from applying a model to make predictions about estimating such result (i.e., training the model is distinct from applying the model to non-training data). With respect to training a model, embodiments of a model are trained using at least some images of the image-based training data set identified at step 401 as well as at least some of the synthetic images generated at step 403. In embodiments, the model is trained using one or more subsets of the imagebased training data set and synthetic images, i.e., training data, which may comprise instances of health records corresponding to the result of interest, such as, for example, the presence of a disease such as glaucoma, as well as images corresponding to an absence of such result of interest, i.e., images ultimately known to not exhibit a disease such as glaucoma.
[0113] In embodiments, training a model using at least a subset of the training data comprises applying an unsupervised learning technique to the model. Unsupervised learning is a machine learning technique known in the art for training a model to, for example, identify or recognize patterns. Unsupervised learning comprises training a model where pre-assigned labels (e.g., regarding the presence of a disease such as glaucoma) are not provided to the model with respect to data used to train the model. That is, in the case of such embodiments of the invention, no labels indicating whether training images correspond to images of subjects known to have (or known not to have) a disease such as glaucoma are provided in connection with training the model. As a result, applying unsupervised learning to train a model entails the model itself discovering patterns among the training data.
[0114] In other embodiments, training a model using at least a subset of the training data comprises applying a semi-supervised learning technique to the model. Semi-supervised learning is a machine learning technique known in the art for training a model to, for example, identify or recognize patterns. Semisupervised learning comprises training a model using both labeled and unlabeled training data.
[0115] In some embodiments, training a model using at least a subset of the training data comprises applying a supervised learning technique to the model. Supervised learning is a machine learning technique known in the art for training a model to, for example, identify or recognize patterns based at least in part on, in this case, images corresponding to a result of interest, such as, for example, images of subjects know to be diagnosed with a disease such as glaucoma as well as images corresponding an absence of the result, such as images corresponding to subjects known not to have the disease such as glaucoma. For purposes of training the model, such different images may be “labeled” based on disease state. Supervised learning comprises training a model using such labeled training data.
[0116] In some embodiments, training and evaluating the model at step 404 involves training and testing a deep learning model for estimating a result using a held out test set and cross validation evaluations.
[0117] Some embodiments apply a round robin training technique to the model at step 404. By round robin training technique, it is meant that input data to the model (i.e., data used to train the model) is divided into multiple partitions such that certain partitions are used to train the model and the remaining partitions are used to generate predictions using the trained model. Such process may be iterated where partitions of data previously used to train the model are subsequently used to generate predictions using the trained model. Round robin training approaches may offer benefits including identifying which data sets used for training result in more accurate predictions.
[0118] In embodiments, training the model at step 404 comprises assessing the accuracy of, or confidence in, predictions made by the model. That is, whether a prediction based on the model that a result is associated with an input image represents a true positive or a false positive and / or whether a prediction based on the model that the result is not associated with the input image represents a true negative or a false negative, as well as how frequently such errors occur.
[0119] In embodiments, the model may be configured to generate information about the confidence of a prediction generated by the model. In such embodiments, training the model may comprise initially applying the model to obtain initial predictions, and using at least a subset of the initial predictions to further train the model, where the subset of initial predictions used to further train the model correspond to higher confidence predictions. That is, in some embodiments, training the model comprises generating a prediction as well as an indication of the degree of confidence that the prediction is accurate.
[0120] In other embodiments, training a deep learning model to estimate a result, such as to estimate the presence of a disease, using an image comprises utilizing one or more of stratified sampling, held out test set, cross validation, hyper parameter tuning or random grid search. In some cases, evaluating the accuracy of estimates of a result, such as estimates of the presence of the disease comprises using any available technique for assessing the accuracy of a model, such as, for example, using one or more of held out test sets and cross validation evaluations. In other cases, embodiments comprise further training the model. In embodiments, further training the deep learning model using training data images comprises training the model using the entirety of the image-based training data set and / or the entirety of available synthetic images.
[0121] Upon completing step 404 of flow diagram 400, the process next moves to step 405 of flow diagram 400.
[0122] At step 405, an experimental image is obtained. In embodiments, the experimental image may be obtained from a population that is not subject to the same data collection bias described above. In embodiments, the experimental image may be an image of a subject not included in the image-based training data set. In embodiments, the experimental image may be an image of a subject with one or more demographic characteristics not reflected in subjects of the image-based training data set. In embodiments, the experimental image may be an image of a subject exhibiting one or more characteristics associated with the plurality of synthetic images. That is, the experimental image may exhibit such characteristics that generating the synthetic images at step 403 and using such images to train the deep learning model at step 404 allowed the deep learning model to more accurately detect the desired result associated with the experimental image. In other words, the experimental image may exhibit such characteristics that had synthetic images not been generated at step 403 and had such images not been used to train the deep learning model at step 404, the deep learning model may not have been able to detect the desired result associated with the experimental image. Such experimental image may be obtained in, for example, a clinical setting, such as a healthcare clinic, in which, for example, a new subject is treated as a patient. Such experimental image may be an image obtained using any convenient imaging modality, as such are described herein, and the imaging modality of the experimental image may be the same as the imaging modality used to obtain the image-based training data set.
[0123] In embodiments in which the deep learning model is trained to estimate the presence of a disease in a subject based on a subject image, the experimental image obtained at step 405 may correspond to a subject who belongs to an at-risk population for the disease. That is, the subject to be evaluated for diagnosis may be selected based on some predetermined criteria. Any applicable predetermined criteria may be applied, and such may vary. In some cases, the at-risk population comprises subjects exhibiting specified symptoms. For example, the specified symptoms may comprise complaints of blurred vision or elevated blood pressure.
[0124] Upon completing step 405 of flow diagram 400, the process next moves to step 406 of flow diagram 400.
[0125] At step 406, the trained deep learning model is used to predict whether a result is present in the experimental image. In embodiments, such prediction may take the form of a percentage reflecting a likelihood that a result is detected, for example.
[0126] In embodiments in which the deep learning model is trained to estimate the presence of a disease in a subject based on a subject image, a subject corresponding to an image, which is estimated by the model to exhibit the disease, may be subjected to confirmation of such diagnosis. In other words, at step 406, further testing may be applied to evaluate whether the deep learning model indicated a true positive or a false positive. Any applicable and / or relevant confirmation may be applied. For example, at step 406, the subject (i.e., the patient) may undergo a biochemical test and / or a clinical examination in order to confirm the diagnosis of the disease. In embodiments, any available relevant biochemical tests and / or clinical examination to confirm a diagnosis may be applied. In the event the diagnosis is not confirmed, further evaluation of the subject may be conducted to confirm whether the deep learning model estimated a false positive. In embodiments, clinically evaluating the presence of the disease comprises clinical evaluation of the subject by a specialist. In some cases, clinically evaluating the presence of the disease comprises conducting biochemical analysis.
[0127] Upon completing step 406 of flow diagram 400, the process ends.
[0128] FIG. 5 illustrates flow diagram 500 for training a deep learning model to detect a result using an enhanced training data set according to embodiments of the present invention. The embodiment of the present invention depicted relates to training a deep learning model to detect a result using an enhanced training data set generated based on existing images in an image-based training data set. Descriptions of aspects of the steps of flow diagram 500 that are already described are not duplicated in connection with flow diagram 500.
[0129] Flow diagram 500 is an exemplary embodiment of the present invention provided for illustrative purposes. While certain explanations and examples of aspects of flow diagram 500 may be presented in the context of healthcare or medicine, such explanations and examples are provided for illustrative purposes, and it is to be understood that the present invention is not limited to methods solely related to healthcare or medicine, but rather may be employed in any context in which it is desired to augment or enhance image-based training data sets such that one or more characteristics of the image-based training data set are not underrepresented when such training data set is used to train a deep learning model and such model is used to detect a result using such an enhanced training data set.
[0130] Flow diagram 500 starts at step 501 . At step 501 , an image-based training data set for a deep learning model is obtained, wherein the image-based training data set comprises a plurality of images.
[0131] Upon completing step 501 of flow diagram 500, the process next moves to step 502 of flow diagram 500.
[0132] At step 502, an autoencoder is applied to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on an image of the first subset of images of the image-based training data set. In embodiments, the first subset of images may be selected according to any convenient protocol, including, for example, randomly selecting images of the image-based training data set or selecting images exhibiting a characteristic, such as, for example, where the subject of the image has received a diagnosis of glaucoma. In embodiments, any convenient autoencoder may be applied. Such autoencoder may be a trained autoencoder. Such trained autoencoder may be trained according to any convenient training protocol.
[0133] Upon completing step 502 of flow diagram 500, the process next moves to step 503 of flow diagram 500.
[0134] At step 503, a deep learning model is trained to detect a result using images from the image-based training data set and the plurality of synthetic images.
[0135] Upon completing step 503 of flow diagram 500, the process next moves to step 504 of flow diagram 500.
[0136] At step 504, an experimental image is obtained.
[0137] Upon completing step 504 of flow diagram 500, the process next moves to step 505 of flow diagram 500.
[0138] At step 505, the trained deep learning model is used to predict whether a result is present in the experimental image.
[0139] Upon completing step 505 of flow diagram 500, the process ends.
[0140] FIG. 6 illustrates a flow diagram 600 for generating an enhanced imagebased training data set for deep learning models according to embodiments of the present invention. The embodiment of the present invention depicted relates to generating an enhanced image-based training data set based on existing images in an image-based training data set. Descriptions of aspects of the steps of flow diagram 600 that are described elsewhere are not duplicated in connection with flow diagram 600.
[0141] Flow diagram 600 is an exemplary embodiment of the present invention provided for illustrative purposes. While certain explanations and examples of aspects of flow diagram 600 may be presented in the context of healthcare or medicine, such explanations and examples are provided for illustrative purposes, and it is to be understood that the present invention is not limited to methods solely related to healthcare or medicine, but rather may be employed in any context in which it is desired to augment or enhance image-based training data sets such that one or more characteristics of the image-based training data set are not underrepresented when such training data set is used to train a deep learning model.
[0142] Flow diagram 600 starts at step 601 . At step 601 , an image-based training data set for a deep learning model is obtained, wherein the image-based training data set comprises a plurality of images, and wherein the plurality of images exhibit a first characteristic at a first frequency. The first frequency may comprise any convenient frequency, such as 100% or 90% or 80% or 70% or 60% or 50% or 40% or 30% or 20% or 10% or 9% or 8% or 7% or 6% or 5% or 4% or 3% or 2% or 1 % or 0.1 % or 0.01 % or 0.001% of images of the imagebased training data set exhibit the first characteristic. That is, in embodiments the first characteristic may be exhibited in, or associated with, 100% or 90% or 80% or 70% or 60% or 50% or 40% or 30% or 20% or 10% or 9% or 8% or 7% or 6% or 5% or 4% or 3% or 2% or 1 % or 0.1 % or 0.01 % or 0.001 % of images of the image-based training data set. For example, the first characteristic may be a first label, e.g., a label that represents a subject diagnosed with a particular disease or condition or a label that represents that a subject is not diagnosed with the particular disease or condition. In other cases, the first characteristic may relate to a condition exhibited in the images, such as, for example, the presence or absence of abnormal tissue in the images.
[0143] Upon completing step 601 of flow diagram 600, the process next moves to step 602 of flow diagram 600.
[0144] At step 602, a first image and a second image are selected from the image-based training data. The first and second images are selected such that both the first and second images share a common characteristic, i.e., the first characteristic. In embodiments, the presence of the first characteristic need not imply the presence of a particular structure or pattern or aspect in an image or a subject thereof but may also reflect, for example, the absence of a particular structure or pattern or aspect in an image of a subject thereof. In embodiments, the first characteristic may relate to a common label of the first and second images. For example, both the first and second images may share a common first characteristic that the subjects associated with each of the first and second images have been diagnosed with the same disease or condition.
[0145] Upon completing step 602 of flow diagram 600, the process next moves to step 603 of flow diagram 600.
[0146] At step 603, a first synthetic image is generated based on the first and second images by combining aspects of the first and second images. In embodiments, the first and second images may be combined to form the first synthetic image in any convenient manner. In some embodiments, the first and second images may be combined such that the first synthetic image comprises certain aspects of the first image and certain aspects of the second image. In other embodiments, the first and second images may be combined such that the first synthetic image comprises aspects based on a combination of aspects of the first and second images. Further, in some embodiments, the first and second images are further combined with one or more images of the image-based training data set (i.e., third image, fourth image, and so on) to form the first synthetic image. In embodiments, the first and second images are combined using an autoencorder. In such embodiments, the first and second images may be encoded using an autoencorder, as described in detail herein, such that the encoder is used to generate first and second representations of the first and second images, respectively. The first and second representations may be, for example, a one-dimensional array, for example. The first and second representations may be combined to form a representation of the first synthetic image. For example, the first and second representations may be combined using a scaled linear combination. In some cases, scale factors are used such that the first synthetic image is based on the first image to the extent of a first factor and is based on the second image to the extent of a second factor. For example, the first synthetic image may be 1% or more, such as 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or 99%, based on the first image and 99%, 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5% or 1 % based on the second image. For example, the first synthetic image may be 50% based on the first image and 50% based on the second image, or in another example, the first synthetic image may be 75% based on the first image and 25% based on the second image. When more than two images are combined to form the first synthetic image, the respective contribution of each of the figures may similarly be scaled as desired. The first and second representations may be combined to form a representation of the first synthetic image by selecting constituent elements from the first representation and the second representation or by adding such elements, averaging such elements or the like. Once the representation of the first synthetic image is generated, e.g., based on a scaled linear combination of the representation of the first image and the representation of the second image, the representation of the first synthetic image is decoded using a decoder to generate the first synthetic image.
[0147] In contrast to step 103 of flow diagram 100, step 103 generates a first synthetic image by combining information from two (or more) images of the image-based training data set.
[0148] Upon completing step 603 of flow diagram 600, the process next moves to step 604 of flow diagram 600.
[0149] At step 604, the image-based training data set and the first synthetic image are combined to generate an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first and second images. Some embodiments of the present invention further comprise generating a plurality of synthetic images based on any combination of two or more images, where such combined images share a first characteristic, and combining the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models.
[0150] Upon completing step 604 of flow diagram 600, the process ends. As described above, when images of the image-based training data are subject images, the subject or subjects or patient or patients may be human. The subject or subjects may be male or female. The subject or subjects may be adult subjects or pediatric subjects or both, ranging, for example, from ten years old to 65 years old. The subject or subjects may have any body mass index and such may vary.
[0151] Autoencoders:
[0152] FIG. 7 depicts aspects of, and operation of, an autoencoder according to embodiments of the present invention, as applied to an exemplary ophthalmologic image. Autoencoder 700 comprises encoder 701 and decoder 702. Encoder 701 is configured to process an image, such as input image 704, into an embedding 703, e.g., an array, such as a one-dimensional array or a two- dimensional array, representation of input image 704, where such embedding 703 is an output of encoder 701 . Decoder 702 is configured to process an embedding, such as embedding 703, into an image, such as reconstructed image 705, where such reconstructed image 705 is an output of decoder 701 . Autoencoder 700 is configured, e.g., trained such that it generates reconstructed image 705 that is identical, or substantially identical to, input image 704, when the embedding 703 is not altered or otherwise modified after encoder 701 generates embedding 703 and before decoder 702 receives embedding 703 as input. As shown in the figure, ideally input 704 and output 705 are substantially identical when embedding 703 is not modified. As described, in embodiments, in order to generate synthetic images, embeddings may be deliberately modified to intentionally introduce differences between an input image and an output image.
[0153] As described, an autoencoder is a powerful generative model that has the ability to learn the underlying distribution of data and generate new data samples based on that distribution. Further details regarding autoencoders are provided in: Kingma, D.P. and Welling, M., 2019, An introduction to variational autoencoders, Foundations and Trends® in Machine Learning, 12(4), pp.307- 392, incorporated herein in its entirety. Autoencoder 600 comprises two components: encoder 601 and decoder 602, both of which can be neural networks or variants thereof as desired and depending on specific applications.
[0154] Embodiments of autoencoders of the present invention, and how such are used in embodiments of the present invention, offer various improvements of existing techniques. As a first improvement over existing techniques, for example, embodiments of the autoencoders of the present invention comprise a multilevel, hierarchical structure, i.e., comprise multiple levels. In embodiments, a lower level of an embodiment of an encoder of an autoencoder is configured to encode small image “patches,” i.e., subsets of an image, into embeddings. Typically, images comprise a group, i.e., a plurality, of patches, such that the lower-level encoder encodes those patches into a group of embeddings. Typically, a higher-level encoder is configured to then further encode this group of embeddings into a final embedding which is a representation of the entire image. As a first improvement over existing techniques, for example, in embodiments, autoencoders are not used to generate new images, i.e., to generate images from scratch. Instead, in embodiments autoencoders are used to synthesize new images based on existing images, i.e., based on an image of an image-based training data set. As described, embodiments of autoencoders transform images into a quantitative format known as an embedding. In embodiments, embeddings comprise a set of numerical values. Such embeddings may be modified by adjusting or altering such numerical values. In embodiments, such numerical values are modified in any convenient manner, including, but not limited to: adding random noise to the embedding or adding patterned changes which follow certain math functions such as sine or cosine functions.
[0155] FIG. 8 depicts aspects of an autoencoder according to embodiments of the present invention, as applied to an exemplary ophthalmologic image. In autoencoding process 800 according to an embodiment, an autoencoder, such as, for example, autoencoder 700, is applied to input image 804. Input image 804 is an ophthalmologic image. The autoencorder encodes 801 input image 804 creating a lower dimensional representation (i.e., a representation or an embedding or an encoding) of input image 804. Such representation 803 of input image 804 may be generated at least in part by treating input image 804 as an RGB array 804A. The output of applying the autoencoder to encode input image 804 is encoding 803, referred to as a latent space representation. Subsequently, the autoencoder is applied to the encoding 803 to create output image 805. Such output image 805 may be generated at least in part by treating output image 805 as an RGB array 805A. In some cases, information is lost in the encoding step 801 , i.e., such that embedding 803 is an incomplete representation of input image 804. In such cases, the decoding step 802 attempts to recreate input image 804 without complete information about input image 804. As a result output image 805 may differ from input image 805, even when such differences may not be noticeable by a human eye. However, such differences, even when not noticeable by a human eye, may benefit the training of a deep learning model by presenting additional variation to the deep learning model, e.g., preventing overtraining of the model.
[0156] FIG. 9 depicts aspects of, and operation of, an autoencoder according to embodiments of the present invention, used to generate a synthetic image based on more than one input images. Autoencoder 900 comprises first encoder 901 A, second encoder 901 B and decoder 902. In some embodiments, a single encoder may be used instead of separate first encoder 901 A and second encoder 901 B. Encoder 901 A is configured to process an image, such as input image 904A, into an embedding 903A, e.g., an array, such as a one-dimensional array or a two-dimensional array, representation of input image 904A, where such embedding 903A is an output of encoder 901 A. Similarly, encoder 901 B is configured to process an image, such as input image 904B, into an embedding 903B, e.g., an array, such as a one-dimensional array or a two-dimensional array, representation of input image 904B, where such embedding 903B is an output of encoder 901 B. As described, while two encoders 901 A, 901 B are depicted, in embodiments, a single encoder may be employed, e.g., as a single encoder algorithm may be employed. Embeddings 903A, 903B are combined to form a combined embedding 903C, where, as a result of combining embeddings 003A, 003B, new embedding 003C reflects aspects of each precursor embedding 003A, 003B. Any convenient technique for combining a plurality of embeddings, such as combining embeddings 003A, 003B may be employed, such as, for example, any convenient simple scaled linear combination of a plurality of embeddings, such as a simple scaled linear combination of embeddings 003A, 903B, may be employed to generate new embeddings 003C. In some cases, new embedding 903C is equal to the sum of a first scale factor times the first embedding 903A and a second scale factor times the second embedding 903B. In some cases, the first and second scale factors sum to one. For example, the third embedding 903C may be set to equal to 0.1 times the first embedding 903A plus 0.9 times the second embedding 903B (i.e., combined_embedding 903C = 0.1 * embedding_one 903A + 0.9 * embedding_two 903B). In another example, the third embedding 903C may be set to equal to 0.2 times the first embedding 903A plus 0.8 times the second embedding 903B (i.e., combined_embedding 903C = 0.2 * embedding_one 903A + 0.8 * embedding_two 903B). In still another example, the third embedding 903C may be set to equal to 0.5 times the first embedding 903A plus 0.5 times the second embedding 903B (i.e., combined_embedding 9030 = 0.5 * embedding_one 903A + 0.5 * embedding_two 903B). In still other cases, such technique may be extended to multiple more than two constituent embeddings by scale factors (where such scale factors sum to 1 ) in order to generate a combined embedding (e.g., 903C) from more than two constituent embeddings (i.e., generate a synthetic image starting from more than two separate images).
[0157] Once new embedding 903C is generated, new embeddings 903C is processed by decoder 902. Decoder 902 is configured to process an embedding, such as embedding 903C, into an image, such as reconstructed image 905, where such reconstructed image 905 is an output of decoder 902. Since embedding 903C is a combination of embeddings 903A, 903B, corresponding to images 904A, 904B, resulting image 905 is therefore a combination of aspects or features of images 904A, 904B. Further, any of embeddings 903A, 903B, 903C may be further processed or manipulated by, e.g., introducing noise into the embedding, to further differentiate or distinguish resulting image 905. As described, in embodiments, in order to generate synthetic images, images 904A, 904 are combined by the encoding / decoding process described in order to intentionally introduce differences from images 904A, 904B to better facilitate training a model, such as a deep learning model. Images 904A, 904B are different retinal images. However, embodiments of the disclosure may be applied to any type of image, e.g., depicting any type of subject matter, and are not limited to retinal images. Preferably, embodiments of the present disclosure are used to combine a plurality of images where all such combined images are members of a same class; e.g., where all combined images depict subjects diagnosed with, or exhibiting, a common disease or condition, or, alternatively, where all combined images depict subjects not diagnosed with, or not exhibiting, a common disease or condition. That is, preferably embodiments of the present disclosure are utilized to combine images of the same class or to combine images exhibiting a common characteristic.
[0158] FIG. 10 panels (a) through (d) present illustrations of autoencoder building blocks according to embodiments of the present invention. FIG. 10, panel (a) depicts a neural net structure of a “resnet contraction block” 1001 , and FIG. 10, panel (b) depicts a structure of a “resnet expansion block” 1002, in each case according to embodiments of the present invention. FIG. 10, panels (c) and (d) depict “resnet block groups” according to embodiments, where each resnet block group comprises of two blocks. FIG. 10, panel (c) depicts a “resnet contraction block group” 1003, and FIG. 10, panel (d) depicts a “resnet expansion block group” 1004.
[0159] As described, embodiments of encoders and decoders of the present invention are designed hierarchically, in some cases, similar to aspects of a ResNet architecture. Further details regarding ResNet architectures are provided in: He, K., Zhang, X., Ren, S. and Sun, J., 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778), incorporated herein in its entirety. Embodiments of encoders and decoders comprise fundamental building blocks known as ResNet blocks. Embodiments of encoders utilize ResNet contraction blocks, while embodiments of decoders employ ResNet expansion blocks.
[0160] FIG. 10, panels (a) and (b) depict exemplary neural network architectures of an expansion block 1001 and a contraction block 1002, respectively. Such blocks are organized into two groups: the resnet contraction group 1003 and the resnet expansion group 1004, each comprising two blocks, as depicted in FIG. 10, panels (c) and (d).
[0161] As described, embodiments of autoencoders of the present invention comprise a dual-level structure. For example, according to an embodiment, each of a Level 1 encoder and decoder are capable of processing images of size, for example, 12 x 12, producing an embedding of size 48. In embodiments, the 12 x 12 refers to the size of an input to an autoencoder. For example, for a first level of an autoencoder, the image patch is 12 X 12 large in width and height. That is, the first level of an autoencoder accepts an input that is a 12 x 12 image patch. In embodiments, the 48 refers to a data size after the autoencoder encodes the input; i.e., 48 is the size of the output. For example, for a first level of an autoencoder, the first level of the autoencoder will encode input that is a 12 X 12 size image into an output that is 48 floating point numbers. FIG. 11 depicts an exemplary neural network architecture of such components, encoder 1 101 and decoder 1102, along with embedding 1 103.
[0162] Since most images (e.g., most images of image-based training data sets of interest in connection with the present invention) exceed a 12 x 12 resolution, embodiments are configured to segment them into so-called 12 x 12 patches. Each patch is then encoded into a separate embedding for further processing. For instance, a 96 x 96 image can be encoded by a Level 1 encoder of an embodiment of an autoencoder into a latent space of dimensions 48 x 8 x 8, where 48 represents a channel size (embedding size). In embodiments, channel size refers to the size of a third dimension of input to an autoencoder. For example, a first layer or first level or Level 1 of an autoencoder encodes an image patch where the channel size may be three, representing each of R (red), G (green) and B (blue) values (i.e., the input to the autoencoder comprises a channel size of three with three numerical values corresponding to R, G and B values, for example). In such embodiment, the second layer or second level or Level 2 of the autoencoder subsequently encodes a group of outputs from the first layer of the autoencoder. For example, given an image that is 120 x 120 in size, since the first layer of the autoencoder encodes a 12 x 12 patch, the first layer output will be 10 x 10 x 48 (in total representing output of the entire 120 x 120 block), where 48 is the output size of the first layer of the autoencoder. This 10 x 10 x 48 output from the first layer of the autoencoder then serves as the input to the second layer of the autoencoder. In this exemplary embodiment, the channel size is 48.
[0163] Additionally, embodiments of the present invention comprise a Level 2 encoder and decoder for processing such latent space. Exemplary architectures of the Level 2 encoder 1201 and decoder 1202 are depicted in FIG. 12, along with embedding 1203.
[0164] Embodiments comprising combined functionality of the Level 1 and Level 2 encoders, 1101 , 1201 , and decoders 1102, 1202, are configured for converting, for example, a 96 x 96 image into an embedding of size 1024. An embodiment, and an illustration of an entire process of such conversion, is illustrated in FIG. 13, in which autoencoder 1300 comprises encoder 1301 and decoder 1302, along with embedding 1303. Such an embodiment of a dual-level structure 1300, offers several advantages over a single-level system, including, for example: (i) improved training efficiency: by segmenting the encoder and decoder into two levels, each can be trained independently, reducing GPU memory requirements, for example; and (ii) improved dynamic composition ability: a specific Level 1 encoder can be paired with various Level 2 encoders, allowing for the encoding of images of different sizes into latent spaces of varying dimensions.
[0165] In some embodiments, training an autoencoder comprises employing a loss function defined based on the discrepancy in pixel values between the input image fed into the encoder and the resultant image output from the decoder. In embodiments autoencoders comprise a framework consistent with a general- purpose image autoencoder, capable of encoding and decoding images across a wide spectrum of types, rather than being restricted to images belonging to a specific category. Consequently, training datasets for use with the autoencoder can be selected to encompass as broad a range of images as desired and / or as feasible. In some cases the image resource, ImageNet, is highly suitable due to its extensive and diverse collection of images. Further details regarding ImageNet are provided in: Deng, J., et al. “ImageNet: A large-scale hierarchical image database,” 2009 IEEE conference on computer vision and pattern recognition, IEEE, 2009, incorporated herein in its entirety. Given the vast quantity of available training data in ImageNet, a conventional approach of segregating data into training and validation subsets need not always be employed in embodiments. Instead, uninterrupted training can be conducted. In some cases, embodiments of autoencoders are exposed to any number, such as approximately 200,000, training iterations, after which embodiments of autoencoders typically demonstrate proficiency in encoding and decoding images beyond those contained within the ImageNet dataset and achieving minimal loss. ImageNet is offered as an exemplary training data set but embodiments of autoencoders of the present invention are not limited to being trained on ImageNet images.
[0166] In embodiments, autoencoder encoders may be trained to map input data to a lower-dimensional latent representation, which may comprise a stochastic compressed representation of the input data. In embodiments, autoencoder decoders, on the other hand, may be trained to generate output images from the latent representation using transpose convolutional neural net. As described, embodiments of autoencoders may be trained using aspects of ImageNet, a general image data bank, without any additional medical images. The trained embodiment of an autoencoder may be used to generate any number of synthetic images based on images of an image-based training data set.
[0167] As described, in embodiments, a key function of an autoencoder is to produce high fidelity images, such as, for example, medical images, for training set enhancement. Other image-processing techniques may be employed. However, in embodiments, an autoencoder may be preferred. For example, in some cases, employing a generative adversarial network (GAN), as opposed to an autoencoder, may not be preferred because training such a network itself is data intensive, and techniques such as diffusion can lead to hallucinations and produce nonsensical images that are counterproductive for generating images for training of deep learning models, such as healthcare Als. Further details regarding generative adversarial networks (GANs) are provided in: Goodfellow, I., et al. “Generative adversarial nets.” Advances in neural information processing systems 27 (2014), incorporated herein in its entirety. Further details regarding techniques such as diffusion are provided in: Ho, J., Jain, A. and Abbeel, P., 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33, pp.6840-6851 , incorporated herein in its entirety. Embodiments of the present invention leverage autoencoders to generate synthetic images whereby controlled variations can be introduced to real-world images, such as real-world medical images. See Kingma, D.P. and Welling, M., 2019. An introduction to variational autoencoders. Foundations and Trends® in Machine Learning, 12(4), pp.307-392. These variations introduced into images go beyond simple rotations, for example. Instead, embodiments of autoencoders first abstract existing images into embeddings (i.e., during an encoding process), before adding transformations through a neural net. The altered embeddings can then be used to generate high fidelity synthetic images of the original image (i.e., during a decoding process). As described, the encoder and decoder of embodiments of an autoencoder may both be implemented as multi-layer ResNets (as shown in FIG. 10) and can be used to encode / decode any image. See He, K., Zhang, X., Ren, S. and Sun, J., 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778).
[0168] FIG. 14, panel (a) depicts an overview of a method of using autoencoder 1400 for generating a synthetic image 1405 based on a subject image 1404, according to an embodiment of the present invention. Autoencoder 1400 is applied to optic disc photos, such as subject image 1404 of an optic disc of a subject, from, for example, a publicly available database. Image 1404 is encoded into embedding 1403a by encoder 1402. A variation or modification, such as noise 1410, is introduced into embedding 1410 by adding a vector of numerical values 1410 to embedding 1403a to produce updated embedding 1403b. Updated embedding 1403b is then input into decoder 1402 to generate synthetic image 1405.
[0169] FIG. 14, panel (b) depicts different representations of various types of noise 1410 that may be added to an embedding of an input image in embodiments of the present invention. FIG. 14, panel (b) shows uniform nose 1410a, in which a fixed value may be added to each entry in an embedding 1403a; step noise 1410b, in which a fixed value may be added to specified entries in an embedding 1403a; and random noise 1410c, in which random values may be added to each entry in an embedding 1403a.
[0170] As described, embodiments of autoencoders, such as autoencoder 1400, first transform an image into a quantitative format known as an embedding, such as embedding 1403a. In embodiments, an embedding comprises a set of numerical values, e.g., a one-dimensional array of floating point numbers. Embodiments of methods of the present invention comprise modifying such embeddings by altering these numerical values. While any convenient modification may be employed, and modifications to the embeddings may be configured such that they lack a specific direction, nonetheless, in embodiments, these modifications may be restricted to ensure the resultant image, such as synthetic image 1405, remains relevant, i.e., relevant for training a deep learning model. For example, modifications may be restricted such that the resulting synthetic image remains substantially recognizable to an expert clinician, in the case of images used in medical contexts. Modifications that introduce excessive alterations may render an image, such as an image for training a deep learning model in the context of medicine, into mere noise, which is not useful in the context of training a deep learning model. Any convenient technique for restricting modifications may be employed, including, for example, restricting the number of numerical values of the embedding that are modified or restricting the relative or absolute amount that the numerical values of the embedding that are modified.
[0171] FIG. 15, panels (a) and (b) depict exemplary autoencoder-generated synthetic images 1505a, 1505b, 1505c and corresponding subject images 1504a, 1504b, 1504c (original images) used to generate such synthetic images.
[0172] FIG. 15, panel (a) depicts three original images 1504a, 1504b, 1504c, i.e., subject images, and, for each original image, three synthetic images 1505a, 1505b, 1505c. Each synthetic image was generated using an autoencoder to introduce one or more variations, as described, into the structures depicted in the original images, i.e., by introducing modifications to embeddings that are generated and processed using an autoencoder. FIG. 15, panel (b) depicts zoomed in views of sections of original images, 1504e, 1504g, and corresponding sections of autoencoder-generated synthetic images, 1505e, 1505g, showing how the autoencoder-generated synthetic images 1505d, 1505f, vary from the original images 1504d, 1504f , i.e., showing the variations introduced into the synthetic images by the autoencoder.
[0173] In the embodiment used to generate the images shown in FIG. 15, panels (a) and (b), an autoencoder added variations to the images that augment the diversity and richness of the training data (i.e., original images 1504a, 1504b, 1504c, 1504d, 1504e, plus synthetic images 1505a, 1505b, 1505c, 1505d, 1505e). The inclusion of these synthetic images in a training set enables deep learning models, such as ViTs, to learn from (i.e., be trained by) a more extensive range of data, leading to a better understanding of the underlying patterns and features relevant to the operation of deep learning models (which, in the case of the images depicted in FIG. 15, panels (a) and (b), relate to glaucoma detection) and ultimately enhancing model performance.
[0174] Deep Learning Models:
[0175] Embodiments of the present invention utilize deep learning models, or comprise image-based training data sets, or generate enhanced image-based training data sets for, deep learning models that are vision transformers (ViTs). Embodiments of ViT architectures can be trained to perform well in the context of medical imaging, such as in the context of glaucoma detection, for example. Details regarding exemplary ViT architectures are described in Hwang, E., et al. “Multi-Dataset Comparison of Vision Transformers and Convolutional Neural Networks for Detecting Glaucomatous Optic Neuropathy from Fundus Photographs.” Bioengineering, 2023, which disclosure is incorporated herein by reference.
[0176] FIG. 16, panel (a) depicts a schematic diagram of an embodiment, in which autoencoder 1600 comprising encoder 1601 and decoder 1602 is employed to generate a synthetic training image 1605 from an original image 1604 by introducing variations into an embedding 1603 representation of original image 1604. Synthetic image 1605 is, in turn, combined with original image 1604, thereby forming enhanced training set 1610. Enhanced training set 1610 is employed to train deep learning model 1620. In the embodiment, deep learning model 1620 is trained to detect a result using enhanced training data set 1610. In the embodiment, the result detected relates to the detection of glaucoma in a subject image, i.e., in an experimental image (not shown) taken of a subject, such as a patient. Such experimental images may differ from images of the enhanced training data set 1610. FIG. 16, panel (b) depicts aspects of a deep learning model, in this case a ViT architecture, according to an embodiment of the present invention.
[0177] Computer Implemented Embodiments
[0178] The various method and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system applying a method according to the present disclosure. The described functionality can be implemented in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the disclosure.
[0179] The various illustrative steps, components, and computing systems (such as devices, databases, interfaces, and engines) described in connection with the embodiments disclosed herein can be implemented or performed by a machine, such as a general purpose processor, a graphics processor unit, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor can be a microprocessor, but in the alternative, the processor can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Although described herein primarily with respect to digital technology, a processor can also include primarily analog components. A computing environment can include any type of computer system, including, but not limited to, a computer system based on a microprocessor, a graphics processor unit, a mainframe computer, a digital signal processor, a portable computing device, a personal organizer, a device controller, and a computational engine within an appliance, to name a few.
[0180] The steps of a method, process, or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module, engine, and associated databases can reside in memory resources such as in RAM memory, PRAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of non-transitory computer-readable storage medium, media, or physical computer storage known in the art. An external storage medium can be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an ASIC. The ASIC can reside in a user terminal. In the alternative, the processor and the storage medium can reside as discrete components in a user terminal.
[0181] As described in detail herein, embodiments of the present invention relate to computer-implemented methods for generating an enhanced image-based training data set for deep learning models. Aspects of the present invention include methods of generating an enhanced image-based training data set for deep learning models, the method comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; selecting a first image from the imagebased training data based at least in part on a first characteristic of the first image; generating a first synthetic image based on the first image by modifying a first aspect of the first image; and combining the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image.
[0182] Aspects of the present invention further include methods of training a deep learning model to detect a result using an enhanced training data set, the method comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; applying an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on an image of the first subset of images of the image-based training data set; training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; obtaining an experimental image; and using the trained deep learning model to predict whether a result is present in the experimental image. SYSTEMS FOR REDUCING DEEP LEARNING DATA REQUIREMENTS
[0183] AS summarized herein, aspects of the present disclosure include systems for generating an enhanced image-based training data set for deep learning models. Systems according to certain embodiments comprise a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to execute steps corresponding to the subject methods described herein.
[0184] In an embodiment, a system for generating an enhanced image-based training data set for deep learning models comprises: a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; select a first image from the image-based training data based at least in part on a first characteristic of the first image; generate a first synthetic image based on the first image by modifying a first aspect of the first image; and combine the imagebased training data set and the first synthetic image into an enhanced imagebased training data set for deep learning models, wherein the enhanced imagebased training data set comprises additional training data associated with the first characteristic of the first image.
[0185] Some embodiments may further comprise a display device, e.g., for displaying results of, or related to, an enhanced image-based training data set for deep learning models or the like or for displaying results of, or related to, training a deep learning model to detect a result using an enhanced training data set. Any convenient display device, such as a liquid crystal display (LCD), lightemitting diode (LED) display, plasma (PDP) display, quantum dot (QLED) display or cathode ray tube display device. The processor and / or memory may be operably connected to the display device, for example, via a wired, such as a Universal Serial Bus (USB) connection, or wireless connection, such as a Bluetooth connection.
[0186] In some instances the systems further include one or more computers for complete automation or partial automation of the methods described herein. In some embodiments, systems include a computer having a computer readable storage medium with a computer program stored thereon.
[0187] In embodiments, the system includes an input module, a processing module and an output module. The subject systems may include both hardware and software components, where the hardware components may take the form of one or more platforms, e.g., in the form of servers, such that the functional elements, i.e., those elements of the system that carry out specific tasks (such as managing input and output of information, processing information, etc.) of the system may be carried out by the execution of software applications on and across the one or more computer platforms represented of the system.
[0188] Systems may include a display and operator input device. Operator input devices may, for example, be a keyboard, mouse, or the like. The processing module includes a processor which has access to a memory having instructions stored thereon for performing the steps of the subject methods. The processing module may include an operating system, a graphical user interface (GUI) controller, a system memory, memory storage devices, and input-output controllers, cache memory, a data backup unit, and many other devices. The processor may be a commercially available processor or it may be one of other processors that are or will become available. The processor executes the operating system and the operating system interfaces with firmware and hardware in a well-known manner, and facilitates the processor in coordinating and executing the functions of various computer programs that may be written in a variety of programming languages, such as Java, Perl, C++, other high level or low level languages, as well as combinations thereof, as is known in the art. The operating system, typically in cooperation with the processor, coordinates and executes functions of the other components of the computer. The operating system also provides scheduling, input-output control, file and data management, memory management, and communication control and related services, all in accordance with known techniques. The processor may be any suitable analog or digital system. In some embodiments, processors include analog electronics. In some embodiments, the processor includes analog electronics which provide feedback control, such as for example negative feedback control.
[0189] The system memory may be any of a variety of known or future memory storage devices. Examples include any commonly available random access memory (RAM), magnetic medium such as a resident hard disk or tape, an optical medium such as a read and write compact disc, flash memory devices, or other memory storage device. The memory storage device may be any of a variety of known or future devices, including a compact disk drive, a tape drive, a removable hard disk drive, or a diskette drive. Such types of memory storage devices typically read from, and / or write to, a program storage medium (not shown) such as, respectively, a compact disk, magnetic tape, removable hard disk, or floppy diskette. Any of these program storage media, or others now in use or that may later be developed, may be considered a computer program product. As will be appreciated, these program storage media typically store a computer software program and / or data. Computer software programs, also called computer control logic, typically are stored in system memory and / or the program storage device used in conjunction with the memory storage device.
[0190] In some embodiments, a computer program product is described comprising a computer usable medium having control logic (computer software program, including program code) stored therein. The control logic, when executed by the processor the computer, causes the processor to perform functions described herein. In other embodiments, some functions are implemented primarily in hardware using, for example, a hardware state machine. Implementation of the hardware state machine so as to perform the functions described herein will be apparent to those skilled in the relevant arts.
[0191] Memory may be any suitable device in which the processor can store and retrieve data, such as magnetic, optical, or solid-state storage devices (including magnetic or optical disks or tape or RAM, or any other suitable device, either fixed or portable). The processor may include a general-purpose digital microprocessor suitably programmed from a computer readable medium carrying necessary program code. Programming can be provided remotely to processor through a communication channel, or previously saved in a computer program product such as memory or some other portable or fixed computer readable storage medium using any of those devices in connection with memory. For example, a magnetic or optical disk may carry the programming, and can be read by a disk writer / reader. Systems of the invention also include programming, e.g., in the form of computer program products, algorithms for use in practicing the methods as described above. Programming according to the present invention can be recorded on computer readable media, e.g., any medium that can be read and accessed directly by a computer. Such media include, but are not limited to: magnetic storage media, such as floppy discs, hard disc storage medium, and magnetic tape; optical storage media such as CD-ROM; electrical storage media such as RAM and ROM; portable flash drive; and hybrids of these categories such as magnetic / optical storage media.
[0192] The processor may also have access to a communication channel to communicate with a user at a remote location. By remote location is meant the user is not directly in contact with the system and relays input information to an input manager from an external device, such as a computer connected to a Wide Area Network (“WAN”), telephone network, satellite network, or any other suitable communication channel, including a mobile telephone (i.e., smartphone).
[0193] In some embodiments, systems according to the present disclosure may be configured to include a communication interface. In some embodiments, the communication interface includes a receiver and / or transmitter for communicating with a network and / or another device. The communication interface can be configured for wired or wireless communication, including, but not limited to, radio frequency (RF) communication (e.g., Radio-Frequency Identification (RFID), Zigbee communication protocols, WiFi, infrared, wireless Universal Serial Bus (USB), Ultra Wide Band (UWB), Bluetooth® communication protocols, and cellular communication, such as code division multiple access (CDMA) or Global System for Mobile communications (GSM).
[0194] In one embodiment, the communication interface is configured to include one or more communication ports, e.g., physical ports or interfaces such as a USB port, an RS-232 port, or any other suitable electrical connection port to allow data communication between the subject systems and other external devices such as a computer terminal (for example, at a physician’s office or in hospital environment) that is configured for similar complementary data communication.
[0195] In one embodiment, the communication interface is configured for infrared communication, Bluetooth® communication, or any other suitable wireless communication protocol to enable the subject systems to communicate with other devices such as computer terminals and / or networks, communication enabled mobile telephones, personal digital assistants, or any other communication devices which the user may use in conjunction with other devices.
[0196] In one embodiment, the communication interface is configured to provide a connection for data transfer utilizing Internet Protocol (IP) through a cell phone network, Short Message Service (SMS), wireless connection to a personal computer (PC) on a Local Area Network (LAN) which is connected to the internet, or WiFi connection to the internet at a WiFi hotspot.
[0197] In one embodiment, the subject systems are configured to wirelessly communicate with a server device via the communication interface, e.g., using a common standard such as 802.1 1 or Bluetooth® RF protocol, or an IrDA infrared protocol. The server device may be another portable device, such as a smart phone, Personal Digital Assistant (PDA) or notebook computer; or a larger device such as a desktop computer, appliance, etc. In some embodiments, the server device has a display, such as a liquid crystal display (LCD), as well as an input device, such as buttons, a keyboard, mouse or touch-screen.
[0198] In some embodiments, the communication interface is configured to automatically or semi-automatically communicate data stored in the subject systems, e.g., in an optional data storage unit, with a network or server device using one or more of the communication protocols and / or mechanisms described herein.
[0199] Output controllers may include controllers for any of a variety of known display devices for presenting information to a user, whether a human or a machine, whether local or remote. If one of the display devices provides visual information, this information typically may be logically and / or physically organized as an array of picture elements. A graphical user interface (GUI) controller may include any of a variety of known or future software programs for providing graphical input and output interfaces between the system and a user, and for processing user inputs. The functional elements of the computer may communicate with each other via system bus. Some of these communications may be accomplished in alternative embodiments using network or other types of remote communications. The output manager may also provide information generated by the processing module to a user at a remote location, e.g., over the Internet, phone or satellite network, in accordance with known techniques. The presentation of data by the output manager may be implemented in accordance with a variety of known techniques. As some examples, data may include SQL, HTML or XML documents, email or other files, or data in other forms. The data may include Internet URL addresses so that a user may retrieve additional SQL, HTML, XML, or other documents or data from remote sources. The one or more platforms present in the subject systems may be any type of known computer platform or a type to be developed in the future, and they may be of a class of computer commonly referred to as servers. However, they may also be a mainframe computer, a workstation, or other computer type, such as a laptop computer. They may be connected via any known or future type of cabling or other communication system including wireless systems, either networked or otherwise. They may be co-located or they may be physically separated. Various operating systems may be employed on any of the computer platforms, possibly depending on the type and / or make of computer platform chosen. Appropriate operating systems include Windows, iOS, Oracle Solaris, Linux, IBM i, Unix, and others. FIG. 17 shows a functional block diagram for one example of a computer system 1700 for practicing methods of the present invention, i.e., a processor operably connected to memory, 1702, for generating an enhanced image-based training data set for deep learning models. A processor and memory 1702 can be configured to implement a variety of processes for generating synthetic images and / or implementing and training a model, for example, in connection with predicting a result.
[0200] An apparatus, 1712 can be configured to acquire training data, such as an image-based training data set. For example, apparatus 1712 may be a remote database, and processor 1702 may be operably connected to such remote database to acquire an image-based training data set for use generating synthetic images for generating an enhanced image-based training data set for deep learning models. A data communication channel can be included between the apparatus 1712 and the processor 1702. An image-based training data set can be provided to the processor 1702 via the data communication channel.
[0201] The processor 1702 can be configured to provide a graphical display, such as one or more images of the image-based training data set or one or more synthetic images or results of training a model using an enhanced training data set to a display device 1706. For example, processor and memory 1702 can be configured to cause a display device 1706 to display one or more diagrams or graphs or probability distributions or images or image excerpts related to results of generating an enhanced imaged-based training data set or training and evaluating a model.
[0202] The display device 1706 can be implemented as a monitor, a tablet computer, a smartphone, or other electronic device configured to present graphical interfaces.
[0203] The processor and memory 1702 can be configured to receive adjustments to configuration settings from a first input device. Such adjustments to configuration settings, when received from the first input device to the processor 1702 can be used to select or adjust aspects of one or more images or one or more image-based training data sets or one or more aspects of training a model. In embodiments, training data may be adjusted in order to avoid overfitting a data set. The first input device can be implemented as a mouse 1710, or the first device can be implemented as the keyboard 1708 or other means for providing an input signal to the processor 1702 such as a touchscreen, a stylus, an optical detector, or a voice recognition system. Some input devices can include multiple inputting functions. In such implementations, the inputting functions can each be considered an input device. For example, mouse 1710 can include a right mouse button and a left mouse button, each of which can generate a triggering event. Such triggering event can cause the processor 1702 to alter the manner in which the data is displayed or in which data is utilized, which portions of the data is actually displayed on the display device 1706, and / or provide input to further processing such as selection of additional training data or evaluation data.
[0204] The processor 1702 can be connected to a storage device 1704. The storage device 1704 can be configured to receive and store image-based training data or synthetic images from the processor 1702. The storage device 1704 can be further configured to allow retrieval of training and / or evaluation data, such as subject images or synthetic images, by the processor 1702.
[0205] The display device 1706 can be further configured to alter the information presented according to input received from the processor 1702 in conjunction with input from the apparatus 1712, the storage device 1704, the keyboard 1708, and / or the mouse 1710.
[0206] In some implementations the processor and memory 1702 can generate a user interface for use with generating an enhanced image-based training data set for a deep learning model or training and evaluating a model for predicting a result. For example, the user interface can include a control for applying certain training data or certain evaluation data to a model.
[0207] FIG. 18 depicts a general architecture of an example computing device 1800 according to certain embodiments. The general architecture of the computing device 1800 includes an arrangement of computer hardware and software components. The computing device 1800 may include many more (or fewer) elements than those shown in FIG. 18. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the computing device 1800 includes a processing unit 1810, a network interface 1820, a computer readable medium 1830, an input / output device interface 1840, a display 1850, and an input device 1860, all of which may communicate with one another by way of a communication bus. The network interface 1820 may provide connectivity to one or more networks or computing systems. The processing unit 1810 may thus receive information and instructions from other computing systems or services via a network. The processing unit 1810 may also communicate to and from memory 1870 and / or a data store 1990 and further provide output information for an optional display 1850 via the input / output device interface 1840. The input / output device interface 1840 may also accept input from the optional input device 1860, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.
[0208] The memory 1870 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 1810 executes in order to implement one or more embodiments. The memory 1870 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer-readable media. The memory 1870 may store an operating system 1872 that provides computer program instructions for use by the processing unit 1810 in the general administration and operation of the computing device 1800. The memory 1870 may further include computer program instructions and other information for implementing aspects of the present disclosure.
[0209] For example, in one embodiment, the memory 1870 includes an obtain an image-based training data set for a deep learning model module 1874, wherein the image-based training data set comprises a plurality of images; a selecting a first image from the image-based training data based at least in part on a first characteristic of the first image module 1876; a generate a first synthetic image based on the first image by modifying a first aspect of the first image module 1878; and a combining the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models module 1880.
[0210] Aspects of the present disclosure further include non-transitory computer readable storage media for generating an enhanced image-based training data set for deep learning models. Non-transitory computer readable storage media according to certain embodiments comprise one or more algorithms corresponding to the subject methods described herein.
[0211] UTILITY
[0212] The subject methods and systems find use in a variety of applications where it is desirable to train a deep learning model in a more timely and accurate manner and / or in a manner that mitigates a bias in the training data and resulting deep learning model trained using such training data than presently available. The subject methods and systems find use in applications, in which it is desirable to obtain additional training data for a deep learning model. In some embodiments, the methods and systems described herein find use in clinical settings such as any clinical setting where traditional, clinical or biochemical testing according to prevailing standards of care may be applied. Embodiments of the methods and systems described herein find use in clinical settings in which images, such as photographic images, MRI images, X-ray images, CT images or ultrasound images may be used in connection with diagnosing or screening for a disease. In other embodiments, the methods and systems described herein find use in remote medicine settings, where specialized medical services capable of providing accurate diagnoses for one or more diseases may be newly enabled by application of the present methods and systems. In addition, the subject methods and systems find use in improving the effectiveness and timeliness of diagnosing disease. Embodiments of methods and systems of the invention find use in addressing concerns around equity of Al models, including, for example, insofar as underrepresented social groups contribute significantly less healthcare data, thus potentially propagating their current lack of access to healthcare into future bias in healthcare Als. See Raju, M., Shanmugam, K.P. and Shyu, C.R., 2023. Application of Machine Learning Predictive Models for Early Detection of Glaucoma Using Real World Data. Applied Sciences, 13(4), p.2445.
[0213] EXAMPLES OF NON-LIMITING ASPECTS OF THE DISCLOSURE
[0214] Aspects, including embodiments, of the present subject matter described herein may be beneficial alone or in combination, with one or more other aspects or embodiments. Without limiting the foregoing description, certain non-limiting aspects of the disclosure numbered 1 -270 are provided below. As will be apparent to those of skill in the art upon reading this disclosure, each of the individually numbered aspects may be used or combined with any of the preceding or following individually numbered aspects. This is intended to provide support for all such combinations of aspects and is not limited to combinations of aspects explicitly provided below:
[0215] Aspect 1 . A computer-implemented method of generating an enhanced image-based training data set for deep learning models, the method comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; selecting a first image from the image-based training data based at least in part on a first characteristic of the first image; generating a first synthetic image based on the first image by modifying a first aspect of the first image; and combining the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image.
[0216] Aspect 2. The method of aspect 1 , wherein generating a first synthetic image based on the first image comprises modifying a representation of a structure present in the first image. Aspect 3. The method of any of the previous aspects, wherein generating a first synthetic image based on the first image comprises introducing a variation in a structure present in the first image.
[0217] Aspect 4. The method of any of the previous aspects, wherein generating a first synthetic image based on the first image comprises introducing random variations into the first image.
[0218] Aspect 5. The method of any of the previous aspects, further comprising: generating a plurality of synthetic images based on the first image by one or more modifications to one or more aspects of the first image; and combining the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models.
[0219] Aspect 6. The method of any of the previous aspects, further comprising: generating a plurality of synthetic images, wherein the synthetic images are generated based on more than one of the plurality of images; and combining the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models.
[0220] Aspect 7. The method of any of the previous aspects, wherein generating a first synthetic image based on the first image comprises introducing a variation in the first image other than mirroring, rotating, smoothing, or contrast reduction of the first image.
[0221] Aspect 8. The method of any of the previous aspects, wherein generating a first synthetic image based on the first image comprises: applying an autoencoder or a diffusion-based model or a generative adversarial network (GAN) to the first image to generate the first synthetic image. Aspect 9. The method of any of the previous aspects, wherein generating a first synthetic image based on the first image comprises applying an autoencoder to the first image to generate the first synthetic image, wherein the autoencoder is a trained autoencoder. Aspect 10. The method of any of aspects 8 to 9, wherein the autoencoder is trained on a second image-based training data set, wherein the second imagebased training data set does not comprise any images of the image-based training data set.
[0222] Aspect 11 . The method of any of aspects 8 to 10, wherein applying an autoencoder to the first image to generate the first synthetic image comprises: introducing a controlled variation of the first image into the first synthetic image.
[0223] Aspect 12. The method of any of aspects 8 to 11 , wherein applying an autoencoder to the first image to generate the first synthetic image comprises: encoding the first image into an embedding; and decoding the embedding into the first synthetic image, wherein the embedding comprises a representation of the first image.
[0224] Aspect 13. The method of any of aspect 12, wherein the embedding comprises a compressed representation of the first image.
[0225] Aspect 14. The method of any of aspects 12 to 13, wherein the embedding comprises a compressed stochastic representation of the first image.
[0226] Aspect 15. The method of any of any of aspects 12 to 14, further comprising: introducing noise into the embedding.
[0227] Aspect 16. The method of aspect 15, wherein the noise introduced into the embedding comprises one or more of uniform noise or step noise or random noise or patterned noise, wherein patterned noise is associated with a mathematical function, optionally comprising one or more of a sine or cosine function.
[0228] Aspect 17. The method of any of aspects 12 to 16, wherein the embedding comprises a linear representation of the first image.
[0229] Aspect 18. The method of any of any of aspects 15 to 17, wherein the noise introduced into the embedding comprises a linear representation of noise values. Aspect 19. The method of any of aspects 8 to 18, wherein the autoencoder comprises an encoder and a decoder. Aspect 20. The method of any of aspects 8 to 19, wherein the autoencoder comprises a dual-level structure.
[0230] Aspect 21 . The method of any of aspects 8 to 20, wherein the autoencoder comprises: an encoder, and a decoder, wherein the encoder comprises a first level encoder configured to encode a subset of the first image, and a second level encoder configured to encode an output of the first level encoder, and the decoder comprises a first level decoder, configured to decode a subset of the first synthetic image, and a second level decoder configured to decode an output of the first level decoder.
[0231] Aspect 22. The method of any of aspects 19 to 21 , wherein the encoder and decoder are modular components.
[0232] Aspect 23. The method of any of aspects 19 to 22, wherein the encoder and decoder are independently trained.
[0233] Aspect 24. The method of any of aspects 19 to 23, wherein the encoder and decoder comprise one or more of neural networks or residual networks.
[0234] Aspect 25. The method of any of aspects 19 to 24, wherein each level of the encoder and decoder are independently trained.
[0235] Aspect 26. The method of any of aspects 19 to 25, wherein each level of the encoder and decoder comprises a neural network.
[0236] Aspect 27. A computer-implemented method of generating an enhanced image-based training data set for deep learning models, the method comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; selecting first and second images from the image-based training data, wherein the first and second images each exhibit a common first characteristic; generating a first synthetic image based on the first and second images by combining the first image with the second image; and combining the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the common characteristic of the first and second images.
[0237] Aspect 28. The method of aspect 27, wherein generating a first synthetic image based on the first and second images comprises combining a representation of the first and second images.
[0238] Aspect 29. The method of aspect 28, wherein combining a representation of the first and second images comprises a scaled linear combination of the representations of the first and second images.
[0239] Aspect 30. The method of any of aspects 27 to 29, wherein generating a first synthetic image based on the first and second images further comprises generating the first synthetic image based on comprises combining the first image, the second image and one or more images from the from the imagebased training data.
[0240] Aspect 31 . The method of any of aspects 27 to 30, wherein generating a first synthetic image based on the first and second images by combining the first image with the second image comprises: applying an autoencoder or a diffusion-based model or a generative adversarial network (GAN) to the first and second images to generate the first synthetic image.
[0241] Aspect 32. The method of any of aspects 27 to 31 , wherein generating a first synthetic image based on the first and second images by combining the first image with the second image comprises applying an autoencoder to the first and second images to generate the first synthetic image, wherein the autoencoder is a trained autoencoder.
[0242] Aspect 33. The method of aspect 32, wherein the autoencoder is trained on a second image-based training data set, wherein the second image-based training data set does not comprise any images of the image-based training data set. Aspect 34. The method of any of aspects 32 to 33, wherein applying an autoencoder to the first and second images to generate the first synthetic image comprises: encoding the first image into a first embedding; encoding the second image into a second embedding; combining the first embedding with the second embedding to form a first synthetic embedding; and decoding the first synthetic embedding into the first synthetic image, wherein the first and second embeddings comprise representations of the first and second images, respectively.
[0243] Aspect 35. The method of aspect 34, wherein the first and second embeddings comprise compressed representations of the first and second images, respectively.
[0244] Aspect 36. The method of any of aspects 34 to 35, wherein the first and second embeddings comprise compressed stochastic representations of the first and second images, respectively.
[0245] Aspect 37. The method of any of aspects 34 to 36, wherein the first embedding comprises a first linear representation, and the second embedding comprises a second linear representation.
[0246] Aspect 38. The method of any of any of aspects 34 to 37, wherein combining the first image with the second image comprises combining the first linear representation with the second linear representation in a scaled linear combination.
[0247] Aspect 39. The method of any of aspects 32 to 38, wherein the autoencoder comprises an encoder and a decoder.
[0248] Aspect 40. The method of any of aspects 32 to 39, wherein the autoencoder comprises a dual-level structure.
[0249] Aspect 41 . The method of any of aspects 32 to 39, wherein the autoencoder comprises: an encoder, and a decoder, wherein the encoder comprises a first level encoder configured to encode a subset of the first and second images, and a second level encoder configured to encode outputs of the first level encoder, and the decoder comprises a first level decoder, configured to decode a subset of the first synthetic image, and a second level decoder configured to decode an output of the first level decoder.
[0250] Aspect 42. The method of aspect 41 , wherein the encoder and decoder are modular components.
[0251] Aspect 43. The method of any of aspects 41 to 42, wherein the encoder and decoder are independently trained.
[0252] Aspect 44. The method of any of aspects 41 to 43, wherein the encoder and decoder comprise one or more of neural networks or residual networks.
[0253] Aspect 45. The method of any of aspects 41 to 44, wherein each level of the encoder and decoder are independently trained.
[0254] Aspect 46. The method of any of aspects 41 to 45, wherein each level of the encoder and decoder comprises a neural network.
[0255] Aspect 47. The method of any of the previous aspects, wherein the plurality of images of the image-based training data set comprises subject images.
[0256] Aspect 48. The method of any of the previous aspects, wherein the plurality of images of the image-based training data set comprises images of biologic tissues.
[0257] Aspect 49. The method of any of the previous aspects, wherein the plurality of images of the image-based training data set corresponds to images of a plurality of subjects.
[0258] Aspect 50. The method of any of the previous aspects, wherein the imagebased training data set comprises images generated by one or more of the following modalities: photographic images, MRI images, X-ray images, CT images or ultrasound images.
[0259] Aspect 51 . The method of any of the previous aspects, wherein at least some of the plurality of images of the image-based training data set exhibit one or more of the following: glaucoma, cancer, a degenerative retinal diseases, an acquired retinal disease, diabetic retinopathy, aneurysm, multiple sclerosis, a spinal cord pathology, stroke, tumor, traumatic brain injury, joint injury or disease, heart disease, vascular disease, a liver disease, Alzheimer’s disease, epilepsy, peripheral nerve compression, renal disease, inflammatory bowel disease, a fracture, arthritis, a lung disease, tuberculosis, pneumonia, chronic obstructive pulmonary disease (COPD), emphysema, pulmonary edema, pneumothorax, a dental disease, a bone infection, scoliosis, osteoporosis, kidney disease, kidney stone, a presence of a foreign object, pulmonary embolism, trauma, stroke, an abdominal disorder, appendicitis, diverticulitis, an ovarian cyst, sinusitis, a thyroid gland disease, pregnancy, gallstone, ascites, pelvic organ prolapse or testicular torsion.
[0260] Aspect 52. The method of any of the previous aspects, wherein the imagebased training data set is a training data set for training a deep learning model. Aspect 53. The method of any of the previous aspects, wherein the imagebased training data set is configured for training a deep learning model to perform one or more of: image classification, object detection or semantic segmentation.
[0261] Aspect 54. The method of any of the previous aspects, wherein the imagebased training data set is configured for training a deep learning model to diagnose or screen for a disease or monitor a disease state or disease progression.
[0262] Aspect 55. The method of any of the previous aspects, wherein the imagebased training data set is a training data set for a convolutional neural network (CNN) or a vision transformer (ViT).
[0263] Aspect 56. The method of any of the previous aspects, wherein the imagebased training data set is a training data set for image-based disease diagnosis or disease monitoring.
[0264] Aspect 57. The method of any of the previous aspects, wherein the imagebased training data set comprises a first subset of images that exhibit a condition or a disease and a second subset of images that do not exhibit the condition of the disease. Aspect 58. The method of any of the previous aspects, wherein the imagebased training data set comprises a bias optionally comprising one or more of a bias related to an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location.
[0265] Aspect 59. The method of any of the previous aspects, wherein the first characteristic optionally comprises one or more of: a diagnosis or a disease state or an injury state or a condition of an image subject. Aspect 60. The method of any of the previous aspects, wherein the first characteristic is a demographic characteristic of the subject optionally comprising one or more of: an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location of an image subject.
[0266] Aspect 61 . The method of any of the previous aspects, wherein the first characteristic is relatively rare in the image-based training data set.
[0267] Aspect 62. The method of any of the previous aspects, wherein the first characteristic is underrepresented in the image-based training data set.
[0268] Aspect 63. The method of any of the previous aspects, wherein it is relatively rare for a subject in the image-based training data set to exhibit the first characteristic.
[0269] Aspect 64. The method of any of the previous aspects, wherein the first characteristic depicts an image aspect of a rare disease.
[0270] Aspect 65. The method of any of the previous aspects, wherein the first or second images correspond to an underrepresented group.
[0271] Aspect 66. The method of any of the previous aspects, wherein an image subject is a member of an underrepresented group.
[0272] Aspect 67. The method of any of the previous aspects, wherein the first characteristic is selected to address a bias in the image-based training data set. Aspect 68. The method of any of the previous aspects, wherein the first characteristic is selected to enhance performance of the deep learning model. Aspect 69. The method of any of the previous aspects, wherein the first characteristic is selected to enhance performance of the deep learning model across a more diverse range of subjects.
[0273] Aspect 70. The method of any of the previous aspects, wherein the image-based training data does not reflect a demographic characteristic of a select population, and the enhanced image-based training data set is generated to reflect the demographic characteristic of the select population.
[0274] Aspect 71 . The method of any of the previous aspects, wherein the enhanced image-based training data set comprises at least one synthetic image corresponding to each image of a plurality of images of the image-based training set.
[0275] Aspect 72. The method of any of the previous aspects, wherein the enhanced image-based training data set comprises at least one synthetic image corresponding to each image of the image-based training set.
[0276] Aspect 73. The method of any of the previous aspects, wherein the enhanced image-based training data set comprises three synthetic images corresponding to each image of the image-based training set.
[0277] Aspect 74. The method of any of the previous aspects, wherein the imagebased training data set reflects a first population of subjects, and wherein the method is a method of generating an enhanced image-based training data set for training a deep learning model for use with individuals that are not adequately represented in the image-based training data set.
[0278] Aspect 75. The method of any of the previous aspects, further comprising: training a deep learning model using the enhanced image-based training data set.
[0279] Aspect 76. The method of any of the previous aspects, further comprising: training a deep learning model for image-based disease diagnosis or disease monitoring using the enhanced image-based training data set.
[0280] Aspect 77. The method of any of the previous aspects, wherein the method is a method of diversifying the image-based training data set. Aspect 78. The method of any of the previous aspects, wherein the method is a method of addressing limitations of the image-based training data set.
[0281] Aspect 79. The method of any of the previous aspects, wherein the method is a method of improving the accuracy of predictions of a deep learning model.
[0282] Aspect 80. The method of any of the previous aspects, wherein the method is a method of addressing underrepresented samples in the image-based training data set.
[0283] Aspect 81 . The method of any of the previous aspects, wherein the method is a method of enhancing underrepresented samples in the image-based training data set.
[0284] Aspect 82. The method of any of the previous aspects, wherein the method is a method of synthetically augmenting underrepresented samples in the imagebased training data set.
[0285] Aspect 83. The method of any of the previous aspects, wherein the method is a method of mitigating an imbalance in the image-based training data set.
[0286] Aspect 84. A computer implemented method of training a deep learning model to detect a result using an enhanced training data set, the method comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; applying an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on an image of the first subset of images of the image-based training data set; training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; obtaining an experimental image; and using the trained deep learning model to predict whether a result is present in the experimental image.
[0287] Aspect 85. A computer implemented method of training a deep learning model to detect a result using an enhanced training data set, the method comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; selecting a first subset of images from the image-based training data based at least in part on a first characteristic of the first subset of images, wherein the first characteristic is associated with a data-collection bias of the image-based training data set; generating synthetic images based on the images of the first subset of images by modifying aspects of each image of the first subset of images such that the synthetic images comprises variations of the first characteristic; training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; obtaining an experimental image; and using the trained deep learning model to predict whether a result is present in the experimental image.
[0288] Aspect 86. A computer-implemented method of generating an enhanced image-based training data set for deep learning models, the method comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images, and wherein the plurality of images exhibit a first characteristic at a first frequency; selecting a first subset of images from the image-based training data based at least in part on a presence of the first characteristic in each image of the first subset of images; generating a first number of synthetic images based on the first subset of images by modifying aspects of one or more images of the first subset of images; combining the image-based training data set and the first number of synthetic images to generate an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image such that the images of the enhanced image-based training data set exhibit the first characteristic at a second frequency. Aspect 87. The method of aspect 66, wherein the image-based training data set corresponds to a first subject population exhibiting the first characteristic at the first frequency, and the enhanced image-based training data set corresponds to a second subject population exhibiting the first characteristic at the second frequency.
[0289] Aspect 88. The method of any of aspects 86 to 87, wherein the first frequency is a biased frequency, and the second frequency is an unbiased frequency.
[0290] Aspect 89. A computer implemented method of training a deep learning model to detect a result using an enhanced training data set, the method comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; applying an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on a combination of at least two images of the first subset of images of the image-based training data set; training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; obtaining an experimental image; and using the trained deep learning model to predict whether a result is present in the experimental image.
[0291] Aspect 90. A system for generating an enhanced image-based training data set for deep learning models, the system comprising: a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; select a first image from the image-based training data based at least in part on a first characteristic of the first image; generate a first synthetic image based on the first image by modifying a first aspect of the first image; and combine the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image.
[0292] Aspect 91 . The system of aspect 90, wherein generating a first synthetic image based on the first image comprises modifying a representation of a structure present in the first image.
[0293] Aspect 92. The system of any of aspects 90 to 91 , wherein generating a first synthetic image based on the first image comprises introducing a variation in a structure present in the first image.
[0294] Aspect 93. The system of any of aspects 90 to 92, wherein generating a first synthetic image based on the first image comprises introducing random variations into the first image.
[0295] Aspect 94. The system of any of aspects 90 to 93, wherein generating a first synthetic image based on the first image comprises introducing a variation in the first image other than mirroring, rotating, smoothing, or contrast reduction of the first image.
[0296] Aspect 95. The system of any of aspects 90 to 94, wherein generating a first synthetic image based on the first image comprises: applying an autoencoder or a diffusion-based model or a generative adversarial network (GAN) to the first image to generate the first synthetic image. Aspect 96. The system of any of aspects 90 to 95, wherein generating a first synthetic image based on the first image comprises applying an autoencoder to the first image to generate the first synthetic image, wherein the autoencoder is a trained autoencoder.
[0297] Aspect 97. The system of any of aspects 95 to 96, wherein the autoencoder is trained on a second image-based training data set, wherein the second imagebased training data set does not comprise any images of the image-based training data set. Aspect 98. The system of any of aspects 95 to 97, wherein applying an autoencoder to the first image to generate the first synthetic image comprises introducing a controlled variation of the first image into the first synthetic image. Aspect 99. The system of any of aspects 95 to 98, wherein applying an autoencoder to the first image to generate the first synthetic image comprises: encoding the first image into an embedding; and decoding the embedding into the first synthetic image, wherein the embedding comprises a representation of the first image.
[0298] Aspect 100. The system of any of aspect 99, wherein the embedding comprises a compressed representation of the first image.
[0299] Aspect 101 . The system of any of aspects 99 to 100, wherein the embedding comprises a compressed stochastic representation of the first image.
[0300] Aspect 102. The system of any of aspects 99 to 101 , wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: introduce noise into the embedding.
[0301] Aspect 103. The system of aspect 102, wherein the noise introduced into the embedding comprises one or more of uniform noise or step noise or random noise or patterned noise, wherein patterned noise is associated with a mathematical function, optionally comprising one or more of a sine or cosine function.
[0302] Aspect 104. The system of any of aspects 99 to 103, wherein the embedding comprises a linear representation of the first image.
[0303] Aspect 105. The system of any of any of aspects 102 to 103, wherein the noise introduced into the embedding comprises a linear representation of noise values. Aspect 106. The system of any of aspects 95 to 105, wherein the autoencoder comprises an encoder and a decoder.
[0304] Aspect 107. The system of any of aspects 95 to 106, wherein the autoencoder comprises a dual-level structure.
[0305] Aspect 108. The system of any of aspects 95 to 107, wherein the autoencoder comprises: an encoder, and a decoder, wherein the encoder comprises a first level encoder configured to encode a subset of the first image, and a second level encoder configured to encode an output of the first level encoder, and the decoder comprises a first level decoder, configured to decode a subset of the first synthetic image, and a second level decoder configured to decode an output of the first level decoder.
[0306] Aspect 109. The system of aspect 108, wherein the encoder and decoder are modular components.
[0307] Aspect 110. The system of any of aspects 108 to 109, wherein the encoder and decoder are independently trained.
[0308] Aspect 111. The system of any of aspects 108 to 110, wherein the encoder and decoder comprise one or more of neural networks or residual networks.
[0309] Aspect 112. The system of any of aspects 95 to 111 , wherein each level of the encoder and decoder are independently trained.
[0310] Aspect 113. The system of any of aspects 95 to 112, wherein each level of the encoder and decoder comprises a neural network.
[0311] Aspect 114. A system for generating an enhanced image-based training data set for deep learning models, the system comprising: a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; select first and second images from the image-based training data, wherein the first and second images each exhibit a common first characteristic; generate a first synthetic image based on the first and second images by combining the first image with the second image; and combine the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image.
[0312] Aspect 115. The system of aspect 114, wherein generating a first synthetic image based on the first and second images comprises combining a representation of the first and second images.
[0313] Aspect 116. The system of aspect 115, wherein combining a representation of the first and second images comprises a scaled linear combination of the representations of the first and second images.
[0314] Aspect 117. The system of any of aspects 114 to 116, wherein generating a first synthetic image based on the first and second images further comprises generating the first synthetic image based on comprises combining the first image, the second image and one or more images from the from the imagebased training data.
[0315] Aspect 118. The system of any of aspects 114 to 117, wherein generating a first synthetic image based on the first and second images by combining the first image with the second image comprises: applying an autoencoder or a diffusion-based model or a generative adversarial network (GAN) to the first and second images to generate the first synthetic image.
[0316] Aspect 119. The system of any of aspects 114 to 118, wherein generating a first synthetic image based on the first and second images by combining the first image with the second image comprises applying an autoencoder to the first and second images to generate the first synthetic image, wherein the autoencoder is a trained autoencoder.
[0317] Aspect 120. The system of aspect 119, wherein the autoencoder is trained on a second image-based training data set, wherein the second image-based training data set does not comprise any images of the image-based training data set. Aspect 121 . The system of any of aspects 119 to 120, wherein applying an autoencoder to the first and second images to generate the first synthetic image comprises: encoding the first image into a first embedding; encoding the second image into a second embedding; combining the first embedding with the second embedding to form a first synthetic embedding; and decoding the first synthetic embedding into the first synthetic image, wherein the first and second embeddings comprise representations of the first and second images, respectively.
[0318] Aspect 122. The system of any of aspect 121 , wherein the first and second embeddings comprise compressed representations of the first and second images, respectively.
[0319] Aspect 123. The system of any of aspects 121 to 122, wherein the first and second embeddings comprise compressed stochastic representations of the first and second images, respectively.
[0320] Aspect 124. The system of any of aspects 121 to 123, wherein the first embedding comprises a first linear representation, and the second embedding comprises a second linear representation.
[0321] Aspect 125. The system of any of any of aspects 121 to 124, wherein combining the first image with the second image comprises combining the first linear representation with the second linear representation in a scaled linear combination.
[0322] Aspect 126. The system of any of aspects 119 to 125, wherein the autoencoder comprises an encoder and a decoder.
[0323] Aspect 127. The system of any of aspects 119 to 126, wherein the autoencoder comprises a dual-level structure.
[0324] Aspect 128. The system of any of aspects 119 to 127, wherein the autoencoder comprises: an encoder, and a decoder, wherein the encoder comprises a first level encoder configured to encode a subset of the first and second images, and a second level encoder configured to encode outputs of the first level encoder, and the decoder comprises a first level decoder, configured to decode a subset of the first synthetic image, and a second level decoder configured to decode an output of the first level decoder.
[0325] Aspect 129. The system of any of aspects 119 to 128, wherein the encoder and decoder are modular components.
[0326] Aspect 130. The system of any of aspects 119 to 129, wherein the encoder and decoder are independently trained.
[0327] Aspect 131 . The system of any of aspects 119 to 130, wherein the encoder and decoder comprise one or more of neural networks or residual networks.
[0328] Aspect 132. The system of any of aspects 119 to 131 , wherein each level of the encoder and decoder are independently trained.
[0329] Aspect 133. The system of any of aspects 119 to 132, wherein each level of the encoder and decoder comprises a neural network.
[0330] Aspect 134. The system of any of aspects 90 to 133, wherein the plurality of images of the image-based training data set comprises subject images.
[0331] Aspect 135. The system of any of aspects 90 to 134, wherein the plurality of images of the image-based training data set comprises images of biologic tissues.
[0332] Aspect 136. The system of any of aspects 90 to 135, wherein the plurality of images of the image-based training data set corresponds to images of a plurality of subjects.
[0333] Aspect 137. The system of any of aspects 90 to 136, wherein the image-based training data set comprises images generated by one or more of the following modalities: photographic images, MRI images, X-ray images, CT images or ultrasound images.
[0334] Aspect 138. The system of any of aspects 90 to 137, wherein at least some of the plurality of images of the image-based training data set exhibit one or more of the following: glaucoma, cancer, a degenerative retinal diseases, an acquired retinal disease, diabetic retinopathy, aneurysm, multiple sclerosis, a spinal cord pathology, stroke, tumor, traumatic brain injury, joint injury or disease, heart disease, vascular disease, a liver disease, Alzheimer’s disease, epilepsy, peripheral nerve compression, renal disease, inflammatory bowel disease, a fracture, arthritis, a lung disease, tuberculosis, pneumonia, chronic obstructive pulmonary disease (COPD), emphysema, pulmonary edema, pneumothorax, a dental disease, a bone infection, scoliosis, osteoporosis, kidney disease, kidney stone, a presence of a foreign object, pulmonary embolism, trauma, stroke, an abdominal disorder, appendicitis, diverticulitis, an ovarian cyst, sinusitis, a thyroid gland disease, pregnancy, gallstone, ascites, pelvic organ prolapse or testicular torsion.
[0335] Aspect 139. The system of any of aspects 90 to 138, wherein the image-based training data set is a training data set for training a deep learning model.
[0336] Aspect 140. The system of any of aspects 90 to 139, wherein the image-based training data set is configured for training a deep learning model to perform one or more of: image classification, object detection or semantic segmentation. Aspect 141 . The system of any of aspects 90 to 140, wherein the image-based training data set is configured for training a deep learning model to diagnose or screen for a disease or monitor a disease state or disease progression.
[0337] Aspect 142. The system of any of aspects 90 to 141 , wherein the image-based training data set is a training data set for a convolutional neural network (CNN) or a vision transformer (ViT).
[0338] Aspect 143. The system of any of aspects 90 to 142, wherein the image-based training data set is a training data set for image-based disease diagnosis or disease monitoring.
[0339] Aspect 144. The system of any of aspects 90 to 143, wherein the image-based training data set comprises a first subset of images that exhibit a condition or a disease and a second subset of images that do not exhibit the condition of the disease.
[0340] Aspect 145. The system of any of aspects 90 to 144, wherein the image-based training data set comprises a bias optionally comprising one or more of a bias related to an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location.
[0341] Aspect 146. The system of any of aspects 90 to 145, wherein the first characteristic optionally comprises one or more of: a diagnosis or a disease state or an injury state or a condition of an image subject. Aspect 147. The system of any of aspects 90 to 146, wherein the first characteristic is a demographic characteristic of a subject optionally comprising one or more of: an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location of an image subject.
[0342] Aspect 148. The system of any of aspects 90 to 147, wherein the first characteristic is relatively rare in the image-based training data set.
[0343] Aspect 149. The system of any of aspects 90 to 148, wherein the first characteristic is underrepresented in the image-based training data set. Aspect 150. The system of any of aspects 90 to 149, wherein it is relatively rare for a subject in the image-based training data set to exhibit the first characteristic. Aspect 151 . The system of any of aspects 90 to 150, wherein the first characteristic depicts an image aspect a rare disease.
[0344] Aspect 152. The system of any of aspects 90 to 151 , wherein the first or second images correspond to an underrepresented group.
[0345] Aspect 153. The system of any of aspects 90 to 152, wherein an image subject is a member of an underrepresented group.
[0346] Aspect 154. The system of any of aspects 90 to 153, wherein the first characteristic is selected to address a bias in the image-based training data set. Aspect 155. The system of any of aspects 90 to 154, wherein the first characteristic is selected to enhance performance of the deep learning model. Aspect 156. The system of any of aspects 90 to 155, wherein the first characteristic is selected to enhance performance of the deep learning model across a more diverse range of subjects.
[0347] Aspect 157. The system of any of aspects 90 to 156, wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: generate a plurality of synthetic images based on the first image by one or more modifications to one or more aspects of the first image; and combine the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models.
[0348] Aspect 158. The system of any of aspects 90 to 157, wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: generate a plurality of synthetic images, wherein the synthetic images are generated based on more than one of the plurality of images; and combine the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models.
[0349] Aspect 159. The system of any of aspects 90 to 158, wherein the image-based training data does not reflect a demographic characteristic of a select population, and the enhanced image-based training data set is generated to reflect the demographic characteristic of the select population.
[0350] Aspect 160. The system of any of aspects 90 to 159, wherein the enhanced image-based training data set comprises at least one synthetic image corresponding to each image of a plurality of images of the image-based training set.
[0351] Aspect 161 . The system of any of aspects 90 to 160, wherein the enhanced image-based training data set comprises at least one synthetic image corresponding to each image of the image-based training set.
[0352] Aspect 162. The system of any of aspects 90 to 161 , wherein the enhanced image-based training data set comprises three synthetic images corresponding to each image of the image-based training set.
[0353] Aspect 163. The system of any of aspects 90 to 162, wherein the image-based training data set reflects a first population of subjects, and wherein the system is a system for generating an enhanced image-based training data set for training a deep learning model for use with individuals that are not adequately represented in the image-based training data set. Aspect 164. The system of any of aspects 90 to 163, wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: train a deep learning model using the enhanced image-based training data set.
[0354] Aspect 165. The system of any of aspects 90 to 164, wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: train a deep learning model for image-based disease diagnosis or disease monitoring using the enhanced image-based training data set.
[0355] Aspect 166. The system of any of aspects 90 to 165, wherein the system is configured to diversify the image-based training data set.
[0356] Aspect 167. The system of any of aspects 90 to 166, wherein the system is configured to address limitations of the image-based training data set.
[0357] Aspect 168. The system of any of aspects 90 to 167, wherein the system is configured to improve the accuracy of predictions of a deep learning model. Aspect 169. The system of any of aspects 90 to 168, wherein the system is configured to address underrepresented samples in the image-based training data set.
[0358] Aspect 170. The system of any of aspects 90 to 169, wherein the system is configured to enhance underrepresented samples in the image-based training data set.
[0359] Aspect 171 . The system of any of aspects 90 to 170, wherein the system is configured to synthetically augment underrepresented samples in the imagebased training data set.
[0360] Aspect 172. The system of any of aspects 90 to 171 , wherein the system is configured to mitigate an imbalance in the image-based training data set.
[0361] Aspect 173. A system for training a deep learning model to detect a result using an enhanced training data set, the system comprising: a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; apply an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on an image of the first subset of images of the image-based training data set; train a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; obtain an experimental image; and use the trained deep learning model to predict whether a result is present in the experimental image.
[0362] Aspect 174. A system for training a deep learning model to detect a result using an enhanced training data set, the system comprising: a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; select a first subset of images from the image-based training data based at least in part on a first characteristic of the first subset of images, wherein the first characteristic is associated with a data-collection bias of the image-based training data set; generate synthetic images based on the images of the first subset of images by modifying aspects of each image of the first subset of images such that the synthetic images comprises variations of the first characteristic; train a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; obtain an experimental image; and use the trained deep learning model to predict whether a result is present in the experimental image.
[0363] Aspect 175. A system for generating an enhanced image-based training data set for deep learning models, the system comprising: a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images, and wherein the plurality of images exhibit a first characteristic at a first frequency; select a first subset of images from the image-based training data based at least in part on a presence of the first characteristic in each image of the first subset of images; generate a first number of synthetic images based on the first subset of images by modifying aspects of one or more images of the first subset of images; combine the image-based training data set and the first number of synthetic images to generate an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image such that the images of the enhanced image-based training data set exhibit the first characteristic at a second frequency.
[0364] Aspect 176. The system of aspect 175, wherein the image-based training data set corresponds to a first subject population exhibiting the first characteristic at the first frequency, and the enhanced image-based training data set corresponds to a second subject population exhibiting the first characteristic at the second frequency.
[0365] Aspect 177. The system of any of aspects 175 to 176, wherein the first frequency is a biased frequency, and the second frequency is an unbiased frequency. Aspect 178. A system for training a deep learning model to detect a result using an enhanced training data set, the system comprising: a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; apply an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on a combination of at least two images of the first subset of images of the image-based training data set; train a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; obtain an experimental image; and use the trained deep learning model to predict whether a result is present in the experimental image.
[0366] Aspect 179. The system of any of aspects 90 to 178, further comprising: a display device operably connected to the processor and memory and configured to display one or more of: aspects of the image-based training data set, one or more synthetic images, the enhanced image-based training data set or one or more result of applying the deep learning model.
[0367] Aspect 180. The system of any of aspects 90 to 179, further comprising: an input device operably connected to the processor and memory and configured to receive input from a user.
[0368] Aspect 181 . The system of aspect 180, wherein the input device is configured to receive a selection of one or more images of the image-based training data set.
[0369] Aspect 182. A non-transitory computer readable storage medium comprising instructions stored thereon, the instructions comprising: algorithm for obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; algorithm for selecting a first image from the image-based training data based at least in part on a first characteristic of the first image; algorithm for generating a first synthetic image based on the first image by modifying a first aspect of the first image; and algorithm for combining the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image.
[0370] Aspect 183. The non-transitory computer readable storage medium of aspect 182, wherein generating a first synthetic image based on the first image comprises modifying a representation of a structure present in the first image. Aspect 184. The non-transitory computer readable storage medium of any of aspects 182 to 183, wherein generating a first synthetic image based on the first image comprises introducing a variation in a structure present in the first image. Aspect 185. The non-transitory computer readable storage medium of any of aspects 182 to 184, wherein generating a first synthetic image based on the first image comprises introducing random variations into the first image.
[0371] Aspect 186. The non-transitory computer readable storage medium of any of aspects 182 to 185, further comprising: generating a plurality of synthetic images based on the first image by one or more modifications to one or more aspects of the first image; and combining the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models.
[0372] Aspect 187. The non-transitory computer readable storage medium of any of aspects 182 to 186, further comprising: generating a plurality of synthetic images, wherein the synthetic images are generated based on more than one of the plurality of images; and combining the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models.
[0373] Aspect 188. The non-transitory computer readable storage medium of any of aspects 182 to 187, wherein generating a first synthetic image based on the first image comprises introducing a variation in the first image other than mirroring, rotating, smoothing, or contrast reduction of the first image.
[0374] Aspect 189. The non-transitory computer readable storage medium of any of aspects 182 to 188, wherein generating a first synthetic image based on the first image comprises: applying an autoencoder or a diffusion-based model or a generative adversarial network (GAN) to the first image to generate the first synthetic image. Aspect 190. The non-transitory computer readable storage medium of any of aspects 182 to 189, wherein generating a first synthetic image based on the first image comprises applying an autoencoder to the first image to generate the first synthetic image, wherein the autoencoder is a trained autoencoder.
[0375] Aspect 191 . The non-transitory computer readable storage medium of aspect 190, wherein the autoencoder is trained on a second image-based training data set, wherein the second image-based training data set does not comprise any images of the image-based training data set.
[0376] Aspect 192. The non-transitory computer readable storage medium of any of aspects 190 to 191 , wherein applying an autoencoder to the first image to generate the first synthetic image comprises: introducing a controlled variation of the first image into the first synthetic image.
[0377] Aspect 193. The non-transitory computer readable storage medium of any of aspects 190 to 192, wherein applying an autoencoder to the first image to generate the first synthetic image comprises: encoding the first image into an embedding; and decoding the embedding into the first synthetic image, wherein the embedding comprises a representation of the first image. Aspect 194. The non-transitory computer readable storage medium of aspect 193, wherein the embedding comprises a compressed representation of the first image.
[0378] Aspect 195. The non-transitory computer readable storage medium of any of aspects 193 to 194, wherein the embedding comprises a compressed stochastic representation of the first image.
[0379] Aspect 196. The non-transitory computer readable storage medium of any of any of aspects 193 to 195, further comprising: algorithm for introducing noise into the embedding.
[0380] Aspect 197. The non-transitory computer readable storage medium of aspect 196, wherein the noise introduced into the embedding comprises one or more of uniform noise or step noise or random noise or patterned noise, wherein patterned noise is associated with a mathematical function, optionally comprising one or more of a sine or cosine function.
[0381] Aspect 198. The non-transitory computer readable storage medium of any of aspects 193 to 197, wherein the embedding comprises a linear representation of the first image.
[0382] Aspect 199. The non-transitory computer readable storage medium of any of any of aspects 193 to 198, wherein the noise introduced into the embedding comprises a linear representation of noise values.
[0383] Aspect 200. The non-transitory computer readable storage medium of any of aspects 189 to 199, wherein the autoencoder comprises an encoder and a decoder.
[0384] Aspect 201 . The non-transitory computer readable storage medium of any of aspects 177 to 188, wherein the autoencoder comprises a dual-level structure. Aspect 202. The non-transitory computer readable storage medium of any of aspects 189 to 200, wherein the autoencoder comprises: an encoder, and a decoder, wherein the encoder comprises a first level encoder configured to encode a subset of the first image, and a second level encoder configured to encode an output of the first level encoder, and the decoder comprises a first level decoder, configured to decode a subset of the first synthetic image, and a second level decoder configured to decode an output of the first level decoder.
[0385] Aspect 203. The non-transitory computer readable storage medium of aspect 202, wherein the encoder and decoder are modular components.
[0386] Aspect 204. The non-transitory computer readable storage medium of any of aspects 202 to 203, wherein the encoder and decoder are independently trained.
[0387] Aspect 205. The non-transitory computer readable storage medium of any of aspects 202 to 204, wherein the encoder and decoder comprise one or more of neural networks or residual networks.
[0388] Aspect 206. The non-transitory computer readable storage medium of any of aspects 189 to 205, wherein each level of the encoder and decoder are independently trained.
[0389] Aspect 207. The non-transitory computer readable storage medium of any of aspects 202 to 206, wherein each level of the encoder and decoder comprises a neural network.
[0390] Aspect 208. A non-transitory computer readable storage medium comprising instructions stored thereon, the instructions comprising: algorithm for obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; algorithm for selecting first and second images from the image-based training data, wherein the first and second images each exhibit a common first characteristic; algorithm for generating a first synthetic image based on the first and second images by combining the first image with the second image; and algorithm for combining the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the common characteristic of the first and second images.
[0391] Aspect 209. The non-transitory computer readable storage medium of aspect 208, wherein generating a first synthetic image based on the first and second images comprises combining a representation of the first and second images. Aspect 210. The non-transitory computer readable storage medium of any of aspects 208 to 209, wherein combining a representation of the first and second images comprises a scaled linear combination of the representations of the first and second images.
[0392] Aspect 211 . The non-transitory computer readable storage medium of any of aspects 208 to 210, wherein generating a first synthetic image based on the first and second images further comprises generating the first synthetic image based on comprises combining the first image, the second image and one or more images from the from the image-based training data.
[0393] Aspect 212. The non-transitory computer readable storage medium of any of aspects 208 to 21 1 , wherein generating a first synthetic image based on the first and second images by combining the first image with the second image comprises: algorithm for applying an autoencoder or a diffusion-based model or a generative adversarial network (GAN) to the first and second images to generate the first synthetic image.
[0394] Aspect 213. The non-transitory computer readable storage medium of any of any of aspects 208 to 212, wherein generating a first synthetic image based on the first and second images by combining the first image with the second image comprises applying an autoencoder to the first and second images to generate the first synthetic image, wherein the autoencoder is a trained autoencoder. Aspect 214. The non-transitory computer readable storage medium of any of aspects 208 to 213, wherein the autoencoder is trained on a second imagebased training data set, wherein the second image-based training data set does not comprise any images of the image-based training data set.
[0395] Aspect 215. The non-transitory computer readable storage medium of any of aspects 208 to 214, wherein algorithm for applying an autoencoder to the first and second images to generate the first synthetic image comprises: algorithm for encoding the first image into a first embedding; algorithm for encoding the second image into a second embedding; algorithm for combining the first embedding with the second embedding to form a first synthetic embedding; and algorithm for decoding the first synthetic embedding into the first synthetic image, wherein the first and second embeddings comprise representations of the first and second images, respectively.
[0396] Aspect 216. The non-transitory computer readable storage medium of aspect 215, wherein the first and second embeddings comprise compressed representations of the first and second images, respectively.
[0397] Aspect 217. The non-transitory computer readable storage medium of any of aspects 215 to 216, wherein the first and second embeddings comprise compressed stochastic representations of the first and second images, respectively.
[0398] Aspect 218. The non-transitory computer readable storage medium of any of aspects 215 to 217, wherein the first embedding comprises a first linear representation, and the second embedding comprises a second linear representation.
[0399] Aspect 219. The non-transitory computer readable storage medium of any of any of aspects 215 to 218, wherein combining the first image with the second image comprises combining the first linear representation with the second linear representation in a scaled linear combination. Aspect 220. The non-transitory computer readable storage medium of any of aspects 212 to 219, wherein the autoencoder comprises an encoder and a decoder.
[0400] Aspect 221 . The non-transitory computer readable storage medium of any of aspects 212 to 220, wherein the autoencoder comprises a dual-level structure. Aspect 222. The non-transitory computer readable storage medium of any of aspects 212 to 221 , wherein the autoencoder comprises: an encoder, and a decoder, wherein the encoder comprises a first level encoder configured to encode a subset of the first and second images, and a second level encoder configured to encode outputs of the first level encoder, and the decoder comprises a first level decoder, configured to decode a subset of the first synthetic image, and a second level decoder configured to decode an output of the first level decoder.
[0401] Aspect 223. The non-transitory computer readable storage medium of aspect 222, wherein the encoder and decoder are modular components.
[0402] Aspect 224. The non-transitory computer readable storage medium of any of aspects 222 to 223, wherein the encoder and decoder are independently trained. Aspect 225. The non-transitory computer readable storage medium of any of aspects 222 to 224, wherein the encoder and decoder comprise one or more of neural networks or residual networks.
[0403] Aspect 226. The non-transitory computer readable storage medium of any of aspects 222 to 225, wherein each level of the encoder and decoder are independently trained.
[0404] Aspect 227. The non-transitory computer readable storage medium of any of aspects 222 to 226, wherein each level of the encoder and decoder comprises a neural network.
[0405] Aspect 228. The non-transitory computer readable storage medium of any of aspects 182 to 227, wherein the plurality of images of the image-based training data set comprises subject images. Aspect 229. The non-transitory computer readable storage medium of any of aspects 182 to 228, wherein the plurality of images of the image-based training data set comprises images of biologic tissues.
[0406] Aspect 230. The non-transitory computer readable storage medium of any of aspects 182 to 229, wherein the plurality of images of the image-based training data set corresponds to images of a plurality of subjects.
[0407] Aspect 231 . The non-transitory computer readable storage medium of any of aspects 182 to 230, wherein the image-based training data set comprises images generated by one or more of the following modalities: photographic images, MRI images, X-ray images, CT images or ultrasound images.
[0408] Aspect 232. The non-transitory computer readable storage medium of any of aspects 182 to 231 , wherein at least some of the plurality of images of the imagebased training data set exhibit one or more of the following: glaucoma, cancer, a degenerative retinal diseases, an acquired retinal disease, diabetic retinopathy, aneurysm, multiple sclerosis, a spinal cord pathology, stroke, tumor, traumatic brain injury, joint injury or disease, heart disease, vascular disease, a liver disease, Alzheimer’s disease, epilepsy, peripheral nerve compression, renal disease, inflammatory bowel disease, a fracture, arthritis, a lung disease, tuberculosis, pneumonia, chronic obstructive pulmonary disease (COPD), emphysema, pulmonary edema, pneumothorax, a dental disease, a bone infection, scoliosis, osteoporosis, kidney disease, kidney stone, a presence of a foreign object, pulmonary embolism, trauma, stroke, an abdominal disorder, appendicitis, diverticulitis, an ovarian cyst, sinusitis, a thyroid gland disease, pregnancy, gallstone, ascites, pelvic organ prolapse or testicular torsion.
[0409] Aspect 233. The non-transitory computer readable storage medium of any of aspects 182 to 232, wherein the image-based training data set is a training data set for training a deep learning model.
[0410] Aspect 234. The non-transitory computer readable storage medium of any of aspects 182 to 233, wherein the image-based training data set is configured for training a deep learning model to perform one or more of: image classification, object detection or semantic segmentation. Aspect 235. The non-transitory computer readable storage medium of any of aspects 182 to 234, wherein the image-based training data set is configured for training a deep learning model to diagnose or screen for a disease or monitor a disease state or disease progression.
[0411] Aspect 236. The non-transitory computer readable storage medium of any of aspects 182 to 235, wherein the image-based training data set is a training data set for a convolutional neural network (CNN) or a vision transformer (ViT).
[0412] Aspect 237. The non-transitory computer readable storage medium of any of aspects 182 to 236, wherein the image-based training data set is a training data set for image-based disease diagnosis or disease monitoring.
[0413] Aspect 238. The non-transitory computer readable storage medium of any of aspects 182 to 237, wherein the image-based training data set comprises a first subset of images that exhibit a condition or a disease and a second subset of images that do not exhibit the condition of the disease.
[0414] Aspect 239. The non-transitory computer readable storage medium of any of aspects 182 to 238, wherein the image-based training data set comprises a bias optionally comprising one or more of a bias related to an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location.
[0415] Aspect 240. The non-transitory computer readable storage medium of any of aspects 182 to 239, wherein the first characteristic optionally comprises one or more of: a diagnosis or a disease state or an injury state or a condition of an image subject. Aspect 241 . The non-transitory computer readable storage medium of any of aspects 182 to 240, wherein the first characteristic is a demographic characteristic of the subject optionally comprising one or more of: an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location of an image subject. Aspect 242. The non-transitory computer readable storage medium of any of aspects 182 to 241 , wherein the first characteristic is relatively rare in the imagebased training data set.
[0416] Aspect 243. The non-transitory computer readable storage medium of any of aspects 182 to 242, wherein the first characteristic is underrepresented in the image-based training data set.
[0417] Aspect 244. The non-transitory computer readable storage medium of any of aspects 182 to 243, wherein it is relatively rare for a subject in the image-based training data set to exhibit the first characteristic.
[0418] Aspect 245. The non-transitory computer readable storage medium of any of aspects 182 to 244, wherein the first characteristic depicts an image aspect of a rare disease.
[0419] Aspect 246. The non-transitory computer readable storage medium of any of aspects 182 to 245, wherein the first or seconds images correspond to an image of an underrepresented group.
[0420] Aspect 247. The non-transitory computer readable storage medium of any of aspects 182 to 246, wherein an image subject is a member of an underrepresented group.
[0421] Aspect 248. The non-transitory computer readable storage medium of any of aspects 182 to 247, wherein the first characteristic is selected to address a bias in the image-based training data set.
[0422] Aspect 249. The non-transitory computer readable storage medium of any of aspects 182 to 248, wherein the first characteristic is selected to enhance performance of the deep learning model.
[0423] Aspect 250. The non-transitory computer readable storage medium of any of aspects 182 to 249, wherein the first characteristic is selected to enhance performance of the deep learning model across a more diverse range of subjects.
[0424] Aspect 251 . The non-transitory computer readable storage medium of any of aspects 182 to 250, wherein the image-based training data does not reflect a demographic characteristic of a select population, and the enhanced image-based training data set is generated to reflect the demographic characteristic of the select population.
[0425] Aspect 252. The non-transitory computer readable storage medium of any of aspects 182 to 251 , wherein the enhanced image-based training data set comprises at least one synthetic image corresponding to each image of a plurality of images of the image-based training set.
[0426] Aspect 253. The non-transitory computer readable storage medium of any of aspects 182 to 252, wherein the enhanced image-based training data set comprises at least one synthetic image corresponding to each image of the image-based training set.
[0427] Aspect 254. The non-transitory computer readable storage medium of any of aspects 182 to 253, wherein the enhanced image-based training data set comprises three synthetic images corresponding to each image of the imagebased training set.
[0428] Aspect 255. The non-transitory computer readable storage medium of any of aspects 182 to 254, wherein the image-based training data set reflects a first population of subjects, and wherein the method is a method of generating an enhanced image-based training data set for training a deep learning model for use with individuals that are not adequately represented in the image-based training data set.
[0429] Aspect 256. The non-transitory computer readable storage medium of any of aspects 182 to 255, further comprising: algorithm for training a deep learning model using the enhanced imagebased training data set.
[0430] Aspect 257. The non-transitory computer readable storage medium of any of aspects 182 to 256, further comprising: algorithm for training a deep learning model for image-based disease diagnosis or disease monitoring using the enhanced image-based training data set. Aspect 258. The non-transitory computer readable storage medium of any of aspects 182 to 258, wherein the non-transitory computer readable storage medium is configured for diversifying the image-based training data set.
[0431] Aspect 259. The non-transitory computer readable storage medium of any of aspects 182 to 258, wherein the non-transitory computer readable storage medium is configured for addressing limitations of the image-based training data set.
[0432] Aspect 260. The non-transitory computer readable storage medium of any of aspects 182 to 259, wherein the non-transitory computer readable storage medium is configured for improving the accuracy of predictions of a deep learning model.
[0433] Aspect 261 . The non-transitory computer readable storage medium of any of aspects 182 to 260, wherein the non-transitory computer readable storage medium is configured for addressing underrepresented samples in the imagebased training data set.
[0434] Aspect 262. The non-transitory computer readable storage medium of any of aspects 182 to 261 , wherein the non-transitory computer readable storage medium is configured for enhancing underrepresented samples in the imagebased training data set.
[0435] Aspect 263. The non-transitory computer readable storage medium of any of aspects 182 to 262, wherein the non-transitory computer readable storage medium is configured for synthetically augmenting underrepresented samples in the image-based training data set.
[0436] Aspect 264. The non-transitory computer readable storage medium of any of aspects 182 to 263, wherein the non-transitory computer readable storage medium is configured for mitigating an imbalance in the image-based training data set.
[0437] Aspect 265. A non-transitory computer readable storage medium for training a deep learning model to detect a result using an enhanced training data set, the method comprising: algorithm for obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; algorithm for applying an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on an image of the first subset of images of the image-based training data set; algorithm for training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; algorithm for obtaining an experimental image; and algorithm for using the trained deep learning model to predict whether a result is present in the experimental image.
[0438] Aspect 266. A non-transitory computer readable storage medium for training a deep learning model to detect a result using an enhanced training data set, the method comprising: algorithm for obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; algorithm for selecting a first subset of images from the image-based training data based at least in part on a first characteristic of the first subset of images, wherein the first characteristic is associated with a data-collection bias of the image-based training data set; algorithm for generating synthetic images based on the images of the first subset of images by modifying aspects of each image of the first subset of images such that the synthetic images comprises variations of the first characteristic; algorithm for training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; algorithm for obtaining an experimental image; and algorithm for using the trained deep learning model to predict whether a result is present in the experimental image.
[0439] Aspect 267. A non-transitory computer readable storage medium for generating an enhanced image-based training data set for deep learning models, the method comprising: algorithm for obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images, and wherein the plurality of images exhibit a first characteristic at a first frequency; algorithm for selecting a first subset of images from the image-based training data based at least in part on a presence of the first characteristic in each image of the first subset of images; algorithm for generating a first number of synthetic images based on the first subset of images by modifying aspects of one or more images of the first subset of images; algorithm for combining the image-based training data set and the first number of synthetic images to generate an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image such that the images of the enhanced image-based training data set exhibit the first characteristic at a second frequency.
[0440] Aspect 268. The non-transitory computer readable storage medium of aspect 267, wherein the image-based training data set corresponds to a first subject population exhibiting the first characteristic at the first frequency, and the enhanced image-based training data set corresponds to a second subject population exhibiting the first characteristic at the second frequency.
[0441] Aspect 269. The non-transitory computer readable storage medium of any of aspects 267 to 268, wherein the first frequency is a biased frequency, and the second frequency is an unbiased frequency. Aspect 270. A non-transitory computer readable storage medium for training a deep learning model to detect a result using an enhanced training data set, the method comprising: algorithm for obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; algorithm for applying an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on a combination of at least two images of the first subset of images of the image-based training data set; algorithm for training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; algorithm for obtaining an experimental image; and algorithm for using the trained deep learning model to predict whether a result is present in the experimental image.
[0442] The following is offered by way of illustration and not by way of limitation.
[0443] EXPERIMENTAL
[0444] Embodiments of methods of the present invention were applied in connection with obtaining experimental results described below. The experimental results described below relate to enhancing image-based training data sets in connection with training deep learning models to detect the presence of glaucoma in a subject that is a patient.
[0445] Synthetic Image Production
[0446] As described, in embodiments, a key function of an autoencoder is to produce high fidelity medical images for training set enhancement. In connection with collecting the presented experimental results, it was decided not to utilize a generative adversarial network (GAN) because training such a network itself is data intensive, and techniques such as diffusion can lead to hallucinations and produce nonsensical images that are counterproductive for the training of healthcare Als. See Goodfellow, I., et al. “Generative adversarial nets.” Advances in neural information processing systems 27 (2014). See also Ho, J., Jain, A. and Abbeel, P., 2020. Denoising diffusion probabilistic models.
[0447] Advances in neural information processing systems, 33, pp.6840-6851 . Instead, an autoencoder was leveraged to generate synthetic images whereby controlled variations were introduced to real-world medical images. See Kingma, D.P. and Welling, M., 2019. An introduction to variational autoencoders. Foundations and Trends® in Machine Learning, 12(4), pp.307-392. These variations go beyond simple rotations. The embodiment of the autoencoder used first abstracts existing images into embeddings (encoding process), before adding transformations through a neural net. The altered embeddings are then used to generate high fidelity synthetic images of the original medical data (decoding process). The embodiments of the encoder and decoder are both multi-layer ResNets and can be used to encode / decode any image. See He, K., Zhang, X., Ren, S. and Sun, J., 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778). As discussed, exemplary ResNet-based architectures are depicted in FIG. 10. For the purposes of collecting the presented experimental results, an embodiment of an autoencoder was applied to optic disc photos from publicly available databases. As discussed, an overview of such an exemplary process is depicted in FIG. 14. Each optic disc photo was used to generate three synthetic images by introducing different types of variations. 100% of the resultant images were recognizable as optic disc photos. One third of the generated synthetic images were randomly selected to expand the training set by 100%, for use towards improving Al performance in the detection of glaucomatous optic neuropathy. Autoencoder enhancement of ViT-based glaucoma detection
[0448] Performance of the embodiment of the autoencoder described was tested on six different publicly available training sets for glaucomatous optic neuropathy, ranging from a minimum of 101 images ( Drishti-GS 1 ) to a maximum of 800 images (REFUGE2), each pre-labeled as positive / glaucomatous images versus negative / control images. A summary of the image-based training data sets utilized is presented in FIG. 19. Embodiments of a deep learning model comprising a ViT were first trained on 80% of each dataset and evaluated on the remaining 20%, confirming performance levels similar to previous works with these datasets in the literature. See Diaz-Pinto, A., et al. “CNNs for automatic glaucoma assessment using fundus images: an extensive validation.” Biomedical engineering online 18 (2019): 1 -19. See also Fan, R., et al. “Detecting glaucoma from fundus photographs using deep learning without convolutions: Transformer for improved generalization.” Ophthalmology science 3.1 (2023): 100233. Each database, i.e., each publicly available training data set, was then enhanced by synthetic images generated by the embodiment of the autoencoder according to the present invention, leading to a 100% expansion of the training set (80% of each dataset). The same ViT models were then retrained using the autoencoder-enhanced datasets and evaluated on the same (all original, unenhanced) 20% validation subset of each database. Specific performance metrics are reported herein, organized by database, in order from smallest dataset to largest.
[0449] Drishti-GS1
[0450] The Drishti-GS1 dataset comprises 101 retinal fundus images, with 70 images representing positive / glaucoma samples and 31 images as negative / control samples. An embodiment of a ViT model trained on Disghti- GS1 showed an AUG of 0.67, a sensitivity of 0.93, and a specificity of 0.40. When enhanced with 80 synthetic images (Drishti-GS1 + autoencoder), the embodiment of the ViT model performance improved to an AUG of 0.77, sensitivity of 0.93, and specificity of 0.60. Results are summarized in FIG. 20. sichoi86-HRF
[0451] The sjchoi86-HRF dataset comprises 401 retinal fundus images, of which 101 are positive / glaucoma samples and 300 are negative / control samples. An embodiment of a ViT model trained with sjchoi86-HRF yielded an AUG of 0.79, sensitivity of 0.67, and specificity of 0.92. When enhanced with 320 synthetic images (sjchoi86-HRF + autoencoder), the embodiment of a ViT model achieved an improved AUG of 0.84, sensitivity of 0.76, and specificity of 0.92. Results are summarized in FIG. 20.
[0452] RIM-ONE
[0453] RIM-ONE comprises 485 retinal fundus images, of which 173 are positive I glaucoma samples and 312 are negative I control samples. An embodiment of a ViT model trained with RIM-ONE showed an AUG, sensitivity, and specificity of 0.88, 0.91 , and 0.85 respectively. When enhanced with 388 synthetic images (RIM-ONE + autoencoder), the embodiment of a ViT model was able to achieve an improved AUG, sensitivity, and specificity of 0.92, 0.91 , and 0.94 respectively. Results are summarized in FIG. 20.
[0454] ORIGA
[0455] The ORIGA dataset comprises 650 retinal fundus images, with 168 images representing positive I glaucoma samples and 482 images serving as negative I control samples. An embodiment of a ViT model trained on ORIGA performed with 0.69 AUG, 0.52 sensitivity, and 0.85 specificity. When enhanced with 520 synthetic images (ORIGA + autoencoder), the embodiment of a ViT model performance improved to an AUC of 0.78, sensitivity of 0.91 , and specificity of 0.65. Results are summarized in FIG. 20.
[0456] ACRIMA
[0457] The ACRIMA dataset comprises 705 retinal fundus images, with 396 images representing positive / glaucoma samples and 309 images as negative / control samples. An embodiment of a ViT model trained on ACRIMA showed an AUG of 0.94, a sensitivity of 1 , and a specificity of 0.88. When enhanced with 564 synthetic training images (ACRIMA + autoencoder), the embodiment of a ViT model performance improved to an AUC of 0.99, sensitivity of 1 , and specificity of 0.98. Results are summarized in FIG. 20.
[0458] Reduction in training data requirement
[0459] The largest publicly available database that was explored is REFUGE2, which comprises 800 retinal fundus images, including 80 positive / glaucoma samples and 720 negative / control samples. A ViT model trained using REFUGE2 resulted in an AUC of 0.95, sensitivity of 0.91 , and specificity of 1 .00. Given this excellent performance, this database was used to further assess the ability of the embodiment to reduce training data requirements. In FIG. 21 , the baseline results were obtained through the application of various image augmentation techniques, such as cropping, flipping, and randomization of image properties. For the data imbalance, down sampling was employed, weighted loss was implemented , and the best results for the baseline were presented. FIG. 21 shows ViT performance using a random 10%, 20%, 30%, 40%, 50%, 60%, 70%, and 80% of the training subset of REFUGE2, resulting in AUCs of 0.62, 0.72, 0.66, 0.50, 0.68, 0.81 , 0.78, and 0.77 respectively. When enhanced with synthetic images from autoencoder, ViT performance improved to AUCs of 0.75, 0.77, 0.83, 0.82, 0.78, 0.84, 0.87, and 0.93 respectively. It's worth noting that when using 80% of training data with synthetic images, the AUC is comparable to that of the full training dataset, suggesting that the introduction of autoencoder and the use of resultant synthetic images are able to reduce the amount of data required for a similar level of ViT performance by 20%.
[0460] Datasets:
[0461] As described, original patient fundus photos were sourced from six public datasets comprising retinal fundus images for glaucoma detection. See Poplin, R., et al. “Prediction of cardiovascular risk factors from retinal fundus photographs via deep learning.” Nature biomedical engineering 2.3 (2018): 158- 164. See also Korot, E., et al. “Predicting sex from retinal fundus photographs using automated deep learning.” Scientific reports 11.1 (2021 ): 10286; Lo, J., et al. “Federated learning for microvasculature segmentation and diabetic retinopathy classification of OCT data.” Ophthalmology Science 1 .4 (2021 ): 100069. Each dataset was independently explored and divided into training and validation sets. Due to the relatively small sizes of the datasets, a separate testing set was not allocated. ACRIMA consists of 705 retinal fundus images, captured from dilated eyes, centered on the optic disc, and selected based on high image quality. See Diaz-Pinto, A., et al. “CNNs for automatic glaucoma assessment using fundus images: an extensive validation.” Biomedical engineering online 18 (2019): 1 -19. Images were annotated by two experienced glaucoma experts. Drishti-GS 1 was specifically designed for automated glaucoma assessment and comprises 101 fundus images. See Sivaswamy, J., Krishnadas, S.R., Joshi, G.D., Jain, M. and Tabish, A. U.S., 2014, April. Drishti- gs: Retinal image dataset for optic nerve head (onh) segmentation. In 2014 IEEE 11 th international symposium on biomedical imaging (ISBI) (pp. 53-56), IEEE. The patient population for this database spanned 40-80 years of age with roughly equal numbers of males and females. All included images were annotated by glaucoma experts. ORIGA contains 650 retinal fundus images collected and annotated by the Singapore Eye Research Institute during a three- year period from 2004 to 2007, involving 149 adult patients, sampled from a larger cross-sectional epidemiological study for visual impairment risk factors among Singapore Malay adults aged from 40 to 80. See Zhang, Z., et al. “Origa- light: An online retinal fundus image database for glaucoma analysis and research.” 2010 Annual international conference of the IEEE engineering in medicine and biology. IEEE, 2010. RIM-ONE comprises 485 retinal fundus images collected from three hospitals in different regions of Spain. See Fumero, F., Alayon, S., Sanchez, J.L., Sigut, J. and Gonzalez-Hernandez, M., 2011 , June. RIM-ONE: An open retinal image database for optic nerve evaluation. In 2011 24th international symposium on computer-based medical systems (CBMS) (pp. 1 -6), IEEE. Images were annotated by five glaucoma experts from these hospitals. The subjects included randomly selected glaucomatous patients and volunteers with healthy eyes. Sjchoi86-HRF contains fundus images from four ophthalmological conditions - normal, cataract, glaucoma, and retinal disease. See sjchoi86: sjchoi86-HRF Database. GitHub. https: / / github.com / yiweichen04 / retina_dataset. Only the normal and glaucoma data are used in this study and together a total of 401 fundus images. REFUGE2 provides 800 color fundus photos, captured from both the right and left eyes, and annotated by experienced ophthalmologists. See Fang, H., et al. “REFUGE2 Challenge: A Treasure Trove for Multi-Dimension Analysis and Evaluation in Glaucoma Screening.” arXiv preprint arXiv:2202.08994 (2022).
[0462] Data Preparation:
[0463] To minimize the presence of redundant information, which may negatively impact deep learning model performance, all images were pre-processed to extract the region around the optic nerve head from each fundus image as shown in FIG. 22, in which pre-processed image 2202 is extracted from original image 2201 . This was achieved using deeplabv3plus, a semantic segmentation model. See Chen, L., et al. “Encoder-decoder with atrous separable convolution for semantic image segmentation.” Proceedings of the European conference on computer vision (ECCV). 2018. Once the region of interest was extracted, a square area centered around the disc was automatically cropped. These extracted regions of interest were then utilized to train deep learning models comprising ViTs. Focusing on these specific areas was intended to improve the model’s ability to identify glaucoma-related features and enhance the accuracy of the automated detection system.
[0464] Autoencoder:
[0465] As described, the autoencoder is a powerful generative model that has the ability to learn the underlying distribution of data and generate new data samples based on that distribution. See Kingma, D.P. and Welling, M., 2019. An introduction to variational autoencoders. Foundations and Trends® in Machine Learning, 12(4), pp.307-392. As described, a schematic of an exemplary autoencoder is shown in FIG. 7. The embodiment of the autoencoder utilized in connection with collecting the experimental results presented here comprises two components: an encoder and a decoder, both of which are neural networks, though such can be variants thereof depending on different specific applications.
[0466] The embodiments of the encoder and decoder employed in connection with collecting these experimental results were designed hierarchically, in some respects similar to the ResNet architecture. See He, K., Zhang, X., Ren, S. and Sun, J., 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778). They are composed of fundamental building blocks known as ResNet blocks. The encoder utilizes ResNet contraction blocks, while the decoder employs ResNet expansion blocks. As described, FIG. 10, panels (a) and (b) depict exemplary neural network architectures of these blocks. These blocks are organized into two groups: the ResNet contraction group and the ResNet expansion group, each comprising two blocks. As described, FIG. 10, panels (c) and (d) depict exemplary groups and blocks thereof.
[0467] The embodiment of the autoencoder employed in connection with collecting the experimental results presented herein features a dual-level structure. The Level 1 encoder and decoder are configured for processing images of size 12 x 12, producing an embedding of size 48 (also referred to as channel size). Such an exemplary architecture is described in connection with FIG. 11. Since most images, including those processed in connection with embodiments used to collect the presented experimental results, exceed a 12 x 12 resolution, such images were segmented into 12 x 12 patches. Each patch is then encoded into a separate embedding for further processing. For instance, a 96 x 96 image can be encoded by the Level 1 encoder into a latent space of dimensions 48 x 8 x 8, where 48 represents the channel size (embedding size). Additionally, embodiments of a Level 2 encoder and decoder were developed for processing the latent space. A schematic of an exemplary architecture of such a Level 2 encoder and decoder is depicted in FIG. 12.
[0468] The combined functionality of the Level 1 and Level 2 encoders and decoders enables converting a 96 x 96 image into an embedding of size 1024. As discussed, exemplary Level 1 and Level 2 encoders and decoders of an embodiment of an autoencoder are depicted in FIG. 13. Such dual-level structure offers several advantages over a single-level system, including, for example, training efficiency: by segmenting the encoder and decoder into two levels, each can be trained independently, reducing the required GPU memory; and dynamic composition ability: a specific Level 1 encoder can be paired with various Level 2 encoders, allowing for the encoding of images of different sizes into latent of varying dimensions.
[0469] In the training of the embodiment of the autoencoder in connection with obtaining the presented results, the employed loss function was defined based on the discrepancy in pixel values between the input image fed into the encoder and the resultant image output from the decoder. The embodiment of the autoencoder developed in connection with collecting the presented experimental results is a general-purpose image autoencoder, capable of encoding and decoding images across a wide spectrum of types, rather than being restricted to images belonging to a specific category. Consequently, the training dataset was selected to encompass as broad a range of images as feasible. In connection with the embodiment presented, it was found that ImageNet is highly suitable due to its extensive and diverse collection of images. See Deng, J., et al. “ImageNet: A large-scale hierarchical image database.” 2009 IEEE conference on computer vision and pattern recognition. IEEE, 2009. Given the vast quantity of available training data in ImageNet, the conventional approach of segregating data into training and validation subsets was not employed. Instead, uninterrupted training was conducted. Following approximately 200,000 training iterations, the autoencoder demonstrated proficiency in encoding and decoding images beyond those contained within the ImageNet dataset, achieving minimal loss. The specific training hyperparameters used in connection with the embodiment discussed are depicted in FIG. 23.
[0470] In the embodiment presented, the encoder of the autoencoder was trained to map input data to a lower-dimensional latent representation, which is a stochastic compressed representation of the data. In the embodiment presented, the decoder, on the other hand, was configured to generate output images from the latent representation using transpose convolutional neural net. The embodiment of the autoencoder developed in connection with the presented experimental results was trained with ImageNet, a general image data bank, without any additional medical images. Using the trained autoencoder, three times the number of training images available in each database were generated, and one third of those synthetic images were randomly selected to achieve 100% expansion of the training subset for each database.
[0471] Vision transformer (ViT) training and evaluation
[0472] In connection with collecting the presented experimental results, a deep learning model comprising a vision transformer architecture was selected due to its outstanding performance for glaucoma detection capability. See Hwang, E., et al. “Multi-Dataset Comparison of Vision Transformers and Convolutional Neural Networks for Detecting Glaucomatous Optic Neuropathy from Fundus Photographs.” Bioengineering, 2023. Each of the six public databases, as described, was split into a training set (80%) and a validation set (20%). The training set alone was used to generate synthetic images using the autoencoder. ViT performance was evaluated using both the original datasets and expanded datasets containing synthetic images. By comparing the performance of the ViTs trained with and without synthetic data, the specific contribution of the synthetic data in enhancing the model's performance was quantified.
[0473] To ensure consistency, the images used were standardized by resizing them to a uniform size of 224 x 224 pixels. Additionally, the pixel values were normalized to a range between 0 and 1 . During training, a batch size of 16 was used and the AdamW optimizer was used with a learning rate of 6e-4 and regularization of 6e-2. These hyperparameters were chosen to optimize the model’s convergence and performance. To compute the loss during training, cross-entropy loss was employed with 0.1 label smoothing.
[0474] The performance of the ViT models was evaluated using the area under the receiver operating characteristic curve (AUG) with 95% confidence intervals (Cl). The sensitivity and specificity of the system were also observed.
[0475] Synthetic images based on combinations of input images
[0476] Embodiments of the present disclosure were applied in connection with obtaining experimental results described below. The experimental results described below relate to enhancing image-based training data sets in connection with training deep learning models, in which a synthetic image was generated based on a combination of input images.
[0477] Dataset acquisition-.
[0478] Fundus photographs in an RP dataset were acquired from patients seen at University of California San Francisco (R.D.) with a familial (pedigree) and / or sequencing-confirmed diagnosis of non-syndromic retinitis pigmentosa, including autosomal dominant (AD), autosomal recessive (AR), X-linked recessive (XLR), or patients with sequencing-confirmed X-linked carrier (XLC) status. A total of 290 color (Topcon and Optos) and 108 auto fluorescent (Optos) fundus photographs were included in the full dataset. For the purpose of determining laterality, only fundus photos containing a visible macula and optic nerve head (ONH) were included. If the patient had multiple imaging dates, only the most recent date (up to August 2024) was included for review. Fundus photos from the most recent date were manually assessed for image quality (excessive blur, artifact, inclusion of macula and ONH in field of view). The highest quality image from each eye was selected for inclusion. Eyes without any images or with no images meeting these quality criteria were excluded. Symptom duration was defined as the length of time in months between patient-reported onset of visual symptoms and imaging date (R.D.). Asymptomatic patients were assigned a symptom duration of 0 months. Statistical analysis was performed with GraphPad Prism software (10.0.3) (https: / / www.graphpad.com / ). Experiments conducted adhered to the tenets of the Declaration of Helsinki and were approved by the LICSF Institutional Review Board, which determined that such retrospective study qualified for a waiver of informed consent.
[0479] Stable Diffusion Variational Autoencoder.
[0480] As described herein, a limitation related to deep learning models is overfitting during training. This limitation is especially pronounced when training data is limited. This leads to the model performing well for training data but failing to yield correct results on the testing data. One reason this happens is that the model is learning characteristics of the training data (e.g., one or more input images) that is irrelevant to the label (e.g., whether a disease or condition is present or exhibited in the test data, e.g., image). To address this limitation, embodiments of the present disclosure may be used to produce more training data using generative artificial intelligence (Al). For medical images, embodiments utilize an autoencoder tool to generate training data. Autoencoders comprise neural networks designed to compress input data into a lower-dimensional representation (sometimes referred to as a latent space or embedding or encoding) before reconstructing it, as illustrated in, for example, FIGS. 7-16. In a traditional autoencoder, the latent space is learned deterministically without any explicit probabilistic structure. This means that while the autoencoder network learns an efficient encoding, it does not enforce any particular distribution on the latent space. As a result, autoencoder models can be sensitive to small fluctuations in input data or model parameters, and the latent representation may become unstructured. A variational autoencoder builds on this by representing each latent attribute as a probability distribution, allowing for controlled variation in generated outputs. In the experimental results presented, a stable diffusion autoencoder was used due to its robustness and extensive validation in the generative Al space. Such autoencoders are well- suited for encoding and decoding a diverse range of images, even those with varying quality or structural complexity.
[0481] An important innovation of certain embodiments of the present disclosure, as illustrated in for example FIG. 9, was combining latent representations from two separate images, e.g., medical images, which have the same label (e.g., images of diseased tissue or images of healthy tissue) into a single latent representation, and then decoding the combined latent representation into a composite image (i.e., the synthetic image). The two precursor images are randomly picked among all images with the same label. The combination of the latent representations is a simple scaled linear combination.
[0482] Vision Transformer (ViT) Training and Evaluation-.
[0483] An 80 / 20 train-test split was applied to the dataset, and two ViT models were trained: one using only the original fundus images and the other incorporating both original and synthetic images. FIG. 24 illustrates the training and validation schema used in connection with collecting experimental results, including the training and generation of synthetic images processes utilized to generate the presented experimental results. Synthetic images were generated from and added to only the training set. The incorporation of synthetic data resulted in a twofold increase in the number of training images. The models’ performance was evaluated based on the area under the receiver operating characteristic curve (AUG). To calculate the AUG of a multi-class model a one- vs-rest approach was used, wherein the AUC was computed for each class against all others. Once each class had an individual AUC, the results were averaged to generate an overall performance metric. This approach ensured that the AUC effectively captured the model’s ability to distinguish between multiple classes. A broad hyperparameter search was conducted, and the optimal configuration was determined to be a pre-trained Vision Transformer employing a transformer encoder architecture, trained on a dataset comprising 14 million images. The use of a pre-trained model facilitated the extraction of rich feature representations learned from large-scale datasets, enhancing the model’s ability to generalize effectively to our class identification.
[0484] To calculate mean AUG, accuracy, recall, and specificity, models were run and performance averaged over five-fold cross-validation.
[0485] Results
[0486] Patient Characteristics
[0487] The final dataset was composed of 290 color and 108 auto fluorescent fundus photographs from 78 RP patients, including 28 autosomal dominant (AD), 20 autosomal recessive (AR), 22 X-linked recessive (XLR), and 8 X-linked carriers (XLC) (Table 1 , Figure 5). FIGS. 25A-25C summarize patient information for these experimental results. Of the 78 patients, 54 patients had fundus photos acquired with ultra-widefield retinal imaging (Optos) and 24 had photos acquired with a traditional fundus camera (Topcon). Patient mean age at time of imaging was 50.2 years (SD 15.5) for AD, 42.9 years (SD 18.7) for AR, 31 .6 years (SD 21 .0) for XLR, and 41 .7 years (SD 12.4) for XLC. Tukey’s multiple comparisons test only revealed a significant difference (adjusted p-value < 0.05) between the AD and XLR group, in line with known earlier onset of symptoms in XLR RP (De Silva 2021 ), while other comparisons were nonsignificant. See De Silva SR et al. The X-linked retinopathies: Physiological insights, pathogenic mechanisms, phenotypic features and novel therapies. Prog Retin Eye Res. 2021 May;82:100898. doi: 10.1016 / j.preteyeres.2020.100898. Epub 2020 Aug 26. PMID: 32860923, incorporated herein in its entirety. Patient mean symptom duration at time of imaging was 25.3 months (SD 17.6) for AD, 18.3 months (SD 18.0) for AR, 18.7 month (SD 15.4) for XLR, and 16.7 months (SD 10.4) for XLC, with no significant between-group differences.
[0488] Synthetic Image Generation /
[0489] A stable diffusion autoencoder, as described herein, was utilized in connection with generating synthetic images as illustrated in FIG. 24. Data Permutation
[0490] It was first examined how modifying input data variables would affect ViT model performance on the original and augmented datasets. In particular, image color was examined, as the dataset included both RGB (color) and monochromatic (autofluorescence) retinal images, and image flipping, a common data augmentation strategy in fundus photo training. For the purpose of testing these two variables, only the Optos ultra wide-field fundus photos were used to minimize confounding. In addition, given the relatively small size of the monochromatic dataset, the autofluorescence photos were further augmented with desaturated color fundus photos (“Greyscale”).
[0491] Model performance was tested on the original and synthetic image- expanded datasets for four conditions: Color with and without OS flipping, and Greyscale with and without OS flipping. Results of are depicted in FIG. 26. No change was observed in the mean AUG with flipping in the color condition in the flipped dataset (pre-expansion 0.51 , post-expansion 0.51 ) or in the non-flipped datasets (pre-expansion 0.51 , post-expansion 0.50), and only minimal improvement in the greyscale condition (Flipped: pre-expansion 0.49, postexpansion 0.52; No flip: pre-expansion 0.52, post-expansion 0.53). Similar trends were observed for accuracy and recall. A decrease in specificity after image expansion in the color flipped condition (pre-expansion 0.88, postexpansion 0.84) and an improvement in specificity without flipping (preexpansion 0.85, post-expansion 0.92) were noted. In contrast, the greyscale conditions demonstrated improvement in specificity with expansion regardless of flipping, though the non-flipped demonstrated better specificity overall (Flipped: pre-expansion 0.70, post-expansion 0.78; No flip: pre-expansion 0.76, postexpansion 0.82). The overall higher specificity without flipping suggests that flipping itself may introduce bias into ViT model performance, as previously found for fundus photo discrimination by CNNs. See Kang TS et al. Asymmetry between right and left fundus images identified using convolutional neural networks. Sci Rep. 2022 Jan 27;12(1 ):1444. doi: 10.1038 / s41598-021 -04323-3. Erratum in: Sci Rep. 2022 Apr 13;12(1 ):6225. doi: 10.1038 / s41598-022-10443-1 . PMID: 35087071 ; PMCID: PMC8795182, incorporated herein by reference. Based upon these results, full model training on the color non-flipped and greyscale non-flipped conditions was explored further.
[0492] Model performance
[0493] The Topcon and Optos fundus photos were combined into two datasets, color and greyscale (autofluorescence + desaturated color). Synthetic image expansion was performed, as shown in, for example FIG. 9, through random patient pairs (with replacement) and generation of combined images at 5 predefined ratios illustrated in FIG. 27, of which one was randomly chosen and the process repeated until the training set was expanded by 100%. The resulting images were generated using different combinations of scalar values for weighting the influence of each of the constituent images used to generate the synthetic image. Specifically, each synthetic image 2701 , 2702, 2703, 2704, 2705 was generated based on a combination of first and second constituent images. Synthetic image 2701 represents a combination of the first and second images with the first image weighted at 90% and the second image weighted at 10% (i.e., where synthetic image 2701 represents a linear combination of an embedding of the first image scaled by 90% and an embedding of the second image scaled at 10%). Synthetic image 2702 represents a combination of the same first and second images with the first image now weighted at 70% and the second image now weighted at 30% (i.e., where synthetic image 2702 represents a linear combination of an embedding of the first image scaled by 70% and an embedding of the second image scaled at 30%). Synthetic image 2703 represents a combination of the same first and second images with the first image now weighted at 50% and the second image now weighted at 50% (i.e., where synthetic image 2703 represents a linear combination of an embedding of the first image scaled by 50% and an embedding of the second image scaled at 50%). Synthetic image 2704 represents a combination of the same first and second images with the first image now weighted at 30% and the second image now weighted at 70% (i.e., where synthetic image 2704 represents a linear combination of an embedding of the first image scaled by 30% and an embedding of the second image scaled at 70%). Synthetic image 2705 represents a combination of the same first and second images with the first image now weighted at 10% and the second image now weighted at 90% (i.e., where synthetic image 2705 represents a linear combination of an embedding of the first image scaled by 10% and an embedding of the second image scaled at 90%).
[0494] It was found that model performance on prediction of inheritance mode (AD, AR, XLR, XLC) dramatically improved with synthetic image augmentation in both the color and greyscale conditions. These results are illustrated in FIGS. 28A-28B. Among the color fundus photos, post-expansion AUG improved to 0.91 (from 0.85), accuracy to 0.58 (from 0.45), recall to 0.58 (from 0.45), and specificity to 0.88 (from 0.83). Among the greyscale fundus photos, postexpansion AUG improved to 0.91 (from 0.78), accuracy to 0.64 (from 0.45), recall to 0.60 (from 0.41 ), and specificity to 0.87 (from 0.80). It is further noted that preexpansion model performance also improves with combination of Optos-imaged patient with Topcon-imaged patients. Overall, these results suggest that ViT model performance demonstrably improves with increased image diversity through expansion of patient image data and synthetic image generation according to an embodiment of the present disclosure.
[0495] Although the foregoing invention has been described in some detail by way of illustration and example for purposes of clarity of understanding, it is readily apparent to those of ordinary skill in the art in light of the teachings of this invention that certain changes and modifications may be made thereto without departing from the spirit or scope of the appended claims.
[0496] Accordingly, the preceding merely illustrates the principles of the invention. It will be appreciated that those skilled in the art will be able to devise various arrangements which, although not explicitly described or shown herein, embody the principles of the invention and are included within its spirit and scope. Furthermore, all examples and conditional language recited herein are principally intended to aid the reader in understanding the principles of the invention and the concepts contributed by the inventors to furthering the art and are to be construed as being without limitation to such specifically recited examples and conditions. Moreover, all statements herein reciting principles, aspects, and embodiments of the invention as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known equivalents and equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims.
[0497] The scope of the present invention, therefore, is not intended to be limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of present invention is embodied by the appended claims. In the claims, 35 U.S.C. § 1 12(f) or 35 U.S.C. § 1 12(6) is expressly defined as being invoked for a limitation in the claim only when the exact phrase “means for” or the exact phrase “step for” is recited at the beginning of such limitation in the claim; if such exact phrase is not used in a limitation in the claim, then 35 U.S.C. § 1 12(f) or 35 U.S.C. § 1 12(6) is not invoked.
Claims
What is claimed is:1 . A computer-implemented method of generating an enhanced imagebased training data set for deep learning models, the method comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; selecting a first image from the image-based training data based at least in part on a first characteristic of the first image; generating a first synthetic image based on the first image by modifying a first aspect of the first image; and combining the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image.
2. The method of claim 1 , wherein generating a first synthetic image based on the first image comprises modifying a representation of a structure present in the first image.
3. The method of any of the previous claims, wherein generating a first synthetic image based on the first image comprises introducing a variation in a structure present in the first image.
4. The method of any of the previous claims, wherein generating a first synthetic image based on the first image comprises introducing random variations into the first image.
5. The method of any of the previous claims, further comprising: generating a plurality of synthetic images based on the first image by one or more modifications to one or more aspects of the first image; andcombining the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models.
6. The method of any of the previous claims, further comprising: generating a plurality of synthetic images, wherein the synthetic images are generated based on more than one of the plurality of images; and combining the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models.
7. The method of any of the previous claims, wherein generating a first synthetic image based on the first image comprises introducing a variation in the first image other than mirroring, rotating, smoothing, or contrast reduction of the first image.
8. The method of any of the previous claims, wherein generating a first synthetic image based on the first image comprises: applying an autoencoder or a diffusion-based model or a generative adversarial network (GAN) to the first image to generate the first synthetic image.
9. The method of any of the previous claims, wherein generating a first synthetic image based on the first image comprises applying an autoencoder to the first image to generate the first synthetic image, wherein the autoencoder is a trained autoencoder.
10. The method of any of claims 8 to 9, wherein the autoencoder is trained on a second image-based training data set, wherein the second image-based training data set does not comprise any images of the image-based training data set.1 1 . The method of any of claims 8 to 10, wherein applying an autoencoder to the first image to generate the first synthetic image comprises: introducing a controlled variation of the first image into the first synthetic image.
12. The method of any of claims 8 to 11 , wherein applying an autoencoder to the first image to generate the first synthetic image comprises: encoding the first image into an embedding; and decoding the embedding into the first synthetic image, wherein the embedding comprises a representation of the first image.
13. The method of any of claim 1 , wherein the embedding comprises a compressed representation of the first image.
14. The method of any of claims 12 to 13, wherein the embedding comprises a compressed stochastic representation of the first image.
15. The method of any of any of claims 12 to 14, further comprising: introducing noise into the embedding.
16. The method of claim 15, wherein the noise introduced into the embedding comprises one or more of uniform noise or step noise or random noise or patterned noise, wherein patterned noise is associated with a mathematical function, optionally comprising one or more of a sine or cosine function.
17. The method of any of claims 12 to 16, wherein the embedding comprises a linear representation of the first image.
18. The method of any of any of claims 15 to 17, wherein the noise introduced into the embedding comprises a linear representation of noise values.
19. The method of any of claims 8 to 18, wherein the autoencoder comprises an encoder and a decoder.
20. The method of any of claims 8 to 19, wherein the autoencoder comprises a dual-level structure.21 . The method of any of claims 8 to 20, wherein the autoencoder comprises: an encoder, and a decoder, wherein the encoder comprises a first level encoder configured to encode a subset of the first image, and a second level encoder configured to encode an output of the first level encoder, and the decoder comprises a first level decoder, configured to decode a subset of the first synthetic image, and a second level decoder configured to decode an output of the first level decoder.
22. The method of any of claims 19 to 21 , wherein the encoder and decoder are modular components.
23. The method of any of claims 19 to 22, wherein the encoder and decoder are independently trained.
24. The method of any of claims 19 to 23, wherein the encoder and decoder comprise one or more of neural networks or residual networks.
25. The method of any of claims 19 to 24, wherein each level of the encoder and decoder are independently trained.
26. The method of any of claims 19 to 25, wherein each level of the encoder and decoder comprises a neural network.
27. A computer-implemented method of generating an enhanced imagebased training data set for deep learning models, the method comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; selecting first and second images from the image-based training data, wherein the first and second images each exhibit a common first characteristic; generating a first synthetic image based on the first and second images by combining the first image with the second image; and combining the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the common characteristic of the first and second images.
28. The method of claim 27, wherein generating a first synthetic image based on the first and second images comprises combining a representation of the first and second images.
29. The method of claim 28, wherein combining a representation of the first and second images comprises a scaled linear combination of the representations of the first and second images.
30. The method of any of claims 27 to 29, wherein generating a first synthetic image based on the first and second images further comprises generating the first synthetic image based on comprises combining the first image, the second image and one or more images from the from the image-based training data.31 . The method of any of claims 27 to 30, wherein generating a first synthetic image based on the first and second images by combining the first image with the second image comprises:applying an autoencoder or a diffusion-based model or a generative adversarial network (GAN) to the first and second images to generate the first synthetic image.
32. The method of any of claims 27 to 31 , wherein generating a first synthetic image based on the first and second images by combining the first image with the second image comprises applying an autoencoder to the first and second images to generate the first synthetic image, wherein the autoencoder is a trained autoencoder.
33. The method of claim 32, wherein the autoencoder is trained on a second image-based training data set, wherein the second image-based training data set does not comprise any images of the image-based training data set.
34. The method of any of claims 32 to 33, wherein applying an autoencoder to the first and second images to generate the first synthetic image comprises: encoding the first image into a first embedding; encoding the second image into a second embedding; combining the first embedding with the second embedding to form a first synthetic embedding; and decoding the first synthetic embedding into the first synthetic image, wherein the first and second embeddings comprise representations of the first and second images, respectively.
35. The method of claim 34, wherein the first and second embeddings comprise compressed representations of the first and second images, respectively.
36. The method of any of claims 34 to 35, wherein the first and second embeddings comprise compressed stochastic representations of the first and second images, respectively.
37. The method of any of claims 34 to 36, wherein the first embedding comprises a first linear representation, and the second embedding comprises a second linear representation.
38. The method of any of any of claims 34 to 37, wherein combining the first image with the second image comprises combining the first linear representation with the second linear representation in a scaled linear combination.
39. The method of any of claims 32 to 38, wherein the autoencoder comprises an encoder and a decoder.
40. The method of any of claims 32 to 39, wherein the autoencoder comprises a dual-level structure.41 . The method of any of claims 32 to 39, wherein the autoencoder comprises: an encoder, and a decoder, wherein the encoder comprises a first level encoder configured to encode a subset of the first and second images, and a second level encoder configured to encode outputs of the first level encoder, and the decoder comprises a first level decoder, configured to decode a subset of the first synthetic image, and a second level decoder configured to decode an output of the first level decoder.
42. The method of claim 41 , wherein the encoder and decoder are modular components.
43. The method of any of claims 41 to 42, wherein the encoder and decoder are independently trained.
44. The method of any of claims 41 to 43, wherein the encoder and decoder comprise one or more of neural networks or residual networks.
45. The method of any of claims 41 to 44, wherein each level of the encoder and decoder are independently trained.
46. The method of any of claims 41 to 45, wherein each level of the encoder and decoder comprises a neural network.
47. The method of any of the previous claims, wherein the plurality of images of the image-based training data set comprises subject images.
48. The method of any of the previous claims, wherein the plurality of images of the image-based training data set comprises images of biologic tissues.
49. The method of any of the previous claims, wherein the plurality of images of the image-based training data set corresponds to images of a plurality of subjects.
50. The method of any of the previous claims, wherein the image-based training data set comprises images generated by one or more of the following modalities: photographic images, MRI images, X-ray images, CT images or ultrasound images.51 . The method of any of the previous claims, wherein at least some of the plurality of images of the image-based training data set exhibit one or more of the following: glaucoma, cancer, a degenerative retinal diseases, an acquired retinal disease, diabetic retinopathy, aneurysm, multiple sclerosis, a spinal cord pathology, stroke, tumor, traumatic brain injury, joint injury or disease, heart disease, vascular disease, a liver disease, Alzheimer’s disease, epilepsy,peripheral nerve compression, renal disease, inflammatory bowel disease, a fracture, arthritis, a lung disease, tuberculosis, pneumonia, chronic obstructive pulmonary disease (COPD), emphysema, pulmonary edema, pneumothorax, a dental disease, a bone infection, scoliosis, osteoporosis, kidney disease, kidney stone, a presence of a foreign object, pulmonary embolism, trauma, stroke, an abdominal disorder, appendicitis, diverticulitis, an ovarian cyst, sinusitis, a thyroid gland disease, pregnancy, gallstone, ascites, pelvic organ prolapse or testicular torsion.
52. The method of any of the previous claims, wherein the image-based training data set is a training data set for training a deep learning model.
53. The method of any of the previous claims, wherein the image-based training data set is configured for training a deep learning model to perform one or more of: image classification, object detection or semantic segmentation.
54. The method of any of the previous claims, wherein the image-based training data set is configured for training a deep learning model to diagnose or screen for a disease or monitor a disease state or disease progression.
55. The method of any of the previous claims, wherein the image-based training data set is a training data set for a convolutional neural network (CNN) or a vision transformer (ViT).
56. The method of any of the previous claims, wherein the image-based training data set is a training data set for image-based disease diagnosis or disease monitoring.
57. The method of any of the previous claims, wherein the image-based training data set comprises a first subset of images that exhibit a condition or adisease and a second subset of images that do not exhibit the condition of the disease.
58. The method of any of the previous claims, wherein the image-based training data set comprises a bias optionally comprising one or more of a bias related to an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location.
59. The method of any of the previous claims, wherein the first characteristic optionally comprises one or more of: a diagnosis or a disease state or an injury state or a condition of an image subject.
60. The method of any of the previous claims, wherein the first characteristic is a demographic characteristic of the subject optionally comprising one or more of: an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location of an image subject.61 . The method of any of the previous claims, wherein the first characteristic is relatively rare in the image-based training data set.
62. The method of any of the previous claims, wherein the first characteristic is underrepresented in the image-based training data set.
63. The method of any of the previous claims, wherein it is relatively rare for a subject in the image-based training data set to exhibit the first characteristic.
64. The method of any of the previous claims, wherein the first characteristic depicts an image aspect of a rare disease.
65. The method of any of the previous claims, wherein the first or second images correspond to an underrepresented group.
66. The method of any of the previous claims, wherein an image subject is a member of an underrepresented group.
67. The method of any of the previous claims, wherein the first characteristic is selected to address a bias in the image-based training data set.
68. The method of any of the previous claims, wherein the first characteristic is selected to enhance performance of the deep learning model.
69. The method of any of the previous claims, wherein the first characteristic is selected to enhance performance of the deep learning model across a more diverse range of subjects.
70. The method of any of the previous claims, wherein the image-based training data does not reflect a demographic characteristic of a select population, and the enhanced image-based training data set is generated to reflect the demographic characteristic of the select population.71 . The method of any of the previous claims, wherein the enhanced imagebased training data set comprises at least one synthetic image corresponding to each image of a plurality of images of the image-based training set.
72. The method of any of the previous claims, wherein the enhanced imagebased training data set comprises at least one synthetic image corresponding to each image of the image-based training set.
73. The method of any of the previous claims, wherein the enhanced imagebased training data set comprises three synthetic images corresponding to each image of the image-based training set.
74. The method of any of the previous claims, wherein the image-based training data set reflects a first population of subjects, and wherein the method is a method of generating an enhanced image-based training data set for training a deep learning model for use with individuals that are not adequately represented in the image-based training data set.
75. The method of any of the previous claims, further comprising: training a deep learning model using the enhanced image-based training data set.
76. The method of any of the previous claims, further comprising: training a deep learning model for image-based disease diagnosis or disease monitoring using the enhanced image-based training data set.
77. The method of any of the previous claims, wherein the method is a method of diversifying the image-based training data set.
78. The method of any of the previous claims, wherein the method is a method of addressing limitations of the image-based training data set.
79. The method of any of the previous claims, wherein the method is a method of improving the accuracy of predictions of a deep learning model.
80. The method of any of the previous claims, wherein the method is a method of addressing underrepresented samples in the image-based training data set.81 . The method of any of the previous claims, wherein the method is a method of enhancing underrepresented samples in the image-based training data set.
82. The method of any of the previous claims, wherein the method is a method of synthetically augmenting underrepresented samples in the imagebased training data set.
83. The method of any of the previous claims, wherein the method is a method of mitigating an imbalance in the image-based training data set.
84. A computer implemented method of training a deep learning model to detect a result using an enhanced training data set, the method comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; applying an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on an image of the first subset of images of the image-based training data set; training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; obtaining an experimental image; and using the trained deep learning model to predict whether a result is present in the experimental image.
85. A computer implemented method of training a deep learning model to detect a result using an enhanced training data set, the method comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; selecting a first subset of images from the image-based training data based at least in part on a first characteristic of the first subset of images,wherein the first characteristic is associated with a data-collection bias of the image-based training data set; generating synthetic images based on the images of the first subset of images by modifying aspects of each image of the first subset of images such that the synthetic images comprises variations of the first characteristic; training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; obtaining an experimental image; and using the trained deep learning model to predict whether a result is present in the experimental image.
86. A computer-implemented method of generating an enhanced imagebased training data set for deep learning models, the method comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images, and wherein the plurality of images exhibit a first characteristic at a first frequency; selecting a first subset of images from the image-based training data based at least in part on a presence of the first characteristic in each image of the first subset of images; generating a first number of synthetic images based on the first subset of images by modifying aspects of one or more images of the first subset of images; combining the image-based training data set and the first number of synthetic images to generate an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image such that the images of the enhanced image-based training data set exhibit the first characteristic at a second frequency.
87. The method of claim 66, wherein the image-based training data set corresponds to a first subject population exhibiting the first characteristic at thefirst frequency, and the enhanced image-based training data set corresponds to a second subject population exhibiting the first characteristic at the second frequency.
88. The method of any of claims 86 to 87, wherein the first frequency is a biased frequency, and the second frequency is an unbiased frequency.
89. A computer implemented method of training a deep learning model to detect a result using an enhanced training data set, the method comprising: obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; applying an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on a combination of at least two images of the first subset of images of the image-based training data set; training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; obtaining an experimental image; and using the trained deep learning model to predict whether a result is present in the experimental image.
90. A system for generating an enhanced image-based training data set for deep learning models, the system comprising: a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; select a first image from the image-based training data based at least in part on a first characteristic of the first image;generate a first synthetic image based on the first image by modifying a first aspect of the first image; and combine the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image.91 . The system of claim 90, wherein generating a first synthetic image based on the first image comprises modifying a representation of a structure present in the first image.
92. The system of any of claims 90 to 91 , wherein generating a first synthetic image based on the first image comprises introducing a variation in a structure present in the first image.
93. The system of any of claims 90 to 92, wherein generating a first synthetic image based on the first image comprises introducing random variations into the first image.
94. The system of any of claims 90 to 93, wherein generating a first synthetic image based on the first image comprises introducing a variation in the first image other than mirroring, rotating, smoothing, or contrast reduction of the first image.
95. The system of any of claims 90 to 94, wherein generating a first synthetic image based on the first image comprises: applying an autoencoder or a diffusion-based model or a generative adversarial network (GAN) to the first image to generate the first synthetic image.
96. The system of any of claims 90 to 95, wherein generating a first synthetic image based on the first image comprises applying an autoencoder to the firstimage to generate the first synthetic image, wherein the autoencoder is a trained autoencoder.
97. The system of any of claims 95 to 96, wherein the autoencoder is trained on a second image-based training data set, wherein the second image-based training data set does not comprise any images of the image-based training data set.
98. The system of any of claims 95 to 97, wherein applying an autoencoder to the first image to generate the first synthetic image comprises introducing a controlled variation of the first image into the first synthetic image.
99. The system of any of claims 95 to 98, wherein applying an autoencoder to the first image to generate the first synthetic image comprises: encoding the first image into an embedding; and decoding the embedding into the first synthetic image, wherein the embedding comprises a representation of the first image.
100. The system of any of claim 99, wherein the embedding comprises a compressed representation of the first image.101 . The system of any of claims 99 to 100, wherein the embedding comprises a compressed stochastic representation of the first image.
102. The system of any of claims 99 to 101 , wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: introduce noise into the embedding.
103. The system of claim 102, wherein the noise introduced into the embedding comprises one or more of uniform noise or step noise or random noise orpatterned noise, wherein patterned noise is associated with a mathematical function, optionally comprising one or more of a sine or cosine function.
104. The system of any of claims 99 to 103, wherein the embedding comprises a linear representation of the first image.
105. The system of any of any of claims 102 to 103, wherein the noise introduced into the embedding comprises a linear representation of noise values.
106. The system of any of claims 95 to 105, wherein the autoencoder comprises an encoder and a decoder.
107. The system of any of claims 95 to 106, wherein the autoencoder comprises a dual-level structure.
108. The system of any of claims 95 to 107, wherein the autoencoder comprises: an encoder, and a decoder, wherein the encoder comprises a first level encoder configured to encode a subset of the first image, and a second level encoder configured to encode an output of the first level encoder, and the decoder comprises a first level decoder, configured to decode a subset of the first synthetic image, and a second level decoder configured to decode an output of the first level decoder.
109. The system of claim 108, wherein the encoder and decoder are modular components.1 10. The system of any of claims 108 to 109, wherein the encoder and decoder are independently trained.1 1 1. The system of any of claims 108 to 110, wherein the encoder and decoder comprise one or more of neural networks or residual networks.1 12. The system of any of claims 95 to 1 11 , wherein each level of the encoder and decoder are independently trained.1 13. The system of any of claims 95 to 1 12, wherein each level of the encoder and decoder comprises a neural network.1 14. A system for generating an enhanced image-based training data set for deep learning models, the system comprising: a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; select first and second images from the image-based training data, wherein the first and second images each exhibit a common first characteristic; generate a first synthetic image based on the first and second images by combining the first image with the second image; and combine the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image.1 15. The system of claim 114, wherein generating a first synthetic image based on the first and second images comprises combining a representation of the first and second images.1 16. The system of claim 115, wherein combining a representation of the first and second images comprises a scaled linear combination of the representations of the first and second images.1 17. The system of any of claims 114 to 116, wherein generating a first synthetic image based on the first and second images further comprises generating the first synthetic image based on comprises combining the first image, the second image and one or more images from the from the imagebased training data.1 18. The system of any of claims 114 to 117, wherein generating a first synthetic image based on the first and second images by combining the first image with the second image comprises: applying an autoencoder or a diffusion-based model or a generative adversarial network (GAN) to the first and second images to generate the first synthetic image.1 19. The system of any of claims 114 to 118, wherein generating a first synthetic image based on the first and second images by combining the first image with the second image comprises applying an autoencoder to the first and second images to generate the first synthetic image, wherein the autoencoder is a trained autoencoder.
120. The system of claim 119, wherein the autoencoder is trained on a second image-based training data set, wherein the second image-based training data set does not comprise any images of the image-based training data set.121 . The system of any of claims 119 to 120, wherein applying an autoencoder to the first and second images to generate the first synthetic image comprises: encoding the first image into a first embedding; encoding the second image into a second embedding;combining the first embedding with the second embedding to form a first synthetic embedding; and decoding the first synthetic embedding into the first synthetic image, wherein the first and second embeddings comprise representations of the first and second images, respectively.
122. The system of any of claim 121 , wherein the first and second embeddings comprise compressed representations of the first and second images, respectively.
123. The system of any of claims 121 to 122, wherein the first and second embeddings comprise compressed stochastic representations of the first and second images, respectively.
124. The system of any of claims 121 to 123, wherein the first embedding comprises a first linear representation, and the second embedding comprises a second linear representation.
125. The system of any of any of claims 121 to 124, wherein combining the first image with the second image comprises combining the first linear representation with the second linear representation in a scaled linear combination.
126. The system of any of claims 119 to 125, wherein the autoencoder comprises an encoder and a decoder.
127. The system of any of claims 119 to 126, wherein the autoencoder comprises a dual-level structure.
128. The system of any of claims 119 to 127, wherein the autoencoder comprises: an encoder, anda decoder, wherein the encoder comprises a first level encoder configured to encode a subset of the first and second images, and a second level encoder configured to encode outputs of the first level encoder, and the decoder comprises a first level decoder, configured to decode a subset of the first synthetic image, and a second level decoder configured to decode an output of the first level decoder.
129. The system of any of claims 119 to 128, wherein the encoder and decoder are modular components.
130. The system of any of claims 119 to 129, wherein the encoder and decoder are independently trained.
131. The system of any of claims 119 to 130, wherein the encoder and decoder comprise one or more of neural networks or residual networks.
132. The system of any of claims 119 to 131 , wherein each level of the encoder and decoder are independently trained.
133. The system of any of claims 119 to 132, wherein each level of the encoder and decoder comprises a neural network.
134. The system of any of claims 90 to 133, wherein the plurality of images of the image-based training data set comprises subject images.
135. The system of any of claims 90 to 134, wherein the plurality of images of the image-based training data set comprises images of biologic tissues.
136. The system of any of claims 90 to 135, wherein the plurality of images of the image-based training data set corresponds to images of a plurality of subjects.
137. The system of any of claims 90 to 136, wherein the image-based training data set comprises images generated by one or more of the following modalities: photographic images, MRI images, X-ray images, CT images or ultrasound images.
138. The system of any of claims 90 to 137, wherein at least some of the plurality of images of the image-based training data set exhibit one or more of the following: glaucoma, cancer, a degenerative retinal diseases, an acquired retinal disease, diabetic retinopathy, aneurysm, multiple sclerosis, a spinal cord pathology, stroke, tumor, traumatic brain injury, joint injury or disease, heart disease, vascular disease, a liver disease, Alzheimer’s disease, epilepsy, peripheral nerve compression, renal disease, inflammatory bowel disease, a fracture, arthritis, a lung disease, tuberculosis, pneumonia, chronic obstructive pulmonary disease (COPD), emphysema, pulmonary edema, pneumothorax, a dental disease, a bone infection, scoliosis, osteoporosis, kidney disease, kidney stone, a presence of a foreign object, pulmonary embolism, trauma, stroke, an abdominal disorder, appendicitis, diverticulitis, an ovarian cyst, sinusitis, a thyroid gland disease, pregnancy, gallstone, ascites, pelvic organ prolapse or testicular torsion.
139. The system of any of claims 90 to 138, wherein the image-based training data set is a training data set for training a deep learning model.
140. The system of any of claims 90 to 139, wherein the image-based training data set is configured for training a deep learning model to perform one or more of: image classification, object detection or semantic segmentation.141 . The system of any of claims 90 to 140, wherein the image-based training data set is configured for training a deep learning model to diagnose or screen for a disease or monitor a disease state or disease progression.
142. The system of any of claims 90 to 141 , wherein the image-based training data set is a training data set for a convolutional neural network (CNN) or a vision transformer (ViT).
143. The system of any of claims 90 to 142, wherein the image-based training data set is a training data set for image-based disease diagnosis or disease monitoring.
144. The system of any of claims 90 to 143, wherein the image-based training data set comprises a first subset of images that exhibit a condition or a disease and a second subset of images that do not exhibit the condition of the disease.
145. The system of any of claims 90 to 144, wherein the image-based training data set comprises a bias optionally comprising one or more of a bias related to an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location.
146. The system of any of claims 90 to 145, wherein the first characteristic optionally comprises one or more of: a diagnosis or a disease state or an injury state or a condition of an image subject.
147. The system of any of claims 90 to 146, wherein the first characteristic is a demographic characteristic of a subject optionally comprising one or more of: an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location of an image subject.
148. The system of any of claims 90 to 147, wherein the first characteristic is relatively rare in the image-based training data set.
149. The system of any of claims 90 to 148, wherein the first characteristic is underrepresented in the image-based training data set.
150. The system of any of claims 90 to 149, wherein it is relatively rare for a subject in the image-based training data set to exhibit the first characteristic.151 . The system of any of claims 90 to 150, wherein the first characteristic depicts an image aspect a rare disease.
152. The system of any of claims 90 to 151 , wherein the first or second images correspond to an underrepresented group.
153. The system of any of claims 90 to 152, wherein an image subject is a member of an underrepresented group.
154. The system of any of claims 90 to 153, wherein the first characteristic is selected to address a bias in the image-based training data set.
155. The system of any of claims 90 to 154, wherein the first characteristic is selected to enhance performance of the deep learning model.
156. The system of any of claims 90 to 155, wherein the first characteristic is selected to enhance performance of the deep learning model across a more diverse range of subjects.
157. The system of any of claims 90 to 156, wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to:generate a plurality of synthetic images based on the first image by one or more modifications to one or more aspects of the first image; and combine the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models.
158. The system of any of claims 90 to 157, wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: generate a plurality of synthetic images, wherein the synthetic images are generated based on more than one of the plurality of images; and combine the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models.
159. The system of any of claims 90 to 158, wherein the image-based training data does not reflect a demographic characteristic of a select population, and the enhanced image-based training data set is generated to reflect the demographic characteristic of the select population.
160. The system of any of claims 90 to 159, wherein the enhanced imagebased training data set comprises at least one synthetic image corresponding to each image of a plurality of images of the image-based training set.161 . The system of any of claims 90 to 160, wherein the enhanced imagebased training data set comprises at least one synthetic image corresponding to each image of the image-based training set.
162. The system of any of claims 90 to 161 , wherein the enhanced imagebased training data set comprises three synthetic images corresponding to each image of the image-based training set.
163. The system of any of claims 90 to 162, wherein the image-based training data set reflects a first population of subjects, and wherein the system is a system for generating an enhanced image-based training data set for training a deep learning model for use with individuals that are not adequately represented in the image-based training data set.
164. The system of any of claims 90 to 163, wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: train a deep learning model using the enhanced image-based training data set.
165. The system of any of claims 90 to 164, wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: train a deep learning model for image-based disease diagnosis or disease monitoring using the enhanced image-based training data set.
166. The system of any of claims 90 to 165, wherein the system is configured to diversify the image-based training data set.
167. The system of any of claims 90 to 166, wherein the system is configured to address limitations of the image-based training data set.
168. The system of any of claims 90 to 167, wherein the system is configured to improve the accuracy of predictions of a deep learning model.
169. The system of any of claims 90 to 168, wherein the system is configured to address underrepresented samples in the image-based training data set.
170. The system of any of claims 90 to 169, wherein the system is configured to enhance underrepresented samples in the image-based training data set.
171. The system of any of claims 90 to 170, wherein the system is configured to synthetically augment underrepresented samples in the image-based training data set.
172. The system of any of claims 90 to 171 , wherein the system is configured to mitigate an imbalance in the image-based training data set.
173. A system for training a deep learning model to detect a result using an enhanced training data set, the system comprising: a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; apply an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on an image of the first subset of images of the image-based training data set; train a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; obtain an experimental image; and use the trained deep learning model to predict whether a result is present in the experimental image.
174. A system for training a deep learning model to detect a result using an enhanced training data set, the system comprising: a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; select a first subset of images from the image-based training data based at least in part on a first characteristic of the first subset of images, wherein the first characteristic is associated with a data-collection bias of the image-based training data set; generate synthetic images based on the images of the first subset of images by modifying aspects of each image of the first subset of images such that the synthetic images comprises variations of the first characteristic; train a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; obtain an experimental image; and use the trained deep learning model to predict whether a result is present in the experimental image.
175. A system for generating an enhanced image-based training data set for deep learning models, the system comprising: a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images, and wherein the plurality of images exhibit a first characteristic at a first frequency; select a first subset of images from the image-based training data based at least in part on a presence of the first characteristic in each image of the first subset of images;generate a first number of synthetic images based on the first subset of images by modifying aspects of one or more images of the first subset of images; combine the image-based training data set and the first number of synthetic images to generate an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image such that the images of the enhanced image-based training data set exhibit the first characteristic at a second frequency.
176. The system of claim 175, wherein the image-based training data set corresponds to a first subject population exhibiting the first characteristic at the first frequency, and the enhanced image-based training data set corresponds to a second subject population exhibiting the first characteristic at the second frequency.
177. The system of any of claims 175 to 176, wherein the first frequency is a biased frequency, and the second frequency is an unbiased frequency.
178. A system for training a deep learning model to detect a result using an enhanced training data set, the system comprising: a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; apply an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on a combination of at least two images of the first subset of images of the image-based training data set;train a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; obtain an experimental image; and use the trained deep learning model to predict whether a result is present in the experimental image.
179. The system of any of claims 90 to 178, further comprising: a display device operably connected to the processor and memory and configured to display one or more of: aspects of the image-based training data set, one or more synthetic images, the enhanced image-based training data set or one or more result of applying the deep learning model.
180. The system of any of claims 90 to 179, further comprising: an input device operably connected to the processor and memory and configured to receive input from a user.181 . The system of claim 180, wherein the input device is configured to receive a selection of one or more images of the image-based training data set.
182. A non-transitory computer readable storage medium comprising instructions stored thereon, the instructions comprising: algorithm for obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; algorithm for selecting a first image from the image-based training data based at least in part on a first characteristic of the first image; algorithm for generating a first synthetic image based on the first image by modifying a first aspect of the first image; and algorithm for combining the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models,wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image.
183. The non-transitory computer readable storage medium of claim 182, wherein generating a first synthetic image based on the first image comprises modifying a representation of a structure present in the first image.
184. The non-transitory computer readable storage medium of any of claims 182 to 183, wherein generating a first synthetic image based on the first image comprises introducing a variation in a structure present in the first image.
185. The non-transitory computer readable storage medium of any of claims 182 to 184, wherein generating a first synthetic image based on the first image comprises introducing random variations into the first image.
186. The non-transitory computer readable storage medium of any of claims182 to 185, further comprising: generating a plurality of synthetic images based on the first image by one or more modifications to one or more aspects of the first image; and combining the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models.
187. The non-transitory computer readable storage medium of any of claims182 to 186, further comprising: generating a plurality of synthetic images, wherein the synthetic images are generated based on more than one of the plurality of images; and combining the image-based training data set and the plurality of synthetic images to form the enhanced image-based training data set for deep learning models.
188. The non-transitory computer readable storage medium of any of claims182 to 187, wherein generating a first synthetic image based on the first image comprises introducing a variation in the first image other than mirroring, rotating, smoothing, or contrast reduction of the first image.
189. The non-transitory computer readable storage medium of any of claims 182 to 188, wherein generating a first synthetic image based on the first image comprises: applying an autoencoder or a diffusion-based model or a generative adversarial network (GAN) to the first image to generate the first synthetic image.
190. The non-transitory computer readable storage medium of any of claims 182 to 189, wherein generating a first synthetic image based on the first image comprises applying an autoencoder to the first image to generate the first synthetic image, wherein the autoencoder is a trained autoencoder.191 . The non-transitory computer readable storage medium of claim 190, wherein the autoencoder is trained on a second image-based training data set, wherein the second image-based training data set does not comprise any images of the image-based training data set.
192. The non-transitory computer readable storage medium of any of claims 190 to 191 , wherein applying an autoencoder to the first image to generate the first synthetic image comprises: introducing a controlled variation of the first image into the first synthetic image.
193. The non-transitory computer readable storage medium of any of claims190 to 192, wherein applying an autoencoder to the first image to generate the first synthetic image comprises: encoding the first image into an embedding; anddecoding the embedding into the first synthetic image, wherein the embedding comprises a representation of the first image.
194. The non-transitory computer readable storage medium of claim 193, wherein the embedding comprises a compressed representation of the first image.
195. The non-transitory computer readable storage medium of any of claims 193 to 194, wherein the embedding comprises a compressed stochastic representation of the first image.
196. The non-transitory computer readable storage medium of any of any of claims 193 to 195, further comprising: algorithm for introducing noise into the embedding.
197. The non-transitory computer readable storage medium of claim 196, wherein the noise introduced into the embedding comprises one or more of uniform noise or step noise or random noise or patterned noise, wherein patterned noise is associated with a mathematical function, optionally comprising one or more of a sine or cosine function.
198. The non-transitory computer readable storage medium of any of claims193 to 197, wherein the embedding comprises a linear representation of the first image.
199. The non-transitory computer readable storage medium of any of any of claims 193 to 198, wherein the noise introduced into the embedding comprises a linear representation of noise values.
200. The non-transitory computer readable storage medium of any of claims 189 to 199, wherein the autoencoder comprises an encoder and a decoder.201 . The non-transitory computer readable storage medium of any of claims 177 to 188, wherein the autoencoder comprises a dual-level structure.
202. The non-transitory computer readable storage medium of any of claims 189 to 200, wherein the autoencoder comprises: an encoder, and a decoder, wherein the encoder comprises a first level encoder configured to encode a subset of the first image, and a second level encoder configured to encode an output of the first level encoder, and the decoder comprises a first level decoder, configured to decode a subset of the first synthetic image, and a second level decoder configured to decode an output of the first level decoder.
203. The non-transitory computer readable storage medium of claim 202, wherein the encoder and decoder are modular components.
204. The non-transitory computer readable storage medium of any of claims 202 to 203, wherein the encoder and decoder are independently trained.
205. The non-transitory computer readable storage medium of any of claims 202 to 204, wherein the encoder and decoder comprise one or more of neural networks or residual networks.
206. The non-transitory computer readable storage medium of any of claims 189 to 205, wherein each level of the encoder and decoder are independently trained.
207. The non-transitory computer readable storage medium of any of claims 202 to 206, wherein each level of the encoder and decoder comprises a neural network.
208. A non-transitory computer readable storage medium comprising instructions stored thereon, the instructions comprising: algorithm for obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; algorithm for selecting first and second images from the image-based training data, wherein the first and second images each exhibit a common first characteristic; algorithm for generating a first synthetic image based on the first and second images by combining the first image with the second image; and algorithm for combining the image-based training data set and the first synthetic image into an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the common characteristic of the first and second images.
209. The non-transitory computer readable storage medium of claim 208, wherein generating a first synthetic image based on the first and second images comprises combining a representation of the first and second images.
210. The non-transitory computer readable storage medium of any of claims 208 to 209, wherein combining a representation of the first and second images comprises a scaled linear combination of the representations of the first and second images.21 1 . The non-transitory computer readable storage medium of any of claims 208 to 210, wherein generating a first synthetic image based on the first and second images further comprises generating the first synthetic image based on comprises combining the first image, the second image and one or more images from the from the image-based training data.
212. The non-transitory computer readable storage medium of any of claims 208 to 21 1 , wherein generating a first synthetic image based on the first and second images by combining the first image with the second image comprises: algorithm for applying an autoencoder or a diffusion-based model or a generative adversarial network (GAN) to the first and second images to generate the first synthetic image.
213. The non-transitory computer readable storage medium of any of any of claims 208 to 212, wherein generating a first synthetic image based on the first and second images by combining the first image with the second image comprises applying an autoencoder to the first and second images to generate the first synthetic image, wherein the autoencoder is a trained autoencoder.
214. The non-transitory computer readable storage medium of any of claims 208 to 213, wherein the autoencoder is trained on a second image-based training data set, wherein the second image-based training data set does not comprise any images of the image-based training data set.
215. The non-transitory computer readable storage medium of any of claims 208 to 214, wherein algorithm for applying an autoencoder to the first and second images to generate the first synthetic image comprises: algorithm for encoding the first image into a first embedding; algorithm for encoding the second image into a second embedding; algorithm for combining the first embedding with the second embedding to form a first synthetic embedding; andalgorithm for decoding the first synthetic embedding into the first synthetic image, wherein the first and second embeddings comprise representations of the first and second images, respectively.
216. The non-transitory computer readable storage medium of claim 215, wherein the first and second embeddings comprise compressed representations of the first and second images, respectively.
217. The non-transitory computer readable storage medium of any of claims 215 to 216, wherein the first and second embeddings comprise compressed stochastic representations of the first and second images, respectively.
218. The non-transitory computer readable storage medium of any of claims 215 to 217, wherein the first embedding comprises a first linear representation, and the second embedding comprises a second linear representation.
219. The non-transitory computer readable storage medium of any of any of claims 215 to 218, wherein combining the first image with the second image comprises combining the first linear representation with the second linear representation in a scaled linear combination.
220. The non-transitory computer readable storage medium of any of claims 212 to 219, wherein the autoencoder comprises an encoder and a decoder.221 . The non-transitory computer readable storage medium of any of claims 212 to 220, wherein the autoencoder comprises a dual-level structure.
222. The non-transitory computer readable storage medium of any of claims 212 to 221 , wherein the autoencoder comprises: an encoder, anda decoder, wherein the encoder comprises a first level encoder configured to encode a subset of the first and second images, and a second level encoder configured to encode outputs of the first level encoder, and the decoder comprises a first level decoder, configured to decode a subset of the first synthetic image, and a second level decoder configured to decode an output of the first level decoder.
223. The non-transitory computer readable storage medium of claim 222, wherein the encoder and decoder are modular components.
224. The non-transitory computer readable storage medium of any of claims 222 to 223, wherein the encoder and decoder are independently trained.
225. The non-transitory computer readable storage medium of any of claims 222 to 224, wherein the encoder and decoder comprise one or more of neural networks or residual networks.
226. The non-transitory computer readable storage medium of any of claims 222 to 225, wherein each level of the encoder and decoder are independently trained.
227. The non-transitory computer readable storage medium of any of claims 222 to 226, wherein each level of the encoder and decoder comprises a neural network.
228. The non-transitory computer readable storage medium of any of claims182 to 227, wherein the plurality of images of the image-based training data set comprises subject images.
229. The non-transitory computer readable storage medium of any of claims182 to 228, wherein the plurality of images of the image-based training data set comprises images of biologic tissues.
230. The non-transitory computer readable storage medium of any of claims182 to 229, wherein the plurality of images of the image-based training data set corresponds to images of a plurality of subjects.231 . The non-transitory computer readable storage medium of any of claims 182 to 230, wherein the image-based training data set comprises images generated by one or more of the following modalities: photographic images, MRI images, X-ray images, CT images or ultrasound images.
232. The non-transitory computer readable storage medium of any of claims182 to 231 , wherein at least some of the plurality of images of the image-based training data set exhibit one or more of the following: glaucoma, cancer, a degenerative retinal diseases, an acquired retinal disease, diabetic retinopathy, aneurysm, multiple sclerosis, a spinal cord pathology, stroke, tumor, traumatic brain injury, joint injury or disease, heart disease, vascular disease, a liver disease, Alzheimer’s disease, epilepsy, peripheral nerve compression, renal disease, inflammatory bowel disease, a fracture, arthritis, a lung disease, tuberculosis, pneumonia, chronic obstructive pulmonary disease (COPD), emphysema, pulmonary edema, pneumothorax, a dental disease, a bone infection, scoliosis, osteoporosis, kidney disease, kidney stone, a presence of a foreign object, pulmonary embolism, trauma, stroke, an abdominal disorder, appendicitis, diverticulitis, an ovarian cyst, sinusitis, a thyroid gland disease, pregnancy, gallstone, ascites, pelvic organ prolapse or testicular torsion.
233. The non-transitory computer readable storage medium of any of claims 182 to 232, wherein the image-based training data set is a training data set for training a deep learning model.
234. The non-transitory computer readable storage medium of any of claims182 to 233, wherein the image-based training data set is configured for training a deep learning model to perform one or more of: image classification, object detection or semantic segmentation.
235. The non-transitory computer readable storage medium of any of claims182 to 234, wherein the image-based training data set is configured for training a deep learning model to diagnose or screen for a disease or monitor a disease state or disease progression.
236. The non-transitory computer readable storage medium of any of claims182 to 235, wherein the image-based training data set is a training data set for a convolutional neural network (CNN) or a vision transformer (ViT).
237. The non-transitory computer readable storage medium of any of claims 182 to 236, wherein the image-based training data set is a training data set for image-based disease diagnosis or disease monitoring.
238. The non-transitory computer readable storage medium of any of claims182 to 237, wherein the image-based training data set comprises a first subset of images that exhibit a condition or a disease and a second subset of images that do not exhibit the condition of the disease.
239. The non-transitory computer readable storage medium of any of claims 182 to 238, wherein the image-based training data set comprises a bias optionally comprising one or more of a bias related to an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location.
240. The non-transitory computer readable storage medium of any of claims 182 to 239, wherein the first characteristic optionally comprises one or more of: a diagnosis or a disease state or an injury state or a condition of an image subject.241 . The non-transitory computer readable storage medium of any of claims 182 to 240, wherein the first characteristic is a demographic characteristic of the subject optionally comprising one or more of: an age or a gender or a race or an ethnicity or a culture or a socioeconomic characteristic or a geographic location of an image subject.
242. The non-transitory computer readable storage medium of any of claims 182 to 241 , wherein the first characteristic is relatively rare in the image-based training data set.
243. The non-transitory computer readable storage medium of any of claims 182 to 242, wherein the first characteristic is underrepresented in the imagebased training data set.
244. The non-transitory computer readable storage medium of any of claims182 to 243, wherein it is relatively rare for a subject in the image-based training data set to exhibit the first characteristic.
245. The non-transitory computer readable storage medium of any of claims 182 to 244, wherein the first characteristic depicts an image aspect of a rare disease.
246. The non-transitory computer readable storage medium of any of claims 182 to 245, wherein the first or seconds images correspond to an image of an underrepresented group.
247. The non-transitory computer readable storage medium of any of claims 182 to 246, wherein an image subject is a member of an underrepresented group.
248. The non-transitory computer readable storage medium of any of claims 182 to 247, wherein the first characteristic is selected to address a bias in the image-based training data set.
249. The non-transitory computer readable storage medium of any of claims 182 to 248, wherein the first characteristic is selected to enhance performance of the deep learning model.
250. The non-transitory computer readable storage medium of any of claims 182 to 249, wherein the first characteristic is selected to enhance performance of the deep learning model across a more diverse range of subjects.251 . The non-transitory computer readable storage medium of any of claims 182 to 250, wherein the image-based training data does not reflect a demographic characteristic of a select population, and the enhanced image-based training data set is generated to reflect the demographic characteristic of the select population.
252. The non-transitory computer readable storage medium of any of claims 182 to 251 , wherein the enhanced image-based training data set comprises at least one synthetic image corresponding to each image of a plurality of images of the image-based training set.
253. The non-transitory computer readable storage medium of any of claims 182 to 252, wherein the enhanced image-based training data set comprises atleast one synthetic image corresponding to each image of the image-based training set.
254. The non-transitory computer readable storage medium of any of claims182 to 253, wherein the enhanced image-based training data set comprises three synthetic images corresponding to each image of the image-based training set.
255. The non-transitory computer readable storage medium of any of claims182 to 254, wherein the image-based training data set reflects a first population of subjects, and wherein the method is a method of generating an enhanced image-based training data set for training a deep learning model for use with individuals that are not adequately represented in the image-based training data set.
256. The non-transitory computer readable storage medium of any of claims182 to 255, further comprising: algorithm for training a deep learning model using the enhanced imagebased training data set.
257. The non-transitory computer readable storage medium of any of claims182 to 256, further comprising: algorithm for training a deep learning model for image-based disease diagnosis or disease monitoring using the enhanced image-based training data set.
258. The non-transitory computer readable storage medium of any of claims 182 to 258, wherein the non-transitory computer readable storage medium is configured for diversifying the image-based training data set.
259. The non-transitory computer readable storage medium of any of claims 182 to 258, wherein the non-transitory computer readable storage medium is configured for addressing limitations of the image-based training data set.
260. The non-transitory computer readable storage medium of any of claims 182 to 259, wherein the non-transitory computer readable storage medium is configured for improving the accuracy of predictions of a deep learning model.261 . The non-transitory computer readable storage medium of any of claims 182 to 260, wherein the non-transitory computer readable storage medium is configured for addressing underrepresented samples in the image-based training data set.
262. The non-transitory computer readable storage medium of any of claims 182 to 261 , wherein the non-transitory computer readable storage medium is configured for enhancing underrepresented samples in the image-based training data set.
263. The non-transitory computer readable storage medium of any of claims 182 to 262, wherein the non-transitory computer readable storage medium is configured for synthetically augmenting underrepresented samples in the imagebased training data set.
264. The non-transitory computer readable storage medium of any of claims 182 to 263, wherein the non-transitory computer readable storage medium is configured for mitigating an imbalance in the image-based training data set.
265. A non-transitory computer readable storage medium for training a deep learning model to detect a result using an enhanced training data set, the method comprising:algorithm for obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; algorithm for applying an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on an image of the first subset of images of the image-based training data set; algorithm for training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; algorithm for obtaining an experimental image; and algorithm for using the trained deep learning model to predict whether a result is present in the experimental image.
266. A non-transitory computer readable storage medium for training a deep learning model to detect a result using an enhanced training data set, the method comprising: algorithm for obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; algorithm for selecting a first subset of images from the image-based training data based at least in part on a first characteristic of the first subset of images, wherein the first characteristic is associated with a data-collection bias of the image-based training data set; algorithm for generating synthetic images based on the images of the first subset of images by modifying aspects of each image of the first subset of images such that the synthetic images comprises variations of the first characteristic; algorithm for training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images;algorithm for obtaining an experimental image; and algorithm for using the trained deep learning model to predict whether a result is present in the experimental image.
267. A non-transitory computer readable storage medium for generating an enhanced image-based training data set for deep learning models, the method comprising: algorithm for obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images, and wherein the plurality of images exhibit a first characteristic at a first frequency; algorithm for selecting a first subset of images from the image-based training data based at least in part on a presence of the first characteristic in each image of the first subset of images; algorithm for generating a first number of synthetic images based on the first subset of images by modifying aspects of one or more images of the first subset of images; algorithm for combining the image-based training data set and the first number of synthetic images to generate an enhanced image-based training data set for deep learning models, wherein the enhanced image-based training data set comprises additional training data associated with the first characteristic of the first image such that the images of the enhanced image-based training data set exhibit the first characteristic at a second frequency.
268. The non-transitory computer readable storage medium of claim 267, wherein the image-based training data set corresponds to a first subject population exhibiting the first characteristic at the first frequency, and the enhanced image-based training data set corresponds to a second subject population exhibiting the first characteristic at the second frequency.
269. The non-transitory computer readable storage medium of any of claims 267 to 268, wherein the first frequency is a biased frequency, and the second frequency is an unbiased frequency.
270. A non-transitory computer readable storage medium for training a deep learning model to detect a result using an enhanced training data set, the method comprising: algorithm for obtaining an image-based training data set for a deep learning model, wherein the image-based training data set comprises a plurality of images; algorithm for applying an autoencoder to a first subset of images of the image-based training data to generate a plurality of synthetic images, wherein each synthetic image is based on a combination of at least two images of the first subset of images of the image-based training data set; algorithm for training a deep learning model to detect a result using images from the image-based training data set and the plurality of synthetic images; algorithm for obtaining an experimental image; and algorithm for using the trained deep learning model to predict whether a result is present in the experimental image.
Citation Information
Patent Citations
Systems and methods to process electronic images for synthetic image generation
US11393574B1
Disentangled representation learning generative adversarial network for pose-invariant face recognition
US11734955B2
Modeling continuous kernels to generate an enhanced digital image from a burst of digital images
US20230237628A1
Multimodality image processing techniques for training image data generation and usage thereof for developing MONO-modality image inferencing models
US20230342427A1