Method and system for selecting embryos
A computationally generated AI model for embryo viability scoring in IVF addresses the subjectivity and cost issues of existing methods, enhancing pregnancy success rates through standardized embryo selection.
Patent Information
- Application Number
- JP2021560476
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-04
- Filing Date
- 2020-04-02
- Publication Date
- 2025-08-27
- Estimated Expiration
- 2040-04-02
AI Technical Summary
The current embryo selection process in IVF is highly subjective and varies significantly among embryologists, leading to low pregnancy success rates, with existing tools like PGS and time-lapse imaging being costly or unreliable.
A computationally generated AI model that estimates embryo viability by training deep learning models on zona pellucida and IZC images, using ensemble methods to combine multiple AI models for accurate embryo scoring.
The AI model provides a standardized and reliable embryo selection method, improving pregnancy success rates by reducing subjectivity and cost barriers of existing technologies.
Smart Images

Figure 0007730152000024 
Figure 0007730152000025 
Figure 0007730152000026
Abstract
Description
[Technical Field]
[0001] This application claims priority to Australian Provisional Patent Application No. 2019901152, entitled "METHOD AND SYSTEM FOR SELECTING EMBRYOS" and filed on April 4, 2019, the contents of which are incorporated herein by reference.
[0002] The present disclosure relates to in vitro fertilization (IVF). In certain aspects, the present disclosure relates to methods of selecting embryos. [Background technology]
[0003] The in vitro fertilization (IVF) process begins with an ovarian stimulation phase, which stimulates egg production. Next, eggs (oocytes) are retrieved from the patient and fertilized in vitro with sperm that penetrate the zona pellucida, a glycoprotein layer surrounding the egg (oocyte), to form a zygote. The embryo develops for approximately five days, after which it forms a blastocyst (formed from the trophoblast, blastocoel, and inner cell mass) suitable for transfer back to the patient. At approximately day five, the blastocyst is still surrounded by the zona pellucida, from which it hatches and then implants into the endometrial wall. Here, the area bounded by the inner surface of the zona pellucida is called the InnerZonal Cavity (IZC). Selection of the best embryo at the time of transfer is crucial to ensuring a positive pregnancy outcome. An embryologist visually evaluates the embryos using a microscope to make this selection. Some clinics record images of the embryos at selected time points, allowing an embryologist to score each embryo based on various metrics and the embryologist's visual evaluation under a microscope. For example, one commonly used scoring system is the Gardner Scale, in which morphological features such as inner cell mass quality, trophectoderm quality, and embryonic developmental progress are assessed and graded according to an alphanumeric scale. The embryologist then selects one (or more) of the embryos, which are then returned to the patient.
[0004] Thus, embryo selection is currently a manual process involving subjective evaluation of embryos by embryologists via visual inspection. One of the major challenges in embryo staging is the high level of subjectivity and intra- and inter-operator variability that exists among embryologists with different skill levels. This means that standardization is difficult even within a single laboratory and impossible across the industry. Thus, the process relies heavily on the expertise of embryologists, and despite their best efforts, IVF success rates remain relatively low (approximately 20%). While the reasons for low pregnancy outcomes are complex, tools for more accurately selecting the most viable embryos are expected to increase the success rate of pregnancy outcomes.
[0005] To date, several tools have been developed to assist embryologists in selecting viable embryos, including preimplantation genetic screening (PGS) or time-lapse photography. However, each approach has significant limitations. PGS involves genetically evaluating several cells from the embryo by performing a biopsy and then screening the extracted cells. While this can be useful for identifying genetic risks that may lead to pregnancy failure, it may also harm the embryo during the biopsy process. It is also expensive and has limited or no availability in many large developing markets, such as China. Another tool that has been considered is the use of time-lapse imaging of the embryonic developmental process. However, this requires expensive and specialized hardware that is cost-prohibitive for many clinics. Furthermore, there is no evidence that it can reliably improve embryo selection. At best, it may help determine whether early-stage embryos will develop to a mature blastocyst, but it has not been demonstrated to reliably predict pregnancy outcome, thus limiting its usefulness for embryo selection.
[0006] Thus, there is a need to provide improved tools to assist embryologists in selecting embryos for implantation, or at least to provide useful alternatives to existing tools and systems. Summary of the Invention
[0007] According to a first aspect, there is provided a method of computationally generating an artificial intelligence (AI) model configured to estimate an embryo viability score from an image, the method comprising: receiving a plurality of images and associated metadata, each image captured during a predetermined time window after in vitro fertilization (IVF), the predetermined time window being within 24 hours, and the metadata associated with the images including at least a pregnancy outcome label; pre-processing each image, including at least segmenting the image to identify zona pellucida regions; generating an artificial intelligence (AI) model configured to generate an embryo viability score from input images by training at least one Zona Deep Learning Model using a deep learning method, the deep learning model comprising training the deep learning model on a set of Zona pellucida images in which zona pellucida regions have been identified, and an associated pregnancy outcome label is used at least to assess the accuracy of the trained model; and Deploying the AI model; Includes:
[0008] In a further embodiment, the set of zona pellucida images includes images in which the areas bounded by the zona pellucida regions are masked.
[0009] In a further aspect, the step of generating the AI model further includes training one or more additional AI models to generate the AI model, each additional AI model being either a computer vision model trained using a machine learning method that uses a combination of one or more computer vision descriptors extracted from images to estimate an embryo viability score, a deep learning model trained on images localized to an embryo including both the zona pellucida region and the IZC region, and a deep learning model trained on a set of intrazona pellucida (IZC) images in which all regions except the IZC are masked; and using an ensemble method to combine at least two of the at least one zona pellucida deep learning model and the one or more additional AI models to generate an AI model embryo viability score from the input image, or using a distillation method to train an AI model using the at least one zona pellucida deep learning model and the one or more additional AI models to generate an AI model embryo viability score.
[0010] In one form, the AI model is generated using an ensemble model that includes selecting at least two contrasting AI models from at least one zona pellucida deep learning model and one or more further AI models, where the selection of the AI models is performed to generate a set of contrasting AI models, and applying a voting strategy to the at least two contrasting AI models that determines how the selected at least two contrasting AI models are combined to generate a result score for the image.
[0011] In a further form, selecting at least two contrasting AI models includes: generating a distribution of embryo viability scores from the set of images for each of the at least one zona pellucida deep learning model and one or more additional AI models; comparing the distributions and discarding models when the associated distribution is too similar to another distribution to select an AI model with a contrasting distribution; Includes:
[0012] In one embodiment, the predetermined time window is a 24-hour timer period beginning on day 5 post-fertilization. In one embodiment, the pregnancy outcome label is a ground truth pregnancy outcome measurement taken within 12 weeks post-embryo transfer. In a further embodiment, the ground truth pregnancy outcome measurement is whether a fetal heartbeat is detected.
[0013] In one form, the method further includes cleaning the plurality of images, where cleaning the plurality of images includes identifying images with likely incorrect pregnancy outcome labels and removing or re-labeling the identified images.
[0014] In a further aspect, the step of cleaning the plurality of images includes estimating the likelihood that the pregnancy outcome label associated with the image is inaccurate, comparing to a threshold, and then rejecting or re-labeling images with a likelihood above the threshold.
[0015] In a further aspect, estimating the likelihood that the pregnancy outcome label associated with the image is inaccurate is performed using multiple AI classification models and k-fold cross-validation, in which the multiple images are divided into k mutually exclusive validation data sets, each of the multiple AI classification models is trained on the combined k-1 validation data sets and then used to classify images in the remaining validation data sets, and the likelihood is determined based on the number of AI classification models that incorrectly classify the pregnancy outcome label of the image.
[0016] In one form, training each AI model or generating the ensemble model includes evaluating the performance of the AI models using multiple metrics including at least one accuracy metric and at least one confidence metric, or one metric that combines accuracy and confidence.
[0017] In one form, the step of pre-processing the image further comprises cropping the image by using deep learning or computer vision methods to localize the embryo within the image.
[0018] In one embodiment, the step of preprocessing the image further includes one or more of padding the image, normalizing the color balance, normalizing the brightness, and scaling the image to a predetermined resolution.
[0019] In one embodiment, padding the image may be performed to create a square aspect ratio for the image. In one embodiment, the method further includes generating one or more augmented images for use in training the AI model. Preparing each image may also include generating the one or more augmented images by making a copy of the image with the changes, or the augmentation may be performed on the image prior to or during training (on the fly). Any number of enhancements may be performed, including rotating the image by 90 degrees, mirror flipping, non-90 degree rotation where a diagonal border is embedded to match the background color, varying the amount of image blur, adjusting the image contrast using an intensity histogram, and applying one or more small random transformations in both horizontal and / or vertical directions: random rotation, JPEG noise, random image resizing, random hue jitter, random brightness jitter, contrast-limited adaptive histogram equalization, random flip / mirror, image sharpening, image embossing, random brightness and contrast, RGB color shift, random hue and saturation, channel shuffle, swapping RGB to BGR or RBG or other, coarse dropout, motion blur, center blur, Gaussian blur, random shift scale rotation (i.e., all three combined).
[0020] In one form, during training of the AI model, one or more augmented images are generated for each image in the training set, and during evaluation of the validation set, the results for the one or more augmented images are combined to generate a single result for the image. The results may be combined using one of the average confidence, median confidence, majority-mean confidence, maximum confidence methods, or other voting strategies for combining model predictions.
[0021] In one embodiment, the step of preprocessing the image may further include annotating the image using one or more feature descriptor models and masking all regions of the image except for regions within a given radius of the descriptor's keypoints. The one or more feature descriptor models may include Gray-Level Co-occurrence Matrix (GLCM) texture analysis, Histogram of Oriented Gradients (HOG), Feature Extraction by Oriented Accelerated Fragment Test (FAST) and Rotational Binary Robust Independent Basic Feature (BRIEF), Binary Robust Invariant Scalable Keypoint (BRISK), Maximum Stable Extremal Region (MSER), or Tracking-Oriented Feature (GFTT) feature detectors.
[0022] In one embodiment, each AI model generates an outcome score, the outcome being an n-ary outcome having n states, and training the AI model includes multiple training / validation cycles, randomly assigning a plurality of images to one of a training set, a validation set, or a blind validation set such that the training dataset includes at least 60% of the images, the validation dataset includes at least 10% of the images, and the blind validation dataset includes at least 10% of the images; after assigning the images to the training set, validation set, and blind validation set, calculating a frequency of each of the n-ary outcome states in each of the training set, validation set, and blind validation set, testing that the frequencies are similar, and discarding the assignments if the frequencies are not similar; and repeating the randomization until randomizations with similar frequencies are obtained.
[0023] In one embodiment, training the computer vision model includes performing multiple training and validation cycles, during each cycle, images are clustered based on computer vision descriptors using an unsupervised clustering algorithm to generate a set of clusters, each image is assigned to a cluster using a distance measure based on the values of the image's computer vision descriptors, and a supervised learning method is used to determine whether a particular combination of these features corresponds to an outcome measure and frequency information for the presence of each computer vision descriptor in the multiple images.
[0024] In one form, the deep learning models may be convolutional neural networks (CNNs), where for an input image, each deep learning model generates an outcome probability.
[0025] In one embodiment, deep learning methods can emphasize a global minimum using a loss function configured to modify the optimization surface. The loss function may include a residual term defined in terms of the network weights, which encodes the collective difference between the predictions from the model and the target outcome for each image and includes it as an additional contribution to a normal cross-entropy loss function.
[0026] In one aspect, the method may be performed on a cloud-based computing system using a web server, a database, and multiple training servers, where the web server receives one or more model training parameters from a user, where the web server initiates a training process on one or more of the multiple training servers including uploading training code to one of the multiple training servers, where the training server requests a plurality of images and associated metadata from a data repository and performs steps of preparing each image, generating a plurality of computer vision models, and generating a plurality of deep learning models, where each training server is configured to periodically save the models to a storage service and accuracy information to one or more log files to allow the training process to be restarted. In a further aspect, the ensemble models may be trained to bias residual inaccuracies to minimize false negatives.
[0027] In one embodiment, the outcome is a binary outcome of viable or non-viable, and randomization may include calculating the frequency of images with a classification of viable and non-viable in each of the training set, validation set, and blind validation set, and testing whether they are similar. In one embodiment, the outcome measure is a measure of embryo viability using the viability classification associated with each image. In one embodiment, each outcome probability may be the probability that the image is viable. In one embodiment, each image may be a phase-contrast image.
[0028] According to a second aspect, there is provided a method of computationally generating an embryo viability score from an image, the method comprising: generating, in a computing system, an artificial intelligence (AI) model configured to generate an embryo viability score from the image according to the method of the first aspect; receiving images captured during a predetermined time window following in vitro fertilization (IVF) from a user via a user interface of the computing system; preprocessing the image according to a preprocessing step used to generate the AI model; providing the preprocessed images to an AI model to obtain an estimate of the embryo viability score; transmitting the embryo viability score to the user via the user interface; Includes.
[0029] According to a third aspect, there is provided a method of obtaining an embryo viability score from an image, the method comprising: uploading, via a user interface, images captured during a predetermined time window after in vitro fertilization (IVF) to a cloud-based artificial intelligence (AI) model configured to generate an embryo viability score from the images, the AI model being generated according to the method of the first aspect; receiving, via a user interface, an embryo viability score from the cloud-based AI model; Includes.
[0030] According to a fourth aspect, there is provided a cloud-based computing system configured to computationally generate an artificial intelligence (AI) model configured to estimate an embryo viability score from an image according to the method of the first aspect.
[0031] According to a fifth aspect, there is provided a cloud-based computing system configured to computationally generate an embryo viability score from an image, the computing system comprising: an artificial intelligence (AI) model configured to generate an embryo viability score from the image, generated according to the method of the first aspect; receiving, from a user via a user interface of the computing system, images captured during a predetermined time window following in vitro fertilization (IVF); Providing images to the AI model to obtain an embryo viability score; and transmitting an embryo viability score to a user via a user interface; Includes.
[0032] According to a sixth aspect, there is provided a computing system configured to generate an embryo viability score from an image, the computing system comprising at least one processor and at least one memory, the at least one memory comprising: receiving images captured during a predetermined time window after in vitro fertilization (IVF); uploading, via a user interface, images captured during a predetermined time window following in vitro fertilization (IVF) to a cloud-based artificial intelligence (AI) model configured to generate an embryo viability score from the images, the AI model being generated according to the method of the first aspect; receiving embryo viability scores from a cloud-based AI model; Displaying embryo viability scores via the user interface; The method includes instructions for configuring at least one processor to: [Brief explanation of the drawings]
[0033] Embodiments of the present disclosure will be discussed with reference to the accompanying drawings. [Figure 1A] 1 is a schematic flow diagram of generating an artificial intelligence (AI) model configured to estimate an embryo viability score from an image, according to one embodiment. [Figure 1B] FIG. 1 is a schematic block diagram of a cloud-based computing system configured to computationally generate and use an AI model configured to estimate an embryo viability score from an image, according to one embodiment. [Figure 2] FIG. 1 is a schematic diagram of an IVF method using an AI model configured to estimate an embryo viability score from images to aid in embryo selection for implantation, according to one embodiment. [Figure 3A] FIG. 1 is a schematic architectural diagram of a cloud-based computing system configured to generate and use an AI model configured to estimate an embryo viability score from an image, according to one embodiment. [Figure 3B] 1 is a schematic flow diagram of a model training process on a training server, according to one embodiment. [Figure 4] FIG. 1 is a schematic diagram of a binary thresholding process for boundary finding on images of human embryos, according to one embodiment. [Figure 5] 1 is a schematic diagram of a method for finding boundaries on an image of a human embryo, according to one embodiment. [Figure 6A] 1 is an example of the use of a geometric active contour (GAC) model applied to fixed regions of an image for image segmentation, according to one embodiment. [Figure 6B] 1 is an example of the use of morphological snakes applied to fixed regions of an image for image segmentation, according to one embodiment. [Figure 6C] FIG. 1 is a schematic architectural diagram of a U-Net architecture for a semantic segmentation model, according to one embodiment. [Figure 6D] FIG. 1 shows images of day 5 embryos. [Figure 6E] FIG. 6E shows a padded version of FIG. 6D that creates a square image. [Figure 6F] FIG. 6E shows a zona pellucida image based on FIG. 6E with the IZC masked, according to one embodiment. [Figure 6G] FIG. 6C shows an IZC image based on FIG. 6E with the zone pellucida and background masked, according to one embodiment. [Figure 7] FIG. 10 is a plot of a gray level co-occurrence matrix (GLCM) showing GLCM correlations of sample feature descriptors: ASM, homogeneity, correlation, contrast, and entropy, calculated for a set of six zona pellucida regions and six cytoplasmic regions, according to a related embodiment. [Figure 8] FIG. 1 is a schematic architecture diagram of a deep learning method including convolutional layers that transform input images into predictions after training, according to one embodiment. [Figure 9] FIG. 10 shows a plot of the accuracy of an embodiment of an ensemble model in determining embryo viability, according to one embodiment. [Figure 10]FIG. 1 shows a bar graph illustrating the accuracy of one embodiment of an ensemble model compared to world-leading embryologists (clinicians) in correctly identifying embryo viability. [Figure 11] FIG. 10 shows a bar graph illustrating the accuracy of one embodiment of an ensemble model compared to a world-leading embryologist (clinician) in accurately identifying embryo viability when the ensemble model's assessment was inaccurate, compared to an embryologist accurately identifying embryo viability when the ensemble model's assessment was inaccurate. [Figure 12] FIG. 1 shows a plot of the distribution of inference scores for viable embryos (successful clinical pregnancies) using one embodiment of an ensemble model when applied to the blind validation dataset from Study 1. [Figure 13] FIG. 1 shows a plot of the distribution of inference scores for non-viable embryos (unsuccessful clinical pregnancies) using one embodiment of the ensemble model when applied to the blind validation dataset from Study 1. [Figure 14] 10 is a histogram of ranks obtained from embryologist scores across the entire blind dataset. [Figure 15] 10 is a histogram of ranks obtained from one embodiment of ensemble model inference across the entire blind dataset. [Figure 16] Histogram of ensemble model inferences before being arranged into rank banding from 1 to 5. [Figure 17] FIG. 10 shows a plot of the distribution of inference scores for viable embryos (successful clinical pregnancies) using the ensemble model when applied to the blind validation dataset from Study 2. [Figure 18] FIG. 10 shows a plot of the distribution of inference scores for non-viable embryos (unsuccessful clinical pregnancies) using the ensemble model when applied to the blind validation dataset from Study 2. [Figure 19]FIG. 10 shows a plot of the distribution of inference scores for viable embryos (successful clinical pregnancies) using the ensemble model when applied to the blind validation dataset from Study 3. [Figure 20] FIG. 10 shows a plot of the distribution of inference scores for non-viable embryos (successful clinical pregnancies) using the ensemble model when applied to the blind validation dataset from Study 3. DETAILED DESCRIPTION OF THE INVENTION
[0034] In the following description, like reference numerals designate like or corresponding parts throughout the figures.
[0035] 1A, 1B, and 2, an embodiment of a cloud-based computing system 1 configured to computationally generate and use an artificial intelligence (AI) model 100 configured to estimate an embryo viability score from a single image of an embryo will now be discussed. This AI model 100 is also referred to as an embryo viability assessment model. FIG. 1A is a schematic flow diagram of the generation of the AI model 100 using the cloud-based computing system 1, according to one embodiment. A plurality of images and associated metadata are received (or acquired) from one or more data sources 101. Each image is captured during a predetermined time window after in vitro fertilization (IVF), such as a 24-hour period beginning on day 5 post-fertilization. The images and metadata may be sourced from an IVF clinic or may be images captured using an optical microscope, including phase-contrast images. The metadata includes a pregnancy outcome label (e.g., a heartbeat detected at the first scan after IVF) and may include various other clinical and patient information.
[0036] The image is then preprocessed (102), which includes segmenting the image to identify zona pellucida regions in the image. Segmentation may also include identifying intrazona pellucida cavities (IZCs) surrounded by the zona pellucida regions. Image preprocessing may also include one or more (or all) of object detection, alpha channel removal, padding, cropping / localization, color balance normalization, brightness normalization, and / or scaling the image to a predetermined resolution, as discussed below. Image preprocessing may also include computing / determining computer vision feature descriptors from the image and performing one or more image augmentations or generating one or more augmented images.
[0037] At least one zona pellucida deep learning model is trained on a set of zona pellucida images (103) to generate an artificial intelligence (AI) model 100 configured to generate an embryo viability score from input images (104). The set of zona pellucida images are images in which the zona pellucida regions have been identified (e.g., during segmentation in step 102). In some embodiments, the set of zona pellucida images are images in which all regions of the images are masked except for the zona pellucida regions (i.e., so the deep learning model is trained only on information from / related to the zona pellucida regions). The pregnancy outcome label is used at least to evaluate the trained model (i.e., to assess accuracy / performance) and may also be used in model training (e.g., by a loss function to drive model optimization). Multiple zona pellucida deep learning models can be trained, and the best-performing model can be selected as the AI model 100.
[0038] In another embodiment, one or more additional AI models are trained on the preprocessed images (106). These may be additional deep learning models trained directly on the embryo images and / or on a set of IZC images in which all regions of the image except the IZC are masked, or computer vision (CV) models trained to combine computer vision features / descriptors generated in the preprocessing step (102) to generate embryo viability scores from the images. Each computer vision model estimates an embryo viability score for an embryo in the image using a combination of one or more computer vision descriptors extracted from the image, and the machine learning method performs multiple training and validation cycles to generate the CV model. Similarly, each deep learning model is trained with multiple training and validation cycles so that each deep learning model learns how to estimate embryo viability scores for embryos in the images. During training, images may be randomly assigned to each of a training set, a validation set, and a blind validation set, and each training and validation cycle includes (further) randomization of multiple images in each of the training set, the validation set, and the blind validation set. That is, the images in each set are randomly sampled for each cycle, so that for each cycle a different subset of images is analyzed or analyzed in a different order. Note, however, that because the images are randomly sampled, this allows for two or more sets to be identical, provided this occurs via a random selection process.
[0039] Next, in step 104, the multiple AI models are combined into a single AI model 100 (107) using ensemble, distillation, or other similar techniques to generate the AU model 100. The ensemble approach involves selecting a model from a set of available models and using a voting strategy that determines how a result score is generated from the individual results of the selected models. In some embodiments, the models are selected to ensure that results are contrasted to generate a distribution of results. They are preferably as independent as possible to ensure a good distribution of results. In distillation, multiple AI models are used as teachers to train a single student model, which becomes the final AI model 100.
[0040] In step 104, a final AI model is selected. This may be one of the zona pellucida deep learning models trained in step 103, or it may be a model obtained using an ensemble, distillation, or similar combination step (step 107), where the training involves at least one zona pellucida deep learning model (from 103) and one or more additional AI models (deep learning and / or CV; step 106). Once the final AI model 100 is generated (104), it is deployed for operational use, for example, on a cloud server configured to receive phase-contrast images of day-5 embryos captured at an IVF clinic using a light microscope, to estimate embryo viability scores from the input images (105). This is further illustrated in FIG. 2 and discussed below. In some embodiments, deployment includes saving or exporting the trained model, such as by writing the model weights and associated model metadata to a file that is transferred to a computing system and uploaded to reproduce the trained model. Deployment may also include moving, copying, or replicating the trained model onto a computing system, such as one or more cloud-based servers or a locally based computer server at an IVF clinic. In one embodiment, deployment may include reconfiguring the computing system on which the AI model was trained to accept new images and generate viability estimates using the trained model, for example, by adding an interface to receive the images, run the trained model on the received images, and send the results back to the source or store the results for later retrieval. The deployed system is configured to receive input images and perform any preprocessing steps used to generate the AI model (i.e., so that new images are preprocessed in the same way as the trained images). In some embodiments, images may be preprocessed (i.e., preprocessed locally) prior to uploading to a cloud system. In some embodiments, preprocessing may be distributed between a local system and a remote (e.g., cloud) system.The developed model is performed or run on the image to generate an embryo viability score, which is then provided to the user.
[0041] FIG. 1B is a schematic block diagram of a cloud-based computing system 1 configured to computationally generate an AI model 100 (i.e., an embryo viability assessment model) configured to estimate an embryo viability score from an image, and then use this AI model 100 to generate an embryo viability score (i.e., an outcome score) that is an estimate (or assessment) of the viability of the received image. Input 10 includes data such as embryo images and pregnancy outcome information (e.g., heartbeat detected on the first ultrasound scan after IVF, live or no live birth, or successful implantation) that can be used to generate a viability classification. This is provided as input to a model creation process 20 that creates and trains AI models. These include a zona pellucida deep learning model (103) and, in some embodiments, additional deep learning and / or computer vision models (106). The models may be trained using a variety of methods and information, including the use of segmented datasets (e.g., zona pellucida images, IZC images, etc.) and pregnancy outcome data. When multiple AI models are trained, the best-performing model may be selected according to some criteria, such as based on pregnancy outcome information, or multiple AI models may be combined using an ensemble model that selects an AI model and generates a result based on a voting strategy, or a distillation method may be used that uses multiple AI models as teachers to train a student AI model, or some other similar method may be used to combine multiple models into a single model. A cloud-based model management and monitoring tool called Model Monitor 21 is used to create (or generate) AI models. It uses a set of linked services, such as Amazon Web Services (AWS), that manages model training, logging, and tracking specific to image analysis. Other similar services on other cloud platforms may also be used. These can use deep learning methods 22, computer vision methods 23, classification methods 24, statistical methods 25, and physics-based models 26.Model generation can also use domain expertise 12 as input, such as from embryologists, computer scientists, scientific / technical literature, etc., on what features to extract and use in computer vision models. The output of the model creation process is an example of an AI model (100), also called a validated embryo assessment model.
[0042] A cloud-based delivery platform 30 is used, providing a user interface 42 to the system for a user 40. This is further illustrated with reference to FIG. 2, which is a schematic diagram of an IVF method 200 that uses a pre-trained AI model to generate an embryo viability score to assist in the selection of embryos for implantation, according to one embodiment. On day 0, retrieved eggs are fertilized (202). They are then cultured in vitro for several days, and then images of the embryos are captured (204), for example, using a phase-contrast microscope. As discussed below, it is generally found that images taken five days after IVF yield better results than images taken on earlier days. Thus, preferably, the model is trained and used on day 5 embryos, although it should be understood that the model may be trained and used on embryos obtained during a particular time window for a particular age. In one embodiment, the time is 24 hours, although other time windows, such as 12, 36, or 48 hours, may also be used. A shorter time window, generally 24 hours or less, is preferred to ensure greater similarity in appearance. In one embodiment, this can be a specific day, a 24-hour window from the beginning of the day (0:00) to the end of the day (23:39), or a specific day, such as day 4 or 5 (a 48-hour window starting from the beginning of day 4). Alternatively, the time window can define a window size and era, such as a 24-hour period centered around day 5 (i.e., day 4.5 to day 5.5). The time window can be variable, with a lower limit, such as at least 5 days. As mentioned above, it is preferred to use images of embryos from a 24-hour time window around day 5, although it should be understood that earlier stage embryos, including images from day 3 or day 4, can be used.
[0043] Typically, several eggs will be fertilized simultaneously, resulting in a set of multiple images to consider which embryos are most suitable for implantation (i.e., most viable). A user uploads the captured images to the platform 30 via a user interface 42, for example, using a "drag-and-drop" function. The user can upload a single image or multiple images, for example, to aid in the selection of which embryos from a set of multiple embryos to consider for implantation. The platform 30 receives one or more images (312) that are stored in a database 36 that includes an image repository. The cloud-based delivery platform includes on-demand cloud servers 32 that can perform image preprocessing (e.g., object detection, segmentation, padding, normalization, cropping, centering, etc.) and then provide the processed images to a trained AI (embryo viability assessment) model 100, which is run on one of the on-demand cloud servers 32 to generate an embryo viability score (314). A report including the embryo viability score is generated (316) and sent or otherwise provided to the user 40, such as via the user interface 42. A user (e.g., an embryologist) receives the embryo viability score via the user interface and can then use the viability score to help decide whether to implant the embryo or which embryo in the set is the best for implantation. The selected embryo is then implanted (205). To aid in further refinement of the AI model, pregnancy outcome data, such as the detection (or non-detection) of a heartbeat in the first ultrasound scan after implantation (usually around 6-10 weeks post-fertilization), can be provided to the system. This allows the AI model to be retrained and updated as more data becomes available.
[0044] Images may be captured using various imaging systems, such as those found in existing IVF clinics. This has the advantage that IVF clinics do not need to purchase new or use specific imaging systems. The imaging system is typically an optical microscope configured to capture single-phase contrast images of the embryo. However, it will be understood that other imaging systems, particularly optical microscope systems using various imaging sensors and image capture technologies, can be used. These may include phase contrast microscopes, polarizing microscopes, differential interference contrast (DIC) microscopes, dark-field microscopes, and bright-field microscopes. Images may be captured using a conventional optical microscope equipped with a camera or image sensor, or the images may be captured by a camera with integrated optics capable of capturing high-resolution or high-magnification images, including smartphone systems. The image sensor may be a CMOS sensor chip or a charge-coupled device (CCD), each with associated electronics. The optics may be configured to collect specific wavelengths or use filters, including bandpass filters, to collect (or reject) specific wavelengths. Some image sensors may be configured to operate or be sensitive to specific wavelengths of light or light with wavelengths beyond the optical range, including infrared (IR) or near-infrared. In some embodiments, the imaging sensor is a multispectral camera that collects images at multiple different wavelength ranges. An illumination system may also be used to illuminate the embryo with light of a specific wavelength, wavelength band, or intensity. Stops and other components may be used to limit or modify illumination to specific portions (or image planes) of the image.
[0045] Additionally, images used in the embodiments described herein may be sourced from video and time-lapse imaging systems. A video stream is a periodic sequence of image frames where the interval between image frames is determined by the capture frame rate (e.g., 24 or 48 frames per second, etc.). Similarly, a time-lapse system captures a sequence of images at a very slow frame rate (e.g., one image per hour) to obtain a sequence of images as the embryo develops (post-fertilization). Thus, it will be understood that the image used in the embodiments described herein may be a single image extracted from a video stream or a time-lapse sequence of images of the embryo. When an image is extracted from a video stream or a time-lapse sequence, the image to be used may be selected as the image having a capture time closest to a reference time point, such as 5.0 or 5.5 days post-fertilization.
[0046] In some embodiments, preprocessing may include image quality assessment so that an image can be rejected if it does not meet the image quality criteria. If the original image does not meet the image quality criteria, additional images can be captured. In embodiments in which images are selected from a video stream or time-lapse sequence, the selected image is the first image that passes the image quality assessment closest to the reference time. Alternatively, a reference time window can be defined along with the image quality criteria (e.g., 30 minutes after the beginning of day 5.0). In this embodiment, the selected image is the selected image with the highest image quality during the reference time window. The image quality criteria used in performing the image quality assessment may be based on pixel color distribution, brightness range, and / or abnormal image characteristics or features indicative of poor image quality or device malfunction. A threshold may be determined by analyzing a reference set of images. This may be based on manual assessment or an automated system that extracts outliers from the distribution.
[0047] The generation of the AI embryo viability assessment model 100 can be further understood with reference to Figure 3A, which is a schematic architectural diagram of a cloud-based computing system 1 configured to generate and use an AI model 100 configured to estimate an embryo viability score from an image, according to one embodiment. Referring to Figure 1B, the AI model generation method is handled by a model monitor 21.
[0048] The model monitor 21 allows a user 40 to provide image data and metadata 14 to a data management platform, including a data repository. Data preparation steps are performed, for example, to move images to specific folders, rename images, and perform preprocessing on the images, such as object detection, segmentation, alpha channel removal, padding, cropping / localization, normalization, and scaling. Feature descriptors can be calculated and augmented images can be generated in advance. However, further preprocessing, including augmentation, can also be performed during training (i.e., on the fly). Images can also undergo image quality assessment to enable rejection of clearly inferior images and capture of replacement images. Similarly, patient records or other clinical data can be processed (prepared) for additional embryo viability classifications (e.g., viable or nonviable) that can be linked or associated with each image for use in training and / or evaluating the AI model. The prepared data, along with the latest versions of the training algorithms, are loaded (16) into a template server 28 on a cloud provider (e.g., AWS). The template server is stored and multiple copies are made across various training server clusters 37, which may be CPU, GPU, ASIC, FPGA, or TPU (tensor processing unit) based, forming training servers 35. The model monitor web server 31 then applies each job submitted by a user 40 to a training server 37 from multiple cloud-based training servers 35. Each training server 35 runs pre-prepared code (from the template server 28) to train an AI model using libraries such as PyTorch, Tensorflow, or equivalent, and can use computer vision libraries such as OpenCV. PyTorch and OpenCV are open-source libraries with low-level commands for building CV machine learning models.
[0049] The training server 37 manages the training process. This may include, for example, dividing images into training, validation, and blind validation sets using a random assignment process. Additionally, during training / validation cycles, the training server 37 may randomize the set of images at the beginning of the cycle so that a different subset of images is analyzed or analyzed in a different order for each cycle. If preprocessing was not previously performed or was incomplete (e.g., during data management), further preprocessing may be performed, including object detection, segmentation, and generation of masked datasets (e.g., zona pellucida-only images, IZC-only images, etc.), calculation / estimation of CV feature descriptors, and generation of data augmentation. Preprocessing may also include padding, normalization, etc., as needed. That is, preprocessing step 102 may be performed prior to training, during training, or some combination (i.e., distributed preprocessing). The number of running training servers 35 can be managed from a browser interface. As training progresses, logging information about the status of training is recorded (62) to a distributed logging service, such as CloudWatch 60. Key patient and accuracy information is also parsed from the logs and stored in a relational database 36. The models are also periodically saved (51) to data storage (e.g., AWS Simple Storage Service (S3) or a similar cloud storage service) 50 so that they can be retrieved and loaded at a later date (e.g., for restart in case of an error or other outage). Email updates about the training server status are sent (44) to users 40 when the training server job is complete or if an error is encountered.
[0050] Within each training cluster 37, numerous processes take place. When the cluster is started via the web server 31, a script automatically runs, which reads prepared images and patient records and launches the specific Pytorch / OpenCV training code requested (71). Input parameters 28 for model training are provided by the user 40 via the browser interface 42 or via a configuration script. The training process 72 is then initiated for the requested model parameters, which can be a lengthy and intensive task. Therefore, to avoid losing progress while training is in progress, logs are periodically saved to a logging (e.g., AWS CloudWatch) service 60, and the current version of the model (as it is being trained) is saved to a data (e.g., S3) storage service for later retrieval and use (51). One embodiment of a schematic flow diagram of the model training process on the training server is shown in FIG. 3B. Various trained AI models are available on the data storage service, and multiple models can be combined using, for example, ensemble, distillation, or similar approaches to incorporate various deep learning models (e.g., PyTorch, etc.) and / or targeted computer vision models (e.g., OpenCV, etc.) to generate a robust AI model 100 that is provided to the cloud-based delivery platform 30.
[0051] The cloud-based delivery platform 30 system then allows users 10 to drag and drop images directly into a web application 34, which prepares the images and passes them through a trained / validated AI model 100 to obtain an embryo viability score, which is immediately returned in a report (as shown in FIG. 2). The web application 34 also allows clinics to store data such as images and patient information in a database 36, generate various reports on the data, and use tools for their organization, group, or specific users, as well as generate audit reports on billing and user accounts (e.g., create users, delete users, reset passwords, change access levels, etc.). The cloud-based delivery platform 30 also allows product administrators to access the system to create new customer accounts and users, reset passwords, as well as access customer / user accounts (including data and screens) to facilitate technical support.
[0052] Various steps and variations in the generation of embodiments of an AI model configured to estimate an embryo viability score from images will now be discussed in further detail. Referring to FIG. 1A , the model was trained using images captured on day 5 post-fertilization (i.e., the 24-hour period from day 5:00:00 to day 5:23:59). Studies on validated models have shown that model performance is significantly improved using images taken on day 5 post-fertilization compared to images taken on day 4 post-fertilization. However, as noted above, effective models can still be developed using shorter time windows, such as 12 hours, or images taken on other days, such as day 3 or day 4, or at least a minimum time post-fertilization, such as day 5 (e.g., an open-ended time window, etc.). Perhaps more important than the exact time window (e.g., day 4 or day 5), is that the images used to train the AI model and the images used for subsequent classification by the trained AI model are taken during similar, preferably the same, time windows (e.g., the same 12- or 24-hour window).
[0053] Prior to analysis, each image undergoes a preprocessing (image preparation) procedure (102), which includes at least segmenting the image to identify the zona pellucida region. Various preprocessing steps or techniques may be applied. These may be performed after adding to the data storage device 14 or during training by the training server 37. In some embodiments, an object detection (localization) module is used to detect and identify the location of the image relative to the embryo. Object detection / localization involves estimating a bounding box that contains the embryo. This can be used for image cropping and / or segmentation. The image may also be padded with a given boundary, and then color-balanced and brightness-normalized. The image is then cropped so that regions outside the embryo are closer to the image boundary. This is achieved using computer vision techniques for boundary selection, including the use of AI object detection models. Image segmentation is a computer vision technique useful for preparing images for a particular model to select relevant regions of interest for model training, such as the zona pellucida and intrazonal cavity (IZC). Images may be masked to generate a zona-only (i.e., cropping the zona pellucida border and masking the IZC) image (see Figure 6F) or an IZC-only (i.e., cropping the IZC border and eliminating the zona pellucida) image (Figure 6G). The background may be left in the image or masked. An embryo viability model may then be trained using only the masked images, e.g., a zona pellucida image masked to include only the zona pellucida and image background and / or an IZC image masked to include only the IZC. Scaling involves rescaling the image to a predefined scale to fit the particular model being trained. Augmentation involves incorporating small modifications to a copy of the image, such as rotating the image to control the orientation of the embryo dish. The use of segmentation prior to deep learning was found to have a significant impact on the performance of deep learning methods. Similarly, augmentation was important for generating robust models.
[0054] A variety of image pre-processing techniques can be used to prepare human embryo images prior to training an AI model. These include: Alpha channel stripping, which involves stripping the image of its alpha channel (if present), for example to remove the transparency map, and ensuring that it is coded in a three-channel format (e.g., RGB, etc.); Pad / bolster each image with a padding border to create a square aspect ratio prior to segmentation, cropping, or boundary finding; this process ensures that image dimensions are consistent, comparable, and compatible with deep learning methods that typically require square images as input, while also ensuring that key components of the image are not cropped; Normalizing RGB (Red-Green-Blue) or grayscale images to a fixed average value for all images, for example this involves taking the average of each RGB channel and dividing each channel by that average, then multiplying each channel by a fixed value of 100 / 255 to ensure that the average value of each image in RGB space is (100,100,100), this step ensures that color bias between images is suppressed and that the brightness of each image is normalized; Image thresholding using binary, Otsu, or adaptive methods, including dilation (opening), erosion (closing), scale gradients, and morphological processing of images using scale masks to extract outer and inner boundaries of shapes; Performing object detection / image cropping to locate the image relative to the embryo and ensure there are no artifacts around the edges of the image, which may be done using an object detector that uses an object detection model (discussed below) trained to estimate a bounding box that contains the embryo (including the zona pellucida); extracting boundary geometric characteristics using an elliptical Hough transform of the image contour, for example, a best ellipse fit from an elliptical Hough transform computed on a binary threshold map of the image, the method working by selecting a hard boundary of the embryo in the image and cropping a square boundary of the new image so that the longest radius of the new ellipse is encompassed by the width and height of the new image and the center of the ellipse is the center of the new image; Zooming an image by ensuring a consistently centered image with a consistent border size around an elliptical region; segmenting the image to identify zona pellucida regions and cytoplasmic intrazona pellucida cavity (IZC) regions, where segmentation may be done by calculating a best fit contour around a non-elliptical image using a geometric active contour (GAC) model or morphological snakes within a given region, where the interior of the snake and other regions may be treated differently depending on the trained model's focus on cytoplasmic (intrazona pellucida cavity) regions or zona pellucida regions, which may harbor the blastocyst; or a semantic segmentation model may be trained to identify a class for each pixel in the image, where in one embodiment the semantic segmentation model is deployed using a U-Net architecture with a ResNet-50 encoder pre-trained to segment the zona pellucida and IZC, and the model is trained using a binary cross-entropy loss function; Annotating an image by selecting a feature descriptor and masking all regions of the image except those within a given radius of the descriptor's keypoints; Resize / scale the entire set of images to a specified resolution; and a tensor transform that involves converting each image into a tensor rather than a visually displayable image, as this data format is more usable by deep learning models; in one embodiment, tensor normalization is obtained from standard pre-trained ImageNet values with mean (0.485, 0.456, 0.406) and standard deviation (0.299, 0.224, 0.225); Includes:
[0055] Figure 4 is a schematic diagram of a binary thresholding process 400 for boundary finding on an image of a human embryo, according to one embodiment. Figure 4 shows eight binary thresholding processes applied to the same image: levels 60, 70, 80, 90, 100, and 110 (images 401, 402, 403, 404, 405, and 406, respectively), adaptive Gaussian 407, and Otsu-Gaussian 408. Figure 5 is a schematic diagram of a boundary finding method 500 on an image of a human embryo, according to one embodiment. The first panel shows an outer boundary 501, an inner boundary 502, and an image 503 with the detected inner and outer boundaries. The inner boundary 502 may correspond approximately to the IZC boundary, and the outer boundary 501 may correspond approximately to the outer edge of the zona pellucida region.
[0056] Figure 6A shows an example 600 of the use of a geometric active contour (GAC) model applied to a fixed region of an image for image segmentation, according to one embodiment. The solid blue line 601 indicates the outer boundary of the zona pellucida region, and the dashed green line 602 indicates the inner boundary that defines the edge of the zona pellucida region and the cytoplasmic (intrazona pellucida cavity or IZC) region. Figure 6B shows an example of the use of morphological snakes applied to a fixed region of an image for image segmentation. Again, the solid blue line 611 indicates the outer boundary of the zona pellucida region, and the dashed green line 612 indicates the inner boundary that defines the edge of the zona pellucida region and the cytoplasmic (inner) region. In this second image, the boundary 612 (defining the cytoplasmic intrazona pellucida cavity region) has an irregular shape with a ridge or protrusion in the lower right quadrant.
[0057] In another embodiment, the object detector uses an object detection model trained to estimate bounding boxes containing germs. The goal of object detection is to identify the largest bounding box that contains all of the pixels associated with the object. This requires the model to model both the object's location and category / label (i.e., what is in the box); therefore, the detection model typically has both an object classifier head and a bounding box regression head.
[0058] One approach is to apply a region-based convolutional neural network (or R-CNN) that uses an expensive search process to search for image patch proposals (potential bounding boxes). These bounding boxes are then used to crop regions of interest. A classification model is then run on the cropped image to classify the content of the image region. This process is complex and computationally expensive. An alternative is Fast CNN, which uses a CNN to propose feature regions rather than searching for image patch proposals. This model uses a CNN to estimate a fixed number of candidate boxes, typically set between 100 and 2000. An even faster alternative approach is Faster RCNN, which uses anchor boxes to limit the search space of required boxes. By default, a standard set of nine anchor boxes (each of a different size) is used. Faster RCNN uses a small network jointly trained to predict feature regions of interest, which can replace expensive region searches and therefore improve runtime compared to R-CNN or Fast CNN.
[0059] For every feature activation coming in, one model is considered an anchor point (red in the image below). For every anchor point, nine anchor boxes (or more or less, depending on the problem) are generated. The anchor boxes correspond to common object sizes in the training dataset. Because there are multiple anchor points with multiple anchor boxes, this results in tens of thousands of region proposals. The proposals are then filtered through a process called Non-Maximal Suppression (NMS), which selects the largest box with a confident smaller box within it. This ensures that there is only one box per object. Because NMS relies on the reliability of each bounding box prediction, a threshold must be set for when objects are considered part of the same object instance. Because the anchor boxes do not fit the object perfectly, the job of the regression head is to predict offsets to these anchor boxes that transform into the best-fit bounding box.
[0060] Detectors can also be specialized, estimating boxes for only a subset of objects, such as only people for a pedestrian detector. Object categories of no interest are encoded as class 0, corresponding to the background class. During training, patches / boxes for the background class are typically randomly sampled from image regions that lack bounding box information. This step allows the model to become invariant to these undesirable objects, e.g., learn to ignore them rather than incorrectly classify them. Bounding boxes are typically represented in two different formats, the most common being (x1, y1, x2, y2), where point p1 = (x1, y1) is the top-left corner of the box and p2 = (x2, y2) is the bottom-right corner. Another common box format is (cx, cy, height, width), where the bounding box / rectangle is encoded as the center point (cx, cy) of the box and the size (height, width) of the box. The detection method uses different encodings / formats depending on the task and situation.
[0061] The regression head may be trained using an L1 loss, and the classification head may be trained using a cross-entropy loss. Objectness loss (is this background or object?) may also be used, and the final loss is calculated as the sum of these losses. Individual losses may also be weighted as follows:
[0062]
number
[0063] As an alternative to GAC segmentation, semantic segmentation may be used. Semantic segmentation is the task of predicting a category or label for every pixel. Tasks like semantic segmentation are called pixel-by-pixel dense prediction tasks because an output is required for every input pixel. Semantic segmentation models are configured differently from standard models because they require full image output. Typically, semantic segmentation (or any dense prediction model) will have an encoding module and a decoding module. The encoding module is responsible for creating a low-dimensional representation (sometimes called a feature representation) of the image. This feature representation is then decoded into the final output image via the decoding module. During training, the predicted label map (for semantic segmentation) is then compared to a ground truth label map that assigns a category to each pixel, and a loss is calculated. The standard loss function for segmentation models is either binary cross-entropy or standard cross-entropy loss (depending on whether the problem is multi-class). These implementations are identical to their image classification cousins, except that the loss is applied pixel-by-pixel (across the image channel dimensions of the tensor).
[0064] Fully convolutional network (FCN)-style architectures are commonly used in the field for general semantic segmentation tasks. In this architecture, a pre-trained model (such as a ResNet) is first used to encode a low-resolution image (approximately 1 / 32 of the original resolution, but could be as low as 1 / 8 if dilated convolutions are used). This low-resolution label map is then upsampled to the original image resolution, and a loss is calculated. The intuition behind predicting the low-resolution label map is that the semantic segmentation mask has very low frequency, avoiding the need for all the extra parameters of a larger decoder. More complex versions of this model exist that use multi-stage upsampling to improve segmentation results. Briefly, the loss is calculated at multiple resolutions in an incremental fashion to refine the prediction at each scale.
[0065] One drawback of this type of model is that when the input data is high-resolution or contains high-frequency information (i.e., smaller / thinner objects), the low-resolution label map cannot capture these small structures (especially if the encoding model does not use dilated convolutions). In a standard encoder / convolutional neural network, the input image / image features are progressively downsampled as the model gets deeper. However, as the image / feature is downsampled, key high-frequency details may be lost. Therefore, to address this, an alternative U-Net architecture can be used that instead uses skip connections between symmetric components of the encoder and decoder. Briefly, every encoding block has a corresponding block in the decoder. Then, the features at each stage are passed to the decoder along with the lowest-resolution feature representation. For each decoding block, the input feature representation is upsampled to match the resolution of its corresponding encoding block. Next, the feature representations from the encoding block and the upsampled low-resolution features are concatenated and passed through a 2D convolutional layer. By concatenating features in this way, the decoder can learn to refine the input in each block, choosing which details (low-resolution details or high-resolution details) to integrate depending on that input.
[0066] An example U-Net architecture 620 is shown in Figure 6C. The key difference between FCN-style models and U-Net-style models is that in FCN models, the encoder is responsible for predicting a low-resolution label map, which is then (possibly progressively) upsampled. U-Net models, on the other hand, do not have fully complete label map predictions until the final layer. Finally, many variations of these models (e.g., hybrids, etc.) exist, trading off the differences between these models. U-net architectures can also use pre-trained weights, such as ResNet-18 or ResNet-50, for use when there is insufficient data to train a model from scratch.
[0067] In some embodiments, segmentation was performed using a U-Net architecture with a pre-trained ResNet-50 encoder trained using binary cross-entropy to identify the zona pellucida region and the intrazona pellucida cavity region. This U-Net architecture-based segmenter generally outperformed active contour-based segmentation, particularly for images with lower image quality. Figures 6D-6F illustrate segmentation according to one embodiment. Figure 6D is an image of a day-5 embryo 630, including a zona pellucida region 631 surrounding an intrazona pellucida cavity (IZC, 632). In this embodiment, the embryo has begun to hatch, and the ISZ has emerged (hatched) from the zona pellucida. The embryo is surrounded by background pixels 633. Figure 6E is a padded image 640 generated from Figure 6D by adding padding pixels 641, 642 to generate a square image that is more easily processed by deep learning methods. FIG. 6F shows a zona pellucida image 650 in which the IZC 652 has been masked to leave the zona pellucida 631 and background pixels 633, and FIG. 6G shows an IZC image 660 in which the zona pellucida and background 661 have been masked, leaving only the IZC region 632. Once segmented, a set of images can be generated in which all regions except the desired region are masked. AI models can then be trained on these particular sets of images. That is, the AI models can be divided into two groups: first, those that involve further image segmentation, and second, those that require entirely unsegmented images. Models trained on images with the IZC masked and the zona pellucida region exposed are designated zona pellucida models. Models trained on images with the zona pellucida masked (designated IZC models) and models trained on full embryo images (i.e., the second group) were also considered for training.
[0068] In one embodiment, to ensure the uniqueness of each image, the name of the new image is set equal to a hash of the original image content as a png (lossless) file so that duplicate records do not bias the results. When executed, the data parser will output images in a multi-threaded manner for any images that do not already exist in the output directory (and will create them if they do not exist), so that lengthy processes can be restarted from the same point if interrupted. The data preparation step may also include processing metadata to eliminate images associated with inconsistent or inconsistent records and identify any erroneous clinical records. For example, a script can be run on a spreadsheet to conform the metadata to a predefined format. This ensures that the data used to generate and train the model is of high quality and has uniform characteristics (e.g., size, color, scale, etc.).
[0069] In some embodiments, the data is cleaned by identifying images with likely inaccurate pregnancy outcome labels (i.e., mislabeled data) and eliminating or relabeling the identified images. In one embodiment, this is done by estimating the likelihood that the pregnancy outcome label associated with the image is inaccurate and comparing that likelihood to a threshold. If the likelihood exceeds the threshold, the image is eliminated or relabeled. Estimating the likelihood that the pregnancy outcome label is inaccurate may be done by using multiple AI classification models and k-fold cross-validation. In this approach, the images are divided into k mutually exclusive validation datasets. Each of the multiple AI classification models is trained on the combined k-1 validation datasets and then used to classify images in the remaining validation datasets. The likelihood is then determined based on the number of AI classification models that incorrectly classify the pregnancy outcome label of the image. In some embodiments, a deep learning model may further be used to learn a likelihood value.
[0070] Once the data has been appropriately preprocessed, it can then be used to train one or more AI models. In one embodiment, the AI model is a deep learning model trained on a set of zona pellucida images in which all regions of the images except the zona pellucida are masked during preprocessing. In one embodiment, multiple AI models are trained and then combined using an ensemble or distillation method. The AI models may be one or more deep learning models and / or one or more computer vision (CV) models. The deep learning models may be trained on whole embryo images, zona pellucida images, or IZC images. The computer vision (CV) models may be generated using machine learning methods using a set of feature descriptors calculated from each image, each individual model configured to estimate an embryo viability score for the embryo in the image, and the AI model combines the selected models to generate an overall embryo viability score returned by the AI model.
[0071] Training is performed using a randomized dataset. Complex image data sets, especially those smaller than approximately 10,000 images, can suffer from uneven distribution, where primary viable or nonviable embryo examples are not evenly distributed throughout the set. Therefore, several (e.g., 20) randomizations of the data are considered at a time and then divided into training, validation, and blind test subsets, as defined below. All randomizations are used on a single training example to determine which one provides the best distribution for training. As a corollary, it is also beneficial to ensure that the ratio of viable to nonviable embryos is the same across all subsets. Embryo images are highly diverse, and therefore, using a uniform distribution of images across the test and training sets ensures that performance can be improved. Therefore, after randomization, the ratio of images classified as viable to images classified as nonviable in each of the training, validation, and blind validation sets is calculated and tested to ensure that the ratios are similar. For example, this may include testing whether the range of the ratios is within a certain variance considering the number of images or below a threshold. If the ranges are not similar, the randomization is discarded and a new randomization is generated and tested until a randomization is obtained whose ratios are similar. More generally, if the outcome is an n-ary outcome with n states, after the randomization is performed the calculation step may include calculating the frequency of each of the n-ary outcome states in each of the training set, validation set, and blind validation set, and testing that the frequencies are similar; if the frequencies are not similar, discarding the assignment and repeating the randomization until a randomization is obtained whose frequencies are similar.
[0072] The training further includes performing multiple training / validation cycles. In each training / validation cycle, each randomization of the total available dataset is divided into typically three separate datasets known as the training, validation, and blind validation datasets. In some variations, four or more datasets can be used, for example, the validation and blind validation datasets can be stratified into multiple sub-test sets of different difficulty.
[0073] The first set is the training dataset, which contains at least 60%, preferably 70-80%, of the images. These images are used by deep learning and computer vision models to create an embryo viability assessment model and accurately identify viable embryos. The second set is the validation dataset, which typically contains approximately (or at least) 10% of the images. This dataset is used to verify or test the accuracy of the model created using the training dataset. Although these images are independent of the training dataset used to create the model, the validation dataset still has a small positive bias in accuracy because it is used to monitor and optimize the progress of model training. Therefore, training tends to target a model that maximizes accuracy for this specific validation dataset, which is not necessarily the best model when applied more generally to other embryo images. The third dataset is the blind validation dataset, which typically contains approximately 10-20% of the images. To address the positive bias using the validation dataset, the third blind validation dataset is used to perform a final, unbiased accuracy assessment of the final model. This validation occurs at the end of the modeling and validation process, when a final model is created and selected. It is important to ensure that the accuracy of the final model is relatively consistent with the validation dataset to ensure that the model can generalize to all embryo images. For the reasons stated above, the accuracy of the validation dataset is likely to be higher than that of the blind validation dataset. The results of the blind validation dataset are a more reliable measure of the model's accuracy.
[0074] In some embodiments, preprocessing the data further includes augmenting the images, where modifications are made to the images. This may be done prior to training or during training (i.e., on the fly). Augmentation may involve augmenting (modifying) the images directly or by creating a copy of the image with small changes. Any number of augmentations can be performed, including 90-degree rotation of the image, mirror flip, non-90-degree rotation where a diagonal border is embedded to match the background color, varying the amount of image blur, adjusting the image contrast using an intensity histogram, and applying one or more small random transformations in both horizontal and / or vertical directions, random rotation, adding JPEG (or compression) noise, random image resizing, random hue jitter, random brightness jitter, contrast-limited adaptive histogram equalization, random flip / mirror, image sharpening, image embossing, random brightness and contrast, RGB color shift, random hue and saturation, channel shuffle, swapping RGB to BGR or RBG or other, coarse dropout, motion blur, center blur, Gaussian blur, random shift-scale rotation (i.e., all three combined). The same set of augmented images may be used for multiple training-validation cycles, or new augmentations may be generated on the fly during each cycle. A further extension used in CV model training is changing the "seed" of the random number generator for extracting feature descriptors. Techniques for obtaining computer vision descriptors have an element of randomness in extracting feature samples. This random number can be changed and included between extensions to provide more robust training for the CV model.
[0075] Computer vision models rely on identifying key image features and representing them in terms of descriptors. These descriptors can encode qualities such as pixel variation, gray level, texture roughness, fixed corner points, or image gradient orientation, and are implemented in OpenCV or similar libraries. By selecting such features to search for in each image, a model can be built by discovering which feature configurations are good indicators of embryo viability. This procedure is best performed by machine learning processes such as random forests or support vector machines, which can segment images in terms of their descriptors from computer vision analysis.
[0076] A variety of computer vision descriptors, encompassing both small-scale and large-scale features, are used and combined with traditional machine learning methods to generate "CV models" for embryo selection. These are optionally later combined with deep learning (DL) models, e.g., into ensemble models, or used in distillation to train student models. Suitable computer vision image descriptors include: Zona pellucida via Hough transform: find inner and outer ellipses to approximate the split of the zona pellucida and the intrazonal cavity, and record the mean and difference of the radii as features; Gray-Level Co-occurrence Matrix (GLCM) texture analysis: detects the roughness of different regions by comparing adjacent pixels within the region. Sample feature descriptors used are: angular second moment (ASM), uniformity, correlation, contrast, and entropy. The selection of regions is obtained by randomly sampling a given number of square subregions of the image of a given size, and the results of each of the five descriptors for each region are recorded as a set of total features; Histogram of Oriented Gradients (HOG): Detects objects and features using scale-invariant feature transformation descriptors and shape context. This method has gained popularity for use in embryology and other medical imaging, but does not constitute a machine learning model in itself; Feature Extraction by Oriented Accelerated Fragmentary Testing (FAST) and Rotational Binary Robust Independent Basic Features (BRIEF) (ORB): an industry standard alternative to SIFT and SURF features, relying on a combination of FAST keypoint detectors (specific pixels) and BRIEF descriptors, modified to include rotational invariance; Binary Robust Invariant Scalable Keypoint (BRISK): A FAST-based detector combined with an assembly of pixel intensity comparisons, which is achieved by sampling each neighborhood around the feature specified by the keypoint; Maximum Stable Extremal Region (MSER): A local morphological feature detection algorithm via the extraction of covariant regions, which are stable connected components with respect to one or more gray level sets extracted from an image; Feature-Friendly Tracking (GFTT): A feature detector that uses an adaptive window size to detect corner textures, identified by using Harris or Shi-Tomasi corner detection and extracting points that exhibit high standard deviations in their spatial intensity profile.
[0077] FIG. 7 is a plot 700 of a gray-level co-occurrence matrix (GLCM) showing GLCM correlations of sample feature descriptors 702: ASM, homogeneity, correlation, contrast, and entropy, computed on a set of six zona pellucida regions (labeled 711 to 716; crosshatched) and six cytoplasm / IZC regions (labeled 721 to 726; dotted lines) in image 701.
[0078] A computer vision (CV) model is constructed by the following method: One (or more) of the computer vision image descriptor techniques described above is selected, and features are extracted from all of the images in the training dataset. These features are arranged into a combined array and then fed into a K-means unsupervised clustering algorithm; this array is called a codebook for "bag of visual words." The number of clusters is a free parameter of the model. The clustered features from this point represent "custom features" used throughout any combination of algorithms, to which each individual image in the validation or test set will be compared. Each image, with its extracted features, is individually clustered. For a given image with its clustered features, its "distance" (in feature space) to each of the clusters in the codebook is measured using a KD tree query algorithm that yields the closest clustered feature. The results from the tree query can then be represented as a histogram, showing the frequency with which each feature occurs in that image. Finally, the question of whether a particular combination of these features corresponds to a measure of embryonic viability needs to be evaluated using machine learning. Here, supervised learning is performed using the histogram and ground truth results. Methods used to obtain the final selection model include random forests or support vector machines (SVMs).
[0079] Multiple deep learning models can also be generated. Deep learning models are based on neural network methods, typically convolutional neural networks (CNNs) consisting of multiple connected layers, with each layer of "neurons" having a nonlinear activation function such as "rectifier" or "sigmoid." In contrast to feature-based methods (i.e., CV models), deep learning and neural networks do not rely on manually designed feature descriptors but instead "learn" features. This allows them to learn "feature representations" tailored to the desired task. These methods are well suited to image analysis because they can pick up both small details and overall morphological shapes to arrive at an overall classification. Various deep learning models are available, each with a different architecture (i.e., different numbers of layers and inter-layer connections), such as residual networks (e.g., ResNet-18, ResNet-50, and ResNet-101), densely connected networks (e.g., DenseNet-121 and DenseNet-161), and other variations (e.g., InceptionV4 and Inception-ResNetV2). Deep learning models can be evaluated based on stability (how stable the accuracy values were relative to the validation set throughout the training process), transferability (how well the accuracy on the training data correlated with the accuracy on the validation set), and predictive accuracy (which model provided the best validation accuracy, total accuracy, and balanced accuracy, defined as the weighted average accuracy across both embryo class types, for both viable and nonviable embryos). Training involves trying different combinations of model parameters and hyperparameters, including input image resolution, optimization algorithm selection, learning rate value and scheduling, momentum value, dropout, and weight initialization (pre-training). A loss function may be defined to evaluate the model's performance; during training, the deep learning model is optimized by varying the learning rate to drive an update mechanism for the network's weight parameters to minimize the objective / loss function.
[0080] Deep learning models may be implemented using various libraries and software languages. In one embodiment, the PyTorch library is used to implement neural networks in the Python language. The PyTorch library also allows tensors to be created that utilize hardware (GPU, TPU) acceleration and includes modules for building multiple layers for neural networks. While deep learning is one of the most powerful techniques for image classification, it can be improved by providing guidance through the use of segmentation or augmentation described above. The use of segmentation prior to deep learning has been found to significantly impact the performance of deep learning methods and aid in the generation of contrasting models. Therefore, preferably, at least some deep learning models are trained on segmented images, such as images in which zona pellucida is identified, or images masked to hide all areas except the zona pellucida region. In some embodiments, the multiple deep learning models include at least one model trained on segmented images and one model trained on images that have not undergone segmentation. Similarly, augmentation is important for generating robust models.
[0081] The effectiveness of the approach is determined by the architecture of the deep neural network (DNN). However, unlike feature descriptor methods, DNNs learn the features themselves through convolutional layers before applying the classifier. That is, without manually incorporating proposed features, DNNs can be used to check existing practices in the literature as well as develop previously unguessed descriptors, especially those that are difficult for the human eye to detect and measure.
[0082] The architecture of a DNN is constrained by the size of the image as input, a hidden layer with the dimensions of a tensor describing the DNN, and a linear classifier with the number of class labels as output. Most architectures utilize multiple downsampling ratios with small (3x3 pixel) filters to capture the concepts of left / right, up / down, and center. A stack of a) convolutional 2d layers, b) rectified linear units (ReLUs), and c) max-pooling layers allows the number of parameters passed through the DNN to remain manageable while allowing the filters to pass over high-level (topological) features of the image and map them onto intermediate and finally microscopic features embedded in the image. The top layer typically includes one or more fully connected neural network layers that act as classifiers, similar to SVMs. A softmax layer is typically used to normalize the resulting tensor so that it has the probability of being the result of the fully connected classifier. Thus, the output of the model is a list of probabilities that the image is either non-viable or viable.
[0083] FIG. 8 is a schematic architecture diagram of a deep learning method including a convolutional layer that converts input images into predictions after training, according to one embodiment. FIG. 8 shows a series of layers based on the RESNET 152 architecture, according to one embodiment. The components are annotated as follows: "CONV" indicates a convolutional 2D layer that calculates the cross-correlation of inputs from the layer below. Each element, or neuron, in a convolutional layer processes inputs only from its receptive field, e.g., 3x3 or 7x7 pixels. This reduces the number of learnable parameters required to describe the layer, allowing for the creation of deeper neural networks than those built from fully connected layers. In these layers, every neuron is connected to every other neuron in the subsequent layer, which is memory intensive and prone to overfitting. Convolutional layers are also spatially invariant, which is useful for processing images where the subject cannot be guaranteed to be precisely centered. "POOL" refers to a max pooling layer, which is a downsampling method whereby only representative neuron weights are selected within a given region, reducing the complexity of the network and also reducing overfitting. For example, for weights within a 4x4 square region of a convolutional layer, the maximum value of each 2x2 corner block is calculated, and these representative values are then used to reduce the size of the square region to a dimension of 2x2. RELU refers to the use of rectified linear units acting as a nonlinear activation function. A common example is when a ramp function is of the following form for an input x from a given neuron:
[0084]
number
[0085] One suitable DNN architecture is ResNet (https: / / ieeexplore.ieee.org / document / 7780459), such as ResNet152, ResNet101, ResNet50, or ResNet-18. ResNet significantly advanced the field in 2016 by using a significantly larger number of hidden layers and by introducing "skip connections," also known as "residual connections." Only the differences from one layer to the next are calculated, which is more time-effective. If little change is detected in a particular layer, that layer is skipped, thus creating a network that adapts very quickly to a combination of small and large features in an image. In particular, ResNet-18, ResNet-50, ResNet-101, DenseNet-121, and DenseNet-161 generally outperformed other architectures. Another suitable DNN architecture is DenseNet (https: / / ieeexplore.ieee.org / document / 8099726), such as DenseNet161, DenseNet201, DenseNet169, and DenseNet121. DenseNet is an evolution of ResNet, where every layer now has a maximum number of skip connections, allowing skipping to any other layer. This architecture requires much more memory and is therefore less efficient, but can show improved performance over ResNet. It is also easy to overtrain / overfit with a large number of model parameters. All model architectures are often combined with methods to control this, especially DenseNet-121 and DenseNet-161. Another suitable DNN architecture is Inception(-ResNet) (https: / / www.aaai.org / ocs / index.php / AAAI / AAAI17 / paper / viewPaper / 14806), such as InceptionV4 and InceptionResNetV2.Inception represents a more complex convolutional unit, whereby instead of simply using fixed-size filters (e.g., 3x3 pixels) as described in Section 3.2, filters of several sizes are computed in parallel (5x5, 3x3, 1x1 pixels) with weights that are free parameters, allowing the neural network to prioritize which filters are most suitable at each layer in the DNN. A development of this type of architecture is to combine it with skip connections in the same way as ResNet, creating Inception-ResNet. In particular, ResNet-18, ResNet-50, ResNet-101, DenseNet-121, and DenseNet-161 generally outperformed other architectures.
[0086] As mentioned above, both computer vision and deep learning methods are trained on pre-processed data using multiple training and validation cycles, which follow the following framework:
[0087] The training data is preprocessed and split into batches (the number of data in each batch is a free model parameter, controlling how fast and how stably the algorithm learns). Augmentation may be performed prior to splitting or during training.
[0088] After each batch, the network weights are adjusted and the running total accuracy is evaluated. In some embodiments, the weights are updated between batches, for example, using gradient accumulation. When all images have been evaluated and one epoch has run, the training set is shuffled (i.e., a new randomization with the set is obtained) and training starts again from the beginning for the next epoch.
[0089] During training, a number of epochs may be performed depending on the size of the dataset, the complexity of the data, and the complexity of the model being trained. The optimal number of epochs typically ranges from 2 to 100, but may be higher depending on the particular case.
[0090] After each epoch, the model is run on a validation set without any training to provide a measure of progress in how accurate the model is and guide the user on whether more epochs should be run or whether more epochs would result in overtraining. The validation set guides the selection of all model parameters or hyperparameters and is therefore not a true blind set. However, it is important that the distribution of images in the validation set is very similar to the final blind test set that will be run after training.
[0091] When reporting validation set results, augmentations can also be included (all) or not (noaug) for each image. Additionally, augmentations for each image may be combined to provide a more robust final result for the image. Several combination / voting strategies can be used, including average confidence (taking the average of the model's inferences across all augmentations), median confidence, population average confidence (taking a population viability rating and providing only the average confidence of those that agree; if a population is not reached, taking the average), maximum confidence, weighted average, population maximum confidence, etc.
[0092] Another method used in the field of machine learning is transfer learning, in which a previously trained model is used as a starting point for training a new model. This is also called pre-training. Pre-training is widely used and allows new models to be built quickly. There are two types of pre-training. One embodiment of pre-training is ImageNet pre-training. Most model architectures are provided with a set of pre-trained weights using the standard image database ImageNet. Although this is not specific to medical images and contains 1000 different types of objects, it provides a way for the model to already learn to identify shapes. The classifier for the 1000 objects is completely removed, and a new classifier for viability replaces it. This type of pre-training outperforms other initialization strategies. Another embodiment of pre-training is custom pre-training, which uses previously trained embryonic models from studies with different outcome sets or different images (PGS instead of viability, or randomly assigned outcomes). These models provide only a small benefit to classification.
[0093] For models that have not undergone pretraining, or for new layers added after pretraining, such as classifiers, weights must be initialized. The initialization method can affect the success of training. For example, setting all weights to 0 or 1 results in very poor performance. A uniform distribution of random numbers or a Gaussian distribution of random numbers also represent commonly used options. These are often combined with regularization methods such as the Xavier or Kaiming algorithms. This addresses the problem that nodes in a neural network can become "trapped" in a particular state by becoming saturated (close to 1) or dead (close to 0), making it difficult to determine the direction in which to adjust the weight associated with that particular neuron. This is particularly prevalent when introducing hyperbolic tangent or sigmoid functions, and is addressed by Xavier initialization.
[0094] In the Xavier initialization protocol, the weights of a neural network are randomized so that each layer's input to the activation function is not too close to either the extremes of saturation or the extremes of dead. However, using ReLU works better, and different initializations offer smaller advantages, such as Kaiming initialization. Kaiming initialization is more suitable when ReLU is used as the nonlinear activation profile for neurons. It effectively achieves the same process as Xavier initialization.
[0095] In deep learning, various free parameters are used to optimize model training on the validation set. One key parameter is the learning rate, which is determined by how much the underlying neuron weights are adjusted after each batch. When training a selection model, overtraining or overfitting the data should be avoided. This occurs when the model has too many parameters to fit and essentially "memorizes" the data, trading generalizability for accuracy on the training or validation set. This is avoided because generalizability is the true measure of whether the model has accurately identified and perfectly fitted the training set, even amidst the noise in the data, the true underlying parameters that indicate embryonic health.
[0096] During the validation and testing phase, the success rate can suddenly drop due to overfitting during the training phase. This can be recovered through various strategies, including slowing or decaying the learning rate (e.g., halving the learning rate every n epochs), or using cosine annealing incorporating batch normalization or the tensor initialization or pretraining methods described above and adding noise such as dropout layers. Batch normalization is used to combat vanishing or exploding gradients and improves the stability of training large models, resulting in improved generalization. Dropout regularization effectively simplifies the network by introducing a random opportunity to set all incoming weights to zero within the rectifier's acceptance range. Introducing noise effectively ensures that the remaining rectifier accurately fits the representation of the data without relying on overspecialization. This allows the DNN to generalize more effectively and become less sensitive to specific values of the network's weights. Similarly, batch normalization improves the training stability of very deep neural networks, allowing for faster learning and better generalization by shifting the input weights to zero mean and unit variance as a precursor to the rectification stage.
[0097] When performing deep learning, the methodology for modifying neuron weights to achieve acceptable classification involves the need to specify an optimization protocol. That is, many techniques need to be specified for a given definition of "accuracy" or "loss" (discussed below), exactly how much weights should be adjusted, and how learning rate values should be used. Suitable optimization techniques include stochastic gradient descent (SGD) with momentum (and / or Nesterov's accelerated gradient method), adaptive gradient with delta (Adadelta), adaptive moment estimation (Adam), root-mean-square propagation (RMSProp), and the limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) algorithm. Of these, SGD-based techniques generally outperformed other optimization techniques. Typical learning rates for phase-contrast microscopy images of human embryos ranged from 0.01 to 0.0001. However, learning rate depends on the batch size, which in turn depends on hardware capacity. For example, larger GPUs allow for larger batch sizes and faster training speeds.
[0098] Stochastic gradient descent (SGD) with momentum (and / or Nesterov's accelerated gradient) represents the simplest and most commonly used optimization algorithm. Gradient descent algorithms typically calculate the gradient (slope) of a given weight's influence on accuracy. While this is slow if the gradient needs to be calculated for the entire dataset to make weight updates, stochastic gradient descent performs updates one at a time, for each training image. While this can introduce fluctuations in the overall target accuracy or loss achieved, it tends to generalize better than other methods because it can dive into new regions of the loss parameter landscape and find a new minimum loss function. SGD performs well against pronounced loss landscapes in difficult problems such as embryo selection. SGD can have trouble navigating asymmetric loss function surface curves that are steeper on one side than the other; this can be compensated for by adding a parameter called momentum. This accelerates SGD in that direction by adding an extra fraction to the weight updates derived from the previous state, helping to dampen high fluctuations in accuracy. A development of this method is to also include an estimate of the position of the weights in the next state; this development is known as Nesterov's accelerated gradient method.
[0099] Adaptive gradients with delta (Adadelta) is an algorithm for adapting the learning rate to the weights themselves, with smaller updates for frequently occurring parameters and larger updates for infrequently occurring features, making it well suited to sparse data. While this can suddenly slow down the learning rate after a few epochs across the entire dataset, adding a delta parameter to limit the window allowed the accumulated past gradients to a constant size. However, this process makes the default learning rate redundant, and the additional degrees of freedom of the free parameter provide some control in finding the best global selection model.
[0100] Adaptive moment estimation (Adam) stores exponentially decaying averages of both past squared and non-squared gradients and incorporates both into the weight updates. This has the effect of providing "friction" to the direction of the weight updates and is suitable for problems with relatively shallow or flat loss minima without large fluctuations. In embryonic selection models, training with Adam tends to perform well on the training set, but often overtrains and is not as suitable as SGD with momentum methods.
[0101] Root-mean-square propagation (RMSProp) is the adaptive gradient optimization algorithm described above, and is nearly identical to Adadelta, except that the update term for the weights divides the learning rate by an exponentially decaying average of the squared gradients.
[0102] The limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) algorithm. Though computationally intensive, the L-BFGS algorithm actually estimates the curvature of the loss landscape, rather than other methods that attempt to compensate for this lack of estimation with additional terms. For small datasets, it tends to outperform Adam, but does not necessarily outperform SGD in terms of speed or accuracy.
[0103] In addition to the above methods, it is also possible to include non-uniform learning rates. That is, the learning rate of the convolutional layers can be specified to be much larger or smaller than the learning rate of the classifier. This is useful in the case of pre-trained models, where changes to the filters below the classifier should be kept "frozen" and the classifier retrained, so that the pre-training is not undone by further retraining.
[0104] Although an optimization algorithm specifies how to update weights given a particular loss or accuracy measure, in some embodiments, loss functions are modified to incorporate distributional effects. These may include cross-entropy (CE) loss, weighted CE, residual CE, inference distributions, or custom loss functions.
[0105] Cross-entropy loss is a commonly used loss function that tends to outperform simple mean squared error between ground truth and predictions. When the network's results are passed through a softmax layer, as is the case here, the distribution of cross-entropy leads to better accuracy. This is naturally to maximize the chances of correctly classifying the input data by de-emphasizing extreme outliers. For an input array, a batch, representing a batch of images, and a class representing viable or non-viable, the cross-entropy loss is:
[0106]
number
[0107]
number
[0108]
number
[0109] If the data has class bias, i.e., more viable examples than non-viable examples (or vice versa), the loss function should be weighted proportionally so that misclassifying elements of the less abundant class is more heavily penalized. This is achieved by pre-multiplying the right-hand side of equation (2) by a factor:
[0110]
number
[0111] In some embodiments, an inferred distribution may be used. While it is important to seek a high level of accuracy in embryo classification, it is also important to seek a high level of transferability in the model. That is, it is often useful to understand the distribution of scores and understand that while high accuracy is an important goal, confidently distinguishing between viable and nonviable embryos is an indicator that the model generalizes well to the test set. Because accuracy on the test set is often used to cite comparisons to clinical benchmarks, such as the accuracy of an embryologist's classification of the same embryos, ensuring generalizability should also be incorporated into the epoch-by-epoch, batch-by-batch evaluation of the model's success.
[0112] In some embodiments, a custom loss function is used. In one embodiment, we customized how the loss function is defined so that the optimization surface is modified to make the global optimum more apparent, thus improving the robustness of the model. To achieve this, a new term that maintains differentiability, called the residual term, defined in terms of the network weights, is added to the loss function. This encodes the collective difference in the predictions from the model and the target outcome for each image and includes it as an additional contribution to the normal cross-entropy loss function. The formula for the residual term is, for N images:
[0113]
number
[0114] This custom loss function considers well-spaced clusters of viable and non-viable embryo scores to be consistent with the improved loss estimate. Note that this custom loss function is not specific to embryo detection applications and can be used in other deep learning models.
[0115] In some embodiments, the models are combined to produce a more robust final AI model 100, i.e., deep learning and / or computer vision models are combined together to contribute to the overall prediction of embryo viability.
[0116] In one embodiment, an ensemble approach is used. First, a well-performing model is selected. Then, each model "votes" on one of the images (using augmentations or otherwise), and the voting strategy that yields the best results is selected. Exemplary voting strategies include maximum confidence, average, majority average, median, mean confidence, median confidence, majority average confidence, weighted average, majority maximum confidence, etc. Once a voting strategy is selected, an evaluation method for the combination of augmentations must also be selected, which, as described above, describes how each of the rotations should be processed by the ensemble. In this embodiment, the final AI model 100 can thus be defined as a collection of AI models trained using deep learning and / or computer vision models, along with a mode that encodes a voting strategy that defines how the results of the individual AI models will be combined, and an evaluation mode that defines how the augmentations, if any, will be combined.
[0117] Model selection was performed so that their results contrasted with each other, i.e., their results were as independent as possible and scores were well distributed. This selection procedure was performed by examining which images in the test set were correctly identified for each model. When comparing two models, if the sets of correctly identified images are very similar or the scores provided by each model are similar to each other for a given image, the models are not considered to be contrasting models. However, if there is little overlap between the two sets of correctly identified images or the scores provided for each image are significantly different from each other, the models are considered to be contrasting. This procedure effectively evaluates whether the distributions of embryo scores in the test set for two different models are similar. The contrast criterion drives model selection with diverse predicted outcome distributions for different input images or segmentations. This method ensures translatability by avoiding the selection of models that performed well only on a specific clinical dataset, thus preventing overfitting. In addition, model selection can also use a diversity criterion. The diversity criterion drives model selection to include different model hyperparameters and configurations. The reason is that in practice, similar model settings may yield similar prediction results and therefore may not be useful for the final ensemble model.
[0118] In one embodiment, this can be done by using a counting approach and specifying a threshold similarity, such as 50%, 75%, or 90% overlapping images in the two sets. In other embodiments, the scores in a set of images (e.g., the viable set) can be summed, the two sets (sums) compared, and ranked similarly if the sum of the two is below a threshold amount. Statistical comparisons can also be used, for example, by considering the number of images in the sets or otherwise comparing the distribution of images in each of the sets.
[0119] In other embodiments, individual AI models can be combined using a distillation method. In this approach, an AI model is used as a teacher model to train a student model. Selection of individual AI models can be done using diversity and contrast criteria, as discussed for ensemble methods. Still other methods of selecting the best model from various models or combining outputs from multiple models into a single output may be used.
[0120] An embodiment of an ensemble-based embryo viability assessment model was created, and two validation (or benchmarking) studies were conducted in an IVF clinic to evaluate the performance of the embryo viability assessment model described herein compared with practicing embryologists. For ease of reference, this will be referred to as the ensemble model. These validation studies showed that the embryo viability assessment model demonstrated an improvement in accuracy of over 30% in identifying embryo viability when directly compared with world-leading embryologists. Thus, this study validates the ability of an embodiment of the ensemble model described herein to inform and support embryologists' selection decisions, which is expected to contribute to improved IVF outcomes for couples.
[0121] The first study was a pilot study conducted at an Australian clinic (Monash IVF), and the second study was conducted across multiple clinics and geographic locations. The studies evaluated the ability of one embodiment of the ensemble-based embryo viability assessment model as described to predict day 5 embryo viability, as measured by clinical pregnancy.
[0122] For each clinical study, each patient in the IVF process may have multiple embryos to select from. The viability of each of these embryos was assessed and scored using one embodiment of the embryo viability assessment model described herein. However, the accuracy of the model can be verified using only embryos that have been implanted and for which the pregnancy outcome (e.g., fetal heartbeat detected in the initial ultrasound scan) is known. Thus, the full dataset contains images of embryos implanted in patients with associated known outcomes against which the accuracy (and therefore performance) of the model can be verified.
[0123] To further enhance validation rigor, some of the images used for validation include embryologist scores for embryo viability. In some cases, embryos scored as "non-viable" may still be implanted if they are still the most preferred embryo option and / or at the patient's request. This data allows for a direct comparison of how well the ensemble model performs compared to the embryologist. Both the ensemble model and embryologist accuracy are measured as the percentage of the number of embryos scored as non-viable and with an unsuccessful pregnancy outcome (true positive) over the number of embryos scored as non-viable and with an unsuccessful pregnancy outcome (true negative), divided by the total number of embryos scored. This approach is used to verify whether the ensemble model performs as well or better when directly compared to the lead embryologist. Note that not all images in the dataset have a corresponding embryologist score.
[0124] To directly compare the accuracy of the selection model with the current manual methods utilized by embryologists, the following interpretation of the embryologist's score for each clinic is used for the degree of growth that is at least blastocyst ("BL" in the Ovation Fertility designation, or "XB" in the Midwest Fertility Specialists designation): Embryos listed as the cell stage (e.g., 10 cells), as compaction from the cell stage to a morula, or as a blastocoel-forming morula (blastocoel cavity less than 50% of total volume at day 5 after IVF) are considered likely to be non-viable.
[0125] Letter grades indicating the quality of the zona pellucida cavity (first letter) and trophectoderm (second letter) are placed into embryo quality bands and identified by the embryologist. A division is then made to indicate whether the embryo is deemed likely non-viable or viable, using Table 1 below. Bands 1 through 3 are considered likely viable, while bands 4 and above are considered likely non-viable. In band 6, if any letter score is worse than a "C," the embryo is considered likely non-viable. In band 7, a score of "1XX" from Midwest Fertility Specialists indicates an early blastocyst with early (large) trophectoderm cells and no discernible zona pellucida cavity, and is considered likely non-viable.
[0126] [Table 1] A set of approximately 20,000 embryo images taken on day 5 after IVF was acquired, along with associated pregnancy and preimplantation genetic screening (PGS) results, and demographic information, including patient age and clinic geographic location. Clinics that contributed data to this study were: Repromed (Adelaide, SA, Australia) as part of Monash IVF Group (Melbourne, VIC, Australia), Ovation Fertility (Austin, TX, USA), San Antonio IVF (San Antonio, TX, USA), Midwest Fertility Specialists (Carmel, IN, USA), Institute for Reproductive Health (Cincinnati, OH, USA), Fertility Associates (Auckland, Hamilton, Wellington, Christchurch, and Dunedin, New Zealand), Oregon Reproductive Medicine (Portland, OR, USA), and Alpha Fertility Centre (Petaling Jaya, Selangor, Malaysia).
[0127] The generation of AI models for use in the testing proceeded as follows. First, various model architectures (or model types) were generated, and each AI model was trained using various settings of model parameters and hyperparameters, including input image resolution, selection of optimization algorithm, learning rate and scheduling, momentum value, dropout, and weight initialization (pre-training). Initial filtering was performed to select models that exhibited stability (stable accuracy throughout the training process), transferability (stable accuracy between the training and validation sets), and predictive accuracy. Predictive accuracy was determined as the weighted average accuracy across both class types of embryos. We investigated which models provided the best validation accuracy, total accuracy, and balanced accuracy for both viable and nonviable embryos. In one embodiment, the use of ImageNet pre-trained weights demonstrated improved performance in these quantities. Evaluation of loss functions showed that weighted CE and residual CE loss functions generally outperformed other models.
[0128] The following models were then divided into two groups: the first group included further image segmentation (identification of the zona pellucida or IZC), and the second group used the entire, unsegmented image (i.e., full embryo models). Models trained on images in which the IZC was masked, exposing the zona pellucida region, were denoted as zona pellucida models. Models trained on images with the zona pellucida masked (denoted as IZC models) and on full embryo images were also considered for training. To provide diversity and maximize performance on the validation set, a group of models was selected encompassing contrasting architectures and preprocessing methods.
[0129] The final ensemble-based AI model was an ensemble of the best-performing individual models, selected based on diversity and contrasting results. Well-performing individual models that exhibited different methodologies or extracted different biases from features obtained through machine learning were combined using various voting strategies based on the reliability of each model. Voting strategies evaluated included average, median, max, majority average vote, max confidence, mean, majority average, median, mean confidence, median confidence, majority average confidence, weighted average, majority max confidence, etc. In one embodiment, the majority average voting strategy is used as a test to see if it outperformed other voting strategies and provided the most stable model across all datasets.
[0130] In this embodiment, the final ensemble-based AI model includes eight deep learning models, four of which are zona pellucida models and four of which are whole embryo models. The configuration of the final model used in this embodiment is: One full-embryo ResNet-152 model trained using SGD with momentum=0.9, CE loss, learning rate 5.0e-5, a stepwise scheduler that halves the learning rate every 3 epochs, a batch size of 32, an input resolution of 224x224, and a dropout value of 0.1; One zona pellucida model, ResNet-152 model, trained using SGD with momentum=0.99, CE loss, learning rate 1.0e-5, a stepwise scheduler dividing the learning rate by 10 every 3 epochs, a batch size of 8, an input resolution of 299x299, and a dropout value of 0.1; Three transparent zone ResNet-152 models trained using SGD with momentum=0.99, CE loss, learning rate 1.0e-5, a stepwise scheduler dividing the learning rate by 10 every 6 epochs, a batch size of 8, an input resolution of 299x299, and a dropout value of 0.1, one of which was trained with a random rotation of any angle; One full-embryo DenseNet-161 model trained using SGD with momentum=0.9, CE loss, learning rate 1.0e-4, a stepwise scheduler that halves the learning rate every 5 epochs, a batch size of 32, an input resolution of 224x224, a dropout value of 0, and additionally trained with random rotations of any angle; One full-embryo DenseNet-161 model trained using SGD with momentum=0.9, CE loss, a learning rate of 1.0e-4, a stepwise scheduler that halves the learning rate every 5 epochs, a batch size of 32, an input resolution of 299x299, and a dropout value of 0; and One full-embryo DenseNet-161 model trained using SGD with momentum=0.9, residual CE loss, learning rate 1.0e-4, a stepwise scheduler that halves the learning rate every 5 epochs, a batch size of 32, an input resolution of 299x299, a dropout value of 0, and additionally trained with random rotations of any angle; is.
[0131] A diagram of the architecture corresponding to ResNet-152, which features prominently in the construction of the final model, is shown in Figure 8. The final ensemble model was then validated and tested on a blind test dataset as described in the Results section.
[0132] Accuracy measures used in assessing model performance on the data included sensitivity, specificity, overall accuracy, distribution of predictions, and comparison with embryologist scoring methods. For the AI model, an embryo viability score of 50% or greater was considered viable, and less than 50% was considered nonviable. Accuracy in identifying viable embryos (sensitivity) was defined as the number of embryos identified by the AI model as viable divided by the total number of known viable embryos that resulted in a positive clinical pregnancy. Accuracy in identifying nonviable embryos (specificity) was defined as the number of embryos identified by the AI model as nonviable divided by the total number of known nonviable embryos that resulted in a negative clinical pregnancy. Overall accuracy of the AI model was determined using a weighted average of sensitivity and specificity, and the percent improvement in accuracy of the AI model compared to the embryologist was defined as the difference in accuracy as a percentage of the original embryologist's accuracy (i.e., (AI_accuracy - embryologist_accuracy) / embryologist_accuracy).
[0133] Pilot study Monash IVF provided the ensemble model with approximately 10,000 embryo images and associated pregnancy and birth data for each image. Additional data provided included patient age, BMI, whether the embryos were freshly implanted or previously cryopreserved, and any fertility-related medical conditions. Data for some of the images included an embryologist's score for embryo viability. Preliminary training, validation, and analysis showed that the model's accuracy was significantly higher for day 5 embryos compared to day 4 embryos. Therefore, all day 4 embryos were removed, leaving approximately 5,000 images. The usable dataset for training and validation was 4,650 images. This initial dataset was split into three separate datasets. An additional 632 images were then provided, which were used as a second blind validation dataset. The final dataset for training and validation included the following: Training dataset: 3892 images; Validation dataset: Of 390 images, 70 (17.9%) had a successful pregnancy outcome, and 149 images contained an embryologist's score for embryo viability; Blind validation dataset 1: Of 368 images, 76 (20.7%) had a successful pregnancy outcome, and 121 images included an embryologist score for embryo viability; and Blind validation dataset 2: Of 632 images, 194 (30.7%) had a successful pregnancy outcome, and 477 images included an embryologist's score for embryo viability.
[0134] Not all images in the dataset have a corresponding embryologist score. The dataset sizes, as well as the subsets that contain embryologist scores, are listed below.
[0135] The ensemble-based AI model was applied to three validation datasets. The overall accuracy results for the ensemble model in identifying viable embryos are shown in Table 2. While the accuracy results for the two blind validation datasets are the primary accuracy metrics, the results for the validation dataset are shown for completeness. Accuracy for identifying viable embryos is calculated as the percentage of viable embryos (i.e., images with successful pregnancy outcomes) that the ensemble model was able to identify as viable (viability score of 50% or higher by the model) divided by the total number of viable embryos in the dataset. Similarly, accuracy for identifying nonviable embryos is calculated as the percentage of nonviable embryos (i.e., images with unsuccessful pregnancy outcomes) that the ensemble model was able to identify as nonviable (viability score of less than 50% by the model) divided by the total number of nonviable embryos in the dataset.
[0136] In the first phase of validation performed at Monash IVF, the trained embryo viability assessment model of the ensemble model was applied to two blind datasets of embryo images with known pregnancy outcomes, totaling 1,000 images (patients) combined. Figure 9 shows a plot 900 of the accuracy of one embodiment of the ensemble model in identifying embryo viability, according to one embodiment. The results show that the ensemble model 910 had an overall accuracy of 67.7% in identifying embryo viability across the two blind validation datasets. Accuracy was calculated by summing the number of embryos identified as viable, leading to a successful outcome, and the number of embryos identified as non-viable, leading to an unsuccessful outcome, and dividing by the total number of embryos. The ensemble model demonstrated a 74.1% accuracy 920 in identifying viable embryos and a 65.3% accuracy 930 in identifying non-viable embryos. This represents a significant improvement in accuracy in this large dataset of embryos already pre-selected by embryologists and implanted in patients, with only 27% resulting in a successful pregnancy outcome.
[0137] To further enhance validation rigor, a subset of images used for validation had an associated embryologist score for embryo viability (598 images). In some cases, an embryo scored as "non-viable" by the embryologist may still be implanted despite a low chance of success if it is considered the most favorable embryo option for the patient and / or at the patient's request. The embryo scores are used as a ground truth for the embryologist's viability assessment, allowing for a direct comparison of the ensemble model's performance against the lead embryologist.
[0138] The worst-case accuracy for blind validation dataset 1 or 2 is 63.2% for identifying viable embryos in blind dataset 1, 57.5% for identifying non-viable embryos in blind dataset 2, and 63.9% for overall accuracy in blind dataset 2.
[0139] Table 3 shows the overall average accuracy across both blind datasets 1 and 2, which is 74.1% for identifying viable embryos, 65.3% for identifying non-viable embryos, and 67.7% for the total accuracy across both viable and non-viable embryos.
[0140] The accuracy values in both tables are high considering that 27% of embryos result in a successful pregnancy outcome and the difficult task of the ensemble model to further classify embryo images that have already been analyzed by embryologists and selected as viable or preferable to other embryos in the same batch.
[0141] [Table 2]
[0142] [Table 3] Table 4 shows the accuracy of the model compared to that of the embryologist. The accuracy values differ from those in the table above because not all embryo images in the dataset have an embryo score; therefore, the results below are accuracy values for each dataset subset. The table shows that the model's accuracy is higher than that of the embryologist in identifying viable embryos. These results are illustrated in the bar graph 1000 of FIG. 10, with the ensemble results 1010 illustrated on the left and the embryologist results 1020 illustrated on the right.
[0143] [Table 4] Table 5 shows the number of times the model was able to correctly identify embryo viability compared to the number of times the embryologist was unable to, and vice versa. The results show that there were fewer instances where the embryologist was accurate and the model was inaccurate compared to the instances where the model was accurate and the embryologist was inaccurate. These results are illustrated in Figure 11. The results further validate the high level of performance and accuracy of the ensemble model's embryo viability assessment model.
[0144] [Table 5] Overall, the ensemble model achieved a total accuracy of 66.7% in identifying embryonic viability, while the embryologist achieved 51% accuracy based on their scoring method ( FIG. 10 ). The 15.7% additional accuracy represents a significant 30.8% performance (accuracy) improvement for the ensemble model compared to the embryologist (p=0.021, n=2, Student's t-test). Specifically, the results show that the ensemble model was able to correctly classify embryonic viability 148 times when the embryologist was inaccurate, while conversely, the embryologist correctly classified embryonic viability only 54 times when the ensemble model was inaccurate. FIG. 11 is a bar graph showing the accuracy of one embodiment of the ensemble model (bar 1110) compared to a world-leading embryologist (clinician) (bar 1120) in accurately identifying embryonic viability when the embryologist's assessment was inaccurate, compared to the embryologist's accuracy when the ensemble model's assessment was inaccurate. These results demonstrate a clear advantage of the ensemble model in identifying viable and nonviable embryos when compared to world-leading embryologists. Further validation tests were performed on embryo images from Ovation Fertility, with similar results.
[0145] Successful validation demonstrated that ensemble modeling approaches and techniques can be applied to embryo images to create a model that can accurately identify viable embryos and ultimately lead to improved IVF outcomes for couples. This model was then further tested in a larger cross-clinic study.
[0146] Cross-Clinic Study In a more general cross-clinic study following the Australian pilot study, over 10,000 embryo images were sourced from multiple demographics. Of these images, over 8,000 could be associated with embryologist scores for embryo viability. For training, each image needs to be labeled as viable or nonviable to enable deep learning and computer vision algorithms to identify patterns and features related to embryo viability.
[0147] In the first cross-clinic study, the available dataset of 2217 images (and associated outcomes) for developing the ensemble model will be split into three subsets in the same manner as the pilot study: a training dataset, a validation dataset, and a blind validation dataset. These studies include data sourced from the following clinics: Ovation Fertility Austin, San Antonio IVF, Midwest Fertility Specialists, and the Institute for Reproductive Health and Fertility Associates NZ. This includes: Training dataset: 1744 images - 886 non-viable, 858 viable; Validation dataset: 193 images - 96 non-viable, 97 viable; and Blind validation dataset 1: 280 images - 139 non-viable, 141 viable; It included.
[0148] After completing the training, validation, and blind validation phases, a second study will be performed on a complete, separate demographic sourced from the Oregon Reproductive Medicine Clinic. This dataset includes: Blind validation dataset 2: 286 images - 106 non-viable, 180 viable; It included.
[0149] The third study looked at EmbryoScope images provided by the clinic: Alpha Fertility Centre: EmbryoScope validation dataset: 62 images - 32 non-viable, 30 viable; Use In creating trained ensemble-based AI models, the same training dataset is used for each model that is trained, so that they can be compared in a consistent manner.
[0150] The final results for the ensemble-based AI model applied to the mixed demographic blind validation dataset are as follows: A summary of the overall accuracy can be seen in Table 6.
[0151] [Table 6] The distribution of inferences, displayed as histograms, is shown in FIGS. 12 and 13. FIG. 12 is a plot 1200 of the distribution of inference scores for viable embryos (successful clinical pregnancies) using an embodiment of an ensemble-based AI model when applied to the blind validation dataset from Study 1. The inferences are normalized between 0 and 1 and can be interpreted as confidence scores. Instances where the model is accurate are marked in boxes filled with a thick downward diagonal line (true positives 1220), and instances where the model is inaccurate are marked in boxes filled with a thin upward diagonal line (false negatives 1210). FIG. 13 is a plot 1300 of the distribution of inference scores for non-viable embryos (unsuccessful clinical pregnancies) using an embodiment of an ensemble-based AI model when applied to the blind validation dataset from Study 1. The inferences are normalized between 0 and 1 and can be interpreted as confidence scores. Examples where the model is accurate are marked in boxes filled with a thick downward diagonal line (true negatives 1320), and examples where the model is inaccurate are marked in boxes filled with a thin upward diagonal line (false positives 1310). There is a clear separation between the two groups. These histograms show a good separation between correctly and incorrectly identified embryo images, providing evidence that the model transfers well to the blind validation set.
[0152] FIG. 13 has a high peak in false positives 1310 (filled boxes with thin upward diagonal lines), which is not as pronounced in the equivalent histogram for false negatives in FIG. 12. The reason for this effect may be due to the presence of patient health factors, such as uterine scarring, that cannot be identified via the embryo image itself. The presence of these factors means that even an ideal embryo may not result in successful implantation. This also limits the upper limit of accuracy in predicting successful clinical pregnancy using embryo image analysis alone.
[0153] In embryo selection, it is widely believed that allowing nonviable embryos to be implanted (false positives) is preferable to endangering potentially healthy embryos (false negatives). Therefore, when obtaining the final ensemble-based AI model that forms the ensemble-based AI model, efforts were made to bias residual imprecision to prioritize minimizing false negatives, if possible. The final model would therefore have higher sensitivity than specificity, i.e., higher accuracy in selecting viable embryos over nonviable embryos. To bias the model to prioritize minimizing false negatives, models were selected for inclusion in the final ensemble-based AI model such that, if possible, the ensemble-based AI model's accuracy on the set of viable embryo images was higher than its accuracy on the set of nonviable embryo images. If a model could not be found that combined together to bias viability accuracy, additional parameters could be supplied during training, which would increase the penalty for incorrectly classifying a viable embryo.
[0154] Although total accuracy is useful for roughly assessing the overall effectiveness of a model, it necessarily averages out the complexities associated with different demographics. Therefore, it is directive to consider the breakdown of results into various major groups, as described below.
[0155] Study 1: Demographic cross-section To investigate the behavior of the ensemble-based AI model, the following demographic groups are considered: First, the accuracy on the dataset provided by Fertility Associates NZ is lower than that of the US-based clinic. This is likely due to the inherent diversity in the data from this clinic, which encompasses many different cities, camera filters, and brightness levels, across which the ensemble-based AI model must average. Further training of the AI on much larger datasets is expected to be able to account for camera diversity by incorporating it into the fine-tuning training dataset. Accuracy including and excluding the NZ data is shown in Tables 7 and 8.
[0156] Due to the small number of images from clinics Midwest Fertility Associates and San Antonio IVF, the sample sizes were individually too small to allow for reliable accuracy measurements, so their results are combined with those from Ovation Fertility Austin in Table 7.
[0157] [Table 7] A study of the effect of patient age on the accuracy of the ensemble-based AI model was also performed and is shown in Table 7. Embryo images corresponding to patients aged 35 years and older were found to be more accurately classified. When the age cutoff was raised to 38 years, accuracy improved again, indicating that the ensemble-based AI model is more sensitive to morphological features that become more pronounced with age.
[0158] [Table 8] We also considered whether the embryos had been treated with a hatching or no-hatching protocol prior to transfer. We found that hatched embryos, which exhibited larger morphological features, were more easily identified by AI than unhatched embryos, although specificity was reduced in the former case. This is likely a result of the fact that ensemble-based AI models trained on a mixed dataset of hatched and unhatched embryos tend to associate successfully hatched embryos with viability.
[0159] Study 1: Comparison of embryologists' rankings A summary of the accuracy of the ensemble-based AI model and embryologists can be seen in Tables 9 and 10 for the same demographic classifications considered in Section 5A. In this study, only embryo images with corresponding embryologist scores are considered.
[0160] The percentage improvement of the ensemble-based AI model compared to the embryologist in accuracy is quoted, as determined by the difference in accuracy as a percentage of the original embryologist's accuracy ((AI_Accuracy - Embryologist_Accuracy) / Embryologist_Accuracy). Although the improvement across the total number of images was 31.85%, the improvement was found to be highly variable across specific demographics, as the improvement factor is highly sensitive to the embryologist's ability on each given dataset.
[0161] In the case of Fertility Associates NZ, embryologists performed significantly better than other demographics, leading to an improvement of only 12.37% when using the ensemble-based AI model. In cases such as Ovation Fertility Austin, where the ensemble-based AI model performed very well, the improvement was as high as 77.71%. The comparison of the performance of the ensemble-based AI model compared to the embryologists is also reflected in the total number of correctly rated images where the comparator rated the same images inaccurately, as seen in the last two columns of both Tables 9 and 10.
[0162] [Table 9] An alternative study comparing the effectiveness of the ensemble-based AI model and the embryologist's assessment can be performed when the embryologist's score has a number or nomenclature that represents the embryo's rank in terms of embryo progression or arrest (number of cells, compaction, morula, blastocoel formation, early blastocyst, complete blastocyst, or hatched blastocyst). A comparison of embryo rank can be performed by dividing the AI inferences into five equal bands (from lowest to highest inference) labeled 1–5 while equating the embryologist's assessment with a numeric score ranging from 1 to 5. When both the ensemble-based AI model and the embryologist's scores are expressed as integers from 1 to 5, a comparison of ranking accuracy can be performed as follows:
[0163] If a given embryo image is given the same rank by the ensemble-based AI model and the embryologist, this is accepted as a match. However, if the ensemble-based AI model gives a higher rank than the embryologist and the ground truth result is recorded as viable, or if the ensemble-based AI model gives a lower rank than the embryologist and the ground truth result is recorded as non-viable, this result is accepted as a model accuracy. Similarly, if the ensemble-based AI model gives a lower rank than the embryologist and the ground truth result is recorded as viable, or if the ensemble-based AI model gives a higher rank and the result is recorded as non-viable, this result is accepted as a model inaccuracy. A summary of the percentage of images rated as match, model accuracy, or model inaccuracy for the same demographic classifications considered above can be seen in Tables 11 and 12. The ensemble-based AI model is considered to have performed well on the dataset if the percentage of model accuracy is high and the percentages of match and model inaccuracy are low.
[0164] [Table 10]
[0165] [Table 11]
[0166] [Table 12] A visual representation of the distribution of ranks obtained from the embryologist and ensemble-based AI model across the full blind dataset for Study 1 can be seen in the histograms of Figures 14 and 15, respectively. Figure 14 is a histogram 1400 of ranks obtained from embryologist scores across the full blind dataset, and Figure 15 is a histogram 1500 of ranks obtained from an embodiment of the ensemble-based AI model inference across the full blind dataset.
[0167] Figures 14 and 15 show different distribution shapes. While the embryologist's scores dominate around a rank value of 3, dropping off sharply for low scores of 1 and 2, the ensemble-based AI model has a more even distribution of scores around values of 2 and 3, with a rank of 4 being the dominant score. Figure 16 is extracted directly from the inferred scores obtained from the ensemble-based AI model, which are shown as a histogram in Figure 13 for comparison. The ranks in Figure 12 are a coarser version of the scores in Figure 13. The finer distribution in Figure 16 shows a clear separation between scores below 50% (predicted non-viable) 1610 and scores above 50% (predicted viable) 1620. This suggests that the ensemble-based AI model offers greater granularity in embryo ranking than standard scoring methods, allowing more decisive selection to be achieved.
[0168] Study 2 - Secondary blind validation In Study 2, embryo images were provided by another clinic, Oregon Reproductive Medicine, and used as secondary blind validation. The total number of images with associated clinical pregnancy outcomes was 286, similar in size to the blind validation dataset in Study 1. The final results for the ensemble-based AI model applied to the mixed-demographic blind validation set can be seen in Table 13. In this blind validation, there was only a decrease in accuracy of 66.43% - 62.64% = 3.49% compared to Study 1, indicating that the model transfers across to the secondary blind set. However, the decrease in accuracy is not uniform across nonviable and viable embryos. Although sensitivity remains stable, specificity decreases. In this study, 183 low-quality images provided by an older (>1 year) Pixelink® camera (that did not meet image quality criteria) were removed before the start of the study to prevent low-quality images from affecting the ensemble-based AI model, which accurately predicts embryo viability.
[0169] [Table 13] To further explore this point, we performed another study in which we successively distorted embryo images by introducing uneven cropping, scaling (blurring), or adding compression noise (e.g., JPEG artifacts). In both cases, we found that the reliability of the ensemble-based AI model predictions decreased as the artifacts increased. Furthermore, we found that the ensemble-based AI model tended to assign a non-viable prediction to distorted images. This makes sense given the ensemble-based AI model's inability to distinguish between images of damaged embryos and damaged images of normal embryos. In both cases, the distortions were identified by the ensemble-based AI model, making it more likely to assign a non-viable prediction to the images.
[0170] As a validation of this analysis, the ensemble-based AI model was applied to only 183 Pixelink camera images removed from the main high-quality image set from Oregon Reproductive Medicine, and the results are shown in Table 14.
[0171] [Table 14] It is clear from Table 14 that for distorted and poor-quality images (i.e., when image quality assessment fails), not only does the ensemble-based AI model perform poorly, but a larger proportion of images are assigned to nonviable predictions. Further analysis of the ensemble-based AI model's behavior in alternative camera settings and how to handle such artifacts to improve results are discussed below. The distributions of inferences, displayed as histograms 1700 and 1800, are shown in Figures 17 and 18. Just as in Study 1, both Figures 17 and 18 show a clear divide between accurate predictions (1720; 1820; boxes filled with thick downward diagonal lines) and inaccurate predictions (1710; 1810; boxes filled with thin upward diagonal lines) for both viable and nonviable embryos. The shapes of the distributions between Figures 17 and 18 are also similar to each other, but the false positive rate is higher than that for false negatives.
[0172] Study 3 - EmbryoScope Validation Study 3 explores the potential performance of ensemble-based AI models on datasets sourced from completely different camera settings. A limited number of EmbryoScope images were obtained from the Alpha Fertility Center with the intention of testing an ensemble-based AI model trained primarily on phase-contrast images. The EmbryoScope images have a clear, bright ring around the embryo resulting from the incubator lamp and a dark area outside this ring, which is not present in the typical phase-contrast images from Study 1. Applying the model to EmbryoScope images without any additional processing resulted in variability in predictions, with a high percentage of images predicted as nonviable, resulting in a high false-negative rate and low sensitivity, as shown in Table 15. However, using computer vision imaging techniques, applying a coarse first pass to approximate the images to their expected morphology resulted in a significant rebalancing of the inferences and improved accuracy.
[0173] [Table 15] Although this dataset is small, it nevertheless provides evidence that computer vision techniques that reduce variability in image morphology can be used to improve the generalizability of ensemble-based AI models. A comparison with embryologists was also performed. Although scores were not provided directly by the Alpha Fertility Center, we found that the conservative assumption that embryos were predicted to be likely viable (to avoid false negatives) resulted in accuracies very similar to those of true embryologists in Study 1. Therefore, by making this assumption, a comparison of the accuracy of the ensemble-based AI model with that of embryologists can also be performed in the same way, as shown in Table 16. In this study, an improvement rate of 33.33% was observed, similar to the overall improvement rate of 31.85% obtained from Study 1.
[0174] [Table 16] Distributions of inferences can also be obtained in this study, as shown in Figures 19 and 20. Figure 19 is a plot 1900 of the distribution of inference scores for viable embryos (successful clinical pregnancies) using an ensemble-based AI model (false negatives 1910, boxes filled with thin upward diagonal lines; true positives 1920, boxes filled with thick downward diagonal lines). Figure 20 is a plot 2000 of the distribution of inference scores for nonviable embryos (successful clinical pregnancies) using an ensemble-based AI model (false negatives 1220, boxes filled with thin upward diagonal lines; true positives 2020, boxes filled with thick downward diagonal lines). Although the limited study size (62 images) does not allow the distributions to be very clear, it can nevertheless be observed that in this case the gap between accurate (1920, 2020) and inaccurate (1910, 2010) predictions for both viable and nonviable embryos is much less clear. This would be expected for images that exhibit additional features as artifacts entirely different from the EmbryoScope camera settings. These additional artifacts effectively add noise to the image, making it more difficult to extract relevant features indicative of embryo health.
[0175] Furthermore, accuracy in the viable category is significantly lower than in the nonviable category, resulting in a high false negative rate. However, we found that this effect was much reduced even after preliminary computer vision processing of the images, providing evidence for improved handling of images from different camera sources. In addition, the addition of EmbryoScope images during subsequent training or fine-tuning phases is also expected to lead to improved performance.
[0176] summary The effectiveness of AI models, including deep learning and computer vision models, for predicting embryo viability based on microscopic images was investigated in an Australian pilot study and three cross-clinic studies to develop a general ensemble-based AI model.
[0177] A pilot study involving a single Australian clinic was able to yield an overall accuracy of 67.7% in identifying embryo viability, 74.1% accuracy for viable embryos, and 65.3% accuracy for nonviable embryos. This represents a 30.8% improvement in embryologist classification rates. The success of these results prompted a more thorough cross-clinic study.
[0178] In three separate cross-clinic studies, we developed, validated, and tested a general AI selection model in various demographics from different clinics across the United States, New Zealand, and Malaysia. Study 1 found that the ensemble-based AI model was able to achieve high accuracy when compared with embryologists from each of the clinics, with an average improvement rate of 31.85% in the cross-clinic blind validation study, similar to the improvement rate in the Australian pilot study. In addition, the distribution of inference scores obtained from the ensemble-based AI model showed a clear gap between accurate and inaccurate predictions for both viable and nonviable embryos, providing evidence that the model transfers correctly to prospective blind datasets.
[0179] A comparative study with embryologist scores was developed to consider the impact of embryo rank ordering. By converting the ensemble-based AI model's inference and the embryologist's ranks to integers between 1 and 5, we were able to directly compare how the ensemble-based AI model performed compared to the embryologist in ranking embryos from most viable to least viable. We found that the ensemble-based AI model again outperformed the embryologist, giving 40.08% of images an improved rank, while only 25.19% of images received a worse rank and 34.73% of images showed no change in their rank.
[0180] The ensemble-based AI model was applied to a second blind validation set, demonstrating accuracy within a few percent of Study 1. The ability of the ensemble-based AI model to function on damaged or distorted images was also evaluated. Images that did not match standard phase-contrast microscopy images, or images with low quality, blurry, compressed, or poorly cropped images, were found to be scored as likely nonviable, reducing the ensemble-based AI model's reliability in predicting embryo images.
[0181] To understand the issues with different camera hardware and how they affect the results of the study, a dataset of EmbryoScope images was obtained, and it was found that the ensemble-based AI model, when simply applied to this dataset, did not reach the high accuracy achieved in the original set in Study 1. However, preliminary data cleaning of the images to address artifacts and reduce noise systematically present in EmbryoScope images significantly improved the results, bringing the accuracy of the ensemble-based AI model much closer to its optimum in Study 1. Due to the ability of ensemble-based AI models to be improved by incorporating larger and more diverse datasets into the training process, and thus fine-tuning the model so that it can self-improve over time, the three studies in this application provide compelling evidence for the effectiveness of AI models as an important and essential tool for robust and consistent assessment of embryo viability in the near future.
[0182] Furthermore, although the above examples use phase-contrast images from an EmbryoScope system and an optical microscope, further testing indicates that the method can be used with images captured using a variety of imaging systems. This testing indicates that the method is robust to a variety of image sensors and images (i.e., more than just fetoscopes and phase-contrast images), including images extracted from video and time-lapse systems. When using images extracted from video and time-lapse systems, a reference capture time point can be defined, and the image extracted from such a system can be the image closest to this reference capture time point or the first image captured after the reference time. Quality assessment can be performed on the images to ensure that the selected images meet minimum quality standards.
[0183] Embodiments of methods and systems have been described for the computational generation of AI models configured to generate embryo viability scores from images using one or more deep learning models. Given a new set of embryo images for training, a new AI model for estimating embryo viability can be generated by segmenting the images and identifying the zona pellucida and IZC regions, which annotate the images to their key morphological components. Next, at least one zona pellucida deep learning model is trained on the zona pellucida mask images. In some embodiments, multiple AI models, including deep learning models and / or computer vision models, are generated, and models that demonstrate stability, transferability from a validation set to a blind test set, and retained predictive accuracy are selected. These AI models can be combined using, for example, an ensemble model, which selects models to be combined based on contrast and diversity criteria using a confidence-based voting strategy. Once a suitable AI model is trained, it can be deployed to estimate the viability of newly collected images. This can be provided as a cloud service, allowing an IVF clinic or embryologist to upload captured images and obtain a viability score to assist in deciding whether to implant an embryo or not, or in the case of multiple embryos available, in selecting which embryo(s) are most likely to be viable. Deployment may involve exporting the model coefficients and model metadata to a file and then loading it onto another computing system to process new images, or reconfiguring the computing system to receive new images and generate viability estimates.
[0184] The implementation of ensemble-based AI models includes many options, and the embodiments described herein include several novel and advantageous features: Image pre-processing steps can be performed, such as segmentation to identify zona pellucida and IZC regions, object detection, image normalization, image cropping, and image cleaning, such as removing old or non-matching images (e.g., images with artifacts).
[0185] In relation to the deep learning model, the use of segmentation to identify the zona pellucida had a significant effect, with the final ensemble-based AI model featuring four zona pellucida models. The final model included an ensemble of eight deep learning AI models, and the additional deep learning models were generally found to outperform the computer vision models. However, useful results can still be produced using a single AI model based on the zona pellucida image, or an ensemble (or similar) AI model that includes a combination of deep learning and CV models. Therefore, the use of several deep learning models in which segmentation is performed prior to deep learning is preferred, helping to create contrasting deep learning models for use in the ensemble-based AI model. Image augmentation was also found to improve robustness. Some architectures that performed well included ResNet-152 and DenseNet-161 (although other variants can also be used). Similarly, stochastic gradient descent generally outperformed all other optimization protocols for varying neuron weights in nearly all tests (followed by Adam). The use of a custom loss function that modified the optimization surface to more clearly reveal the global optimum improved robustness. Randomizing the pre-training dataset, specifically checking that the dataset distribution was even (or similar) across the test and training sets, was also found to have a significant effect. Images of viable embryos are highly diverse, so checking randomization provides robustness to the effects of diversity. Using a selection process to choose contrasting models (i.e., their results are as independent as possible and scores are well distributed) to build ensemble-based AI models also improved performance. This can be assessed by examining the overlap in the sets of viable images for the two models. Prioritizing the reduction of false negatives (i.e., data cleansing) also helps improve accuracy.As described herein, in the case of embryo viability assessment models, models using images taken on day 5 after in vitro fertilization outperformed models obtained using images taken earlier (e.g., day 4 or earlier).
[0186] AI models using computer vision and deep learning methods could be generated using one or more of these advantageous features and applied to other image sets besides embryos. Referring to FIG. 1, embryo model 100 could be replaced with another model, trained, and used on other image data, whether medical in nature or not. The method may also be more general to deep learning-based models, including ensemble-based deep learning models. These could be trained and implemented using systems such as those illustrated in FIGS. 3A and 3B and described above.
[0187] Models trained as described herein can be usefully deployed to classify new images, thereby assisting embryologists in making implantation decisions and thus increasing success rates (i.e., pregnancies). Extensive testing of one embodiment of an ensemble-based AI model was conducted, in which the ensemble-based AI model was configured to generate embryo viability scores for embryos from images of embryos taken on day 5 after IVF. The testing demonstrated that the model clearly distinguished between viable and nonviable embryos (see FIG. 13), and Tables 10-12 and FIGS. 14-16 illustrate the model's outperformance compared to embryologists. Notably, as illustrated in the study above, one embodiment of the ensemble-based AI model was found to have high accuracy in identifying both viable (74.1%) and nonviable (65.3%) embryos, significantly outperforming experienced embryologists in assessing image viability by more than 30%.
[0188] Those skilled in the art will understand that information and signals may be represented using any of a variety of technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.
[0189] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software or instructions, middleware, platforms, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.
[0190] The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two, including in a cloud-based system. For a hardware implementation, processing may be implemented within one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, or other electronic units designed to perform the functions described herein, or a combination thereof. Various middleware and computing platforms may also be used.
[0191] In some embodiments, the processor module includes one or more central processing units (CPUs) or graphics processing units (GPUs) configured to perform some of the method steps. Similarly, a computing device may include one or more CPUs and / or GPUs. The CPU may include an input / output interface, an arithmetic logic unit (ALU), and a control unit and program counter element that communicate with input and output devices via the input / output interface. The input / output interface may include a network interface and / or a communication module for communicating with a comparable communication module in another device using a predetermined communication protocol (e.g., Bluetooth, Zigbee, IEEE 802.15, IEEE 802.11, TCP / IP, UDP, etc.). The computing device may include a single CPU (core) or multiple CPUs (multi-core), or multiple processors. The computing device is typically a cloud-based computing device using a GPU cluster, but may also be a parallel processor, vector processor, or distributed computing device. Memory is operatively coupled to the one or more processors and may include RAM and ROM components and may be provided internal or external to the device or processor module. The memory may be used to store an operating system and additional software modules or instructions. One or more processors may be configured to load and execute the software modules or instructions stored in the memory.
[0192] A software module, also known as a computer program, computer code, or instructions, may comprise numerous source code or object code segments or instructions and may reside in any computer-readable medium, such as RAM memory, flash memory, ROM memory, EPROM memory, registers, a hard disk, a removable disk, a CD-ROM, a DVD-ROM, a Blu-ray disc, or any other form of computer-readable medium. In some aspects, a computer-readable medium may include a non-transitory computer-readable medium (e.g., a tangible medium, etc.). Additionally, in other aspects, a computer-readable medium may include a transitory computer-readable medium (e.g., a signal, etc.). Combinations of the above should also be included within the scope of computer-readable media. In another aspect, a computer-readable medium may be integral to a processor. The processor and the computer-readable medium may reside in an ASIC or related device. The software codes may be stored in a memory unit, and the processor may be configured to execute them. The memory unit may be implemented within the processor or external to the processor, in which case it may be communicatively coupled to the processor via various means known in the art.
[0193] Furthermore, it should be appreciated that modules and / or other suitable means for performing the methods and techniques described herein can be downloaded and / or otherwise obtained by a computing device. For example, such a device can be coupled to a server to facilitate the transfer of means for performing the methods described herein. Alternatively, the various methods described herein can be provided via storage means (e.g., RAM, ROM, physical storage media such as a compact disk (CD) or floppy disk, etc.), such that a computing device can obtain the various methods when the storage means is coupled to or provided to the device. Furthermore, any other suitable technology for providing the methods and techniques described herein to a device can be utilized.
[0194] The methods disclosed herein comprise one or more steps or actions for achieving the described method. The steps and / or actions of the methods may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions can be modified without departing from the scope of the claims.
[0195] Throughout this specification and the appended claims, unless the context requires otherwise, the word "comprises" and variations such as "comprises" will be understood to mean the inclusion of a stated integer or group of integers, but not the exclusion of any other integer or group of integers.
[0196] The reference herein to any prior art is not, and should not be taken as, any form of acknowledgment that such prior art forms part of the common general knowledge.
[0197] It will be appreciated by those skilled in the art that the present disclosure is not limited in its use to the particular application or applications described. Nor is the present disclosure limited in its preferred embodiments with respect to the particular elements and / or features described or depicted herein. It will be appreciated that the present disclosure is not limited to the disclosed embodiment or embodiments, but is capable of many rearrangements, modifications, and substitutions without departing from the scope set forth and defined by the appended claims.
Claims
1. 1. A method for computationally generating an artificial intelligence (AI) model configured to estimate an embryo viability score from an image, the method comprising: receiving a plurality of images and associated metadata associated with a plurality of patients, each image being an image of an embryo captured during a predetermined time window following in vitro fertilization (IVF) of eggs retrieved from a patient in the plurality of patients, the predetermined time window being within 24 hours; after each image is captured, the embryo is transferred into a respective patient; and pregnancy outcome information is obtained regarding whether the transfer results in a successful pregnancy in the respective patient, the metadata associated with the images including at least a pregnancy outcome label determined using the pregnancy outcome information; pre-processing each image, including at least segmenting the image to identify zone pellucida regions in each image; generating an artificial intelligence (AI) model by training at least one zona pellucida deep learning model using a deep learning method, wherein the AI model, when deployed, is configured to estimate an embryo viability score from input images of embryos captured during a predetermined time window following IVF of eggs retrieved from a patient, wherein the embryo viability score is an estimate of the likelihood that implantation of the embryo will result in a successful pregnancy in the respective patient; and training each of the at least one zona pellucida deep learning model includes training the deep learning model on the set of preprocessed images in which the zona pellucida region is identified, wherein the training includes performing multiple training and validation cycles on the deep learning model to generate the zona pellucida deep learning model, wherein each image is first assigned to one of a training set, a validation set, or a blind validation set; during a training validation cycle, the images in the training set are used to train the deep learning model to estimate embryo viability scores for the images in the training set; the validation set is then used to evaluate the accuracy of the generated deep learning model on the training set using pregnancy outcome labels associated with the images in the validation set; during or at the end of a training validation cycle, one or more model weights are adjusted; and after the multiple training validation cycles, the accuracy of the zona pellucida deep learning model is evaluated using pregnancy outcome labels associated with the images in the blind validation set; developing the AI model; A method comprising:
2. 2. The method of claim 1, wherein segmenting the image to identify clear zone regions in the image further comprises masking each image, wherein in each image, areas bounded by clear zone regions in the image are masked.
3. 3. The method of claim 1, wherein the step of generating the AI model further comprises: training one or more additional AI models to generate the AI model, each of which is either a computer vision model trained using a machine learning method that uses a combination of one or more computer vision descriptors extracted from images to estimate an embryo viability score; a deep learning model trained on images localized to an embryo including both the zona pellucida region and the intrazona pellucida cavity (IZC) region; or a deep learning model trained on a set of IZC images in which all regions except the intrazona pellucida cavity (IZC) are masked; and using an ensemble method to combine at least two of the at least one zona pellucida deep learning model and the one or more additional AI models to generate an embryo viability score for the AI model from an input image, or using a distillation method to train an AI model using the at least one zona pellucida deep learning model and the one or more additional AI models to generate an embryo viability score for the AI model.
4. 4. The method of claim 3, wherein the AI model is generated using an ensemble model, the method comprising: selecting at least two contrasting AI models from the at least one zona pellucida deep learning model and the one or more further AI models, wherein the selection of AI models is performed to generate a set of contrasting AI models; and applying a voting strategy to the at least two selected contrasting AI models that determines how the at least two selected contrasting AI models are combined to generate a result score for the image.
5. The method of claim 4, wherein generating the ensemble model includes evaluating the performance of the AI model using multiple metrics including at least one accuracy metric and at least one confidence metric, or one metric that combines accuracy and confidence.
6. The method described in claim 4 or 5, wherein the ensemble model is trained to bias residuals to minimize false negatives.
7. Selecting at least two contrasting AI models includes: generating a distribution of embryo viability scores from a set of images for each of the at least one zona pellucida deep learning model and the one or more further AI models; comparing the distributions and discarding models when the associated distribution is too similar to another distribution to select an AI model with a contrasting distribution; The method of claim 4, comprising:
8. 8. The method of any one of claims 1 to 7, wherein the predetermined time window is a 24 hour period beginning on the fifth day after fertilization.
9. 9. The method of claim 1, wherein the pregnancy outcome label is a ground truth pregnancy outcome measurement taken within 12 weeks after embryo transfer.
10. 10. The method of claim 9, wherein the ground truth pregnancy outcome measure is whether a fetal heartbeat is detected.
11. 11. The method of claim 1, further comprising cleaning the plurality of images, wherein cleaning the plurality of images comprises identifying images with likely incorrect pregnancy outcome labels and excluding or re-labeling the identified images.
12. 12. The method of claim 11, wherein cleaning the plurality of images comprises estimating the likelihood that a pregnancy outcome label associated with an image is inaccurate, comparing to a threshold, and then rejecting or re-labeling images with a likelihood above the threshold.
13. 13. The method of claim 12, wherein estimating the likelihood that a pregnancy outcome label associated with an image is incorrect is performed using multiple AI classification models and k-fold cross-validation, in which a plurality of images are divided into k mutually exclusive validation data sets, each of the multiple AI classification models trained on the combined k-1 validation data sets and then used to classify images in the remaining validation data sets, and the likelihood is determined based on the number of AI classification models that incorrectly classify the pregnancy outcome label of the image.
14. 14. The method of any one of claims 3 to 13, wherein training each AI model comprises evaluating the performance of the AI model using multiple metrics including at least one accuracy metric and at least one confidence metric, or one metric combining accuracy and confidence.
15. 15. The method of any one of claims 1 to 14, wherein pre-processing the image further comprises cropping the image by localizing an embryo in the image using deep learning or computer vision methods.
16. 16. The method of claim 1, wherein pre-processing the image further comprises one or more of padding the image, normalizing color balance, normalizing brightness, and scaling the image to a predetermined resolution.
17. 17. The method of claim 1, further comprising generating one or more augmented images for use in training an AI model.
18. 18. The method of any one of claims 1 to 17, wherein the augmented image is generated by applying one or more of rotation, reflection, resizing, blurring, contrast variation, jittering, or random compression noise to the image.
19. 19. The method of claim 17 or 18, wherein during training of an AI model, one or more augmented images are generated for each image in the training set, and during evaluation of the validation set, results for the one or more augmented images are combined to generate a single result for the image.
20. 20. The method of claim 1, wherein the step of preprocessing the image further comprises annotating the image with one or more feature descriptor models and masking all regions of the image except for regions within a given radius of keypoints of the descriptors.
21. 21. The method of claim 1, wherein each AI model generates an outcome score, in which an outcome is an n-ary outcome having n states, and training the AI model includes a plurality of training-validation cycles, further comprising: randomly assigning the plurality of images to one of a training set, a validation set, or a blind validation set such that a training dataset includes at least 60% of the images, a validation dataset includes at least 10% of the images, and a blind validation dataset includes at least 10% of the images; after assigning the images to the training set, the validation set, and the blind validation set, calculating a frequency of each of the n-ary outcome states in each of the training set, the validation set, and the blind validation set; testing that the frequencies are similar; and if the frequencies are not similar, discarding the assignments; and repeating the randomization until randomizations with similar frequencies are obtained.
22. 4. The method of claim 3, wherein training a computer vision model includes performing a plurality of training-validation cycles, during each cycle, the images are clustered based on the computer vision descriptors using an unsupervised clustering algorithm to generate a set of clusters, each image is assigned to a cluster using a distance measure based on the values of the computer vision descriptors for the image, and a supervised learning method is used to determine whether a particular combination of these features corresponds to an outcome measure and frequency information of the presence of each computer vision descriptor in the plurality of images.
23. 23. The method of any one of claims 1 to 22, wherein each deep learning model is a convolutional neural network (CNN), and for an input image, each deep learning model generates an outcome probability.
24. 24. The method of any one of claims 1 to 23, wherein the deep learning method emphasizes a global optimum using a loss function configured to modify an optimization surface.
25. 25. The method of claim 24, wherein the one or more model weights are network weights of the deep learning model, and the loss function includes a residual term defined in terms of the network weights, the residual term encoding the collective difference between predictions from the model and a target outcome for each image, and including it as an additional contribution to a normal cross-entropy loss function.
26. 26. The method of any one of claims 1 to 25, wherein the method is performed on a cloud-based computing system using a web server, a database, and multiple training servers, wherein the web server receives one or more model training parameters from a user, the web server initiates a training process on one or more of the multiple training servers including uploading training code to one of the multiple training servers, the training server requests the multiple images and associated metadata from a data repository, and performs steps of preparing each image, generating multiple computer vision models, and generating multiple deep learning models, and each training server is configured to periodically save the models to a storage service and accuracy information to one or more log files to allow the training process to be restarted.
27. 27. The method of any one of claims 1 to 26, wherein the embryo viability score is a binary outcome of viable or non-viable.
28. 28. A method according to any preceding claim, wherein each image is a phase contrast image.
29. 1. A method for computationally generating an embryo viability score from an image, comprising: generating, in a computing system, an artificial intelligence (AI) model configured to generate an embryo viability score from the images according to the method of any one of claims 1 to 28; receiving, from a user via a user interface of the computing system, images of embryos captured during a predetermined time window following in vitro fertilization (IVF) of eggs retrieved from a patient; pre-processing the images according to the pre-processing steps used to generate the AI model; providing the pre-processed images to the AI model to obtain an estimate of the embryo viability score, wherein the embryo viability score is an estimate of the likelihood that transfer of the embryo will result in a successful pregnancy in the respective patient; transmitting the embryo viability score to the user via the user interface; A method comprising:
30. 1. A method for obtaining an embryo viability score from an image, comprising: uploading, via a user interface, images of embryos captured during a predetermined time window following in vitro fertilization (IVF) of eggs retrieved from a patient to a cloud-based artificial intelligence (AI) model configured to generate an embryo viability score from the images, the AI model being generated according to the method of any one of claims 1 to 28; receiving, via the user interface, an embryo viability score from the cloud-based AI model, wherein the embryo viability score is an estimate of the likelihood that transfer of the embryo will result in a successful pregnancy in the respective patient; A method comprising:
31. 29. A cloud-based computing system configured to computationally generate an artificial intelligence (AI) model configured to estimate an embryo viability score from images of embryos captured during a predetermined time window following in vitro fertilization (IVF) of eggs retrieved from a patient, according to the method of any one of claims 1 to 28, wherein the embryo viability score is an estimate of the likelihood that transfer of the embryo will result in a successful pregnancy in the respective patient.
32. 1. A cloud-based computing system configured to computationally generate an embryo viability score from an image, the computing system comprising:
29. An artificial intelligence (AI) model configured to generate an embryo viability score from an image, generated according to the method of any one of claims 1 to 28. receiving, from a user via a user interface of the computing system, images of embryos captured during a predetermined time window following in vitro fertilization (IVF) of eggs retrieved from a patient; providing the image to the AI model to obtain an embryo viability score, wherein the embryo viability score is an estimate of the likelihood that transfer of the embryo will result in a successful pregnancy in the respective patient; and transmitting the embryo viability score to the user via the user interface; 1. A computing system comprising:
33. 1. A computing system configured to generate an embryo viability score from an image, the computing system including at least one processor and at least one memory, the at least one memory comprising: receiving images of embryos captured during a predetermined time window following in vitro fertilization (IVF) of eggs retrieved from a patient; uploading, via a user interface, images captured during a predetermined time window after the in vitro fertilization (IVF) procedure to a cloud-based artificial intelligence (AI) model configured to generate an embryo viability score from the images, the AI model being generated according to the method of any one of claims 1 to 28. receiving an embryo viability score from the cloud-based AI model, wherein the embryo viability score is an estimate of the likelihood that transfer of the embryo will result in a successful pregnancy in the respective patient; displaying the embryo viability score via the user interface. and instructions for configuring the at least one processor to:
Citation Information
Patent Citations
Preparation of aqueous finishing oil for guide oiling and feeding thereof
JP1989014310A
Apparatus, method, and system for image-based human embryonic cell classification
JP2016509845A
Automated image analysis to assess reproductive potential of human oocytes and pronuclear embryos
US20190042958A1