Method and system for performing non-invasive genetic testing using artificial intelligence (AI) models

An AI-based method for non-invasive aneuploidy screening in IVF addresses the limitations of traditional PGT-A by using computer vision and deep learning to rapidly and accurately identify chromosomal abnormalities in embryos, enhancing IVF efficiency and safety.

JP7760165B2Active Publication Date: 2025-10-27ASUTETABUKUYUUGEN
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2022518893
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-25
Filing Date
2020-09-25
Publication Date
2025-10-27
Estimated Expiration
2040-09-25

AI Technical Summary

Technical Problem

Current preimplantation genetic screening (PGT-A) methods for IVF are invasive, time-consuming, and prone to inaccuracies due to embryonic mosaicism, delaying conception and posing unknown long-term effects on embryonic development.

Method used

A non-invasive AI-based method using computer vision and deep learning to analyze embryo images for aneuploidy, employing chromosome group-specific AI models and ensemble techniques to provide rapid and accurate aneuploidy screening.

Benefits of technology

The AI method offers rapid, non-invasive, and accurate aneuploidy screening, reducing delays and uncertainties associated with traditional PGT-A, while improving the reliability and safety of embryo selection for IVF.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007760165000026
    Figure 0007760165000026
  • Figure 0007760165000027
    Figure 0007760165000027
  • Figure 0007760165000028
    Figure 0007760165000028
Patent Text Reader

Abstract

An artificial intelligence (AI)-based computational system is used to noninvasively estimate the presence of a range of aneuploidies and mosaicism in images of preimplantation embryos. Aneuploidies and mosaicisms with similar risks of adverse events are grouped, and training images are labeled with their groups. A separate AI model is trained for each group using the same training dataset. The separate models are then combined, for example, by using an ensemble or distillation approach, to develop a model that can identify a wide range of aneuploidy and mosaicism risks. AI models for groups are created by training multiple models, including binary models, hierarchical models, and multi-class models. In particular, hierarchical models are generated by assigning quality labels to images. At each layer, the training set is divided into the best quality images and other images. The model in that layer is trained on the best quality images, and the other images are passed on to the next layer, and the process is repeated (thus dividing the remaining images into the next best quality images and other images). The final model can then be used to noninvasively identify aneuploidies and mosaicisms and the associated risks of adverse events from preimplantation embryo images.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Priority document This application claims priority to Australian Provisional Patent Application No. 2019903584, entitled "Method and System for performing non-invasive genetic testing using an Artificial Intelligence (AI) Model," filed September 25, 2019, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates to artificial intelligence (AI), including image classification based on computer vision and deep learning. In a particular aspect, the present disclosure relates to a computational AI method for non-invasively identifying aneuploidy in embryos for in vitro fertilization (IVF). [Background technology]

[0003] Unless suffering from side effects such as cytotoxicity caused by radiation or a genetic condition / congenital disease, human cells contain 23 pairs of chromosomes (46 in total). In these cases, one or more of the chromosomes can be modified, either in whole or in part. This can have far-reaching and long-term health effects on the developing embryo, continuing into adulthood, and there is a high level of value in understanding whether a patient exhibits such chromosomal abnormalities or whether they are carriers of chromosomal mutations that predispose their children to such disorders, so that they can be appropriately treated. While prospective parents may have one or many genetic predispositions, it is not possible to predict in advance whether their offspring will actually exhibit one or more genetic abnormalities.

[0004] One commonly used assisted reproductive technology (ART) involves testing embryos after fertilization and performing genetic sequencing to measure the embryo's genetic health and further classify it as "euploid" (genetically typical) or "aneuploid" (exhibiting a genetic mutation).

[0005] This screening technique is particularly prominent in the IVF process, where embryos are fertilized outside the body and then reimplanted into the future mother within approximately three to five days of fertilization. This is often a decision made by the patient in consultation with their IVF physician as part of a process to help diagnose any pregnancy complications the couple may be experiencing, or to diagnose any disease risks early and select against those risks.

[0006] This screening process, known as preimplantation genetic screening (PGS), or preimplantation embryo aneuploidy testing (PGT-A), has many characteristics that make it less than ideal, but it still remains the most viable option currently available in the fertility industry for obtaining genetic information about embryos.

[0007] The greatest risk factor in performing PGT-A is its invasive nature, which typically requires the removal of a small number of cells from a developing embryo (using one of various biopsy techniques). The long-term effects of this technique on embryonic development are unknown and have not been fully characterized. Furthermore, all embryos undergoing PGT-A must be transported to and from the laboratory where the biopsy is attempted, resulting in a delay of several days or weeks for the clinic to receive the results. This means that the "time to conception," a key measure of IVF success, is extended, and all such embryos must undergo freezing. Modern freezing techniques, such as vitrification, have significantly improved in recent years compared to "slow freezing" in terms of embryo survival, and are now commonly practiced among many IVF clinics, even when PGT-A is performed. The reason behind this is to allow the future mother's hormone levels to rebalance after the stimulation of hyperovulation to increase the chances of embryo implantation.

[0008] It is unclear whether modern vitrification techniques are harmful to embryos. The use of vitrification and PGT-A is widespread and widely accepted, especially in the United States, where PGT-A is routinely performed and most embryos undergo this process, providing genetic data for clinics and patients.

[0009] An additional challenge to PGT-A performance is embryonic "mosaicism." This term refers to the possibility that the chromosomal profile of individual cells collected during a biopsy may not represent the entire embryo, even at early cell division stages of embryonic development. That is, a mosaic embryo is a mixture of euploid (chromosomally normal) and aneuploid (chromosomal extra / deletion / alteration) cells, and multiple different aneuploids may be present in different cells (including cases where all cells are aneuploid and there are no euploid cells in the embryo). As a result, PGT-A results obtained from different cells from the same embryo may not agree with each other. The overall accuracy and reliability of such PGT-A tests are reduced because there is no way to assess whether a biopsy is representative.

[0010] Thus, there is a need to provide improved methods for conducting genetic screening of embryos, or at least to provide useful alternatives to existing methods. [Prior art documents] [Non-patent literature]

[0011] [Non-Patent Document 1] Griffiths AJF, Miller JH, Suzuki DT, et al., “An Introduction to Genetic Analysis”, 7th edition, New York: WH Freeman; 2000 Summary of the Invention

[0012] According to a first aspect of the present invention, there is provided a method of computationally generating an aneuploidy screening artificial intelligence (AI) model for screening embryo images for the presence of aneuploidy, the method comprising: determining a plurality of chromosome group labels, each group including one or more different aneuploidies including different genetic mutations or chromosomal abnormalities; creating a training dataset from the first set of images, each image comprising an image of an embryo captured after in vitro fertilization and labeled with one or more chromosome group labels, each label indicating whether at least one aneuploidy associated with a respective chromosome group is present in at least one cell of the embryo, the training dataset comprising images labeled with each of the chromosome groups; creating a test dataset from the second set of images, each image comprising an image of an embryo obtained after in vitro fertilization and labeled with one or more chromosome group labels, each label indicating whether at least one aneuploidy associated with a respective chromosome group is present, the test dataset comprising images labeled with each of the chromosome groups; separately training at least one chromosome group AI model for each chromosome group using the training dataset for training all models, each chromosome group AI model being trained to identify morphological features in images labeled with an associated chromosome group label; and / or training at least one multi-group AI model on the training data, each multi-group AI model being trained to independently identify morphological features in images labeled with each of the associated chromosome group labels and to generate a multi-group output on the input image indicating the presence or absence of at least one aneuploidy associated with each of the chromosome groups in the image; Using the test dataset, select the best chromosome group AI model or the best multi-group AI model for each chromosome group; and deploying the selected AI model to screen embryo images for the presence of one or more aneuploidies; Includes.

[0013] In one embodiment, the step of separately training at least one chromosome group AI model for each chromosome group includes training a hierarchical layered model and / or training at least one multi-group AI model, wherein training the hierarchical layered model includes: The method may include training a hierarchical sequence layer model, where in each layer, images associated with one chromosome group are given a first label and trained against a second set of images, the second set of images being grouped based on a maximum level of quality, and in each sequential layer, the second set of images is a subset of images from the second set in the previous layer that have a quality lower than the maximum quality of the second set in the previous layer.

[0014] In a further aspect, training a hierarchical model includes: assigning a quality label to each image in the plurality of images, the set of quality labels comprising a hierarchical set of quality labels including at least "viable euploid embryo," "euploid nonviable embryo," "non-severe aneuploid embryo," and "severe aneuploid embryo"; training a top-layer model by splitting the training set into a first quality dataset labeled "viable euploid embryos" and another dataset containing all other images, and training the model on the images labeled with the chromosome group and the images in the first quality dataset; Sequentially training one or more intermediate layer models, where in each intermediate layer, a next quality level dataset is created from selecting images with the highest quality labels in the other datasets, and the models are trained on the chromosome group labeled images and images in the next quality dataset; and training a base layer model on the chromosome group labeled images and images in other datasets from previous layers; may include:

[0015] In a further aspect, after training the first base-level model for the first chromosome group, training a hierarchical model for each other chromosome group includes training the other chromosome groups on other datasets used to train the first base-level model.

[0016] In a further aspect, the step of separately training at least one chromosome group AI model for each chromosome group may further include training one or more binary models for each chromosome group: The method includes labeling images in a training data set with a label that matches the chromosome group with a present label and labeling all other images in the training set with an absent label, and training a binary model using the present and absent labels to generate a binary output for the input image indicating whether a chromosomal abnormality associated with the chromosome group is present in the image.

[0017] In a further aspect, the hierarchical models are each binary models.

[0018] In one embodiment, each chromosome group label further includes multiple mutually exclusive aneuploidy classes, the probabilities of the aneuploidy classes within the chromosome group sum to 1, and the AI ​​model is a multi-class AI model trained to estimate the probability of each aneuploidy class within the chromosome group. In a further embodiment, the aneuploidy classes may include "loss," "gain," "duplication," "deletion," and "normal."

[0019] In one aspect, the method comprises: The method may further include a step of generating an ensemble model for each chromosome group, wherein the step of generating an ensemble model for each chromosome group includes: training a plurality of final models, each of the plurality of final models being based on the best chromosome group AI model for a respective group, each of the plurality of final models being trained on a training dataset having a different set of initial conditions and image orderings; Combining multiple trained final models according to an ensemble majority voting strategy; Includes.

[0020] In one aspect, the method comprises: The method may further include creating a distilled model for each chromosome group, the creating a distilled model for each chromosome group comprising: training a plurality of teacher models, each of the plurality of teacher models based on the best chromosome group AI model for a respective group, each of the plurality of teacher models being trained on at least a portion of the training dataset having a different set of initial conditions and image orderings; training a student model using multiple trained teacher models on a training dataset using a distilled loss function; Includes.

[0021] In one aspect, the method comprises: receiving a plurality of images, each image comprising an image of an embryo obtained after in vitro fertilization and one or more aneuploidy results; dividing the plurality of images into a first image set and a second image set, and assigning one or more chromosome group labels to each image based on one or more associated aneuploidy results, wherein the first image set and the second image set have similar proportions of each of the chromosome group labels; It may further include:

[0022] In one embodiment, each group includes multiple different aneuploidies with similar risks of adverse events. In a further embodiment, the multiple chromosome group labels include at least a low-risk group and a high-risk group. In a further embodiment, the low-risk group includes at least chromosomes 1, 3, 4, 5, 17, 19, 20, and "47,XYY," and the high-risk group includes at least chromosomes 13, 16, 21, and "45,X," "47,XXY," and "47,XXX."

[0023] In one form, the image may be captured within 3 to 5 days after fertilization.

[0024] In one embodiment, the relative proportions of each of the chromosome groups in the test data set are similar to the relative proportions of each of the chromosome groups in the training data set.

[0025] According to a second aspect of the present invention there is provided a method of computationally generating an estimate of the presence of one or more aneuploidies in an image of an embryo, the method comprising: generating an aneuploidy screening AI model in a computing system according to the method of the first aspect; receiving an image containing an embryo captured after in vitro fertilization from a user via a user interface of the computing system; providing the image to an aneuploidy screening AI model to obtain an estimate of the presence of one or more aneuploidies in the image; and transmitting a report regarding the presence of one or more aneuploidies in the image to the user via the user interface; Includes.

[0026] According to a third aspect of the present invention there is provided a method of obtaining an estimate of the presence of one or more aneuploidies in an image of an embryo, the method comprising: uploading, via a user interface, images captured during a predetermined time window after in vitro fertilization (IVF) to a cloud-based artificial intelligence (AI) model configured to generate an estimate of the presence of one or more aneuploidies in the images, the AI ​​model being created according to the method of the first aspect; receiving, via a user interface, an estimate of the presence of one or more aneuploidies in the image of the embryo; Includes.

[0027] According to a fourth aspect of the present invention, there is provided a cloud-based computing system configured to computationally create an aneuploidy screening artificial intelligence (AI) model configured according to the method of the first aspect.

[0028] According to a fifth aspect of the present invention there is provided a cloud-based computing system configured to computationally generate an estimate of the presence of one or more aneuploidies in an image of an embryo, the computing system comprising: and one or more computing servers comprising one or more processors and one or more memories configured to store an aneuploidy screening artificial intelligence (AI) model configured to generate an estimate of the presence of one or more aneuploidies in an image of an embryo, wherein the aneuploidy screening artificial intelligence (AI) model is created according to the method of the first aspect, wherein the one or more computing servers: receiving an image from a user via a user interface of a computing system; providing the image to an aneuploidy screening artificial intelligence (AI) model to obtain an estimate of the presence of one or more aneuploidies in the image; and transmitting a report regarding the presence of one or more aneuploidies in the image to the user via the user interface; It is structured as follows.

[0029] According to a sixth aspect of the present invention there is provided a computing system configured to generate an estimate of the presence of one or more aneuploidies in an image of an embryo, the computing system comprising: at least one processor; receiving images captured during a predetermined time window after in vitro fertilization (IVF); uploading, via a user interface, images captured during a predetermined time window following in vitro fertilization (IVF) to a cloud-based artificial intelligence (AI) model configured to generate an estimate of the presence of one or more aneuploidies in images of an embryo, the AI ​​model being created according to the method of any one of claims 1 to 13; receiving, via a user interface, an estimate of the presence of one or more aneuploidies in the image of the embryo; and displaying, via a user interface, an estimate of the presence of one or more aneuploidies in the embryo image; at least one memory containing instructions for configuring the at least one processor to Includes. [Brief explanation of the drawings]

[0030] Embodiments of the present disclosure will be discussed with reference to the accompanying drawings. [Figure 1A] 1 is a flow diagram of a method for computationally creating an aneuploidy screening artificial intelligence (AI) model for screening embryo images for the presence of aneuploidy, according to one embodiment. [Figure 1B] 1 is a flow diagram of a method for computationally generating an estimate of the presence of one or more aneuploidies in an image of an embryo using a trained aneuploidy screening AI model, according to one embodiment. [Figure 2A] 1 is a flow diagram of steps for training a binary model, according to one embodiment. [Figure 2B] 1 is a flow diagram of steps for training a hierarchical model, according to one embodiment. [Figure 2C] 1 is a flow diagram of steps for training a multi-class model, according to one embodiment. [Figure 2D] 1 is a flow diagram of steps for selecting the best chromosome group AI model, according to one embodiment. [Figure 3] FIG. 1 is a schematic architectural diagram of a cloud-based computing system configured to computationally create and use an aneuploidy screening AI model, according to one embodiment. [Figure 4] FIG. 1 is a schematic diagram of an IVF method using an aneuploidy screening AI model to assist in the selection of embryos for implantation, according to one embodiment. [Figure 5A] 1 is a schematic flow diagram of the creation of an aneuploidy screening model using a cloud-based computing system, according to one embodiment. [Figure 5B] 1 is a schematic flow diagram of a model training process on a training server, according to one embodiment. [Figure 5C]FIG. 1 is a schematic architecture diagram of a deep learning method including convolutional layers that transform input images into predictions after training, according to one embodiment. [Figure 6A] FIG. 10 shows a plot of the confidence of a chromosome 21 AI model for detecting aneuploid chromosome 21 embryos in a blind test set, according to one embodiment, where low confidence estimates are indicated by bars filled with diagonal forward slashes on the left and high confidence estimates are indicated by bars filled with diagonal backslashes on the right. [Figure 6B] FIG. 10 shows a plot of the confidence of the chromosome 21 AI model in detecting euploid viable embryos in a blind test set, according to one embodiment, where low confidence estimates are indicated by bars filled with diagonal forward slashes on the left and high confidence estimates are indicated by bars filled with diagonal backslashes on the right. [Figure 7A] FIG. 10 shows a plot of the confidence of a chromosome 16 AI model for detecting aneuploid chromosome 16 embryos in a blind test set, according to one embodiment, where low confidence estimates are indicated by bars filled with diagonal forward slashes on the left and high confidence estimates are indicated by bars filled with diagonal backslashes on the right. [Figure 7B] FIG. 10 shows a plot of the confidence of the chromosome 16 AI model in detecting euploid viable embryos in a blind test set, according to one embodiment, where low confidence estimates are indicated by bars filled with diagonal forward slashes on the left and high confidence estimates are indicated by bars filled with diagonal backslashes on the right. [Figure 8A]FIG. 1 shows a plot of the confidence of the chromosome severe group (14, 16, 18, 21, and 45,X) AI model in detecting aneuploidies in chromosomes 14, 16, 18, 21, and 45,X in a blind test set, according to one embodiment, where low confidence estimates are indicated by bars filled with diagonal forward slashes on the left and high confidence estimates are indicated by bars filled with diagonal backslashes on the right. [Figure 8B] FIG. 10 shows a plot of the confidence of a chromosomal severity group (14, 16, 18, 21, and 45,X) AI model detecting euploid viable embryos in a blind test set, according to one embodiment, where low confidence estimates are indicated by bars filled with diagonal forward slashes on the left and high confidence estimates are indicated by bars filled with diagonal backslashes on the right. DETAILED DESCRIPTION OF THE INVENTION

[0031] In the following description, like reference numerals designate like or corresponding parts throughout the figures.

[0032] Embodiments of non-invasive methods for screening embryos for the presence / possibility of aneuploidies (genetic mutations) are described. These aneuploidies, or genetic mutations, result in alterations, deletions, or additional copies of portions of chromosomes or even entire chromosomes. Often, these chromosomal abnormalities result in subtle (and sometimes obvious) changes in the appearance of chromosomes in an image of the embryo. Embodiments of the method use computer vision-based artificial intelligence (AI) / machine learning models to detect the presence of aneuploidies (i.e., chromosomal abnormalities) entirely based on morphological data extracted from phase-contrast microscopy images of embryos (or images of similar embryos). The AI ​​models often use computer vision techniques to detect subtle morphological features in embryo images to estimate the probability or likelihood of the presence (or absence) of various aneuploidies. This estimate / information can then be used to aid in implantation decisions or the selection of embryos for invasive PGT-A testing.

[0033] The system is non-invasive (i.e., it works solely on microscopic images) and allows analysis within seconds of image collection by uploading images using a cloud-based user interface, which has the advantage of analyzing the images using a pre-trained AI model on the cloud-based server to quickly return an estimate of the likelihood of aneuploidy (or a specific aneuploidy) to the clinician.

[0034] Figure 1A is a flow chart of a method 100 for computationally creating an aneuploidy screening artificial intelligence (AI) model for screening embryo images for the presence of aneuploidy. Figure 1B is a flow chart of a method 110 for computationally generating an estimate of the presence of one or more aneuploidies in an embryo image using a trained aneuploidy screening AI model (i.e., one created according to Figure 1A).

[0035] For the 22 types of non-sex chromosomes, the types of chromosomal abnormalities considered include: complete gains, complete losses, deletions (segmental deletions within a chromosome), and duplications (segmental duplications within a chromosome) compared to the normal chromosome structure. For the sex chromosomes, the types of abnormalities considered include: deletions (segmental deletions within a chromosome), duplications (segmental duplications within a chromosome), complete losses (45,X), and three complete gains (47,XXX, 47,XXY, and 47,XYY) compared to the normal XX or XY chromosomes.

[0036] Embryos can also exhibit mosaicism, in which different cells within the embryo have different sets of chromosomes. That is, an embryo may contain one or more euploid cells and one or more aneuploid cells (i.e., have one or more chromosomal abnormalities). Furthermore, multiple aneuploidies may exist, with different cells having different aneuploidies (e.g., one cell may have a deletion on chromosome 1, while another cell may have an X gain, such as 47,XXX). In some extreme cases, every cell within a mosaic embryo exhibits aneuploidy (i.e., no euploid cells exist). Thus, an AI model can be trained to detect aneuploidy in one or more cells of an embryo, thus detecting the presence of mosaicism.

[0037] The output of the AI ​​model can be expressed as a probability of an outcome, such as an aneuploidy risk score, or as an embryo viability score. Embryo viability and aneuploidy risk are understood to be complementary terms. For example, when they are probabilities, the sum of embryo viability risk and aneuploidy risk can be 1. That is, both measure the likelihood of an adverse event, such as the risk of miscarriage or a serious genetic disorder. Therefore, this result will be referred to as the aneuploidy risk / embryo viability score. The result may also be the probability of being in a particular risk category of an adverse event, such as very low risk, low risk, medium risk, high risk, or very high risk. Each risk category may include at least one, and typically more, specific chromosomal abnormality group with similar probabilities of the adverse event. For example, very low risk may be no detectable aneuploidy, a low-risk group may be aneuploidy / mosaicism in chromosomes 1, 3, 10, 12, and 19, and a medium-risk group may be aneuploidy / mosaicism in chromosomes 4, 5, and 47,XXY. Likelihood can be expressed as a score on a predefined scale, a probability between 0 and 1.0, or as a strict classification such as a strict binary classification (aneuploidy present or absent) or into one of several groups (low risk, medium risk, high risk, very high risk).

[0038] In step 101, the inventors define multiple chromosome group labels. Each group contains one or more different aneuploidies, including different gene mutations or chromosomal abnormalities. Different aneuploidies / gene mutations affect the embryo differently and lead to different chromosomal abnormalities. Within a chromosome group label, a separate mosaicism category may be defined, which, if present, may be low, medium, or high, indicating that the same embryo may exhibit different types of chromosomal abnormalities. The level of severity (risk) of mosaicism can also take into account the number of cells exhibiting mosaicism and / or the type of aneuploidy present. Thus, a chromosome group label can include not only the number of affected chromosomes, but also whether mosaicism (to some extent) exists. This allows for a more detailed description of the aneuploidy progression level or genetic health of the embryo. As described in Table 1 below, the level of severity of mosaicism present depends on the level of severity of the chromosomes involved in the mosaicism. In addition, the level of severity may be related to the number of cells exhibiting mosaicism. Based on clinical evidence, e.g., based on the results of PGT-A testing and pregnancy, it is possible to group different aneuploidies / chromosomal abnormalities based on the risk and severity of adverse events, and thus give them a priority for implantation (aneuploidy risk / embryo viability score). Table 1 lists the number and types of chromosomal abnormalities in spontaneous abortions and live births in 100,000 pregnancies from Non-Patent Document 1.

[0039] [Table 1-1]

[0040] [Table 1-2] Using Table 1, or similar data from other clinical studies, aneuploidies can be grouped based on risk level. The highest risk aneuploidies are considered the lowest priority for implantation and the highest priority for identification to avoid post-implantation adverse events. For example, a first low-risk group consisting of chromosomes 1, 3, 4, 5, 17, 19, 20, and 47,XYY could be formed based on spontaneous abortions occurring at less than 100 per 100,000 pregnancies. A medium-risk group consisting of chromosomes 2 and 6-12 could be defined based on spontaneous abortions occurring at less than 200 (and greater than 100) per 100,000 pregnancies. A high-risk group consisting of chromosomes 14, 15, 18, and 22 could be defined based on spontaneous abortions occurring at more than 200 per 100,000 pregnancies. A final very high-risk group, consisting of chromosomes 13, 16, 21, and 45,X, 47,XXY, and 47,XXX, can be defined based on the number of spontaneous abortions exceeding 1,000 per 100,000 pregnancies or known to result in births with adverse health outcomes. Other divisions can also be used; for example, the first group can be divided into chromosomes 1, 3, 10, 12, 19, and 20, and a second, slightly higher-risk group consisting of chromosomes 4, 5, and 47,XYY. Chromosomes can also be classified separately based on complete addition (trisomy), normal pair (disomy), and complete deletion (monosomy). For example, chromosome 3 (disomy) can be in a different group than chromosome 3 (trisomy). Generally, trisomy (complete addition) is considered high-risk and is avoided.

[0041] A chromosome group may include a single chromosome or a subset of chromosomes, e.g., with a similar risk profile or below a risk threshold. A chromosome group can define a particular type or class of mosaicism, such as the type of chromosome and the number of aneuploid cells in an embryo. These chromosomes then become the focus of building an AI / machine learning model that will identify morphological features associated with alterations to that chromosome. In one embodiment, each image is labeled using a class label based on the grouping outlined above based on implantation priority / risk profile, e.g., the risks listed in Table 1 (e.g., embryo images in the "low risk" group may be given class label 1, and embryo images in the "medium risk" group may be given class label 2). Note that the above groupings are merely exemplary; in other embodiments, other clinical risk profiles or other clinical data or risk factors can be used to define (different) chromosome groups and assign chromosome group labels to images. As noted above, embryos may exhibit mosaicism, in which different cells within the embryo have different sets of chromosomes, such that the (mosaic) embryo is a mixture of euploid (chromosomally normal) and aneuploid (with extra / missing / altered chromosomes) cells. Risk groups may therefore be established based on the presence of mosaicism and the type and number / extent of aneuploidy present. In some embodiments, risk may be based on the most severe aneuploidy present in the embryo (even if present in only a single cell). In other embodiments, a threshold number of low-risk aneuploidies can be established, after which embryos are reclassified as higher risk depending on the amount of aneuploidy present (i.e., the number of aneuploidies exceeds the threshold).

[0042] In step 102, we create a training dataset 120 from the first set of images. Each image includes an image of an embryo captured after in vitro fertilization and is labeled with one or more chromosome group labels. Each label indicates whether at least one aneuploidy associated with the respective chromosome group is present. The training dataset is configured to include images labeled with each of the chromosome groups, so that the model is exposed to each of the chromosome groups to be detected. Furthermore, it should be noted that individual embryos / images may have multiple different aneuploidies and therefore be labeled with and included in multiple chromosome groups.

[0043] Similarly, in step 103, we create a test dataset 140 from the second set of images. Again, each image includes an image of an embryo taken after in vitro fertilization and is labeled with one or more chromosome group labels. Each label indicates whether at least one aneuploidy associated with the respective chromosome group is present. Similar to the training dataset 120, the test dataset 140 includes images labeled with each of the chromosome groups.

[0044] The training set 120 and test set 140 can be created using images of embryos for which PGT-A results and / or pregnancy outcomes (e.g., for implanted embryos) are available, which can be used to label the images. Typically, the images are phase-contrast microscopy images of embryos captured 3-5 days after in vitro fertilization (IVF). Such images are routinely captured during the IVF procedure to assist embryologists in determining which embryos to select for implantation. However, it will be understood that other microscopy images captured at other times under other lighting conditions or magnification ranges can be used. In some embodiments, a time-lapse sequence of images can be used, for example, by combining / concatenating a series of images into a single image that is analyzed by the AI ​​model. Typically, the pool of available images is divided into a large training set with approximately 90% of the images and a small (remaining 10%) blind holdout test set; i.e., the test data set is not used to train the model. A small percentage, such as 10%, of the training set 120 can also be allocated to a validation data set. Preferably, the relative proportions of each of the chromosome groups in the test dataset 120 are similar to the relative proportions of each of the chromosome groups in the training dataset 140 (e.g., within 10%, preferably within 5% or less).

[0045] In step 104, we separately train at least one chromosome group AI model for each chromosome group, using the same training dataset 120 to train all models. Each chromosome group AI model is trained to identify morphological features in images labeled with an associated chromosome group label. Additionally or alternatively, at least one multi-group AI model can be trained on the training data. Each multi-group AI model is trained to independently identify morphological features in images labeled with each of the associated chromosome group labels and generate a multi-group output for the input image to indicate the presence or absence of at least one aneuploidy associated with each of the chromosome groups in the image.

[0046] In step 105, we then use the test dataset to select the best chromosome group AI model or the best multi-group AI model for each of the chromosome groups (depending on which was created in step 104). In some embodiments, the final selected model will be used to create further ensemble or knowledge distillation models. In step 106, we deploy the selected AI model to screen embryo images for the presence of one or more aneuploidies.

[0047] Therefore, one approach to building an aneuploidy screening AI model capable of detecting / predicting a wide range of chromosomal deletions is to overcome this problem and train individual, targeted AI models for a subset of chromosomal deletions, and then combine the separate AI models to detect a wider set of chromosomal deletions. As mentioned above, each chromosomal group will be treated independently of the others. Each embryo (or embryo image) may have multiple chromosomal deletions, and a single image may be associated with multiple chromosomal groups. That is, a mosaic embryo in which different cells have different aneuploidies may have multiple group labels corresponding to each aneuploidy present in the embryo. In either case, a complete training dataset will be utilized to create a machine learning model (or ensemble / distilled model). This will be repeated multiple times for each chromosomal group of interest (as limited by the quality and total size of the dataset so that machine learning models can be created), allowing for the creation of multiple models covering different chromosomal alterations from the same training dataset. In some cases, models established from the same "base" model may be very similar to each other, but in the final layer, separate classifiers will correspond to each chromosome considered. In other cases, models can treat only one chromosome individually and are combined together using ensemble or distillation methods. These scenarios are discussed below with regard to selecting the best model.

[0048] In one embodiment, separately training at least one chromosome group AI model 103 for each chromosome group includes training one or more of a binary model, a hierarchical (multi-layer) model, or a single multi-group model for each chromosome group, and then using a test data set to select the best model for that chromosome group. This is further illustrated in Figures 2A-2D and discussed below. Figures 2A, 2B, and 2C show a flowchart of training a binary model 137, a hierarchical model 138, and a multi-group model 139, according to one embodiment. Figure 2D is a flowchart of selecting the best chromosome group AI model 146 using a test data set 140, according to one embodiment.

[0049] 2A is a flow diagram of step 104a of training a binary model for a chromosome group. Images in a training set 120 are labeled with a label that matches the i-th chromosome group, such as a "present" label (or "Yes" or 1), to create an i-th chromosome group image 121. Next, we label all other images in the training set with an "absent" label (or "No" or 0) to create a set of all other images 122. Next, we train a binary model 131 using the presence and absence labels (i.e., in the i-th chromosome group 121 and all other images 122) so that the binary model 127 will produce a binary output in the input image indicating whether a chromosomal abnormality associated with the i-th chromosome group is present (or absent) in the image. The presence label will typically indicate presence in at least one cell (e.g., in the case of mosaicism), but could also indicate presence in a threshold number of cells or in all cells in the embryo.

[0050] FIG. 2B is a flow diagram of step 104b of training a layer model of a hierarchical sequence for a chromosome group (i-th chromosome group 121), which we will refer to as a hierarchical model (or multi-layer model). In this embodiment, this involves training a layer binary model of a hierarchical sequence (although, as discussed below, the requirements for a binary model can be relaxed). At each layer, images associated with the i-th chromosome group 121 are assigned a first label and trained against a second set of images from the training dataset 120, which are grouped based on a maximum level of quality (quality group). At each sequence layer, the second set of images used in training is a subset of images from the second set in the previous layer, and has a quality lower than the maximum quality of the second set in the previous layer. That is, we first divide the dataset into images labeled (associated with) the i-th chromosome group and the remaining images in the dataset. The remaining images in the dataset are assigned a quality label. Then, at each level, we divide the current dataset into a first group corresponding to the highest quality level remaining in the dataset and the remaining lower quality images (i.e., the remaining image group). At the next level, the previous remaining group is further divided into a first group corresponding to the highest quality level in the (remaining) dataset and the remaining lower quality images (which become the updated remaining image group). This is repeated until we reach the lowest level.

[0051] That is, the second group (or second classification) includes embryo images of varying levels of quality based on genetic integrity (chromosomal deletions, including mosaicism) and viability (a viable embryo, when implanted in a patient, will result in a pregnancy and be considered "good"). The rationale behind the hierarchical model approach is that embryo images considered to be of high quality will likely have the least abnormalities and the highest quality morphological features in the image (i.e., "look like the best embryo") and therefore will have the greatest morphological differences / distinctions compared to embryo images containing chromosomal deletions (i.e., "look bad or have abnormal features"), thus enabling AI algorithms to better detect and predict morphological features between these two (extreme) image classifications.

[0052] In the embodiment shown in FIG. 2B , we assign a quality label to each image in the training dataset 120, which includes a hierarchical set of quality labels that can be used to divide the training dataset 120 into different quality subsets. Each image has a single quality label. In this embodiment, these include "viable euploid embryos" 123, "euploid nonviable embryos" 125, "non-severe aneuploid embryos" 127, and "severe aneuploid embryos" 129. In other embodiments, these may be risk categories of adverse events, such as very low risk, low risk, medium risk, high risk, or very high risk, or simply low risk, medium risk, and high risk. The training dataset is then divided into the i-th chromosome group of interest and the remaining images, as outlined above. Next, we divide the training set into a first-quality dataset with the label "viable euploid embryo" 123 and another dataset 124 containing all other images, and further train a top-layer binary model 132 by training a binary model on images labeled with the i-th chromosome group 121 and images in the first-quality dataset 123. Next, we sequentially train one or more, in this case two, intermediate-layer binary models 132-133, where in each intermediate layer, a next-level quality dataset is created by selecting images with the highest-quality label in the other dataset, and a binary model is trained on images labeled with the chromosome group and images in the next-level quality dataset. Thus, we select a "euploid nonviable embryo" 125 from the other lower-quality images 124 in the top layer and train a first intermediate-layer model 133 on images from the i-th chromosome group 121 and the "euploid nonviable embryo" 125. The remaining lower quality images 126 include "non-severe aneuploid embryos" 127 and "severe aneuploid embryos" 129. In the next hidden layer, we again extract images of the next quality level, namely "non-severe aneuploid embryos" 127, and train another hidden layer model 133 on images from the i-th chromosome group 121 and "non-severe aneuploid embryos" 127.The remaining other low-quality images 128 now include "severe aneuploid embryos" 129. Next, we train a binary substratum model 135 on the images labeled with the i-th chromosome group 121 and on other images in the dataset from the previous layer, i.e., the images 129 of the "severe aneuploid embryos" 129. The output of this step is a trained hierarchical (binary) model 138 that produces a binary output on the input image to indicate whether a chromosomal abnormality associated with the i-th chromosome group is present (or not) in the image. This can be repeated many times by varying the number of layers / quality labels (e.g., from 5 (very low risk, low risk, medium risk, high risk, very high risk) to 3 (low risk, medium risk, high risk)) to create many different models for comparison / selection.

[0053] In some embodiments, after training the first binary base-level model on the first chromosome group, we then reuse the images of "severe aneuploid embryos" 129 to train each other chromosome group. That is, training the hierarchical models includes training the other chromosome groups on the other dataset ("severe aneuploid embryos" 129) used to train the first binary base-level model. In some embodiments, we may skip the intermediate layers and simply use the top-layer model and the base-layer model (in this case, the base layer is trained on images with multiple quality levels, but not on "viable euploid embryos" 123).

[0054] In the above example, the model is a single-label model, where each label is simply present / absent (with a certain probability). However, in another embodiment, the model can be a multi-class model, where each independent label contains multiple independent aneuploidy classes. For example, if the group labels are "severe," "moderate," or "mild," the labels can each have aneuploidy classes such as "loss," "gain," "duplication," "deletion," and "normal." Note that because the group labels are independent, the confidence in the type of aneuploidy in one chromosome group does not affect the confidence of the model in the aneuploidy in another chromosome group; for example, they can be high or low confidence. The classes within a label are mutually exclusive, such that the probabilities of different classes within a label sum to 1 (e.g., you cannot have both a loss and a gain in the same chromosome). Therefore, the output for each group is not a binary / yes output, but a list of probabilities for each class. Furthermore, because the labels are independent, different labels / chromosome groups can have different binary / multi-class classes; for example, some may be binary (euploid, aneuploid) and others may be multi-class ("loss," "gain," "duplication," "deletion," "normal"). That is, the model is trained to estimate the probability of each aneuploidy class within a label. That is, if there are m classes, the output is a set of mn yes / no (presence / absence, or 1 / 0 values) outcomes (e.g., in a list or similar data structure), the probabilities for the classes, and the overall probability of the label.

[0055] In another embodiment, the layer models of the hierarchical sequence may be a mixture of multi-class and binary models, or may all be multi-class models. That is, with reference to the above discussion, we may replace one or more (or all) of the binary models with multi-class models, trained on a set of chromosome groups in addition to quality labels. In this regard, the complete layer models of the hierarchical sequence cover all available group labels available in the training set. However, each model in the sequence may be trained only on a subset of one or more chromosome groups. In this manner, a top-level model can be trained on a large data set and the data set can be triaged into one or more predicted outcomes. Subsequent models in the sequence are then trained on a subset of data related to one of the outcomes, further classifying this set into finer subgroups. This process can be repeated multiple times to create a sequence of models. By repeating this process and varying not only which levels use binary models and which levels use multi-class models, but also the number of different quality labels, different models can be trained on different subsets of the training data set.

[0056] In another embodiment, the model can be trained as a multigroup model (single or multi-class). That is, if there are n chromosome labels / groups, rather than training a model separately for each group, we train a single multigroup model that simultaneously estimates each of the n group labels in a single pass through the data. Figure 2C is a flow diagram of step 104c, training a single multigroup model 136 for a chromosome group (i-th chromosome group 121). We then train the single multigroup model 136 on the training data 120 using present and absent labels (if binary) for each of the chromosome groups 120 and all other images 122 to create the multigroup model 139. When presented with an input image, the multigroup model 139 generates a multigroup output indicating whether at least one aneuploidy associated with each of the chromosome groups is present or absent in the image. That is, if there are n chromosome groups, the output is a set of n yes / no (presence / absence, or 1 / 0 values) results (e.g., in a list or similar data structure). In the multi-class case, there are additional probability estimates for each of the classes in the chromosome groups. Note that the architecture of the specific model (e.g., the configuration of convolutional and pooling layers) is similar to that for the single-chromosome group model described above, but differs in the final output layer because, rather than binary or multi-class classification for a specific chromosome group, the output layer must generate independent estimates for each chromosome group. This change in the output layer effectively changes the optimization problem and therefore provides different performance / results compared to the multiple single-chromosome group models described above, thus providing more diversity in the models / results, which may aid in finding the optimal overall AI model. In addition, a multi-group model does not need to estimate / classify all chromosome groups; instead, we could train several (M>1) multi-group models, each estimating / classifying a different subset of chromosome groups. For example, if there are n chromosome groups, we can train M multi-group models separately, where each model estimates k groups simultaneously, and M, k, and n are integers such that n=Mk.However, each multigroup model does not have to estimate / classify the same number of chromosome groups; for example, we may train M multigroup models and each multigroup model may estimate / classify k chromosome groups. m When classifying the chromosome groups together,

[0057]

number

[0058] In another embodiment, the multigroup model can also be used in a hierarchical model approach, as illustrated in FIG. 2B. That is, instead of training a separate hierarchical model for each chromosome group (e.g., n hierarchical models for n chromosome groups), we can train a single multigroup model for all chromosome groups using a hierarchical approach, where the dataset is successively divided based on the quality level of the images remaining in the dataset (and we train a new multigroup model at each layer). In addition, we can use a hierarchical approach to train several (M>1) multigroup models, where each multigroup model classifies a different subset of chromosome groups, such that every group is classified by one of the M multigroup models. Furthermore, each multigroup model can be a binary model or a multiclass model.

[0059] FIG. 2D is a flow diagram of step 105 for selecting the best chromosome group AI model or best multigroup model for the i-th chromosome group. The test dataset has images across all of the euploid and aneuploid categories selected for training the models. We obtain a test dataset 140 and provide (unlabeled) images as input to each of the binary model 137, hierarchical model 128, and multigroup model 139 from FIGS. 2A-2D. We obtain test results 141 for the binary model, 142 for the hierarchical model, and 143 for the multigroup model, and compare model results 144 using the i-th chromosome group label 145. The best-performing model 146 is then selected using a selection criterion, e.g., based on calculating one or more metrics and comparing the models to each other using one or more of the metrics. The metric may be chosen from (but is not limited to) a list of commonly accepted performance metrics such as: overall accuracy, average accuracy, F1 score, average class accuracy, precision, recall, Log Loss, or a custom reliability or loss metric such as those described below (e.g., Equation (7)). The performance of the models on a set of validation images is measured with respect to the metric, and then the best-performing model is selected accordingly. These models may be further sorted by a secondary metric, and this process may be repeated multiple times until a final model or a selection of models (to create an ensemble model, if desired) is obtained.

[0060] The best-performing model can be further refined using ensemble or knowledge distillation methods. In one embodiment, an ensemble model for each chromosome group can be created by training multiple final models, where each of the multiple final models is based on the best chromosome group AI model for its respective group (or multiple groups if a multi-group model is selected), and each of the multiple final models is trained on a training dataset with a different set of initial conditions and image orderings. The final ensemble model is obtained by combining multiple trained final models according to an ensemble voting strategy, combining models that exhibit contrasting or complementary behavior according to their performance on one or more indicators from the list above.

[0061] In one embodiment, a distilled model is created for each chromosome group. This involves training multiple teacher models, each based on the best chromosome group AI model for its respective group (or groups if a multi-group model is selected), and each of the multiple teacher models is trained on at least a portion of a training dataset with a different set of initial conditions and image orderings. We then train a student model using the multiple trained teacher models on the training dataset using a distillation loss function.

[0062] This is repeated for each group of chromosomes to create a comprehensive aneuploidy screening AI model 150. Once the aneuploidy screening AI model 150 is trained, it can be deployed in a computational system to provide real-time (or near real-time) screening results. Figure 1B is a flow diagram of a method 110 for computationally generating estimates of the presence of one or more aneuploidies in images of embryos using a trained aneuploidy screening AI model, according to one embodiment.

[0063] In step 111, we create an aneuploidy screening AI model 150 in a computing system according to method 100 above. In step 112, we receive images containing embryos captured after in vitro fertilization from a user via a user interface of the computing system. In step 113, we provide the images to the aneuploidy screening AI model 150 to obtain an estimate of the presence of one or more aneuploidies in the images. Then, in step 114, we send a report regarding the presence of one or more aneuploidies in the images to the user via the user interface.

[0064] An associated cloud-based computing system may also be provided that is configured to computationally create an aneuploidy screening artificial intelligence (AI) model 150 configured according to the training method 100 and further to estimate the presence of one or more aneuploidies in an image of an embryo (including in at least one cell in the case of mosaicism, or in all cells of the embryo) (method 110), as further illustrated in Figures 3, 4, 5A, and 5B.

[0065] FIG. 3 is a schematic architecture diagram of a cloud-based computing system 1 configured to computationally create an aneuploidy screening AI model 150 and then use this model to generate a report with an estimate of the presence of one or more aneuploidies in received embryo images. Input 10 includes data, such as embryo images and outcome information (presence of one or more aneuploidies, birth or nonbirth, or successful implantation), that can be used to generate a label (classification). This is provided as input to a model creation process 20, which creates computer vision and deep learning models to analyze the input images and combine them to create an aneuploidy screening AI model. This may also be referred to as an aneuploidy screening artificial intelligence (AI) model or an aneuploidy screening AI model. A cloud-based model management and monitoring tool 21, called Model Monitor, is used to create (or generate) the AI ​​model. It uses a set of linked services, such as Amazon Web Services (AWS), that manage the model and manage the training, logging, and tracking of the model, which are specific to image analysis. Other similar services on other cloud platforms may also be used. These can use deep learning methods 22, computer vision methods 23, classification methods 24, statistical methods 25, and physics-based models 26. Modeling can also use domain expertise 12 as input, such as from embryologists, computer scientists, scientific / technical literature, etc., such as on what features to extract and use in the computer vision model. The output of the modeling process is an example of an aneuploidy screening AI model, which in this embodiment is a validated aneuploidy screening (or embryo evaluation) AI model 150. Other aneuploidy screening AI models 150 can be created using other image data with associated outcome data.

[0066] A cloud-based delivery platform 30 is used, providing a user interface 42 to the system for a user 40. This is further illustrated with reference to FIG. 4, which is a schematic diagram of an IVF method 200 that uses an aneuploidy screening AI model 150 to assist in the selection of embryos for implantation, or which embryos to reject or subject to invasive PGD testing, according to one embodiment. On day 0, retrieved eggs are fertilized (202). They are then cultured in vitro for several days, after which images of the embryos are captured (204), for example, using a phase-contrast microscope. Preferably, the model is trained and used on images of embryos captured on the same day or during a specific time window referenced to a specific epoch. In one embodiment, the time is 24 hours, although other time windows, such as 12, 36, or 48 hours, can also be used. A shorter time window, generally 24 hours or less, is preferred to ensure greater similarity in appearance. In one embodiment, this can be a specific day, a 24-hour window starting from the beginning of the day (0:00) to the end of the day (23:39), or a specific day, such as day 4 or 5 (a 48-hour window starting from the beginning of day 4). Alternatively, the time window can define a window size and epoch, such as a 24-hour period centered around day 5 (i.e., day 4.5 to day 5.5). The time window can be variable, with a lower limit, such as at least 5 days. As mentioned above, it is preferred to use images of embryos from a 24-hour time window around day 5, although it should be understood that earlier stage embryos, including images from day 3 or day 4, can be used.

[0067] Typically, several eggs will be fertilized simultaneously, and therefore multiple images will be obtained in order to consider which embryos are best (i.e., most viable) for implantation (which may include identifying which embryos should be eliminated due to a high risk of serious defects). A user uploads the captured images to platform 30 via user interface 42, for example, using a "drag-and-drop" function. The user can upload a single image or multiple images, for example, to aid in the selection of which embryos from a set of multiple embryos will be considered for implantation (or rejected). Platform 30 receives one or more images (312) that are stored in database 36, which includes an image repository. The cloud-based delivery platform includes on-demand cloud servers 32 configured to perform image preprocessing (e.g., object detection, segmentation, padding, normalization, cropping, centering, etc.) and then provide the processed images to a trained AI (aneuploidy screening) model 150, which runs on one of the on-demand cloud servers 32 and analyzes the images to generate an aneuploidy risk / embryo viability score (314). A report of the model results, such as the likelihood of the presence of one or more aneuploidies, or a binary call (use / no use), or other information derived from the model, is generated (316) and is sent or otherwise provided to a user 40, such as via a user interface 42. A user (e.g., an embryologist, etc.) receives the aneuploidy risk / embryo viability score and report via the user interface and can then use the report (likelihood) to help decide whether to implant the embryo or which embryo in the set is the best for implantation. The selected embryo is then implanted (205). To help further improve the AI ​​model, pregnancy outcome data, such as the detection (or non-detection) of a heartbeat on the first ultrasound scan after implantation (usually around 6-10 weeks after fertilization), or aneuploidy results from PGT-A testing, can be provided to the system, allowing the AI ​​model to be retrained and updated as more data becomes available.

[0068] Images may be captured using various imaging systems, such as those found in existing IVF clinics. This has the advantage that IVF clinics do not need to purchase new or use specific imaging systems. The imaging system is typically an optical microscope configured to capture single-phase contrast images of the embryo. However, it will be understood that other imaging systems, particularly optical microscope systems using various imaging sensors and image capture technologies, can be used. These may include phase contrast microscopes, polarizing microscopes, differential interference contrast (DIC) microscopes, dark-field microscopes, and bright-field microscopes. Images may be captured using a conventional optical microscope equipped with a camera or image sensor, or the images may be captured by a camera with integrated optics capable of capturing high-resolution or high-magnification images, including smartphone systems. The image sensor may be a CMOS sensor chip or a charge-coupled device (CCD), each with associated electronics. The optics may be configured to collect specific wavelengths or use filters, including bandpass filters, to collect (or reject) specific wavelengths. Some image sensors may be configured to operate or be sensitive to specific wavelengths of light or light with wavelengths beyond the optical range, including infrared (IR) or near-infrared. In some embodiments, the imaging sensor is a multispectral camera that collects images in multiple different wavelength ranges. An illumination system may also be used to illuminate the embryo with light of a specific wavelength, wavelength band, or intensity. Stops and other components may be used to limit or modify illumination to specific portions (or image planes) of the image.

[0069] Additionally, images used in the embodiments described herein may be sourced from video and time-lapse imaging systems. A video stream is a periodic sequence of image frames, with the spacing between image frames determined by the capture frame rate (e.g., 24 or 48 frames per second, etc.). Similarly, a time-lapse system captures a sequence of images at a very slow frame rate (e.g., one image per hour) to obtain a sequence of images as the embryo develops (post-fertilization). Thus, it will be understood that the image used in the embodiments described herein may be a single image extracted from a video stream or time-lapse sequence of images of the embryo. If an image is extracted from a video stream or time-lapse sequence, the image to be used may be selected as the image having a capture time closest to a reference time point, such as 5.0 or 5.5 days post-fertilization.

[0070] In some embodiments, preprocessing may include image quality assessment so that an image can be rejected if it does not meet the image quality criteria. If the original image does not meet the image quality criteria, additional images can be captured. In embodiments in which images are selected from a video stream or time-lapse sequence, the selected image is the first image that passes the image quality assessment closest to the reference time. Alternatively, a reference time window can be defined along with the image quality criteria (e.g., 30 minutes after the beginning of day 5.0). In this embodiment, the selected image is the selected image with the highest image quality during the reference time window. The image quality criteria used in performing the image quality assessment may be based on pixel color distribution, brightness range, and / or abnormal image characteristics or features indicative of poor image quality or device malfunction. A threshold may be determined by analyzing a reference set of images. This may be based on manual assessment or an automated system that extracts outliers from the distribution.

[0071] The creation of the aneuploidy screening AI model 150 can be further understood with reference to Figure 5A, which is a schematic flow diagram of the creation of the aneuploidy screening AI model 150 using a cloud-based computing system 1 configured to create and use an AI model 100 configured to estimate the presence of aneuploidy (including mosaicism) in an image, according to one embodiment. With reference to Figure 5B, this creation method is handled by a model monitor 21.

[0072] The model monitor 21 allows users 40 to provide image data and metadata 14 to a data management platform, including a data repository. Data preparation steps are performed, for example, to move images to specific folders, rename images, and perform preprocessing on the images, such as object detection, segmentation, alpha channel removal, padding, cropping / localization, normalization, and scaling. Feature descriptors can be calculated and augmented images can be generated in advance. However, further preprocessing, including augmentation, can also be performed during training (i.e., on the fly). Images can also undergo image quality assessment to enable rejection of clearly inferior images and capture of replacement images. Similarly, patient records or other clinical data can be processed (prepared) for additional viability classifications (e.g., viable or nonviable, presence and type of aneuploidy, etc.), which can be linked or associated with each image for use in training machine learning and deep learning models. The prepared data, along with the latest versions of the training algorithms, are loaded (16) into a template server 28 on a cloud provider (e.g., AWS). The template server is stored and multiple copies are made across different training server clusters 37 (which may be CPU, GPU, ASIC, FPGA, or TPU (tensor processing unit) based) that form the training servers 35.

[0073] The model monitor web server 31 then applies each job submitted by a user 40 to a training server 37 from multiple cloud-based training servers 35. Each training server 35 runs pre-prepared code (from the template server 28) to train an AI model using libraries such as PyTorch, Tensorflow, or equivalent, and may also use computer vision libraries such as OpenCV. PyTorch and OpenCV are open-source libraries with low-level commands for building CV machine learning models.

[0074] The training server 37 manages the training process. This may include, for example, dividing images into training, validation, and blind validation sets using a random assignment process. Additionally, during training / validation cycles, the training server 37 may randomize the set of images at the beginning of the cycle so that a different subset of images is analyzed or analyzed in a different order for each cycle. If preprocessing was not previously performed or was incomplete (e.g., during data management), further preprocessing may be performed, including object detection, segmentation, and generation of a masked dataset (e.g., IZC-only images), calculation / estimation of CV feature descriptors, and generation of data augmentation. Preprocessing may also include padding, normalization, etc., as needed. That is, preprocessing step 102 may be performed prior to training, during training, or some combination (i.e., distributed preprocessing). The number of running training servers 35 can be managed from a browser interface. As training progresses, logging information about the status of training is recorded (62) to a distributed logging service, such as CloudWatch 60. Key patient and accuracy information is also parsed from the logs and stored in a relational database 36. The models are also periodically saved (51) to data storage (e.g., AWS Simple Storage Service (S3) or a similar cloud storage service) 50 so that they can be retrieved and loaded at a later date (e.g., for restart in case of an error or other outage). Email updates about the training server status are sent (44) to users 40 when the training server job is complete or if an error is encountered.

[0075] Within each training cluster 37, numerous processes take place. When the cluster is started via the web server 31, a script automatically runs, which reads prepared images and patient records and launches the specific Pytorch / OpenCV training code requested (71). Input parameters 28 for model training are provided by the user 40 via the browser interface 42 or via a configuration script. The training process 72 is then initiated for the requested model parameters, which can be a lengthy and intensive task. Therefore, to avoid losing progress while training is in progress, logs are periodically saved to a logging (e.g., AWS CloudWatch) service 60 (62), and the current version of the model (as it is being trained) is saved to a data (e.g., S3) storage service for later retrieval and use (51). One embodiment of a schematic flow diagram of the model training process on the training server is shown in FIG. 5B. Various trained AI models are available on data storage services, and multiple models can be combined using, for example, ensemble, distillation, or similar approaches to incorporate various deep learning models (e.g., PyTorch, etc.) and / or targeted computer vision models (e.g., OpenCV, etc.) to create a more robust aneuploidy screening AI model 100 that is provided to the cloud-based delivery platform 30.

[0076] The cloud-based delivery platform (or system) 30 then allows the user 10 to drag and drop images directly into a web application 34, which prepares the image and passes it through a trained / validated aneuploidy screening AI model 30 to obtain a viability score (or aneuploidy risk), which is immediately returned in a report (as illustrated in FIG. 4 ). The web application 34 also allows clinics to store data such as images and patient information in a database 36, generate various reports on the data, and use tools for their organization, group, or specific users, as well as generate audit reports on billing and user accounts (e.g., create users, delete users, reset passwords, change access levels, etc.). The cloud-based delivery platform 30 also allows product administrators to access the system to create new customer accounts and users, reset passwords, and access customer / user accounts (including data and screens) to facilitate technical support.

[0077] Various steps and variations in the creation of embodiments of an AI model configured to estimate aneuploidy risk / embryo viability scores from images will now be discussed in further detail. Referring to FIG. 3 , the model was trained using images captured on day 5 post-fertilization (i.e., the 24-hour period from day 5:00:00 to day 5:23:59). However, as noted above, an effective model can still be developed using a shorter time window, such as 12 hours, a longer time window of 48 hours, or even a zero time window (i.e., an open-ended window). Additional images can be obtained on other days, such as days 1, 2, 3, or 4, or a minimum time after fertilization (e.g., an open-ended time window), such as at least day 3 or at least day 5. However, it is generally preferred (but not necessarily required) that the images used for training the AI ​​model and the images used for subsequent classification by the trained AI model be taken during similar, preferably the same, time window (e.g., the same 12-, 24-, or 48-hour time window).

[0078] Prior to analysis, each image undergoes a preprocessing (image preparation) procedure. Various preprocessing steps or techniques may be applied. These may be performed after addition to the data store 14 or during training by the training server 37. In some embodiments, an object detection (localization) module is used to detect and identify the location of the image relative to the embryo. Object detection / localization involves estimating a bounding box containing the embryo. This can be used for image cropping and / or segmentation. The image may also be padded with a given boundary, and then color-balanced and brightness-normalized. The image is then cropped so that the region outside the embryo is closer to the image boundary. This is achieved using computer vision techniques for boundary selection, including the use of AI object detection models.

[0079] Image segmentation is a computer vision technique useful for preparing images for a particular model to select relevant regions for model training that will be of interest, such as the intrazona pellucida cavity (IZC), individual cells within the embryo (i.e., cell boundaries to help identify mosaicism), or other regions such as the zona pellucida. As outlined above, mosaicism occurs when different cells in an embryo have different sets of chromosomes. That is, a mosaic embryo is a mixture of euploid (chromosomally normal) and aneuploid (with extra / missing / altered chromosomes) cells, and multiple different aneuploidies may be present, and in some cases, no euploid cells may be present. Segmentation can be used to identify IZCs or cell boundaries, thus segmenting the embryo into individual cells. In some embodiments, multiple masked (expanded) images of the embryo are generated, each masked except for a single cell. Images may also be masked to eliminate the zona pellucida and background, thus generating an image of only the IZC, or these may remain in the image. An aneuploidy AI model may then be trained using masked images, such as IZC images masked to include only IZCs, or images masked to identify individual cells in the embryo. Scaling involves rescaling the image to a predefined scale to fit the particular model being trained. Augmentation involves incorporating small changes to a copy of the image, such as rotating the image to control the orientation of the embryo dish. The use of segmentation prior to deep learning was found to have a significant impact on the performance of the deep learning method. Similarly, augmentation was important for creating robust models.

[0080] Various image pre-processing techniques can be used to prepare embryo images for analysis, for example to ensure that images are standardized. For example: Alpha channel stripping, which involves stripping the image of its alpha channel (if present), for example to remove the transparency map, and ensuring that it is coded in a three-channel format (e.g., RGB, etc.); Pad / bolster each image with a padding border to create a square aspect ratio prior to segmentation, cropping, or boundary finding; this process ensures that image dimensions are consistent, comparable, and compatible with deep learning methods that typically require square images as input, while also ensuring that key components of the image are not cropped; Normalizing RGB (Red-Green-Blue) or grayscale images to a fixed average value for all images, for example this involves taking the average of each RGB channel and dividing each channel by that average, then multiplying each channel by a fixed value of 100 / 255 to ensure that the average value of each image in RGB space is (100,100,100), this step ensures that color bias between images is suppressed and that the brightness of each image is normalized; Image thresholding using binary, Otsu, or adaptive methods, including dilation (opening), erosion (closing), scale gradients, and morphological processing of images using scale masks to extract outer and inner boundaries of shapes; Performing object detection / cropping of the image to locate the image relative to the embryo and ensure there are no artifacts around the edges of the image, which may be done using an object detector that uses an object detection model (discussed below) trained to estimate a bounding box that contains the main features of the image, such as the embryo (IZC or zona pellucida), so that the image is well-centered and the embryo is cropped; extracting boundary geometric characteristics using an elliptical Hough transform of the image contour, for example, a best ellipse fit from an elliptical Hough transform computed on a binary threshold map of the image, the method working by selecting a hard boundary of the embryo in the image and cropping a square boundary of the new image so that the longest radius of the new ellipse is encompassed by the width and height of the new image and the center of the ellipse is the center of the new image; Zooming an image by ensuring a consistently centered image with a consistent border size around an elliptical region; Segmenting the image to identify intrazona pellucida (IZC) regions, zona pellucida regions, and / or cell boundaries, where segmentation may be performed by calculating a best-fit contour around a non-elliptical image using a geometric active contour (GAC) model or morphological snakes within a given region, and the interior and other regions of the snake may be treated differently depending on the trained model's focus on intrazona pellucida (IZC) or cells within the blastocyst, which may contain the blastocyst; or a semantic segmentation model may be trained to identify a class for each pixel in the image, for example, a semantic segmentation model deployed using a U-Net architecture with a ResNet-50 encoder pre-trained to segment the background, zona pellucida, and IZC, or to segment cells within the IZC, and training this model using a binary cross-entropy loss function; Annotating an image by selecting a feature descriptor and masking all regions of the image except those within a given radius of the descriptor's keypoints; Resize / scale the entire set of images to a specified resolution; a tensor transform that involves converting each image into a tensor rather than a visually displayable image, as this data format is more usable by deep learning models, and in one embodiment, tensor normalization is obtained from standard pre-trained ImageNet values, e.g., with mean (0.485, 0.456, 0.406) and standard deviation (0.299, 0.224, 0.225); Examples include:

[0081] In another embodiment, the object detector uses an object detection model trained to estimate bounding boxes containing germs. The goal of object detection is to identify the largest bounding box that contains all of the pixels associated with the object. This requires the model to model both the object's location and category / label (i.e., what is in the box); therefore, the detection model typically has both an object classifier head and a bounding box regression head.

[0082] One approach is to apply a region-based convolutional neural network (or R-CNN) that uses an expensive search process to search for image patch proposals (potential bounding boxes). These bounding boxes are then used to crop regions of interest. A classification model is then run on the cropped image to classify the content of the image region. This process is complex and computationally expensive. An alternative is Fast CNN, which uses a CNN to propose feature regions rather than searching for image patch proposals. This model uses a CNN to estimate a fixed number of candidate boxes, typically set between 100 and 2000. An even faster alternative approach is Faster RCNN, which uses anchor boxes to limit the search space of required boxes. By default, a standard set of nine anchor boxes (each of a different size) is used. Faster RCNN uses a small network jointly trained to predict feature regions of interest, which can replace expensive region searches and therefore improve runtime compared to R-CNN or Fast CNN.

[0083] For every feature activation coming in, one model is considered the anchor point (red in the image below). For every anchor point, nine anchor boxes are generated (or more or less, depending on the problem). The anchor boxes correspond to common object sizes in the training dataset. Because there are many anchor points with many anchor boxes, this results in tens of thousands of region proposals. The proposals are then filtered through a process called Non-Maximal Suppression (NMS), which selects the largest box with a confident smaller box within it. This ensures that there is only one box per object. Because NMS relies on the reliability of each bounding box prediction, a threshold must be set for when objects are considered part of the same object instance. Because the anchor boxes do not fit the object perfectly, the job of the regression head is to predict offsets to these anchor boxes that transform into the best-fit bounding box.

[0084] Detectors can also be specialized, estimating boxes for only a subset of objects, such as only people for a pedestrian detector. Object categories of uninteresting interest are encoded as class 0, corresponding to the background class. During training, patches / boxes for the background class are typically randomly sampled from image regions that lack bounding box information. This step allows the model to become invariant to these undesirable objects, e.g., learn to ignore them rather than incorrectly classify them. Bounding boxes are typically represented in two different formats, the most common being (x1, y1, x2, y2), where point p1 = (x1, y1) is the top-left corner of the box and p2 = (x2, y2) is the bottom-right corner. Another common box format is (cx, cy, height, width), where the bounding box / rectangle is encoded as the center point (cx, cy) of the box and the size (height, width) of the box. The detection method uses different encodings / formats depending on the task and situation.

[0085] The regression head may be trained using an L1 loss, and the classification head may be trained using a cross-entropy loss. Objectness loss (is this background or object?) may also be used, and the final loss is calculated as the sum of these losses. Individual losses may also be weighted as follows:

[0086]

number

[0087] As an alternative to GAC segmentation, semantic segmentation may be used. Semantic segmentation is the task of predicting a category or label for every pixel. Tasks like semantic segmentation are called pixel-by-pixel dense prediction tasks because an output is required for every input pixel. Semantic segmentation models are configured differently from standard models because they require full image output. Typically, semantic segmentation (or any dense prediction model) will have an encoding module and a decoding module. The encoding module is responsible for creating a low-dimensional representation (sometimes called a feature representation) of the image. This feature representation is then decoded into the final output image via the decoding module. During training, the predicted label map (for semantic segmentation) is then compared to a ground truth label map that assigns a category to each pixel, and a loss is calculated. The standard loss function for segmentation models is either binary cross-entropy or standard cross-entropy loss (depending on whether the problem is multi-class). These implementations are identical to their image classification cousins, except that the loss is applied pixel-by-pixel (across the image channel dimensions of the tensor).

[0088] Fully convolutional network (FCN)-style architectures are commonly used in the field for general semantic segmentation tasks. In this architecture, a pre-trained model (such as ResNet) is first used to encode a low-resolution image (approximately 1 / 32 of the original resolution, but could be as low as 1 / 8 if dilated convolutions are used). This low-resolution label map is then upsampled to the original image resolution, and a loss is computed. The intuition behind predicting the low-resolution label map is that the semantic segmentation mask is very low frequency, avoiding the need for all the extra parameters of a larger decoder. More complex versions of this model exist that use multi-stage upsampling to improve segmentation results. Briefly, a loss is computed at multiple resolutions in an incremental fashion to refine the prediction at each scale.

[0089] One drawback of this type of model is that when the input data is high-resolution or contains high-frequency information (i.e., smaller / thinner objects), the low-resolution label map cannot capture these small structures (especially if the encoding model does not use dilated convolutions). In a standard encoder / convolutional neural network, the input image / image features are progressively downsampled as the model gets deeper. However, as the image / feature is downsampled, key high-frequency details may be lost. Therefore, to address this, an alternative U-Net architecture can be used that instead uses skip connections between symmetric components of the encoder and decoder. Briefly, every encoding block has a corresponding block in the decoder. Then, the features at each stage are passed to the decoder along with the lowest-resolution feature representation. For each decoding block, the input feature representation is upsampled to match the resolution of its corresponding encoding block. Next, the feature representations from the encoding block and the upsampled low-resolution features are concatenated and passed through a 2D convolutional layer. By concatenating features in this way, the decoder can learn to refine the input at each block, choosing which details (low-resolution details or high-resolution details) to integrate depending on the input. The key difference between FCN-style models and U-Net-style models is that in FCN models, the encoder is responsible for predicting a low-resolution label map and then (possibly progressively) upsampling it. On the other hand, U-Net models do not have fully complete label map predictions until the final layer. Finally, many variants of these models (e.g., hybrids, etc.) exist, trading off the differences between these models. The U-net architecture can also use pre-trained weights, such as ResNet-18 or ResNet-50, for use when there is insufficient data to train a model from scratch.

[0090] In some embodiments, segmentation was performed using a U-Net architecture with a pre-trained ResNet-50 encoder trained using binary cross-entropy to identify zona pellucida regions, intrazona pellucida cavity regions, and / or cell boundaries. Once segmented, an image set can be generated in which all regions except the desired region are masked. AI models can then be trained on these specific image sets. That is, AI models can be divided into two groups: first, those that involve further image segmentation, and second, those that require entirely unsegmented images. Models trained on images in which the IZC was masked and the zona pellucida region was exposed are referred to as zona pellucida models. Models trained on images with the zona pellucida masked (referred to as IZC models) and models trained on full embryo images (i.e., the second group) were also considered for training.

[0091] In one embodiment, to ensure the uniqueness of each image, the name of the new image is set equal to a hash of the original image content as a png (lossless) file so that duplicate records do not bias the results. When executed, the data parser will output images in a multi-threaded manner for any images that do not already exist in the output directory (and will create them if they do not exist), so that lengthy processes can be restarted from the same point if interrupted. The data preparation step may also include processing metadata to eliminate images associated with inconsistent or inconsistent records and identify any erroneous clinical records. For example, a script can be run on a spreadsheet to conform the metadata to a predefined format. This ensures that the data used to create and train the model is of high quality and has uniform characteristics (e.g., size, color, scale, etc.).

[0092] Once the data is properly prepared, it can then be used to train an AI model, as described above. In one embodiment, multiple computer vision (CV) models are created using machine learning methods, and multiple deep learning models are created using deep learning methods. The deep learning models may be trained on full embryo images or a set of masked images. The computer vision (CV) models may be created using machine learning methods using a set of feature descriptors calculated from each image, with each individual model configured to estimate a likelihood, such as an aneuploidy risk / embryo viability score, of the embryo in the image. The AI ​​model combines selected models to generate an overall aneuploidy risk / embryo viability score, or similar overall likelihood or hard classification. Models created for individual chromosome groups can be improved using ensemble and knowledge distillation techniques. Training is performed using a randomized dataset. Complex image data sets may suffer from uneven distribution, particularly when the dataset is smaller than approximately 10,000 images, where primary viable or nonviable embryo examples are not evenly distributed throughout the set. Thus, several (e.g., 20) randomizations of the data are considered at a time and then divided into training, validation, and blind test subsets, as defined below. All randomizations are used on a single training example to determine which one provides the best distribution for training. As a corollary, it is also beneficial to ensure that the ratio of viable to nonviable embryos is the same across all subsets. Embryo images are highly diverse, and therefore using a uniform distribution of images across the test and training sets ensures that performance can be improved. Thus, after randomization, the ratio of images with a viable classification to images with a nonviable classification in each of the training, validation, and blind validation sets is calculated and tested to ensure that the ratios are similar. For example, this may include testing whether the range of the ratios is within a certain variance given the number of images or is below a threshold.If the ranges are not similar, the randomization is discarded and a new randomization is generated and tested until a randomization is obtained whose ratios are similar. More generally, if the outcome is an n-ary outcome with n states, after the randomization is performed the calculation step may include calculating the frequency of each of the n-ary outcome states in each of the training set, validation set, and blind validation set, and testing that the frequencies are similar; if the frequencies are not similar, discarding the assignment and repeating the randomization until a randomization is obtained whose frequencies are similar.

[0093] The training further includes performing multiple training / validation cycles. In each training / validation cycle, each randomization of the total available dataset is divided into typically three separate datasets known as the training, validation, and blind validation datasets. In some variations, four or more datasets can be used, for example, the validation and blind validation datasets can be stratified into multiple sub-test sets of different difficulty.

[0094] The first set is the training dataset, which contains at least 60%, preferably 70-80%, of the images. These images are used by deep learning and computer vision models to create an aneuploidy screening AI model and accurately identify viable embryos. The second set is the validation dataset, which typically contains approximately (or at least) 10% of the images. This dataset is used to verify or test the accuracy of the model created using the training dataset. Although these images are independent of the training dataset used to create the model, the validation dataset still has a small positive bias in accuracy because it is used to monitor and optimize the progress of model training. Therefore, training tends to target models that maximize the accuracy of this specific validation dataset, which is not necessarily the best model when applied more generally to other embryo images. The third dataset is the blind validation dataset, which typically contains approximately 10-20% of the images. To address the positive bias using the validation dataset, the third blind validation dataset is used to perform a final, unbiased accuracy assessment of the final model. This validation occurs at the end of the modeling and validation process, when a final model is created and selected. It is important to ensure that the accuracy of the final model is relatively consistent with the validation dataset to ensure that the model can generalize to all embryo images. For the reasons stated above, the accuracy of the validation dataset is likely to be higher than that of the blind validation dataset. The results of the blind validation dataset are a more reliable measure of the model's accuracy.

[0095] In some embodiments, preprocessing the data further includes augmenting the images, where modifications are made to the images. This may be done prior to training or during training (i.e., on the fly). Augmentation may involve augmenting (modifying) the images directly or by creating a copy of the image with small changes. Any number of augmentations can be performed, including rotating, mirroring, flipping, non-90 degree rotation of the image where a diagonal border is embedded to match the background color, varying the amount of image blur, adjusting the image contrast using an intensity histogram, and applying one or more small random transformations in both horizontal and / or vertical directions, random rotation, adding JPEG (or compression) noise, random image resizing, random hue jitter, random brightness jitter, contrast-limited adaptive histogram equalization, random flip / mirror, image sharpening, image embossing, random brightness and contrast, RGB color shift, random hue and saturation, channel shuffle, swapping RGB to BGR or RBG or other, coarse dropout, motion blur, center blur, Gaussian blur, random shift-scale rotation (i.e., all three combined). The same set of augmented images may be used for multiple training-validation cycles, or new augmentations may be generated on the fly during each cycle. A further extension used in CV model training is changing the "seed" of the random number generator for extracting feature descriptors. Techniques for obtaining computer vision descriptors have an element of randomness in extracting feature samples. This random number can be changed and included between extensions to provide more robust training for the CV model.

[0096] Computer vision models rely on identifying key image features and representing them in terms of descriptors. These descriptors can encode qualities such as pixel variation, gray level, texture roughness, fixed corner points, or image gradient orientation, and are implemented in OpenCV or similar libraries. By selecting such features to search for in each image, a model can be built by discovering which feature configurations are good indicators of aneuploidy / embryo viability. This procedure is best performed by machine learning processes such as random forests or support vector machines, which can segment images in terms of their descriptors from computer vision analysis.

[0097] A variety of computer vision descriptors, encompassing both small-scale and large-scale features, are used and combined with traditional machine learning methods to generate "CV models" for identifying aneuploidy and mosaicism. These can then optionally be combined with deep learning (DL) models, e.g., into ensemble models, or used in distillation to train student models. Suitable computer vision image descriptors include: Zona pellucida via Hough transform: find inner and outer ellipses to approximate the split of the zona pellucida and the intrazonal cavity, and record the mean and difference of the radii as features; Gray-Level Co-occurrence Matrix (GLCM) texture analysis: detects the roughness of different regions by comparing adjacent pixels within the region. Sample feature descriptors used are: angular second moment (ASM), uniformity, correlation, contrast, and entropy. The selection of regions is obtained by randomly sampling a given number of square subregions of the image of a given size, and the results of each of the five descriptors for each region are recorded as a set of total features; Histogram of Oriented Gradients (HOG): Detects objects and features using scale-invariant feature transformation descriptors and shape context. This method has gained popularity for use in embryology and other medical imaging, but does not constitute a machine learning model in itself; Feature Extraction by Oriented Accelerated Fragmentary Testing (FAST) and Rotational Binary Robust Independent Basic Features (BRIEF) (ORB): an industry standard alternative to SIFT and SURF features, relying on a combination of FAST keypoint detectors (specific pixels) and BRIEF descriptors, modified to include rotational invariance; Binary Robust Invariant Scalable Keypoint (BRISK): A FAST-based detector combined with an assembly of pixel intensity comparisons, which is achieved by sampling each neighborhood around the feature specified by the keypoint; Maximum Stable Extremal Region (MSER): A local morphological feature detection algorithm via the extraction of covariant regions, which are stable connected components with respect to one or more gray level sets extracted from an image; Feature-Friendly Tracking (GFTT): A feature detector that uses an adaptive window size to detect corner textures, identified by using Harris or Shi-Tomasi corner detection and extracting points that exhibit high standard deviations in their spatial intensity profile.

[0098] A computer vision (CV) model is constructed by the following method: One (or more) of the computer vision image descriptor techniques described above is selected, and features are extracted from all of the images in the training dataset. These features are arranged into a combined array and then fed into a K-means unsupervised clustering algorithm; this array is called a codebook for "bag of visual words." The number of clusters is a free parameter of the model. The clustered features from this point represent "custom features" used throughout any combination of algorithms, to which each individual image in the validation or test set will be compared. Each image, with its extracted features, is individually clustered. For a given image with its clustered features, its "distance" (in feature space) to each of the clusters in the codebook is measured using a KD tree query algorithm that yields the closest clustered feature. The results from the tree query can then be represented as a histogram, showing the frequency with which each feature occurs in that image. Finally, the question of whether a particular combination of these features corresponds to a measure of aneuploidy risk / embryo viability needs to be evaluated using machine learning. Here, supervised learning is performed using the histogram and ground truth results. Methods used to obtain the final selection model include random forests or support vector machines (SVMs).

[0099] Multiple deep learning models can also be created. Deep learning models are based on neural network methods, typically convolutional neural networks (CNNs) consisting of multiple connected layers, where each layer of "neurons" has a nonlinear activation function such as "rectifier" or "sigmoid." In contrast to feature-based methods (i.e., CV models), deep learning and neural networks do not rely on manually designed feature descriptors but instead "learn" features. This allows them to learn "feature representations" tailored to the desired task.

[0100] These methods are suitable for image analysis because they can pick up both small details and overall morphological shapes to arrive at an overall classification. Various deep learning models are available, each with a different architecture (i.e., a different number of layers and inter-layer connections), such as residual networks (e.g., ResNet-18, ResNet-50, and ResNet-101), densely connected networks (e.g., DenseNet-121 and DenseNet-161), and other variations (e.g., InceptionV4 and Inception-ResNetV2). Deep learning models can be evaluated based on stability (how stable the accuracy values ​​were on the validation set throughout the training process), transferability (how well the accuracy on the training data correlated with the accuracy on the validation set), and prediction accuracy (which model provided the best validation accuracy, total accuracy, and balanced accuracy, defined as the weighted average accuracy across both embryo class types, for both viable and nonviable embryos). Training involves trying different combinations of model parameters and hyperparameters, including input image resolution, optimization algorithm selection, learning rate value and scheduling, momentum value, dropout, and weight initialization (pre-training). A loss function may be defined to evaluate the model's performance, and during training, the deep learning model is optimized by varying the learning rate to drive an update mechanism for the network's weight parameters to minimize the objective / loss function.

[0101] Deep learning models may be implemented using various libraries and software languages. In one embodiment, the PyTorch library is used to implement neural networks in the Python language. The PyTorch library also allows tensors to be created that take advantage of hardware (GPU, TPU) acceleration and includes modules for building multiple layers for neural networks. While deep learning is one of the most powerful techniques for image classification, it can be improved by providing guidance through the use of segmentation or augmentation described above. The use of segmentation prior to deep learning has been found to significantly impact the performance of deep learning methods and aid in the creation of contrasting models. Therefore, preferably, at least some deep learning models are trained on segmented images, such as images in which IZCs or cell boundaries have been identified, or images that have been masked to exclude regions outside the IZCs or cell boundaries. In some embodiments, the multiple deep learning models include at least one model trained on segmented images and one model trained on images that have not undergone segmentation. Similarly, augmentation was important for creating robust models.

[0102] The effectiveness of the approach is determined by the architecture of the deep neural network (DNN). However, unlike feature descriptor methods, DNNs learn the features themselves through convolutional layers before applying the classifier. That is, without manually incorporating proposed features, DNNs can be used to check existing practices in the literature as well as develop previously unguessed descriptors, especially those that are difficult for the human eye to detect and measure.

[0103] The architecture of a DNN is constrained by the size of the image as input, a hidden layer with the dimensions of a tensor describing the DNN, and a linear classifier with the number of class labels as output. Most architectures utilize multiple downsampling ratios with small (3x3 pixel) filters to capture the concepts of left / right, up / down, and center. A stack of a) convolutional 2d layers, b) rectified linear units (ReLUs), and c) max-pooling layers allows the number of parameters passed through the DNN to remain manageable while allowing the filters to pass over high-level (topological) features of the image and map them onto intermediate and finally microscopic features embedded in the image. The top layer typically includes one or more fully connected neural network layers that act as classifiers, similar to SVMs. A softmax layer is typically used to normalize the resulting tensor so that it has the probability of being the result of the fully connected classifier. Thus, the output of the model is a list of probabilities that the image is either nonviable or viable. The range of AI architectures may be based on neural network architectures such as ResNet varieties (18, 34, 50, 101, 152), Wide ResNet varieties (50-2, 101-2), ResNeXt varieties (50-32x4d, 1-1-32x8d), DenseNet varieties (121, 161, 169, 201), Inception (v4), Inception-ResNet (v2), and EffficientNet varieties (b0, b1, b2, b3).

[0104] FIG. 5C is a schematic architecture diagram of an AI model 151, according to one embodiment, including a series of layers based on the RESNET 152 architecture that converts input images into predictions. These include a 2D convolutional layer, annotated as "CONV" in FIG. 5C, which calculates the cross-correlation of inputs from the layer below. Each element, or neuron, in a convolutional layer processes inputs only from its receptive field, e.g., 3x3 or 7x7 pixels. This reduces the number of learnable parameters required to describe the layer, allowing for the creation of deeper neural networks than those built from fully connected layers. In these layers, every neuron is connected to every other neuron in the subsequent layer, which is memory intensive and prone to overfitting. Convolutional layers are also spatially invariant, which is useful for processing images where the subject cannot be guaranteed to be precisely centered. The AI ​​architecture in FIG. 5C further includes a max-pooling layer, annotated as "POOL" in FIG. 5C, which is a downsampling method that selects only representative neuron weights within a given region, reducing network complexity and overfitting. For example, for weights within a 4x4 square region of a convolutional layer, the maximum value of each 2x2 corner block is calculated, and these representative values ​​are then used to reduce the size of the square region to a dimension of 2x2. The architecture may also include the use of a rectified linear unit to act as a nonlinear activation function. As a common example, a ramp function may be of the following form for an input x from a given neuron:

[0105]

number

[0106] The final layer at the end of the network, after the input has passed through all of the convolutional layers, is typically a fully connected (FC) layer, which acts as a classifier. This layer takes the final input and outputs an array of dimensions equal to the number of classification categories. For example, for two categories, such as "aneuploidy present" and "aneuploidy absent," the final layer would output an array of length 2, indicating the proportion of input images with features that align with each category. A final softmax layer is often added, which converts the final numbers in the output array into a matching percentage between 0 and 1, adding both together to sum to 1, so that the final output can be interpreted as a confidence limit for the image's classification into one of the categories.

[0107] One suitable DNN architecture is ResNet (and its variants; see https: / / ieeexplore.ieee.org / document / 7780459), such as ResNet152, ResNet101, ResNet50, or ResNet-18. ResNet significantly advanced the field in 2016 by using a significantly larger number of hidden layers and by introducing "skip connections," also known as "residual connections." Only the differences from one layer to the next are calculated, which is more time-effective; if little change is detected in a particular layer, that layer is skipped, thus creating a network that adapts very quickly to a combination of small and large features in an image.

[0108] Another suitable DNN architecture is the DenseNet variant (https: / / ieeexplore.ieee.org / document / 8099726), which includes DenseNet161, DenseNet201, DenseNet169, and DenseNet121. DenseNet is an evolution of ResNet, where every layer now has a maximum number of skip connections and can skip to any other layer. This architecture requires much more memory and is therefore less efficient, but can show improved performance over ResNet. It is also easy to overtrain / overfit with a large number of model parameters. All model architectures are often combined with methods to control this.

[0109] Another suitable DNN architecture is Inception(-ResNet) (https: / / www.aaai.org / ocs / index.php / AAAI / AAAI17 / paper / viewPaper / 14806), such as InceptionV4 and InceptionResNetV2. Inception represents a more complex convolutional unit, whereby instead of simply using fixed-size filters (e.g., 3x3 pixels) as described in Section 3.2, filters of several sizes are computed in parallel (5x5, 3x3, 1x1 pixels) with weights that are free parameters, allowing the neural network to prioritize which filters are most suitable at each layer in the DNN. A development of this kind of architecture is to combine it with skip connections in the same way as ResNet, creating Inception-ResNet. As mentioned above, both computer vision and deep learning methods are trained on pre-processed data using multiple training and validation cycles, which follow the following framework:

[0110] The training data is preprocessed and split into batches (the number of data in each batch is a free model parameter, controlling how fast and how stably the algorithm learns). Augmentation may be performed prior to splitting or during training.

[0111] After each batch, the network weights are adjusted and the running total accuracy is evaluated. In some embodiments, the weights are updated between batches, for example, using gradient accumulation. When all images have been evaluated and one epoch has run, the training set is shuffled (i.e., a new randomization with the set is obtained) and training starts again from the beginning for the next epoch.

[0112] During training, a number of epochs may be performed depending on the size of the dataset, the complexity of the data, and the complexity of the model being trained. The optimal number of epochs typically ranges from 2 to 100, but may be higher depending on the specific case. After each epoch, the model is run on a validation set without any training to provide a progress measure in how accurate the model is and guide the user on whether more epochs should be performed or whether more epochs would result in overtraining.

[0113] The validation set guides the selection of all model parameters or hyperparameters and is therefore not a truly blind set, but it is important that the distribution of images in the validation set is very similar to the final blind test set that will be run after training.

[0114] When reporting validation set results, augmentations can also be included (all) or not (noaug) for each image. Additionally, augmentations for each image may be combined to provide a more robust final result for the image. Several combination / voting strategies can be used, including average confidence (taking the average of the model's inferences across all augmentations), median confidence, population mean confidence (taking a population viability assessment and providing only the average confidence of those that agree; if a majority is not reached, taking the average), maximum confidence, weighted average, population maximum confidence, etc.

[0115] Another method used in the field of machine learning is transfer learning, in which a previously trained model is used as a starting point for training a new model. This is also called pre-training. Pre-training is widely used and allows new models to be built quickly. There are two types of pre-training. One embodiment of pre-training is ImageNet pre-training. Most model architectures are provided with a set of pre-trained weights using the standard image database ImageNet. Although this is not specific to medical images and contains 1000 different types of objects, it provides a way for the model to already learn to identify shapes. The classifier for the 1000 objects is completely removed, and a new classifier for viability replaces it. This type of pre-training outperforms other initialization strategies. Another embodiment of pre-training is custom pre-training, which uses previously trained embryonic models from studies with different outcome sets or different images (PGS instead of viability, or randomly assigned outcomes). These models provide only a small benefit to classification.

[0116] For models that have not undergone pretraining, or for new layers added after pretraining, such as classifiers, weights must be initialized. The initialization method can affect the success of training. For example, setting all weights to 0 or 1 results in very poor performance. A uniform distribution of random numbers or a Gaussian distribution of random numbers also represent commonly used options. These are often combined with regularization methods such as the Xavier or Kaiming algorithms. This addresses the problem that nodes in a neural network can become "trapped" in a particular state by becoming saturated (close to 1) or dead (close to 0), making it difficult to determine the direction in which to adjust the weight associated with that particular neuron. This is particularly prevalent when introducing hyperbolic tangent or sigmoid functions, and is addressed by Xavier initialization.

[0117] In the Xavier initialization protocol, the weights of a neural network are randomized so that each layer's input to the activation function is not too close to either the extremes of saturation or the extremes of dead. However, using ReLU works better, and different initializations offer smaller advantages, such as Kaiming initialization. Kaiming initialization is more suitable when ReLU is used as the nonlinear activation profile for neurons. It effectively achieves the same process as Xavier initialization.

[0118] In deep learning, various free parameters are used to optimize model training on the validation set. One key parameter is the learning rate, which is determined by how much the underlying neuron weights are adjusted after each batch. When training a selection model, overtraining or overfitting the data should be avoided. This occurs when the model has too many parameters to fit and essentially "memorizes" the data, trading generalizability for accuracy on the training or validation set. This is avoided because generalizability is the true measure of whether the model has accurately identified and perfectly fitted the training set, even amidst the noise in the data, the true underlying parameters that indicate embryonic health.

[0119] During the validation and testing phase, the success rate can suddenly drop due to overfitting during the training phase. This can be recovered through various strategies, including slowing or decaying the learning rate (e.g., halving the learning rate every n epochs), or using cosine annealing incorporating batch normalization or the tensor initialization or pretraining methods described above and adding noise such as dropout layers. Batch normalization is used to combat vanishing or exploding gradients and improves the stability of training large models, resulting in improved generalization. Dropout regularization effectively simplifies the network by introducing a random opportunity to set all incoming weights to zero within the rectifier's acceptance range. Introducing noise effectively ensures that the remaining rectifier accurately fits the representation of the data without relying on overspecialization. This allows the DNN to generalize more effectively and become less sensitive to specific values ​​of the network's weights. Similarly, batch normalization improves the training stability of very deep neural networks, allowing for faster learning and better generalization by shifting the input weights to zero mean and unit variance as a precursor to the rectification stage.

[0120] When performing deep learning, the methodology for modifying neuron weights to achieve acceptable classification involves the need to specify an optimization protocol. That is, many techniques need to be specified for a given definition of "accuracy" or "loss" (discussed below), exactly how much weights should be adjusted, and how learning rate values ​​should be used. Suitable optimization techniques include stochastic gradient descent (SGD) with momentum (and / or Nesterov's accelerated gradient method), adaptive gradient with delta (Adadelta), adaptive moment estimation (Adam), root-mean-square propagation (RMSProp), and the limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) algorithm. Of these, SGD-based techniques generally outperformed other optimization techniques. Typical learning rates for phase-contrast microscopy images of human embryos ranged from 0.01 to 0.0001. However, learning rate depends on the batch size, which in turn depends on hardware capacity. For example, larger GPUs allow for larger batch sizes and faster training speeds.

[0121] Stochastic gradient descent (SGD) with momentum (and / or Nesterov's accelerated gradient) represents the simplest and most commonly used optimization algorithm. Gradient descent algorithms typically calculate the gradient (slope) of a given weight's influence on accuracy. While this is slow if the gradient needs to be calculated for the entire dataset to make weight updates, stochastic gradient descent performs updates one at a time, for each training image. While this can introduce fluctuations in the overall target accuracy or loss achieved, it tends to generalize better than other methods because it can dive into new regions of the loss parameter landscape and find a new minimum loss function. SGD performs well against pronounced loss landscapes in difficult problems such as embryo selection. SGD can have trouble navigating asymmetric loss function surface curves that are steeper on one side than the other; this can be compensated for by adding a parameter called momentum. This accelerates SGD in that direction by adding an extra fraction to the weight updates derived from the previous state, helping to dampen high fluctuations in accuracy. A development of this method is to also include an estimate of the position of the weights in the next state; this development is known as Nesterov's accelerated gradient method.

[0122] Adaptive gradients with delta (Adadelta) is an algorithm for adapting the learning rate to the weights themselves, with smaller updates for frequently occurring parameters and larger updates for infrequently occurring features, making it well suited to sparse data. While this can suddenly slow down the learning rate after a few epochs across the entire dataset, adding a delta parameter to limit the window allowed the accumulated past gradients to a certain size. However, this process makes the default learning rate redundant, and the additional degrees of freedom of the free parameter provide some control in finding the best global selection model.

[0123] Adaptive moment estimation (Adam) stores exponentially decaying averages of both past squared and non-squared gradients and incorporates both into the weight updates. This has the effect of providing "friction" to the direction of the weight updates and is suitable for problems with relatively shallow or flat loss minima without large fluctuations. In embryonic selection models, training with Adam tends to perform well on the training set, but often overtrains and is not as suitable as SGD with momentum methods.

[0124] Root-mean-square propagation (RMSProp) is the adaptive gradient optimization algorithm described above, and is nearly identical to Adadelta, except that the update term for the weights divides the learning rate by an exponentially decaying average of the squared gradients.

[0125] The limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) algorithm. Though computationally intensive, the L-BFGS algorithm actually estimates the curvature of the loss landscape, rather than other methods that attempt to compensate for this lack of estimation with additional terms. For small datasets, it tends to outperform Adam, but does not necessarily outperform SGD in terms of speed or accuracy.

[0126] In addition to the above methods, it is also possible to include non-uniform learning rates. That is, the learning rate of the convolutional layers can be specified to be much larger or smaller than the learning rate of the classifier. This is useful in the case of pre-trained models, where changes to the filters below the classifier should be kept "frozen" and the classifier retrained, so that the pre-training is not undone by further retraining.

[0127] Although an optimization algorithm specifies how to update weights given a particular loss or accuracy measure, in some embodiments, loss functions are modified to incorporate distributional effects. These may include cross-entropy (CE) loss, weighted CE, residual CE, inference distributions, or custom loss functions.

[0128] Cross-entropy loss is a commonly used loss function that tends to outperform simple mean squared error between ground truth and predictions. When the network's results are passed through a softmax layer, as is the case here, the distribution of cross-entropy leads to better accuracy. This is naturally to maximize the chances of correctly classifying the input data by de-emphasizing extreme outliers. For an input array, a batch, representing a batch of images, and classes representing viable or non-viable, the cross-entropy loss is:

[0129]

number

[0130]

number

[0131]

number

[0132] If the data has a class bias, i.e., more viable examples than non-viable examples (or vice versa), the loss function should be weighted proportionally so that misclassifying elements of the less abundant class is more heavily penalized. This is achieved by pre-multiplying the right-hand side of equation (2) by a factor:

[0133]

number

[0134] In some embodiments, an inferred distribution may be used. While it is important to seek a high level of accuracy in embryo classification, it is also important to seek a high level of transferability in the model. That is, it is often useful to understand the distribution of scores and understand that while high accuracy is an important goal, confidently distinguishing between viable and nonviable embryos is an indicator that the model generalizes well to the test set. Because accuracy on the test set is often used to cite comparisons to important clinical benchmarks, such as the accuracy of an embryologist's classification of the same embryos, ensuring generalizability should also be incorporated into the epoch-by-epoch, batch-by-batch evaluation of the model's success.

[0135] In some embodiments, a custom loss function is used. In one embodiment, we customized how the loss function is defined so that the optimization surface is modified to make the global optimum more apparent, thus improving the robustness of the model. To achieve this, a new term that maintains differentiability, called the residual term, defined in terms of the network weights, is added to the loss function. This encodes the collective difference in the predictions from the model and the target outcome for each image and includes it as an additional contribution to the normal cross-entropy loss function. The formula for the residual term is, for N images:

[0136]

number

[0137] This custom loss function considers well-spaced clusters of viable and non-viable embryo scores to be consistent with the improved loss estimate. Note that this custom loss function is not specific to embryo detection applications and can be used in other deep learning models.

[0138] In some embodiments, a custom confidence-based loss function is used. This is a weighted loss function with two variants (linear and non-linear). In both cases, it is intended to encode the separation of scores as contributions to the loss function, but in a different way from the above by integrating the difference between classes in predicted scores together as weights in the loss function. The larger the difference, the greater the loss reduction. This loss function will help drive the predictive model to magnify the difference between the two classes and increase the model's confidence in the results. For confidence weighting, the binary target label of the i-th input sample specifies the ground truth class.

[0139]

number

[0140]

number

[0141]

number

[0142]

number

[0143]

number

[0144] For the standard log-softmax function, we calculate p as follows: t (log(pt) is obtained in the loss function as the standard cross-entropy loss function).

[0145]

number

[0146]

number

[0147]

number

[0148]

number

[0149]

number

[0150]

number

[0151] In some embodiments, the models are combined to create a more robust final AI model 100, i.e., deep learning and / or computer vision models are combined together to contribute to the overall prediction of aneuploidy.

[0152] In one embodiment, an ensemble approach is used. First, a well-performing model is selected. Then, each model "votes" on one of the images (using augmentations or otherwise), and the voting strategy that yields the best results is selected. Exemplary voting strategies include maximum confidence, average, majority average, median, average confidence, median confidence, majority average confidence, weighted average, majority maximum confidence, etc. Once a voting strategy is selected, an evaluation method for the combination of augmentations must also be selected, which, as described above, describes how each of the rotations should be processed by the ensemble. In this embodiment, the final AI model 100 can thus be defined as a collection of trained AI models using deep learning and / or computer vision models, along with a mode that encodes a majority voting strategy that determines how the results of the individual AI models will be combined, and an evaluation mode that determines how the augmentations, if any, will be combined.

[0153] Model selection can be performed so that their results contrast with each other, i.e., so that their results are as independent as possible and the scores are well distributed. This selection procedure is performed by examining which images in the test set were correctly identified for each model. When comparing two models, if the sets of correctly identified images are very similar or the scores provided by each model are similar to each other for a given image, the models are not considered to be contrasting models. However, if there is little overlap between the two sets of correctly identified images or the scores provided for each image are significantly different from each other, the models are considered to be contrasting. This procedure effectively evaluates whether the distributions of embryo scores in the test set for two different models are similar. The contrast criterion drives model selection with diverse predicted outcome distributions for different input images or segmentations. This method ensures translatability by avoiding the selection of models that performed well only on a specific clinical dataset, thus preventing overfitting. In addition, model selection can also use a diversity criterion. The diversity criterion drives model selection to include different model hyperparameters and configurations. The reason is that in practice, similar model settings may yield similar prediction results and therefore may not be useful for the final ensemble model.

[0154] In one embodiment, this can be done by using a counting approach and specifying a threshold similarity, such as 50%, 75%, or 90% overlapping images in the two sets. In other embodiments, the scores in a set of images (e.g., the viable set) can be summed, the two sets (sums) compared, and ranked similarly if the sum of the two is below a threshold amount. Statistical comparisons can also be used, for example, by considering the number of images in the sets or otherwise comparing the distribution of images in each of the sets.

[0155] Another approach in AI and machine learning is known as "knowledge distillation" (abbreviated to "distillation") or "student-teacher" models, in which distributions of weight parameters obtained from one (or many) models (one or more teachers) are used to inform the weight updates of another model (the student) via the student model's loss function. We use the term "distillation" to describe the process of training a student model using one or more teacher models. The idea behind this procedure is to train the student model to mimic a set of one or more teacher models. The intuition behind this process is that the teacher models have a subtle but important relationship between their predicted output probabilities (soft labels) that is not present in the original predicted probabilities (hard labels) obtained directly from the model's results in the absence of distributions from one or more teacher models.

[0156] First, a set of one or more teacher models is trained on the dataset of interest. The teacher models may be of any neural network or model architecture, and may even be completely different architectures from each other or from the student models. They may share the exact same dataset or may have disjoint or overlapping subsets of the original dataset. Once the teacher models are trained, the student models are trained using a distillation loss function to mimic the output of the teacher models. The distillation process begins by first applying the teacher models to a dataset made available to both the teacher and student models, known as a "transfer dataset." The transfer dataset may be a holdout blind dataset drawn from the original dataset, or it may be the original dataset itself. Furthermore, the transfer dataset does not need to be fully labeled; i.e., some of the data is not associated with known outcomes. This removal of the label restriction allows the size of the dataset to be artificially increased. Next, the student models are applied to the transfer dataset. The output probabilities of the teacher model (soft labels) are compared to the output probabilities of the student model via a divergence measure function, such as KL-Divergence or a "relative entropy" function calculated from the distributions. A divergence measure is an accepted mathematical method for measuring the "distance" between two probability distributions. The divergence measure is then summed with a standard cross-entropy classification loss function, resulting in a loss function that effectively minimizes the classification loss, improving model performance and simultaneously improving the divergence of the student model from the teacher model. Typically, the soft-label matching loss (the divergence component of the new loss) and the hard-label classification loss (the original component of the loss) are weighted relative to each other (introducing an extra adjustable parameter to the training process) to control the contribution of each of the two terms in the new loss function.

[0157] A model can be defined by its network weights. In some embodiments, this may involve exporting or saving a checkpoint file or model file using an appropriate function in the machine learning code / API. The checkpoint file may be a file created by the machine learning code / library with a defined format that can be exported and then read back (reloaded) using standard functions (e.g., ModelCheckpoint() and load_weights()) provided as part of the machine learning code / API. The file format may be directly transmitted or copied (e.g., via ftp or a similar protocol), or may be serialized and transmitted using JSON, YAML, or a similar data transfer protocol. In some embodiments, additional model metadata, such as model accuracy, number of epochs, etc., which can further characterize the model or otherwise aid in building another model (e.g., a student model) on another node / server, may be exported / saved and transmitted along with the network weights.

[0158] Embodiments of the method may be used to create AI models for obtaining estimates of the presence of one or more aneuploidies in images of embryos. These may be implemented in a cloud-based computing system configured to computationally create aneuploidy screening artificial intelligence (AI) models. Once the models are created, they can be deployed in a cloud-based computing system configured to computationally generate estimates of the presence of one or more aneuploidies in images of embryos. In this system, the cloud-based computing system includes a previously created (trained) aneuploidy screening artificial intelligence (AI) model, and the computing system is configured to receive images provided to the aneuploidy screening artificial intelligence (AI) model from a user via a user interface of the computing system to obtain estimates of the presence of one or more aneuploidies in the images. A report regarding the presence of one or more aneuploidies in the images is provided to the user via the user interface. Similarly, a computing system may be located at a clinic or similar location where images are obtained and configured to generate estimates of the presence of one or more aneuploidies in images of embryos. In this embodiment, the computing system includes at least one processor and at least one memory, the at least one memory including instructions for configuring the processor to receive images captured during a predetermined time window after in vitro fertilization (IVF) and further upload, via a user interface, the images captured during the predetermined time window after IVF to a cloud-based artificial intelligence (AI) model configured to generate an estimate of the presence of one or more aneuploidies in the images of the embryo, the estimate of the presence of one or more aneuploidies in the images of the embryo being received via the user interface and displayed by the user interface.

[0159] result Results demonstrating the AI ​​model's ability to isolate morphological features corresponding to specific chromosomes or groups of chromosomes purely from phase-contrast microscopy images are presented below. This includes a series of example studies focusing on some of the most severe chromosomal defects (i.e., high risk of post-implantation adverse events) according to Table 1. In the first three cases, simplified examples are constructed to illustrate whether there are morphological features corresponding to specific chromosomal abnormalities. This is done by including only affected chromosomes and euploid viable embryos. These simplified examples provide evidence that it is feasible to create a holistic model based on combining separate models, each focusing on a different chromosomal defect / gene defect. Additional examples were also created using the chromosome groups containing the aneuploidies listed in Table 1.

[0160] The first study was conducted to evaluate whether the AI ​​model could detect the difference between euploid, viable embryos and embryos containing any abnormalities, including chromosome 21 (including mosaic embryos), which is associated with Down syndrome. Results of the trained model on a blind dataset of 214 images achieved an overall accuracy of 71.0%.

[0161] Prior to AI model training, a blind test set with representation of all chromosomes considered to be a serious health risk if involved in aneuploidy (from Table 1) and images of viable euploids was constrained to be used as a common test set for the trained model. The total number of images involved in the study is shown in Table 2:

[0162] [Table 2] The accuracy results on the test set are as follows: - Embryos with some abnormality on chromosome 21: 76.47% (52 / 68 correctly identified); and - Viable euploid embryos: 68.49% (100 / 146 correctly identified).

[0163] The resulting distributions are shown in Figures 6A and 6B for aneuploid and euploid viable embryos, respectively. Aneuploid distribution 600 shows a small set of chromosome 21 abnormal embryos 610 (bars filled with diagonal forward slashes on the left) that were incorrectly identified (overlooked) by the AI ​​model on the left, and a larger set of chromosome 21 abnormal embryos 620 (bars filled with diagonal backslashes on the right) that were correctly identified by the AI ​​model. Similarly, euploid distribution 630 shows a small set of chromosome 21 normal embryos 640 (bars filled with diagonal forward slashes on the left) that were incorrectly identified (overlooked) by the AI ​​model on the left, and a larger set of chromosome 21 normal embryos 650 (bars filled with diagonal backslashes on the right) that were correctly identified by the AI ​​model. In both plots, aneuploid and euploid embryo images are well separated with a clear clustering of their ploidy state outcome scores as provided by the AI ​​model.

[0164] In the same manner as in the study of chromosome 21, this methodology was repeated for chromosome 16, alterations of which have been associated with autism. The total number of images involved in this study is shown in Table 3.

[0165] [Table 3] The accuracy results on the test set are as follows: - Embryos with some abnormality on chromosome 16: 70.21% (33 / 47 correctly identified); and - Viable euploid embryos: 73.97% (108 / 146 correctly identified).

[0166] The resulting distributions are shown in Figures 7A and 7B for aneuploid and euploid viable embryos, respectively. Aneuploid distribution 700 shows a small set of chromosome 16 abnormal embryos 710 (bars filled with diagonal forward slashes on the left) that were incorrectly identified (overlooked) by the AI ​​model, and a larger set of chromosome 16 abnormal embryos 750 (bars filled with diagonal backslashes on the right) that were correctly identified by the AI ​​model. Similarly, euploid distribution 730 shows a small set of chromosome 16 normal embryos 740 (bars filled with diagonal forward slashes on the left) that were incorrectly identified (overlooked) by the AI ​​model, and a larger set of chromosome 16 normal embryos 750 (bars filled with diagonal backslashes on the right) that were correctly identified by the AI ​​model.

[0167] As a third case study, this methodology is repeated for chromosome 13, which is associated with Patau syndrome. The total number of images involved in this study is shown in Table 4.

[0168] [Table 4] The accuracy results are as follows: - Embryos with some abnormality on chromosome 13: 54.55% (24 / 44 correctly identified); and - Viable euploid embryos: 69.13% (103 / 149 correctly identified).

[0169] Although the accuracy for this particular chromosome is lower than for chromosomes 21 and 16, it is expected that different chromosomes will have different levels of confidence in identifying images corresponding to their specific associated aneuploidies for a given dataset size. That is, each genetic abnormality will exhibit different visible characteristics, and therefore, some abnormalities will be more easily detectable than others. However, as with most machine learning systems, increasing the size and diversity of the training dataset is expected to maximize the model's ability to detect the presence of specific chromosomal abnormalities. As a result, a combined approach that can assess multiple aneuploidies separately, all at once, can provide a useful overall picture of genetic abnormalities associated with embryos with different levels of confidence, depending on the rarity of the cases incorporated into the training.

[0170] As a fourth case study, this methodology was used in chromosome group analysis, where viable euploid embryos were included with a chromosome group of chromosome alterations considered "severe" (according to Table 1), including chromosomes 13, 14, 16, 18, 21, and 45,X. For the purposes of this example, both mosaicism and non-mosaicism are included, and all types of chromosome alterations are included. The total number of images involved in this study is shown in Table 5.

[0171] [Table 5] The accuracy results are as follows: - Embryos with any severe chromosomal abnormality: 54.95% (50 / 91 correctly identified); and - Viable euploid embryos: 64.41% (38 / 59 correctly identified).

[0172] The resulting distributions are shown in Figures 8A and 8B for aneuploid and euploid viable embryos, respectively. Aneuploid distribution 800 shows a small set of aneuploid / abnormal embryos 810 (bars filled with diagonal forward slashes on the left) that were incorrectly identified (overlooked) by the AI ​​model on the left as being in the chromosomal severe group, and a larger set of aneuploid / abnormal embryos 820 (bars filled with diagonal backslashes on the right) that were correctly identified by the AI ​​model as being in the chromosomal severe group. Similarly, euploid distribution 830 shows a small set of normal euploid embryos 840 (bars filled with diagonal forward slashes on the left) that were incorrectly identified (overlooked) by the AI ​​model on the left as being chromosomal severe, and a larger set of normal euploid embryos 850 (bars filled with diagonal backslashes on the right) that were correctly identified by the AI ​​model.

[0173] Although this accuracy for chromosome groups is lower than for individual chromosomes, it is expected that certain combinations of grouped chromosomes or morphological bases of similar severity can be identified corresponding to their specific associated aneuploidies for a given dataset size. That is, each genetic abnormality will exhibit different visible characteristics, and therefore, some abnormalities are expected to be more easily detectable than others. However, as with most machine learning systems, increasing the size and diversity of the training dataset is expected to maximize the model's ability to detect the presence of specific chromosomal abnormalities. As a result, a combined approach that can evaluate multiple aneuploidies separately and all at once can provide a useful overall picture of genetic abnormalities associated with embryos with different levels of confidence, depending on the rarity of the cases incorporated into the training.

[0174] These four studies demonstrate that AI / machine learning and computer vision techniques can independently identify morphological features associated with abnormalities in chromosomes 21, 16, and 13, and in combined chromosome groups.

[0175] Each AI model is able to detect, with some degree of reliability, morphological features associated with specific severe chromosomal abnormalities. Histograms of scores associated with ploidy state provided by the selected models show reasonable separation between euploid and aneuploid embryo images.

[0176] Morphological features associated with chromosomal abnormalities can potentially be subtle and complex, making it challenging to effectively discover these patterns by training on small datasets. Although this study shows a strong correlation between embryo morphology in images and chromosomal abnormalities, it is expected that greater accuracy will be achieved with a much larger and more diverse dataset for training AI models.

[0177] These studies demonstrate the feasibility of building a general aneuploidy assessment model based on combining separate models, each focusing on a different chromosomal abnormality. Such a more general aneuploidy assessment model could incorporate a broader variety of both severe and mild chromosomal abnormalities, as outlined in Table 1 or as determined according to clinical practice. That is, in contrast to conventional systems that typically simply aggregate all aneuploidies (and mosaicism) and give a presence / absence call, the present system improves performance by dividing the problem into independent chromosomal groups and training these models separately for each group before combining the individual models to enable the detection of a broad range of chromosomal abnormalities. By dividing the problem into smaller chromosomal groups and then training multiple different models, each trained in a different way or with different configurations or architectures (e.g., hierarchical, binary, multi-class, multi-group), a variety of models are created, each effectively solving a different optimization problem and thus yielding different results for the input image. This diversity then allows the optimal model to be selected. In addition, this approach is designed to identify mosaicism not currently detectable by invasive screening methods. During IVF cycles, embryos are a precious and limited resource. Current success rates (in terms of viable pregnancies) are low, and the financial and emotional costs of additional cycles are high. Therefore, providing improved, non-invasive aneuploidy assessment tools based on defining chromosomal groups, such as based on the severity of adverse events, would provide more nuanced benefits for clinicians and patients. This would allow for more informed decisions to be made, particularly in the difficult situation where all available embryos (for the current cycle) exhibit aneuploidy or mosaicism, and therefore allow clinicians and patients to balance potential risks and make more informed selection decisions about which embryos to implant.

[0178] Several embodiments are discussed, including hierarchical and binary models, as well as single-group or multi-group models. In particular, hierarchical models can be used to train AI models by providing quality labels to embryo images. In this embodiment, a hierarchical sequence of layer models can be created, with a separate layer model for each chromosome group. At each layer, images are separated based on quality, and the best-quality images are used to train the models at that layer. That is, at each layer, the training set is divided into the best-quality images and other images. The models at that layer are trained on the best-quality images, and the other images are passed to the next layer, and the process is repeated (thus separating the remaining images into the next-best-quality images and other images). The models in a hierarchical model can be all binary models, all multi-class models, or a combination of both binary and multi-class models across multiple layers. In addition, this hierarchical training method can also be used to train multi-group models. The rationale behind the hierarchical model approach is that embryo images deemed high quality are likely to have the highest quality morphological features in images with minimal abnormalities (i.e., "best embryo-like") and therefore the greatest morphological disparity / difference compared to embryo images containing chromosomal defects (i.e., "appear to have poor or abnormal features"). This therefore enables AI algorithms to better detect and predict morphological features between these two (extreme) classifications of images. This process can be repeated multiple times with different numbers of layers / quality labels to create a set of hierarchical models. Many independent hierarchical models are created for each chromosome group, and from this set of hierarchical models, the best hierarchical model can be selected. This may be based on a quality metric, or ensemble or distillation techniques may be used.

[0179] In some embodiments, a set of binary models may be created for each chromosome group, or for one or more multigroup models that classify all of the chromosome groups (or at least a large number of chromosome groups). Similar to multigroup models, including hierarchical multigroup models, many different sets of binary models and many multiclass models may be created. These provide further diversity in AI models. Once a set of candidate models is created, they can be used to create a final AI model to identify each of the chromosome groups in an image. This can be further refined or created using ensemble, distillation, or other similar methods to train a final single model based on multiple models. Once a final model is selected, it can then be deployed to classify new images during IVF and thus aid in the selection of embryo(s) for implantation, for example, by identifying and eliminating embryos with high risk or by identifying embryos with the lowest risk of aneuploidy.

[0180] Therefore, the methodology developed in combination with research on chromosomal abnormalities can be used as a prescreening tool to characterize embryo images prior to preimplantation genetic diagnosis (PGD), or to provide a series of high-level genetic analyses to complement clinics that cannot use readily available PGD technology. For example, if the images suggest a high probability / confidence of the presence of harmful chromosomal abnormalities, the embryos can be discarded or subjected to invasive (and high-risk) PGD techniques so that only embryos considered to be low-risk are implanted.

[0181] Those skilled in the art will appreciate that information and signals may be represented using any of a variety of technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.

[0182] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software or instructions, middleware, platforms, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.

[0183] The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two, including in a cloud-based system. For a hardware implementation, processing may be implemented within one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, or other electronic units designed to perform the functions described herein, or a combination thereof. Various middleware and computing platforms may also be used.

[0184] In some embodiments, the processor module includes one or more central processing units (CPUs), graphics processing units (GPUs), and / or tensor processing units (TPUs) configured to perform some of the method steps. Similarly, a computing device may include one or more CPUs, GPUs, and / or TPUs. The CPU may include an input / output interface, an arithmetic logic unit (ALU), and a control unit and program counter element that communicate with input and output devices via the input / output interface. The input / output interface may include a network interface and / or a communication module for communicating with a comparable communication module in another device using a predetermined communication protocol (e.g., IEEE 802.11, IEEE 802.15, TCP / IP, UDP, etc.). The computing device may include a single CPU (core) or multiple CPUs (multi-core), or multiple processors. The computing device is typically a cloud-based computing device using a GPU or TPU cluster, but may also be a parallel processor, vector processor, or distributed computing device. Memory is operatively coupled to the one or more processors and may include RAM and ROM components and may be provided internal or external to the device or processor module. The memory can be used to store an operating system and additional software modules or instructions. The one or more processors may be configured to load and execute the software modules or instructions stored in the memory.

[0185] A software module, also known as a computer program, computer code, or instructions, may comprise numerous source code or object code segments or instructions and may reside in any computer-readable medium, such as RAM memory, flash memory, ROM memory, EPROM memory, registers, a hard disk, a removable disk, a CD-ROM, a DVD-ROM, a Blu-ray disc, or any other form of computer-readable medium. In some aspects, a computer-readable medium may include a non-transitory computer-readable medium (e.g., a tangible medium, etc.). Additionally, in other aspects, a computer-readable medium may include a transitory computer-readable medium (e.g., a signal, etc.). Combinations of the above should also be included within the scope of computer-readable media. In another aspect, a computer-readable medium may be integral to a processor. The processor and the computer-readable medium may reside in an ASIC or related device. The software codes may be stored in a memory unit, and the processor may be configured to execute them. The memory unit may be implemented within the processor or external to the processor, in which case it may be communicatively coupled to the processor via various means known in the art.

[0186] Furthermore, it should be appreciated that modules and / or other suitable means for performing the methods and techniques described herein can be downloaded and / or otherwise obtained by a computing device. For example, such a device can be coupled to a server to facilitate the transfer of means for performing the methods described herein. Alternatively, the various methods described herein can be provided via storage means (e.g., RAM, ROM, physical storage media such as a compact disk (CD) or floppy disk, etc.), such that a computing device can obtain the various methods when the storage means is coupled to or provided to the device. Furthermore, any other suitable technique for providing the methods and techniques described herein to a device can be utilized.

[0187] The methods disclosed herein comprise one or more steps or actions for achieving the described method. The steps and / or actions of the methods may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions can be modified without departing from the scope of the claims.

[0188] Throughout this specification and the appended claims, unless the context requires otherwise, the word "comprises" and variations such as "comprises" will be understood to mean the inclusion of a stated integer or group of integers, but not the exclusion of any other integer or group of integers. The reference to any prior art in this specification is not, and should not be construed as, any form of acknowledgment that such prior art forms part of the common general knowledge.

[0189] It will be appreciated by those skilled in the art that the present invention is not limited in its use to the particular applications described. Nor is the present invention limited in its preferred embodiments with respect to the specific elements and / or features described or depicted herein. It will be appreciated that the present invention is not limited to the disclosed embodiment or embodiments, but is capable of many rearrangements, modifications, and substitutions without departing from the scope of the present invention as set forth and defined by the appended claims.

Claims

1. 1. A method for computationally generating an aneuploidy screening artificial intelligence (AI) model for screening embryo images for the presence of aneuploidy, comprising: determining a plurality of chromosome group labels, each group including one or more different aneuploidies including different genetic mutations or chromosomal abnormalities; creating a training dataset from a first set of images, each image comprising an image of an embryo captured after in vitro fertilization and labeled with one or more chromosome group labels, each label indicating whether at least one aneuploidy associated with a respective chromosome group is present in at least one cell of the embryo, the training dataset comprising images labeled with each of the chromosome groups; creating a test dataset from a second set of images, each image comprising an image of an embryo obtained after in vitro fertilization and labeled with one or more chromosome group labels, each label indicating whether at least one aneuploidy associated with a respective chromosome group is present, the test dataset comprising images labeled with each of the chromosome groups; training at least one chromosome group AI model separately for each chromosome group using the training dataset for training all models, each chromosome group AI model being trained to identify morphological features in images labeled with an associated chromosome group label; and / or training at least one multi-group AI model on the training dataset, each multi-group AI model being trained to independently identify morphological features in images labeled with each of its associated chromosome group labels and to generate a multi-group output for an input image indicating the presence or absence of at least one aneuploidy associated with each of the chromosome groups in the image; using the test dataset to select a best chromosome group AI model or a best multi-group AI model for each of the chromosome groups; deploying the selected AI model to screen embryo images for the presence of one or more aneuploidies; A method comprising:

2. The step of separately training at least one chromosome group AI model for each chromosome group and / or the step of training the at least one multi-group AI model comprises training a hierarchical model, and training the hierarchical model comprises:

2. The method of claim 1, comprising training hierarchical sequence layer models, wherein in each layer, images associated with one chromosome group are given a first label and trained on a second set of images, the second set of images being grouped based on a maximum level of quality, and wherein in each sequence layer, the second set of images is a subset of images from the second set in a previous layer and has a quality lower than the maximum quality of the second set in the previous layer.

3. Training the hierarchical model comprises: assigning a quality label to each image in the training dataset, the set of quality labels comprising a hierarchical set of quality labels including at least "viable euploid embryo", "euploid nonviable embryo", "non-severe aneuploid embryo", and "severe aneuploid embryo"; training a top-layer model by splitting the training dataset into a first quality dataset labeled "viable euploid embryos" and another dataset containing all other images, and training a model on the images labeled with the chromosome group and on images in the first quality dataset; training one or more intermediate layer models in sequence, where in each intermediate layer, a next quality level dataset is created from selecting images with the highest quality labels in the other datasets, and models are trained on images labeled with the chromosome groups and on images in the next quality dataset; and training a base layer model on the chromosome-labeled images and images in other datasets from previous layers; The method of claim 2 , comprising:

4. 4. The method of claim 3, wherein after training a first base-level model for a first chromosome group, training a hierarchical model for each other chromosome group comprises training the other chromosome groups on other datasets used to train the first base-level model.

5. wherein the step of separately training at least one chromosome group AI model for each chromosome group further comprises training one or more binary models for each chromosome group; labeling images in the training dataset with a label that matches the chromosome group with a present label and labeling all other images in the training dataset with an absent label, and training a binary model using the present label and the absent label to produce a binary output for an input image indicating whether a chromosomal abnormality associated with the chromosome group is present in the image; The method of claim 2 , comprising:

6. The method of claim 2 , wherein the hierarchical models are each binary models.

7. 2. The method of claim 1, wherein each chromosome group further comprises multiple mutually exclusive aneuploidy classes, the probabilities of the aneuploidy classes within a chromosome group sum to 1, and one or more of the AI ​​models is a multi-class AI model trained to estimate the probability of each aneuploidy class within a chromosome group.

8. The method of claim 7, wherein the aneuploidy classes include "deletion," "gain," "duplication," "deletion," and "normal."

9. The method further includes a step of creating an ensemble model for each chromosome group, wherein the step of creating an ensemble model for each chromosome group includes: training a plurality of final models, each of the plurality of final models based on the best chromosome group AI model for a respective group, each of the plurality of final models being trained on the training dataset having a different set of initial conditions and image orderings; and combining a plurality of the trained final models according to an ensemble voting strategy; The method of claim 1 , comprising:

10. The method further includes a step of creating a distillation model for each chromosome group, wherein the step of creating a distillation model for each chromosome group includes: training a plurality of teacher models, each of the plurality of teacher models being based on the best chromosome group AI model for a respective group, each of the plurality of teacher models being trained on at least a portion of the training dataset having a different set of initial conditions and image orderings; and training a student model using a plurality of trained teacher models on the training dataset using a distilled loss function; The method of claim 1 , comprising:

11. receiving a plurality of images, each image comprising an image of an embryo obtained after in vitro fertilization and one or more aneuploidy results; dividing the plurality of images into the first set of images and the second set of images and assigning one or more chromosome group labels to each image based on the one or more associated aneuploidy results, wherein the first set of images and the second set of images have similar proportions of each of the chromosome group labels; The method of claim 1 further comprising:

12. 10. The method of claim 1, wherein each group comprises multiple different aneuploidies with similar risks of adverse events.

13. The method of claim 12 , wherein the plurality of chromosome group labels comprises at least a low risk group and a high risk group.

14. 14. The method of claim 13, wherein the low-risk group includes at least chromosomes 1, 3, 4, 5, 17, 19, 20, and 47,XYY, and the high-risk group includes at least chromosomes 13, 16, 21, and 45,X, 47,XXY, and 47,XXX.

15. The method of claim 1 , wherein the image is captured within 3 to 5 days after fertilization.

16. 2. The method of claim 1, wherein the relative proportions of each of the chromosome groups in the test data set are similar to the relative proportions of each of the chromosome groups in the training data set.

17. 1. A method for computationally generating an estimate of the presence of one or more aneuploidies in an image of an embryo, comprising:

17. A method for generating an aneuploidy screening AI model in a computing system according to the method of any one of claims 1 to 16, receiving an image containing an embryo captured after in vitro fertilization from a user via a user interface of the computing system; providing the image to the aneuploidy screening AI model to obtain an estimate of the presence of one or more aneuploidies in the image; transmitting a report to the user via the user interface regarding the presence of one or more aneuploidies in the image; A method comprising:

18. 17. A cloud-based computing system comprising one or more computing devices including one or more processors and one or more memories, the cloud-based computing system configured to computationally create an aneuploidy screening artificial intelligence (AI) model configured according to the method of any one of claims 1 to 16.

Citation Information

Patent Citations

  • Chromosome abnormality detection model, chromosome abnormality detection system, and chromosome abnormality detection method

    CN110265087A

  • Artificial Intelligence and Devices for Diagnosis, Screening, Prevention, and Treatment of Maternal-Fetal Conditions

    JP2007521854A

  • Information processor, information processing method, and program

    JP2012043156A

  • Imaging and evaluation of embryos, oocytes, and stem cells

    JP2013502233A

  • Image diagnosis system, image diagnosis program and image diagnosis method for fertilized egg

    JP2019097425A