Methods and systems for non-invasive genetic testing using artificial intelligence (AI) models
The AI model for aneuploidy screening, generated through computation, solves the invasiveness and reliability issues of existing embryo gene screening methods, enabling rapid and accurate aneuploidy screening of embryos, reducing invasive procedures on embryos and prolonging pregnancy time.
Patent Information
- Application Number
- CN202080081475.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-25
- Filing Date
- 2020-09-25
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2040-09-25
AI Technical Summary
Existing embryo genetic screening methods such as PGT-A are highly invasive, have uncertain long-term effects, low reliability of results, and prolong pregnancy time, and the chromosome test results of mosaic embryos are inconsistent.
An artificial intelligence (AI) model for aneuploidy screening, generated through computation, identifies aneuploids in embryo images using training and testing datasets. It provides a non-invasive method for embryo gene screening by using hierarchical or multi-group AI models.
It enables rapid and accurate screening for embryonic aneuploidy, reduces invasive procedures on embryos, improves the reliability and efficiency of screening, and reduces the risk of prolonged pregnancy.
Smart Images

Figure CN114846507B_ABST
Abstract
Description
[0001] Priority document
[0002] This application claims priority to Australian Provisional Patent Application No. 2019903584, filed on September 25, 2019, entitled “Method and System for Performing Non-invasive Genetic Testing Using an Artificial Intelligence (AI) Model,” the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to artificial intelligence (AI), including image classification based on computer vision and deep learning. In a particular form, the invention relates to a computational AI method for non-invasively identifying aneuploidy in embryos used for in vitro fertilization (IVF). Background Technology
[0004] Human cells contain 23 pairs of chromosomes (46 in total), unless adversely affected by factors such as radiation or cell damage caused by genetic conditions / congenital diseases. In these cases, one or more chromosomes may be completely or partially altered. This can have extensive and long-term (lasting into adulthood) health effects on embryonic development, making it crucial to understand whether a patient exhibits this chromosomal abnormality or is a carrier of the chromosomal variant (making their child susceptible to the disease) so that they can receive adequate treatment. While prospective parents may have one or more genetic predispositions, it is impossible to predict in advance whether their offspring will actually exhibit one or more genetic abnormalities.
[0005] A common assisted reproductive technology (ART) involves testing the embryo, after fertilization, and performing gene sequencing to assess the genetic health of the embryo and classify it as "euploid" (genetically normal) or "aneuploid" (exhibiting gene alterations).
[0006] This screening technique is particularly important in the IVF process (where embryos are fertilized in vitro and then re-implanted into the expectant mother approximately 3 to 5 days after fertilization). This is typically a decision made by the patient in consultation with their IVF doctor as part of a process that assists in diagnosing potential fertility complications the couple may experience or in early detection of any disease risks and selection to prevent them.
[0007] This screening process, known as preimplantation genetic screening (PGS) or preimplantation aneuploidy gene testing (PGT-A), has many features that make it less than ideal; however, it remains the most viable option for obtaining genetic information about embryos in the fertility industry.
[0008] The biggest risk factor for PGT-A is that the test is highly invasive, as it typically requires the removal of a small number of cells from the developing embryo (using one of a range of biopsy techniques) before testing can be performed. The long-term effects of this technique on embryonic development are uncertain and not fully characterized. Furthermore, all embryos undergoing PGT-A require round trips to the laboratory where the biopsy is performed, and results are delayed at the clinic by days or weeks. This means that the “time to pregnancy” (a key measure of IVF success) is prolonged, and all such embryos must be frozen. Because modern freezing techniques such as vitrification have significantly improved embryo viability compared to “slow freezing” in recent years, many IVF clinics now use this technique, even when performing PGT-A. The logic behind this is to allow the prospective mother's hormone levels to rebalance after excessive ovulation stimulation, increasing the likelihood of embryo implantation.
[0009] It remains unclear whether modern vitrification techniques are harmful to embryos. Due to the popularity and widespread acceptance of vitrification and PGT-A, especially in the United States where PGT-A is performed routinely, most embryos undergo this process to obtain genetic data for clinics and patients.
[0010] Another issue with PGT-A performance is due to embryonic "mosaicism." This term means that the chromosome profile of a single cell collected in a biopsy may not be representative of the entire embryo, even during early cell division stages of embryonic development. In other words, a mosaic embryo is a mixture of euploid (chromosomally normal) and aneuploid (chromosomal excess / deletion / modification) cells, and multiple distinct aneuploidies may exist in different cells (including cases where all cells are aneuploid and no euploid cells are present in the embryo). Therefore, PGT-A results from different cells from the same embryo can be contradictory. Because the representativeness of the biopsy cannot be assessed, the overall accuracy / reliability of such PGT-A tests is reduced.
[0011] Therefore, there is a need to provide improved methods for embryo genetic screening, or at least useful alternatives to existing methods. Summary of the Invention
[0012] According to a first aspect of the present invention, a method is provided for computationally generating an artificial intelligence (AI) model for screening aneuploidy, the AI model being used to screen for the presence of aneuploidy in embryo images, the method comprising:
[0013] Define multiple chromosome set tags, where each set includes one or more different aneuploids, which contain different gene alterations or chromosomal abnormalities;
[0014] A training dataset is generated from a first set of images, wherein each image includes an image of an embryo captured after in vitro fertilization and is labeled with one or more chromosome set tags, each tag indicating the presence of at least one aneuploidy associated with the corresponding chromosome set in at least one cell of the embryo, and the training dataset includes images labeled with each chromosome set;
[0015] A test dataset is generated from a second set of images, wherein each image includes an image of an embryo captured after in vitro fertilization and is labeled with one or more chromosome set tags, each tag indicating the presence of at least one aneuploidy associated with the corresponding chromosome set, and the test dataset includes images labeled with each chromosome set;
[0016] At least one chromosome group AI model is trained for each chromosome group using the training dataset used to train all models, wherein each chromosome group AI model is trained to identify morphological features in images labeled with the relevant chromosome group, and / or at least one multi-group AI model is trained on the training data, wherein each multi-group AI model is trained to independently identify morphological features in images labeled with each relevant chromosome group, to generate a multi-group output for the input image to indicate whether at least one aneuploidy associated with each chromosome group exists in the image;
[0017] Use the test dataset to select the best chromosome set AI model for each chromosome set, or a best multi-set AI model; and
[0018] The selected AI model is deployed to screen for the presence of one or more aneuploidies in embryo images.
[0019] In one form, the step of training at least one chromosome set AI model for each chromosome set includes: training a hierarchical model and / or training at least one multi-set AI model, wherein training the hierarchical model includes:
[0020] The hierarchical sequence of training hierarchical models is provided, wherein in each layer, images associated with chromosome sets are assigned a first label and trained on a second set of images, wherein the second set of images is grouped based on the highest quality level, and in each sequential layer the second set of images is a subset of the second set of images from the previous layer whose quality is lower than the highest quality of the second set of images in the previous layer.
[0021] In another form, training a hierarchical model includes:
[0022] A quality label is assigned to each of the plurality of images, wherein the quality label set includes a graded quality label set, which includes at least “viable euploid embryos”, “non-viable euploid embryos”, “mild aneuploid embryos” and “severe aneuploid embryos”.
[0023] The top-level model was trained by dividing the training set into a first quality dataset labeled “viable euploid embryos” and another dataset containing all other images, and the model was trained on images labeled with chromosome sets and on images within the first quality dataset.
[0024] One or more intermediate layer models are trained sequentially, wherein in each intermediate layer, images with the highest quality labels are selected from another dataset to generate the next quality-level dataset, and the model is trained on images labeled with the chromosome set and images within the next quality-level dataset; and
[0025] The base layer model is trained on images labeled with the chromosome set and images from the other datasets from the previous layer.
[0026] In another form, after training a first base-level model for a first chromosome set, training a hierarchical model for each other chromosome set includes training the other chromosome sets for the other datasets used to train the first base-level model.
[0027] In another form, the step of training at least one chromosome set AI model for each chromosome set may further include:
[0028] Train one or more binary models for each chromosome set, including:
[0029] Images in the training dataset that have labels matching the chromosome set are labeled with "existence" and all other images in the training set are labeled with "non-existence". A binary model is trained using the "existence" and "non-existence" labels to generate a binary output about the input image to indicate whether there is a chromosomal abnormality in the image that is associated with the chromosome set.
[0030] In another form, both hierarchical and layered models are binary models.
[0031] In one form, each chromosome set label also includes multiple mutually exclusive aneuploid categories, wherein the sum of the probabilities of the aneuploid categories within the chromosome set is 1, and the AI model is a multi-class AI model trained to estimate the probability of each aneuploid category within the chromosome set. In another form, the aneuploid categories may include: “missing,” “insertion,” “duplication,” “deleted,” and “normal.”
[0032] In one form, the method may also include:
[0033] Generate an ensemble model for each chromosome set, including:
[0034] Train multiple final models, each of which is based on the best chromosome set AI model for its corresponding group, and each of the multiple final models is trained on a training dataset with different initial condition sets and image rankings; and
[0035] Multiple trained final models are combined according to the ensemble voting strategy.
[0036] In one form, the method may also include:
[0037] Generate distillation models for each chromosome set, including:
[0038] Train multiple teacher models, each of which is based on the best chromosome set AI model for its corresponding group, and each of the multiple teacher models is trained on at least a portion of a training dataset with different initial condition sets and image rankings; and
[0039] The student model is trained on the training dataset using multiple trained teacher models with the distillation loss function.
[0040] In one form, the method may also include:
[0041] Receive multiple images, each image including an embryo image taken after in vitro fertilization and one or more aneuploid results;
[0042] The plurality of images are divided into a first group of images and a second group of images, and one or more chromosome group labels are assigned to each image based on one or more associated aneuploid results, wherein the first group of images and the second group of images have each chromosome group label in a similar proportion.
[0043] In one form, each group comprises multiple distinct aneuploids with similar risks of adverse outcomes. In another form, the multiple chromosome set tags include at least a low-risk group and a high-risk group. In yet another form, the low-risk group includes at least chromosomes 1, 3, 4, 5, 17, 19, 20, and “47,XYY”, and the high-risk group includes at least chromosomes 13, 16, 21, “45,X”, “47,XXY”, and “47,XXX”.
[0044] In one form, images can be captured 3 to 5 days after fertilization.
[0045] In one form, the relative proportion of each chromosome group in the test dataset is similar to the relative proportion of each chromosome group in the training dataset.
[0046] According to a second aspect of the present invention, a method is provided for computationally generating an estimate of the presence or absence of one or more aneuploidies in an embryo image, the method comprising:
[0047] In the computing system, an aneuploid screening AI model is generated according to the method in the first aspect;
[0048] Images of embryos captured after in vitro fertilization are received from the user via the user interface of the computing system;
[0049] The image is provided to the aneuploid screening AI model to obtain an estimate of whether one or more aneuploids exist in the image; and
[0050] The system sends a report to the user via the user interface regarding the presence of one or more aneuploids in the image.
[0051] According to a third aspect of the present invention, a method is provided for obtaining an estimate of the presence of one or more aneuploids in an embryo image, the method comprising:
[0052] Images captured during a predetermined time window following in vitro fertilization (IVF) are uploaded to a cloud-based artificial intelligence (AI) model via a user interface. The AI model is used to generate estimates about the presence of one or more aneuploids in the images, wherein the AI model is generated according to the method of the first aspect.
[0053] The user interface receives an estimate of the presence of one or more aneuploids in the embryo image.
[0054] According to a fourth aspect of the invention, a cloud-based computing system is provided for computationally generating an artificial intelligence (AI) model for aneuploid screening, the model being configured according to the method of the first aspect.
[0055] According to a fifth aspect of the invention, a cloud-based computing system is provided for computationally generating estimates regarding the presence of one or more aneuploidies in an embryo image, wherein the computing system comprises:
[0056] One or more computing servers, including one or more processors and one or more memories, the memories being used to store an aneuploidy screening artificial intelligence (AI) model, the aneuploidy screening AI model being used to generate estimates about the presence of one or more aneuploids in an embryo image, wherein the aneuploidy screening AI model is generated according to the method of the first aspect, and the one or more computing servers are used for:
[0057] Images are received from the user via the user interface of the computing system;
[0058] The image is provided to the aneuploid screening artificial intelligence (AI) model to obtain an estimate of whether one or more aneuploids exist in the image; and
[0059] The system sends a report to the user via the user interface regarding the presence of one or more aneuploids in the image.
[0060] According to a sixth aspect of the invention, a computing system is provided for generating an estimate of the presence of one or more aneuploids in an embryo image, wherein the computing system includes at least one processor and at least one memory, the memory including instructions for causing the at least one processor to perform the following operations:
[0061] Receive images captured within a predetermined time window after in vitro fertilization (IVF);
[0062] The images captured within a predetermined time window after in vitro fertilization (IVF) are uploaded to a cloud-based artificial intelligence (AI) model via a user interface. The AI model is used to generate an estimate of the presence of one or more aneuploids in the embryo images, wherein the AI model is generated according to the method of the first aspect.
[0063] Receive, via the user interface, an estimate of the presence of one or more aneuploids in the embryo image; and
[0064] The estimation results regarding the presence of one or more aneuploids in the embryo image are displayed via the user interface. Attached Figure Description
[0065] Embodiments of the invention will be discussed with reference to the accompanying drawings, in which:
[0066] Figure 1A This is a flowchart of a method for computationally generating an artificial intelligence (AI) model according to one embodiment, the AI model being used to screen for aneuploidy in embryo images;
[0067] Figure 1B This is a flowchart of a method, according to one embodiment, for using a trained aneuploid screening AI model to computationally generate an estimate of the presence of one or more aneuploids in an image of an embryo;
[0068] Figure 2A This is a flowchart of the training steps of a binary model according to one embodiment;
[0069] Figure 2B This is a flowchart of the training steps of a hierarchical and layered model according to one embodiment;
[0070] Figure 2C This is a flowchart of the training steps for a multi-class model according to one embodiment;
[0071] Figure 2D This is a flowchart of the steps for selecting the optimal chromosome set AI model according to one embodiment;
[0072] Figure 3 This is a schematic architecture of a cloud-based computing system according to one embodiment for computationally generating and using AI models for screening aneuploids;
[0073] Figure 4 This is a schematic diagram of an IVF procedure that uses an aneuploid screening AI model to help select embryos for implantation, according to one embodiment.
[0074] Figure 5A This is a schematic flowchart illustrating the generation of an aneuploid screening model using a cloud-based computing system according to one embodiment.
[0075] Figure 5B This is a schematic flowchart of a model training process on a training server according to one embodiment;
[0076] Figure 5C This is a schematic architecture diagram of a deep learning method including convolutional layers according to one embodiment, whereby the convolutional layers convert the input image into a prediction after training;
[0077] Figure 6A This is a confidence graph of the chromosome 21 AI model for detecting aneuploid chromosome 21 embryos in a blind test set according to one embodiment, wherein low confidence estimates are shown in the diagonally filled bar on the left and high confidence estimates are shown in the diagonally filled bar on the right.
[0078] Figure 6B This is a confidence graph of an AI model for detecting chromosome 21 of euploid viable embryos in a blind test set according to one embodiment, wherein low confidence estimates are shown in the diagonally filled bars on the left and high confidence estimates are shown in the diagonally filled bars on the right.
[0079] Figure 7A This is a confidence graph of the AI model for detecting aneuploid chromosome 16 embryos in a blind test set according to one embodiment, wherein low confidence estimates are shown in the diagonally filled bars on the left and high confidence estimates are shown in the diagonally filled bars on the right.
[0080] Figure 7B This is a confidence graph of an AI model for detecting chromosome 16 of euploid viable embryos in a blind test set according to one embodiment, wherein low confidence estimates are shown in the diagonally filled bars on the left and high confidence estimates are shown in the diagonally filled bars on the right.
[0081] Figure 8A This is a confidence plot of an AI model for detecting aneuploidy in chromosomes 14, 16, 18, 21, and 45,X in a blind test set according to one embodiment, wherein low confidence estimates are shown in the diagonally filled bar on the left and high confidence estimates are shown in the diagonally filled bar on the right.
[0082] Figure 8B This is a confidence plot of an AI model for detecting severe chromosome groups (14, 16, 18, 21, and 45,X) in viable euploid embryos in a blind test set according to one embodiment, wherein low confidence estimates are shown in the diagonally filled bars on the left and high confidence estimates are shown in the diagonally filled bars on the right.
[0083] In the following description, the same reference numerals denote the same or corresponding parts throughout the drawings. Detailed Implementation
[0084] Examples of a non-invasive method for screening for the presence or likelihood of aneuploidy (genetic alterations) in embryos are described. These aneuploidies, or gene alterations, result in modifications, deletions, or extra copies of parts or even entire chromosomes. In many cases, these chromosomal abnormalities cause subtle (sometimes pronounced) changes in the appearance of chromosomes in embryonic images. Examples of this method utilize a computer vision-based artificial intelligence (AI) / machine learning model, relying entirely on morphological data extracted from phase-contrast microscopy images (or similar images) of the embryo, to detect the presence of aneuploidy (i.e., chromosomal abnormalities). The AI model uses computer vision techniques to detect typically subtle morphological features in embryonic images to estimate the probability or likelihood of a range of aneuploidies being present (or absent). This estimate / information can then be used to aid in making implantation decisions or determining which embryos to select for invasive PGT-A testing.
[0085] The system's advantage is that it is non-invasive (i.e., works purely on microscope images) and can perform analysis within seconds of image collection by uploading images through a cloud-based user interface. This user interface analyzes the images on a cloud-based server using a previously trained AI model to quickly return estimates of the probability of aneuploidy (or specific aneuploidy) to clinicians.
[0086] Figure 1A This is a flowchart of a method 100 for computationally generating an artificial intelligence (AI) model for screening aneuploidy in embryo images. Figure 1B It is used to screen for aneuploidy using a trained AI model (i.e., based on...). Figure 1A A flowchart of a method 110 for generating an estimate of the presence of one or more aneuploids in an embryo image using a computational AI model.
[0087] For the 22 non-sex chromosomes, the types of chromosomal abnormalities considered include: complete insertion, complete loss, deletion (partial, within the chromosome), and duplication (partial, within the chromosome) compared to the normal chromosome structure. For sex chromosomes, the types of abnormalities considered include: deletion (partial, within the chromosome), duplication (partial, within the chromosome), and complete loss: “45,X”. Complete insertions, compared to normal XX or normal XY chromosomes, include three types: “47,XXX”, “47,XXY”, and “47,XYY”.
[0088] Embryos can also exhibit mosaicism, where different cells within the embryo have different sets of chromosomes. That is, an embryo can contain one or more euploid cells and one or more aneuploid cells (i.e., cells with one or more chromosomal abnormalities). Furthermore, multiple aneuploidies may exist, with different cells exhibiting different aneuploidy patterns (e.g., one cell may have a deletion on chromosome 1, while another cell may have an X insertion, such as 47,XXX). In some extreme cases, every cell in a mosaic embryo may exhibit aneuploidy (i.e., there are no euploid cells). Therefore, AI models can be trained to detect aneuploidy in one or more cells of an embryo, thereby detecting the presence of mosaicism.
[0089] The output of an AI model can be represented as the probability of an outcome, such as an aneuploidy risk score or an embryo viability score. It should be understood that embryo viability and aneuploidy risk are complementary terms. For example, if they were probabilities, the sum of embryo viability risk and aneuploidy risk might be 1. That is, both can measure the likelihood of an adverse outcome, such as the risk of miscarriage or a serious genetic disease. Therefore, we refer to the outcome as an aneuploidy risk / embryo viability score. An outcome can also be the probability of being in a specific risk category of an unfavorable outcome, such as very low risk, low risk, intermediate risk, high risk, and very high risk. Each risk category can include a group (at least one, and often more) of specific chromosomal abnormalities with similar probabilities of adverse outcomes. For example, very low risk might be no aneuploidy detected, the low risk group might be aneuploidy / mosaic in chromosomes 1, 3, 10, 12, and 19, the intermediate risk group might be aneuploidy / mosaic in chromosomes 4, 5, and 47,XYY, etc. Probability can be represented as a score on a predefined scale, a probability from 0 to 1.0, or a hard classification, such as a hard binary classification (presence / absence of aneuploids) or a hard classification into one of several groups (low risk, medium risk, high risk, very high risk).
[0090] In step 101, we defined multiple chromosomal set tags. Each set contains one or more different aneuploidies, including different genetic alterations or chromosomal abnormalities. Different aneuploidies / genetic alterations have different effects on the embryo, leading to different chromosomal abnormalities. Within the chromosomal set tags, individual mosaic categories can be defined, which, if present, can be low, intermediate, or high, indicating that the same embryo can exhibit different types of chromosomal abnormalities. The severity (risk) level of the mosaic can also be considered in terms of the number of cells involved in the mosaic and / or the type of aneuploidy present. Therefore, chromosomal set tags can include not only the number of chromosomes affected but also the presence (to some extent) of mosaicism. This allows for a more refined description of the progressive level of aneuploidy or genetic health in the embryo. The severity level of the present mosaicism depends on the severity of the chromosomes involved in the mosaicism, as described in Table 1 below. Moreover, the severity level may be related to the number of cells exhibiting mosaicism. Based on clinical evidence, such as based on PGT-A testing and pregnancy outcomes, different aneuploidies / chromosomal abnormalities can be grouped according to the risk and severity of adverse outcomes, thereby assigning priority for implantation (aneuploidy risk / embryo viability score). Table 1 lists the number and types of chromosomal abnormalities in 100,000 cases of spontaneous abortion and live birth, from Introduction to Genetic Analysis, 7th edition, Griffiths AJF, Miller JH, Suzuki DT, et al., New York: WH Freeman; 2000.
[0091] Table 1
[0092] The number and types of chromosomal abnormalities in spontaneous abortions and live births in 100,000 pregnancies, from *Introduction to Genetic Analysis*, 7th ed., Griffiths AJF, Miller JH, Suzuki DT, et al., New York: WH Freeman; 2000. In cases defined by sex chromosome redundancy or deletion, the format is “47” (or “45”) followed by the sex chromosome.
[0093]
[0094]
[0095] Table 1 or similar data from other clinical studies can be used to group aneuploidy based on risk levels. Those with the highest risk are considered the lowest priority for metastasis and the highest priority for avoiding adverse post-implantation outcomes. For example, we can form a first low-risk group consisting of chromosomes 1, 3, 4, 5, 17, 19, 20, and "47,XYY" based on the number of spontaneous abortions (less than 100 per 100,000 pregnancies). An intermediate-risk group consisting of chromosomes 2, 6-12 can be defined based on the number of spontaneous abortions (less than 200 per 100,000 pregnancies, but more than 100). A high-risk group consisting of chromosomes 14, 15, 18, and 22 can be defined based on the number of spontaneous abortions (more than 200 per 100,000 pregnancies). The highest risk group can be defined based on the number of spontaneous abortions (more than 1,000 per 100,000 pregnancies) or live births known to have adverse health effects, consisting of chromosomes 13, 16, 21, "45,X", "47,XXY", and "47,XXX". Other classification methods can also be used, such as a first group including chromosomes 1, 3, 10, 12, 19, and 20, and a second, slightly higher-risk group including chromosomes 4, 5, and "47,XYY". Chromosomes can also be classified separately according to complete addition (trisomy), normal pairing (disomy), and complete deletion (monosomy). For example, chromosome 3 (disomy) could be in a different group than chromosome 3 (trisomy). Trisomy (complete addition) is generally considered high-risk and should be avoided.
[0096] Chromosome sets can include a single chromosome or a subset of chromosomes (e.g., a subset of chromosomes with similar risk distributions or below a risk threshold). Chromosome sets can define a specific type or category of mosaicism, such as the type of chromosomes and the count of aneuploid cells in an embryo. These chromosomes then become the focus of building an AI / machine learning model that identifies morphological features associated with that chromosomal modification. In one embodiment, each image is labeled using category labels based on implantation priority / risk distribution, for example, based on the grouping methods listed above, which are based on the risks listed in Table 1 (e.g., embryo images in the "low-risk" group could be assigned category label 1, embryo images in the "intermediate-risk" group could be assigned category label 2, etc.). Note that the grouping methods above are illustrative only; in other embodiments, other clinical risk distributions or other clinical data or risk factors can be used to define (different) chromosome sets and assign chromosome set labels to images. As mentioned above, embryos may exhibit mosaicism, where different cells in the embryo have different chromosome sets, and therefore (mosaic) embryos are a mixture of euploid (normal chromosomes) and aneuploid cells (chromosomal redundancy / deletion / modification). Therefore, risk groups can be defined based on the presence of mosaics and the type and number / range of aneuploidy present. In some embodiments, risk can be based on the most severe aneuploidy present in the embryo (even if it is only present in a single cell). In other embodiments, a threshold number of low-risk aneuploidy can be defined, and then if there are too many aneuploidy (i.e., the number of aneuploidy exceeds the threshold), the embryo is reclassified as higher risk.
[0097] In step 102, we generate a training dataset 120 from the first set of images. Each image includes an image of an embryo captured after in vitro fertilization and is labeled with one or more chromosome set tags. Each tag indicates the presence of at least one aneuploidy associated with the corresponding chromosome set. The training dataset is configured to include images labeled with each chromosome set such that the model is exposed to each chromosome set to be detected. It should also be noted that a single embryo / image may have multiple different aneuploidies, and therefore be labeled with multiple chromosome sets and included in multiple chromosome sets.
[0098] Similarly, in step 103, we generate test dataset 140 from the second set of images. Again, each image includes an image of an embryo taken after in vitro fertilization and labeled with one or more chromosome set tags. Each tag indicates the presence of at least one aneuploidy associated with the corresponding chromosome set. Like training dataset 120, test dataset 140 includes images labeled with each chromosome set.
[0099] Training set 120 and test set 140 can be generated from images of embryos from which PGT-A results and / or pregnancy results (e.g., for implanted embryos) can be obtained, which can be used to label the images. Typically, the images will be phase-contrast microscopic images of embryos captured 3–5 days after in vitro fertilization (IVF). Such images are usually captured during IVF to help embryologists decide which embryos to select for implantation. However, it should be understood that other microscopic images captured at other times, under other lighting conditions, or at other magnifications can be used. In some embodiments, time-lapse sequences of images can be used, for example, by combining / concatenating a series of images into a single image that is analyzed by the AI model. Typically, the available image pool will be divided into a large training set containing approximately 90% of the images and a small (remaining 10%) blinded test set; that is, the test dataset is not used to train the model. A small portion of training set 120, such as 10%, can also be allocated to the validation dataset. Preferably, the relative proportion of each chromosome set in the test dataset 120 is similar to the relative proportion of each chromosome set in the training dataset 140 (e.g., within 10%, preferably within 5% or lower).
[0100] In step 104, we use the same training dataset 120 used to train all models to train at least one chromosome group AI model for each chromosome group. Each chromosome group AI model is trained to recognize morphological features in images labeled with the relevant chromosome group. As an additional or alternative, at least one multi-group AI model can be trained on the training data. Each multi-group AI model is trained to independently recognize morphological features in images labeled with each relevant chromosome group to generate multiple sets of outputs for the input image, indicating the presence of at least one aneuploidy associated with each chromosome group in the image.
[0101] Then, in step 105, we use the test dataset to select the best chromosome set AI model for each chromosome set or the best multi-set AI model (depending on what was generated in step 104). In some embodiments, the finally selected model will be used to generate further ensemble or knowledge distillation models. In step 106, we deploy the selected AI model to screen for the presence of one or more aneuploidies in the embryo images.
[0102] Therefore, the approach to building an AI model for aneuploidy screening that can detect / predict a wide range of chromosomal defects is to decompose the problem, train a single target AI model for a subset of chromosomal defects, and then combine the individual AI models to detect a larger range of chromosomal defects. As mentioned above, each chromosome set will be treated as independent of each other. Each embryo (or embryo image) may have multiple chromosomal defects, and a single image may be associated with multiple chromosome sets. That is, a mosaic embryo (where different cells have different aneuploidies) may have multiple set labels corresponding to each aneuploidy present in the embryo. In each case, the complete training dataset will be used to create a machine learning model (or ensemble / distillation model). This process is repeated multiple times for each chromosome set of interest (limited by the quality and total size of the dataset capable of creating machine learning models) so that multiple models covering different chromosomal modifications can be created from the same training dataset. In some cases, these models may be very similar to each other, all built from the same "base" model, but with separate classifiers in the final layer corresponding to each chromosome considered. In other cases, the models may only process one chromosome separately, and these models may be combined using an ensemble or distillation approach. These scenarios will be discussed below in conjunction with the selection of the best model.
[0103] In one embodiment, the step of training at least one chromosome group AI model 103 for each chromosome group includes: training one or more of a binary model, a hierarchical (multi-layered) model, or a single multi-group model for each chromosome group, and then selecting the best model for that chromosome group using a test dataset. Figures 2A to 2D This will be further explained and discussed below. Figure 2A , 2B Figures 1 and 2C illustrate flowcharts of the training steps for a binary model 137, a hierarchical model 138, and a multi-group model 139 according to one embodiment. Figure 2D This is a flowchart of the steps for selecting the best chromosome set AI model 146 using a test dataset 140 according to one embodiment.
[0104] Figure 2AThis is a flowchart of step 104a, training a binary model for chromosome sets. Images in the training set 120 are labeled with a label matching the i-th chromosome set, such as a "present" label (or "yes", or 1), to create the i-th chromosome set 121 for the images. Then, all other images in the training set are labeled with a "not present" label (or "no", or 0), to create a set 122 of all other images. The binary model is then trained (step 131) using both the "present" and "not present" labels (i.e., on the i-th chromosome set 121 and all other images 122), such that the binary model 127 generates a binary output about the input image indicating whether the i-th chromosome set associated with the chromosomal abnormality exists (or does not exist) in the image. The "present" label typically indicates presence in at least one cell (e.g., if there is mosaicism), but can also indicate a threshold number of cells in the embryo, or presence in all cells.
[0105] Figure 2B This is a flowchart of step 104b of training a hierarchical model for a hierarchical sequence for chromosome set (i-th chromosome set 121), which we refer to as a hierarchical hierarchical model (or multi-layer model). In this embodiment, it includes a hierarchical sequence for training a hierarchical binary model (although the requirements for the binary model can be relaxed as described below). At each layer, images associated with the i-th chromosome set 121 are assigned a first label and trained on a second set of images from training dataset 120, where the second set of images is grouped based on the highest quality level (quality group). At each sequential layer, the second set of images used for training is a subset of the second set of images in the previous layer, with a quality lower than the highest quality of the second set of images in the previous layer. That is, we first divide the dataset into: images labeled with the i-th chromosome set (associated with the i-th chromosome set label), and the remaining images in the dataset. The remaining images in the dataset are assigned a quality label. Then at each level, we divide the current dataset into: a first group corresponding to the highest remaining quality level in the dataset, and the remaining lower quality images (i.e., the residual group of images). At the next level, the previous residual groups are further divided into: a first group corresponding to the highest quality level within the (residual) dataset, and the remaining lower quality images (which become the updated residual image group). This process is repeated until we reach the lowest level.
[0106] That is, the second group (or second category) includes embryo images of different quality levels based on genetic integrity (including chromosomal defects such as mosaicism) and viability (if implanted in a patient, a viable embryo leads to pregnancy and is considered "good"). The basic principle behind the hierarchical model approach is that an embryo image considered high-quality is likely to have the highest quality morphological features (i.e., "looks like the best embryo") among the images with the fewest abnormalities, and therefore will have the greatest morphological difference compared to an embryo image containing chromosomal defects (i.e., "looks bad or has abnormal features"), thus enabling AI algorithms to better detect and predict morphological features between these two (extreme) categories of images.
[0107] exist Figure 2BIn the illustrated embodiment, we assign quality labels to each image in the training dataset 120, where the quality label set comprises a set of quality label gradations that can be used to partition the training dataset 120 into subsets of different qualities. Each image has a quality label. In this embodiment, they include “viable euploid embryos” 123, “non-viable euploid embryos” 125, “mild aneuploid embryos” 127, and “severe aneuploid embryos” 129. In other embodiments, they can be risk categories of adverse outcomes, such as very low risk, low risk, moderate risk, high risk, or very high risk, or simply low risk, moderate risk, and high risk. The training dataset is then partitioned as described above: the i-th chromosome group of interest, and the remaining images. We then train the top-level binary model 132 by partitioning the training set into a first quality dataset 123 labeled “viable euploid embryos” and another dataset 124 containing all other images, and training a binary model on the images labeled the i-th chromosome group 121 and the images in the first quality dataset 123. Then, we sequentially train one or more (two in this example) intermediate layer binary models 132, 133, where in each intermediate layer, the next quality level dataset is generated by selecting images with labels having the highest quality labels in another dataset, and the binary model is trained on images labeled with chromosome sets and images in the next quality dataset. Therefore, we select “euploid non-viable embryos” 125 from the other lower quality images 124 in the top layer, and train the first intermediate layer model 133 on images from the i-th chromosome set 121 and “euploid non-viable embryos” 125. The remaining other lower quality images 126 include “non-severe aneuploid embryos” 127 and “severe aneuploid embryos” 129. In the next intermediate layer, the next quality level image is extracted again, namely “non-severe aneuploid embryos” 127, and another intermediate layer model 133 is trained on images from the i-th chromosome set 121 and “non-severe aneuploid embryos” 127. The remaining lower-quality images 128 now include “severe aneuploid embryos” 129. We then train a binary base layer model 135 on the image labeled with the i-th chromosome set 121 and images from another dataset (i.e., “severe aneuploid embryos” 129) from the previous layer. The output of this step is a trained hierarchical (binary) model 138 that generates a binary output about the input image to indicate the presence (or absence) of a chromosomal abnormality associated with the i-th chromosome set in the image. This can be repeated multiple times to generate several different models for comparison / selection, including by varying the number of layers / quality labels (e.g., from 5: (very low risk, low risk, intermediate risk, high risk, very high risk) to 3 (low risk, intermediate risk, high risk)).
[0108] In some embodiments, after training a first binary basal-level model for a first chromosome set, each of the other chromosome sets is trained again using images of “severe aneuploid embryos” 129. That is, training the hierarchical model involves training the other chromosome sets against the other datasets (“severe aneuploid embryos” 129) used to train the first binary basal-level model. In some embodiments, we may skip the intermediate layers and simply use the top-level and basal-level models (in which case the basal layers are trained on images with multiple quality levels, but not on “euploid viable embryos” 123).
[0109] In the example above, the model is a single-label model, where each label is simply "present" / "absent" (with a probability). However, in another embodiment, the model can also be a multi-class model, where each independent label includes multiple independent aneuploid categories. For example, if the group labels are "severe," "moderate," or "mild," each label can have aneuploid categories such as ("missing," "insertion," "duplication," "deleted," "normal"). It's important to note that group labels are independent, so the confidence level of an aneuploidy in one chromosome set does not affect the model's confidence level for aneuploidies in another chromosome set; for example, they can all be high confidence or all low confidence. Categories within a label are mutually exclusive, so the sum of probabilities for different categories within a label is 1 (e.g., you cannot have both a loss and an insertion on the same chromosome). Therefore, the output for each group is a list of probabilities for each category, rather than a binary / "yes" output. Furthermore, since the labels are independent, different labels / chromosome groups can have different binary / multi-class categories. For example, some might be binary (euploid, aneuploid), while others might be multi-class ("missing", "insertion", "duplication", "deleted", "normal"). That is, the model is trained to estimate the probability of each aneuploid category within a label. In other words, if there are m categories, the output for those categories will be a set of mn yes / no (existent / non-existent, or 1 / 0 value) probabilities (e.g., in a list or similar data structure), along with the overall probability of that label.
[0110] In another embodiment, the hierarchical sequence of the hierarchical model can be a mixture of multi-class and binary models, or it can be entirely multi-class models. That is, referring to the discussion above, we can replace one or more (or all) binary models with multi-class models, which can also be trained on a set of chromosome groups in addition to quality labels. In this respect, the complete hierarchical sequence of the hierarchical model covers all available group labels within the training set. However, each model in the sequence can be trained only on a subset of one or more chromosome groups. In this way, a top-level model can be trained on a large dataset and classify the dataset into one or more predicted outcomes. Subsequent models in the sequence are then trained on a subset of data related to one of these outcomes, further classifying this set into more refined subgroups. This process can be repeated multiple times to create a series of models. By repeating this process and changing which levels use binary models, which levels use multi-class models, and the number of different quality labels, a series of models can be trained on different subsets of the training dataset.
[0111] In another embodiment, the model can be trained as a multi-group model (single-class or multi-class). That is, if there are n chromosome labels / groups, instead of training a model for each group separately, a single multi-group model is trained that simultaneously estimates each of the n group labels in a single pass through the data. Figure 2CThis is a flowchart of step 104c, which describes the training of a single multi-set model 136 for a chromosome set (the i-th chromosome set 121). Then, the single multi-set model 136 is trained on training data 120 using presence and absence labels (in the binary case) for each chromosome set 120 and all other images 122, to generate a multi-set model 139. When an input image is presented, the multi-set model 139 generates multiple sets of outputs indicating the presence or absence of at least one aneuploid associated with each chromosome set in the image. That is, if there are n chromosome sets, the output will be a set of n yes / no (present / absent, or 1 / 0 values) results (e.g., in a list or similar data structure). In the multi-class case, there will be additional probability estimates for each class within the chromosome set. Note that the specific model architecture (e.g., the configuration of convolutional and pooling layers) is similar to the single-chromosome model discussed above, but the final output layer will differ because, for a particular chromosome set, which is neither binary nor multi-class classification, the output layer must generate independent estimates for each chromosome set. This change in the output layer effectively alters the optimization problem, thus offering different performance / results compared to the multiple single-chromosome models discussed above, providing further diversity in the model / results, which may help in finding the optimal overall AI model. Moreover, multi-set models do not need to estimate / classify all chromosome sets; instead, we can train several (M>1) multi-set models, each estimating / classifying a different subset of chromosome sets. For example, if there are n chromosome sets, we can train M multi-set models separately, where each model simultaneously estimates k sets, where M, k, and n are integers, and n = Mk. However, note that each multi-set model does not need to estimate / classify the same number of chromosome sets; for example, if we train M multi-set models and each multi-set model jointly classifies k... m If there are 3 chromosome sets, then n = ∑ m=1..M k m .
[0112] In another embodiment, multiple sets of models can also be used. Figure 2B In the hierarchical model approach shown, instead of training a hierarchical model separately for each chromosome group (e.g., training n hierarchical models for n chromosome groups), we can use a hierarchical approach to train a single multi-group model for all chromosome groups, where the dataset is continuously partitioned based on the quality level of the remaining images within the dataset (we train a new multi-group model at each layer). Furthermore, we can use a hierarchical approach to train several (M>1) multi-group models, where each multi-group model classifies a different subset of the chromosome group, such that all groups are classified by one of the M multi-group models. Moreover, each multi-group model can be a binary model or a multi-class model.
[0113] Figure 2DThis is a flowchart of step 105, which selects the best chromosome set AI model for the i-th chromosome set or the best multi-set model. The test dataset contains images of all euploid and aneuploid categories selected for training the model. We use the test dataset 140 and provide (unlabeled) images as input from... Figures 2A to 2D Each of the binary model 137, the hierarchical model 128, and the multi-group model 139 is used. We obtain the test results 141 for the binary model, 142 for the hierarchical model, and 143 for the multi-group model, and compare the model results 144 using the i-th chromosome group label 145. The best-performing model 146 is then selected using selection criteria, such as calculating one or more metrics and comparing the models with each other using one or more of these metrics. The metric can be selected from a list of recognized performance metrics, such as (but not limited to): total accuracy, balanced accuracy, F1 score, average class accuracy, precision, recall, log loss, or a custom confidence or loss metric, as described below (e.g., Equation (7)). The performance of the model on a set of validation images is measured according to the metrics, and the best-performing model is selected accordingly. These models can be further ranked using secondary metrics, and this process is repeated multiple times until a final model is obtained or a model is selected (if necessary, for creating an ensemble model).
[0114] The best-performing model can be further refined using ensemble or knowledge distillation methods. In one embodiment, an ensemble model for each chromosome set can be generated by training multiple final models, each based on the best chromosome set AI model for the corresponding set (or multiple sets, if multiple sets are selected), and each trained on a training dataset with different initial condition sets and image rankings. The final ensemble model is obtained by combining the multiple trained final models according to an ensemble voting strategy, combining models that exhibit contrasting or complementary behavior based on their performance on one or more of the metrics listed above.
[0115] In one embodiment, a distillation model is generated for each chromosome set. This involves training multiple teacher models, each based on the best chromosome set AI model for the corresponding set (or multiple sets, if multiple set models are selected), and each teacher model being trained on at least a portion of the training dataset with different initial conditions and image sorting sets. A student model is then trained on the training dataset using the multiple trained teacher models with a distillation loss function.
[0116] These operations are repeated for each chromosome set to generate a total aneuploidy screening AI model 150. Once the aneuploidy screening AI model 150 is trained, it can be deployed on a computing system to provide real-time (or near real-time) screening results. Figure 1B This is a flowchart of a method 110 according to one embodiment of using a trained aneuploid screening AI model to computationally generate an estimate of the presence of one or more aneuploids in an embryo image.
[0117] In step 111, an aneuploidy screening AI model 150 is generated in the computing system according to the method 100 described above. In step 112, an image containing an embryo captured after in vitro fertilization is received from the user via the user interface of the computing system. In step 113, the image is provided to the aneuploidy screening AI model 150 to obtain an estimate of the presence of one or more aneuploids in the image. Then, in step 114, a report regarding the presence of one or more aneuploids in the image is sent to the user via the user interface.
[0118] A related cloud-based computing system may also be provided, which is used to computationally generate an artificial intelligence (AI) model 150 for aneuploidy screening, configured to be generated according to training method 100, to estimate the presence of one or more aneuploids in an embryo image (including whether they are present in at least one cell in the case of mosaicism, or whether they are present in all cells of the embryo) (method 110). Figure 3 , 4 Further explanation is provided in 5A and 5B.
[0119] Figure 3This is a schematic architecture of a cloud-based computing system 1, used to computationally generate an aneuploidy screening AI model 150, which is then used to generate a report with estimates of the presence of one or more aneuploids in received embryo images. Input 10 includes data such as embryo images and outcome information that can be used to generate labels (classifications) (presence of one or more aneuploids, live birth, or successful implantation, etc.). This is fed as input to a model creation process 20 that creates computer vision and deep learning models, which are combined to generate an aneuploidy screening AI model to analyze the input images. This can also be referred to as an aneuploidy screening artificial intelligence (AI) model or an aneuploidy screening AI model. A cloud-based model management and monitoring tool (which we call model monitor 21) is used to create (or generate) the AI model. It uses a suite of linked services, such as Amazon Web Services (AWS), which manages the training, logging, and tracking of the model in relation to image analysis and model specificity. Other similar services on other cloud platforms can also be used. These services can utilize deep learning methods 22, computer vision methods 23, classification methods 24, statistical methods 25, and physics-based models 26. Model generation can also use domain expertise 12 (e.g., domain expertise from embryologists, computer scientists, scientific / technical literature, etc., such as what features to extract and use in computer vision models) as input. The output of the model creation process is an instance of an aneuploidy screening AI model, which in this embodiment is a validated aneuploidy screening (or embryo evaluation) AI model 150. Other aneuploidy screening AI models 150 can be generated using other image data with relevant result data.
[0120] A cloud-based delivery platform 30 is used, which provides user interface 42 for user 40 to access the system. (See reference) Figure 4 To further illustrate this point, Figure 4This is a schematic diagram of an IVF procedure 200 according to one embodiment, which uses an aneuploidy screening AI model 150 to help select embryos for implantation, or to select which embryos to reject, or to select which embryos to undergo invasive PGT-A testing. On day 0, the retrieved oocytes are fertilized (202). They are then cultured in vitro for several days, and images of the embryos are captured, for example, using a phase-contrast microscope (204). Preferably, the model is trained and used to reference images of embryos captured on the same day or during a specific time window at a particular epoch. In one embodiment, the time is 24 hours, but other time windows, such as 12 hours, 36 hours, or 48 hours, may also be used. Generally, a smaller time window of 24 hours or less is preferred to ensure greater similarity in appearance. In one embodiment, it may be a specific day (a 24-hour window from the beginning (0:00) to the end (23:39) of the day), or a specific number of days such as day 4 or day 5 (a 48-hour window starting from day 4). Alternatively, the time window can define the window size and epoch, such as a 24-hour window centered on day 5 (i.e., day 4.5 to day 5.5). The time window can be open and have a lower limit, such as at least 5 days. As mentioned above, while it is best to use embryonic images within a 24-hour time window around day 5, it is understood that earlier embryos, including images from day 3 or day 4, can be used.
[0121] Typically, several eggs are fertilized simultaneously, resulting in multiple images to consider which embryo is best suited for implantation (i.e., most viable) (this may include identifying embryos to be excluded due to a high risk of severe defects). Users upload captured images to platform 30 via user interface 42, for example, using a drag-and-drop function. Users can upload single or multiple images, for example, to help select (or reject) which embryo from a set of multiple embryos considered for implantation. Platform 30 receives one or more images 312 stored in a database 36 that includes an image repository. The cloud-based delivery platform includes on-demand cloud servers 32 that can perform image preprocessing (e.g., object detection, segmentation, padding, normalization, cropping, centering, etc.) and then feed the processed images to a trained AI (aneuploidy) screening model 150, which runs on one of the on-demand cloud servers 32, to analyze the images to generate an aneuploidy risk / embryo viability score 314. A report (316) is generated of model results (e.g., the probability of one or more aneuploidies, or a binary call (use / not use), or other information obtained from the model), and is sent, for example, via user interface 42 or otherwise provided to user 40. The user (e.g., an embryologist) receives the aneuploidy risk / embryo viability score and report via the user interface, and can then use the report (probability) to help decide whether to implant the embryo, or which embryo in the set is most likely to be implanted. The selected embryos are then implanted. To help further refine the AI model, pregnancy outcome data can be provided to the system, such as the detection (or non-detection) of a heartbeat in the first ultrasound scan after implantation (typically around 6-10 weeks post-fertilization), or aneuploidy results from a PGT-A test. This allows the AI model to be retrained and updated as more data becomes available.
[0122] Images can be captured using a range of imaging systems, such as those already in existing IVF clinics. This eliminates the need for IVF clinics to purchase new or specialized imaging systems. Imaging systems are typically optical microscopes used to capture single-phase images of embryos. However, it should be understood that other imaging systems can be used, particularly optical microscope systems employing a range of imaging sensors and image capture techniques. These can include phase-contrast microscopy, polarization microscopy, differential interference contrast (DIC) microscopy, dark-field microscopy, and bright-field microscopy. Images can be captured using a conventional optical microscope equipped with a camera or image sensor, or using a camera with an integrated optical system (including smartphone systems) capable of capturing high-resolution or high-magnification images. The image sensor can be a CMOS sensor chip or a charge-coupled device (CCD), each with associated electronics. The optical system can be used to collect specific wavelengths or to collect (or exclude) specific wavelengths using filters including bandpass filters. Some image sensors can be used to operate on or be sensitive to light of specific wavelengths or wavelengths outside the optical range, including infrared (IR) or near-infrared. In some embodiments, the imaging sensor is a multispectral camera that collects images across multiple different wavelength ranges. The lighting system can also be used to illuminate the embryo with light of a specific wavelength, band, or intensity. Stop and other components can be used to limit or modify the illumination of certain parts of the image (or image plane).
[0123] Furthermore, the images used in the embodiments described herein can be from video and time-lapse imaging systems. A video stream is a periodic sequence of image frames, where the interval between image frames is defined by the capture frame rate (e.g., 24 or 48 frames / second). Similarly, a time-lapse system captures image sequences at a very slow frame rate (e.g., 1 image / hour) to obtain image sequences during embryonic development (post-fertilization). Therefore, it can be understood that the images used in the embodiments described herein can be single images extracted from a video stream or time-lapse sequences of images of the embryo. In the case of extracting images from a video stream or a time-lapse sequence, the images to be used can be selected from those captured at times closest to a reference time point (e.g., 5.0 days or 5.5 days post-fertilization).
[0124] In some embodiments, preprocessing may include image quality assessment such that an image that fails the quality assessment can be excluded. If the original image fails the quality assessment, another image can be captured. In embodiments where images are selected from a video stream or time-lapse sequence, the selected image is the first image that passes the quality assessment and is closest to a reference time. Alternatively, a reference time window (e.g., 30 minutes after the start of day 5.0) and image quality criteria may be defined. In this embodiment, the selected image is the image with the highest quality during the selected reference time window. The image quality criteria used to perform the quality assessment may be based on pixel color distribution, brightness range, and / or anomalous image characteristics or features indicating poor quality or device malfunction. Thresholds may be determined by analyzing a reference set of images. This may be based on manual assessment or an automated system that extracts outliers from the distribution.
[0125] You can refer to Figure 5A To further understand the generation of the aneuploid screening AI model 150, Figure 5A This is a schematic flowchart illustrating the generation of an aneuploid screening model 150 using a cloud-based computing system 1 according to one embodiment. This AI model 150 is used to estimate the presence of aneuploids (including mosaicks) in an image. Reference Figure 5B The generation method is handled by the model monitor 21.
[0126] Model monitor 21 allows user 40 to provide image data and metadata to a data management platform, including a data repository (step 14). Data preparation steps are performed, such as moving images to specific folders and renaming and preprocessing them (e.g., object detection, segmentation, alpha channel removal, padding, cropping / localization, normalization, scaling, etc.). Feature descriptors can also be computed, and enhanced images can be pre-generated. However, additional preprocessing, including enhancement, can also be performed during training (i.e., in-run). Image quality assessment can also be performed to allow rejection of significantly inferior images and to allow the capture of replacement images. Similarly, patient records or other clinical data are processed (prepared) to add embryo viability classification (e.g., viable or non-viable), which is linked or associated with each image to enable its use in training machine learning and deep learning models. The prepared data is loaded (step 16) onto a cloud provider's (e.g., AWS) template server 28, which has the latest version of the training algorithm. The template server is saved and multiple copies are made on a series of training server clusters 37, which can be CPU, GPU, ASIC, FPGA or TPU (Tensor Processing Unit) based training server clusters, which form (local) training servers 35.
[0127] Then, for each job submitted by user 40, the model monitor web server 31 requests training servers 37 from multiple cloud-based training servers 35. Each training server 35 uses libraries such as PyTorch, Tensorflow, or equivalents to run pre-prepared code (from template server 28) for training the AI model, and can use computer vision libraries such as OpenCV. PyTorch and OpenCV are open-source libraries with low-level commands for building CV machine learning models.
[0128] Training server 37 manages the training process. This may include, for example, dividing images into training, validation, and blind validation sets using a random assignment process. Furthermore, during training and validation cycles, training server 37 may randomize the image set at the beginning of each cycle to allow analysis of different image subsets in each cycle, or to analyze different image subsets in different orders. Additional preprocessing may be performed if no preprocessing was previously performed or the preprocessing was incomplete (e.g., during data management), including object detection, segmentation, and generating masked datasets (e.g., IZC images only), calculating / estimating CV feature descriptors, and generating data augmentations. Preprocessing may also include padding, normalization, etc., as needed. That is, preprocessing step 102 can be performed before training, during training, or in some combination (i.e., distributed preprocessing). The number of running training servers 35 can be managed from a browser interface. As training progresses, logging information about the training status is recorded on a distributed logging service, such as CloudWatch 60 (step 62). Key patient and accuracy information is also parsed from the logs and saved to relational database 36. The models are also periodically saved to a data storage device (e.g., AWS Simple Storage Service (S3) or a similar cloud storage service) 50 (step 51) so that they can be retrieved and loaded later (e.g., in case of an error or other stoppage). If the training server's job completes or encounters an error, an email update on the status of the training server is sent to user 40 (step 44).
[0129] Multiple processes occur within each training cluster 37. Once the cluster is started via web server 31, the script runs automatically, reads the prepared images and patient records, and begins the requested specific PyTorch / OpenCV training code 71. The input parameters for model training 28 are provided by user 40 via browser interface 42 or via configuration script. The training process 72 is then initiated for the requested model parameters, which can be a lengthy and intensive task. Therefore, to prevent loss of progress during training, logs are periodically saved to a logging service (e.g., AWS Cloudwatch) 60 (step 62), and the current version of the model (at training time) is saved to a data storage service (e.g., S3) 51 (step 51) for later retrieval and use. Figure 5B An embodiment of a schematic flowchart illustrating the model training process on a training server is shown. By accessing a series of trained AI models on a data storage service, multiple models can be combined, for example using ensembles, distillation, or similar methods, to merge a series of deep learning models (such as PyTorch) and / or target computer vision models (such as OpenCV), generating a more robust aneuploid screening AI model 100 provided to a cloud-based delivery platform 30.
[0130] Then, the cloud-based delivery platform 30 system allows user 10 to drag and drop images directly onto the web application 34. The web application 34 prepares the images and passes them to the trained / validated aneuploidy screening AI model 30 to obtain an embryo viability score (or aneuploidy risk), which is immediately returned in a report (e.g., ...). Figure 4 (As shown). Web application 34 also allows medical institutions to store data such as images and patient information in database 36, create various reports on the data, create audit reports on tool usage for their organization, group, or specific users, and manage billing and user accounts (e.g., creating users, deleting users, resetting passwords, changing access levels, etc.). Cloud-based delivery platform 30 also allows product administrators to access the system to create new customer accounts and users, reset passwords, and access customer / user accounts (including data and screens) for technical support.
[0131] The various steps and variations in the embodiments for generating an AI model to estimate aneuploidy risk / embryo viability scores from images will now be discussed in more detail. (References) Figure 3The model was trained using images captured 5 days post-fertilization (i.e., the 24-hour period from 00:00 to 23:59 on day 5). However, as mentioned above, it is still possible to develop an effective model using shorter time windows (e.g., 12 hours), longer time windows (e.g., 48 hours), or even no time window (i.e., open). More images can be taken on other days, such as day 1, 2, 3, or 4, or during the shortest post-fertilization period, such as at least 3 days or at least 5 days (e.g., open time windows). However, it is generally preferred (but not absolutely necessary) that the images used for training the AI model and subsequently classified by the trained AI model are taken during similar and preferably the same time window (e.g., the same 12, 24, or 48-hour time window).
[0132] Before analysis, each image undergoes preprocessing (image preparation). A series of preprocessing steps or techniques can be applied. This can be performed after being added to data storage 14 or during training on training server 37. In some embodiments, an object detection (localization) module is used to detect and localize images on embryos. Object detection / localization includes estimating bounding boxes containing embryos. This can be used for image cropping and / or segmentation. Images can also be filled with given boundaries, and then color balance and brightness normalized. The image is then cropped so that the outer region of the embryo is close to the image boundary. This is achieved by using computer vision techniques for boundary selection, including the use of AI object detection models.
[0133] Image segmentation is a computer vision technique used to prepare images for certain models to select relevant regions of interest for model training, such as intracavitary spaces (IZCs), individual cells within the embryo (i.e., cell boundaries to assist in mosaic identification), or other regions like the zona pellucida. As mentioned above, mosaicism occurs when different cells in an embryo have different sets of chromosomes. That is, a mosaic embryo is a mixture of euploid (normal chromosomes) and aneuploid cells (chromosomal excess / deletion / modification), and there may be multiple aneuploid cells with different aneuploidy, and in some cases, no euploid cells may be present. Segmentation can be used to identify IZCs or cell boundaries, thereby segmenting the embryo into individual cells. In some embodiments, multiple masked (enhanced) images of the embryo are generated, where each image except for individual cells is masked. Images can also be masked to generate only images of IZCs, thus excluding the zona pellucida and background, or these images can remain in the image. The masked images (e.g., IZC images masked to contain only IZCs, or masked to identify individual cells in the embryo) can then be used to train an aneuploid AI model. Scaling involves resizing an image to a predefined scale to fit the specific model being trained. Augmentation involves making minor changes to image copies, such as rotating the image to control the orientation of the embryonic disc. Using segmentation before deep learning has a significant impact on the performance of deep learning methods. Similarly, augmentation is crucial for generating robust models.
[0134] A range of image preprocessing techniques can be used to prepare embryo images for analysis, such as ensuring image standardization. Examples include:
[0135] Alpha channel stripping: This involves stripping the alpha channel (if present) from the image to ensure it is encoded in a 3-channel format (e.g., RGB), such as by removing the transparency map;
[0136] Fill / Enhance: Before segmentation, cropping, or boundary finding, each image is filled / enhanced with fill borders to generate square aspect ratios. This process ensures consistent and comparable image dimensions and compatibility with deep learning methods, which typically require square-sized images as input, while also ensuring that critical components of the image are not cropped.
[0137] Normalization: Normalizes an RGB (red, green, blue) or grayscale image to a fixed average across all images. For example, this involves taking the average of each RGB channel and dividing each channel by its average. Then, each channel is multiplied by a fixed value of 100 / 255 to ensure that the average of each image in the RGB space is (100, 100, 100). This step ensures that color deviations between images are suppressed and that the brightness of each image is normalized.
[0138] Thresholding: Image thresholding is performed using binary, Otsu, or adaptive methods. This includes morphological processing of the image using dilation (opening), erosion (closing), and scaling gradients, and using scaling masks to extract the external and internal boundaries of the shape;
[0139] Object detection / cropping: Perform object detection / cropping on the image to locate the embryo and ensure that there are no artifacts at the image edges. This can be done using an object detector that uses an object detection model (discussed below) that is trained to estimate bounding boxes containing the main features of the image, such as the embryo (IZC or zona pellucida), so that the image is a well-centered and cropped embryo;
[0140] Extraction: Geometric properties of the boundaries are extracted using the elliptical Hough transform of the image contour, for example, based on the best ellipse fit of the elliptical Hough transform calculated on a binary thresholded map of the image. This method selects the hard boundaries of the embryo in the image and crops the square boundaries of the new image so that the longest radius of the new ellipse is enclosed by the width and height of the new image, and the center of the ellipse is the center of the new image.
[0141] Zooming: Scaling an image by ensuring a consistently centered image with a consistent boundary size around the elliptical region;
[0142] Segmentation: Images are segmented to identify intracranial cytoplasmic zone (IZC) regions, zona pellucida regions, and / or cell boundaries. Segmentation can be performed by computing the best-fit contour around a non-elliptical image using a geometrically active contour (GAC) model or a morphological snake within a given region. The snake's interior and other regions can be processed differently based on the focal point of the trained model on the intracranial cytoplasmic zone (IZC), which may contain blastocysts or cells within blastocysts. Alternatively, a semantic segmentation model can be trained that identifies the category of each pixel in the image. For example, a semantic segmentation model can be developed using a U-Net architecture with a pre-trained ResNet-50 encoder and trained using a binary cross-entropy loss function to segment the background, zona pellucida, and IZC, or to segment cells within the IZC.
[0143] Note: The image is annotated by selecting feature descriptors and masking all regions of the image (except for the region within a given radius of the descriptor keypoints);
[0144] Resizing / scaling: Resizes / scales the entire image set to a specified resolution; and
[0145] Tensor transformation: This involves converting each image into a tensor instead of a visually displayed image, as this data format is more suitable for deep learning models. In one embodiment, tensor normalization is obtained from standard pre-trained ImageNet values using the mean (0.485, 0.456, 0.406) and standard deviation (0.299, 0.224, 0.225).
[0146] In another embodiment, the object detector uses an object detection model trained to estimate bounding boxes containing embryos. The goal of object detection is to identify the largest bounding box containing all pixels associated with the object. This requires the model to model both the object's location and its category / label (i.e., the content within the box), so the detection model typically includes an object classifier head and a bounding box regression head.
[0147] One approach is to use a region convolutional neural network (or R-CNN) with an expensive search process to search for image inpainting schemes (potential bounding boxes). These bounding boxes are then used to crop the image region of interest. The cropped image is then passed through a classification model to classify the content of the image region. This process is complex and computationally expensive. Another approach is Fast CNN, which uses a CNN to propose feature regions instead of searching for image inpainting schemes. This model uses a CNN to estimate a fixed number of candidate boxes, typically between 100 and 2000. A faster alternative is Faster R-CNN, which uses anchor boxes to constrain the search space of the desired boxes. By default, a standard set of 9 anchor boxes (each of different sizes) is used. Faster R-CNN uses a small network that jointly learns to predict the feature regions of interest, resulting in faster runtime compared to R-CNN or Fast CNN because the expensive region search is replaced.
[0148] For each feature activation emerging from the backside, a model is considered an anchor point (red in the image below). For each anchor point, nine (or more, or fewer, depending on the problem) anchor boxes are generated. The anchor boxes correspond to common object sizes within the training dataset. With multiple anchor points and multiple anchor boxes, thousands of region schemes are generated. These schemes are then filtered through a process called Non-Maximum Suppression (NMS), which selects the largest box containing the smaller bounding box. This ensures that each object has only one box. Since NMS relies on the confidence of each bounding box prediction, a threshold must be considered for when to consider an object as part of the same object instance. Because the anchor boxes cannot perfectly fit the objects, the regression head's job is to predict the offsets of these anchor boxes, thus transforming them into best-fit bounding boxes.
[0149] The detector can also specifically estimate boxes for only a subset of objects; for example, a pedestrian detector might only estimate boxes for people. Object categories that are not of interest are encoded into class 0, corresponding to the background class. During training, patches / boxes for the background class are typically randomly sampled from image regions that do not contain bounding box information. This step allows the model to remain invariant for unwanted objects; for example, it can learn to ignore them instead of misclassifying them. Bounding boxes are typically represented in two different formats: the most common is (x1, y1, x2, y2), where point p1 = (x1, y1) is the top-left corner of the box and p2 = (x2, y2) is the bottom-right corner. Another common box format is (cx, cy, height, width), where the bounding box / rectangle is encoded as the center point (cx, cy) and the box size (height, width). Different detection methods will use different encoding / formats depending on the task and context.
[0150] The regression head can be trained using L1 loss, and the classification head can be trained using cross-entropy loss. Object-specific loss (whether it's background or object) can also be used. The final loss is the sum of these losses. Individual losses can also be weighted, for example:
[0151] Loss = λ1 Regression Loss + λ2 Classification Loss + λ3 Objectivity Loss
[0152] (loss=λ1regression_loss+λ2classification_loss+λ3objectness_loss) (1)
[0153] In one embodiment, a faster RNN-based embryo detection model was used. In this embodiment, approximately 2000 images were manually labeled with ground truth bounding boxes. The boxes were labeled so that the entire embryo (including the zona pellucida) was within the bounding box. If more than one embryo was present (also known as a double embryo transfer), both embryos were labeled so that the model could distinguish between double and single embryo transfers. Because it was impossible to coordinate which embryo was which in a double embryo transfer, the model was configured to raise a usage error when a double embryo transfer was detected. Models with multiple "lobes" were labeled as single embryos.
[0154] As an alternative to GAC segmentation, semantic segmentation can be used. Semantic segmentation is a task that attempts to predict the category or label for each pixel. Tasks like semantic segmentation are called pixel-dense prediction tasks because each input pixel requires an output. The setup of semantic segmentation models differs from standard models because they require the complete image output. Typically, semantic segmentation (or any dense prediction model) has an encoding module and a decoding module. The encoding module is responsible for creating a low-dimensional representation of the image (sometimes called a feature representation). This feature representation is then decoded into the final output image by the decoding module. During training, the predicted label map (used for semantic segmentation) is compared with the ground truth label map that assigns a category to each pixel, and a loss is computed. The standard loss function for segmentation models is binary cross-entropy or standard cross-entropy loss (depending on whether the problem is multi-class). These implementations are consistent with their image classification counterparts, except that the loss is applied pixel-wise (across the entire image channel dimension of the tensor).
[0155] Fully Convolutional Network (FCN) style architectures are commonly used in the domain of general semantic segmentation tasks. In this architecture, a low-resolution image is first encoded using a pre-trained model (such as ResNet) (approximately 1 / 32 of the original resolution, but potentially 1 / 8 if expanded convolutions are used). This low-resolution label map is then upsampled to the original image resolution, and a loss is computed. The intuition behind predicting the low-resolution label map is that semantic segmentation masks are very infrequent and do not require all the extra parameters of a larger decoder. More complex versions of this model exist that use multi-level upsampling to improve segmentation results. Simply put, the loss is computed progressively across multiple resolutions to refine the predictions at each scale.
[0156] One drawback of this model is that if the input data is high-resolution or contains high-frequency information (i.e., smaller / thinner objects), the low-resolution label mapping will fail to capture these smaller structures (especially when the encoding model does not use extended convolutions). In standard encoder / convolutional neural networks, the input image / image features are gradually downsampled as the model progresses. However, because the images / features are downsampled, key high-frequency details may be lost. Therefore, to address this issue, an alternative U-Net architecture can be used instead of using skip connections between the symmetric components of the encoder and decoder. Simply put, each encoded block has a corresponding block in the decoder. The features of each stage are then passed to the decoder along with the lowest-resolution feature representation. For each decoded block, the input feature representation is upsampled to match the resolution of its corresponding encoded block. The feature representation of the encoded block and the upsampled low-resolution features are then concatenated and passed through a 2D convolutional layer. By concatenating features in this way, the decoder can learn to refine the input for each block, selecting which details (low-resolution or high-resolution details) to integrate based on the input. The main difference between FCN-style and U-Net-style models is that in FCN models, the encoder is responsible for predicting a low-resolution label map, which is then upsampled (potentially progressively). However, U-Net models don't have a fully accurate label map prediction until the very last layer. Ultimately, there are many variations of these models that can weigh the differences between them (e.g., hybrid models). The U-Net architecture can also use pre-trained weights, such as ResNet-18 or ResNet-50, for situations where there isn't enough data to train the model from scratch.
[0157] In some embodiments, segmentation is performed using a U-net architecture with a pre-trained dynamic ResNet-50 encoder trained with binary cross-entropy to identify the zona pellucida region, intraluminal region, and / or cell boundaries. Once segmented, an image set is generated in which all regions except the desired region are masked. The AI model can then be trained on these specific image sets. That is, the AI models can be divided into two groups: the first group includes models that incorporate additional image segmentation, and the second group requires the entire unsegmented image. The model trained on images with masked IZC, exposing the zona pellucida region, is called the zona pellucida model. During training, models trained on images with masked zona pellucida (called IZC models) and models trained on images of the complete embryo (i.e., the second group) are also considered.
[0158] In one embodiment, to ensure the uniqueness of each image so that copies of records do not deviate from the results, the name of the new image is set to a hash equal to the content of the original image, as a PNG (lossless) file. During runtime, for any image not present in the output directory (if it doesn't exist, it will be created), the data parser outputs the image in a multi-threaded manner, so if this is a lengthy process, it can be restarted from the same point even if interrupted. The data preparation step may also include processing metadata to remove images associated with inconsistent or contradictory records and to identify any erroneous clinical records. For example, a script can be run on a spreadsheet to integrate metadata into a predefined format. This ensures that the data used to generate and train the model is of high quality and has uniform characteristics (e.g., size, color, scale, etc.).
[0159] Once the data is ready, it can be used to train the AI models discussed above. In one embodiment, multiple computer vision (CV) models are generated using machine learning methods and multiple deep learning models are generated using deep learning methods. The deep learning models can be trained on full embryo images or masked image sets. The computer vision (CV) models can be generated using machine learning methods, using a set of feature descriptors computed from each image. Each individual model is configured to estimate probabilities, such as aneuploidy risk / embryo viability score in an embryo image, and the AI models combine selected models to produce an overall aneuploidy risk / embryo viability score, or similar overall probability or hard classification. Models generated on a single chromosome set can be improved using ensemble and knowledge distillation techniques. Training is performed using a random dataset. Complex image datasets can exhibit uneven distributions, especially if the dataset contains fewer than 10,000 images, where the distribution of key viable or non-viable embryos in the set is uneven. Therefore, consider randomizing the data several times (e.g., 20 times) and then dividing it into training subsets, validation subsets, and blind test subsets as defined below. All randomizations are applied to a single training example to determine which exhibits the optimal distribution for training. As an inference, it is also beneficial to ensure that the ratio of viable to non-viable embryos is the same within each subset. Embryo images are highly diverse, so ensuring a uniform distribution of images between the test and training sets can improve performance. Therefore, after performing randomization, the ratio of images with viable classifications to images with non-viable classifications within each training, validation, and blind validation set is calculated and tested to ensure similar ratios. For example, this might include tests such as whether the range of the ratio is less than a threshold or within a certain variance range considering the number of images. If the ranges are not similar, the randomization is discarded, and new randomizations are generated and tested until randomizations with similar ratios are obtained. More generally, if the result is an n-gram result with n states, the computational steps after performing randomization may include calculating the frequency of each n-gram result state within each training, validation, and blind validation set, and testing whether the frequencies are similar. If the frequencies are not similar, the assignment is discarded and randomization is repeated until randomizations with similar frequencies are obtained.
[0160] Training also includes executing multiple training and validation cycles. In each training and validation cycle, each randomization of the total available dataset typically splits it into three separate datasets, called the training dataset, the validation dataset, and the blind validation dataset. In some variations, more than three datasets can be used; for example, the validation dataset and the blind validation dataset can be stratified into multiple sub-test sets of varying difficulty.
[0161] The first dataset is the training dataset, comprising at least 60%, preferably 70-80%, of the images. Deep learning and computer vision models use these images to create embryo viability assessment models to accurately identify viable embryos. The second dataset is the validation dataset, typically comprising about (or at least) 10% of the images. This dataset is used to validate or test the accuracy of the model created using the training dataset. Although these images are independent of the training dataset used to create the model, the validation dataset still has a small positive bias in accuracy because it is used to monitor and optimize the progress of model training. Therefore, training often aims to maximize the accuracy of the model on this particular validation dataset, which may not necessarily be the optimal model when applied more generally to other embryo images. The third dataset is the blind validation dataset, typically comprising about 10-20% of the images. To address the positive bias of the validation dataset mentioned above, a final unbiased accuracy evaluation of the final model is performed using the third blind validation dataset. This validation occurs at the end of the modeling and validation process, i.e., when the final model is created and selected. It is important to ensure that the accuracy of the final model is relatively consistent with the validation dataset to ensure that the model is applicable to all images. For the reasons mentioned above, the accuracy of validation datasets may be higher than that of blind validation datasets. Results from blind validation datasets are a more reliable measure of model accuracy.
[0162] In some embodiments, data preprocessing also includes image enhancement, wherein the images are modified. This can be performed before or during training (i.e., on the fly). Enhancement can include directly enhancing (changing) the image or by making copies of the image with minor changes. Any number of enhancements can be performed as follows: different numbers of 90-degree rotations of the image, mirror flips, non-90-degree rotations (where diagonal boundaries are filled to match the background color), image blurring, adjusting image contrast using intensity bars, and applying one or more small random translations in the horizontal and / or vertical directions, random rotations, JPEG or compressed noise, random image resizing, random tone jittering, random brightness jittering, contrast-limited adaptive bar equalization, random flips / mirrors, image sharpening, image embossing, random brightness and contrast, RGB color shifts, random hue and saturation, channel shuffling: switching RGB to BGR or RBG or others, coarse dropout, motion blur, median blur, Gaussian blur, random offset scaling rotation (i.e., all three combined). The same set of enhanced images can be used for multiple training and validation epochs, or new enhancements can be generated on the fly in each epoch. Another enhancement for training CV models is changing the "seed" of the random number generator used to extract feature descriptors. Techniques for obtaining computer vision descriptors incorporate random elements when extracting feature samples. This randomness can be modified and included in the enhancement to provide more robust training for the CV model.
[0163] Computer vision models rely on identifying key features of an image and representing them with descriptors. These descriptors can encode qualities such as pixel variations, grayscale, texture roughness, fixed corner points, or image gradient directions; they are implemented in OpenCV or similar libraries. By selecting the features to search for in each image, a model can be built by discovering which arrangement of features is a good indicator of embryonic viability. This process is best implemented using machine learning processes (such as random forests or support vector machines), which are capable of separating images from computer vision analysis based on their descriptions.
[0164] A range of computer vision descriptors, including small and large features, were used in conjunction with traditional machine learning methods to generate “CV models” for embryo selection. Optionally, these can be later combined with deep learning (DL) models to form, for example, ensemble models or used for distillation to train student models. Suitable computer vision image descriptors include:
[0165] Transparent bands obtained through Hough transform: find inner and outer ellipses to approximate the separate transparent bands and band cavities, and record the average and difference of the radii as features;
[0166] Gray-Level Co-occurrence Matrix (GLCM) texture analysis: This method detects roughness in different regions by comparing adjacent pixels within a region. The sample feature descriptors used are: Angular Second Moment (ASM), homogeneity, correlation, contrast, and entropy. The region is selected by randomly sampling a given number of square sub-regions of a given size from the image, and the results of each of the five descriptors for each region are recorded as a feature set.
[0167] Oriented Gradient Bar Chart (HOG): Detects objects and features using a scale-invariant feature transform descriptor and shape context. This method is primarily used in embryology and other medical imaging, but it does not constitute a machine learning model itself.
[0168] Oriented features derived from Accelerated Segmentation Test (FAST) and Rotation Binary Robust Independent Basic Features (BRIENT) (ORB): Industry-standard alternatives to SIFT and SURF features, which rely on a combination of fast keypoint detectors (pixel-specific) and short descriptors, and have been modified to include rotation invariance;
[0169] Binary Robust Invariant Scalable Keypoint (BRISK): A fast detector that combines a set of pixel intensity comparisons by sampling each neighborhood around a keypoint’s specified features;
[0170] Maximum Stable Extreme Region (MSER): A local morphological feature detection algorithm that extracts covariant regions, which are stable connected components associated with one or more grayscale sets extracted from an image.
[0171] Good Tracking Features (GFTT): A feature detector that uses an adaptive window size to detect corner textures, identifies corners using Harris or Shi-Tomasi corner detection, and extracts points that show a high standard deviation in their spatial intensity profile.
[0172] Computer vision (CV) models are constructed using the following method. One (or more) computer vision image descriptor techniques listed above are chosen, and features are extracted from all images within the training dataset. These features are arranged into a combined array and then fed to a K-means unsupervised clustering algorithm; this array is called the codebook and serves as a “visual word bag.” The number of clusters is a free parameter of the model. The clustered features from this point represent “custom features” used by the algorithm in combination, which are compared to each individual image within the validation or test set. Features are extracted for each image, and clusters are performed separately. For a given image with clustered features, the “distance” to each cluster in the codebook is measured (in feature space) using a KD-tree query algorithm, which yields the nearest clustered features. The result of the tree query can then be represented as a bar chart showing the frequency of each feature in the image. Finally, machine learning is used to evaluate whether a particular combination of these features corresponds to a measure of embryo viability. Here, the bar chart and the ground truth are used to perform supervised learning. Methods used to obtain the final selected model include random forests or support vector machines (SVMs).
[0173] Multiple deep learning models can also be generated. Deep learning models are based on neural network methods, typically convolutional neural networks (CNNs) consisting of multiple connected layers. Compared to feature-based methods (i.e., CV models), where each "neuron" contains a non-linear activation function such as a "rectifier" or "sigmoid"), deep learning and neural networks "learn" features instead of relying on hand-designed feature descriptors. This allows them to learn "feature representations" tailored to the desired task.
[0174] These methods are suitable for image analysis because they can extract small details and overall morphology for holistic classification. Therefore, a variety of deep learning models can be used, each with a different architecture (i.e., different numbers of layers and inter-layer connections), such as residual networks (e.g., ResNet-18, ResNet-50, and ResNet-101), densely connected networks (e.g., DenseNet-121 and DenseNet-161), and other variants (e.g., Inception V4 and Inception-ResNetV2). Deep learning models can be evaluated based on stability (stability of the accuracy values on the validation set during training), transferability (the degree of correlation between the accuracy on the training data and the accuracy on the validation set), and prediction accuracy (which models provide the best validation accuracy, including, for both viable and non-viable embryos: total combined accuracy, and balanced accuracy, which is defined as the weighted average accuracy of the two types of embryos). Training involves trying different combinations of model parameters and hyperparameters, including input image resolution, optimizer selection, learning rate and scheduling, momentum value, dropout, and weight initialization (pre-training). A loss function can be defined to evaluate the model's performance, and during training, the deep learning model can be optimized by changing the learning rate to drive the update mechanism of the network weight parameters, thereby minimizing the objective / loss function.
[0175] Deep learning models can be implemented using various libraries and software languages. In one embodiment, the PyTorch library is used to implement neural networks in Python. The PyTorch library also allows the creation of tensors that leverage hardware acceleration (GPU, TPU) and includes modules for building multilayer neural networks. While deep learning is one of the most powerful techniques for image classification, it can be improved by using segmentation or augmentation as described above. Research has found that using segmentation before deep learning has a significant impact on the performance of deep learning methods and helps generate contrasting models. Therefore, preferably, at least some deep learning models are trained on segmented images (e.g., images that identify IZCs or cell boundaries, or images that are masked to exclude regions outside of IZCs or cell boundaries). In some embodiments, multiple deep learning models include at least one model trained on a segmented image and one model trained on an unsegmented image. Similarly, augmentation is also important for generating robust models.
[0176] The effectiveness of a method is determined by the architecture of the deep neural network (DNN). However, unlike feature descriptor methods, DNNs learn the features themselves throughout the convolutional layers before using a classifier. That is, without manually adding proposed features, DNNs can be used to examine existing practices in the literature and to develop previously unused descriptors, especially those that are difficult for the human eye to detect and measure.
[0177] The architecture of a DNN is constrained by the size of the image as input, the hidden layers (with tensor dimensions describing the DNN), and the linear classifier (outputting the number of class labels). Most architectures employ numerous downsampling rates, using small (3×3 pixel) filters to capture left-right, top-bottom, and center concepts. The stacking of (a) 2D convolutional layers, (b) rectified linear units (ReLU), and (c) max-pooling layers allows the number of parameters through the DNN to remain solvable while allowing filters to map high-level (topological) features of the image to intermediate and final micro-features embedded within it. The top layer typically comprises one or more fully connected neural network layers that act as classifiers, similar to an SVM. Typically, a softmax layer is used to normalize the resulting tensor to the probabilities after incorporating the fully connected classifier. Therefore, the model's output is a list of probabilities for an image to be either inactive or active. A range of AI architectures can be based on neural network architectures such as ResNet variants (18, 34, 50, 101, 152), Wide ResNet variants (50-2, 101-2), ResNeXt variants (50-32x4d, l1-32x8d), DenseNet variants (121, 161, 169, 201), Inception (v4), Inception-ResNet (v2), and EfficientNet variants (b0, bl, b2, b3).
[0178] Figure 5C This is a schematic architecture diagram of an AI model 151 comprising a series of layers based on the RESNET152 architecture, according to one embodiment, which transforms an input image into a prediction. It includes two-dimensional convolutional layers... Figure 5C The symbol "CONV" calculates the cross-correlation of inputs from the layers below. Each element or neuron in a convolutional layer processes only input from its receptive field, such as 3×3 or 7×7 pixels. This reduces the number of learnable parameters needed to describe the layer and allows for the formation of deeper neural networks than those constructed with fully connected layers, where each neuron is connected to every other neuron in the next layer. This is highly memory-dense and prone to overfitting. Convolutional layers are also spatially translation-invariant, which is very useful for processing images where the subject cannot be guaranteed to be precisely centered. Figure 5C The AI architecture also includes a max pooling layer, in Figure 5CThe term "POOL" indicates a downsampling method that selects only representative neuron weights within a given region to reduce network complexity and minimize overfitting. For example, for weights within a 4×4 square region of a convolutional layer, the maximum value is calculated for each 2×2 corner block, and then these representative values are used to reduce the dimensionality of the square region to 2×2. The architecture may also include the use of calibrated linear units as non-linear activation functions. As a common example, the ramp function takes the form of the input x from a given neuron, similar to neuron activation in biology:
[0179] f(x)=max(0,x) (2)
[0180] After the input passes through all convolutional layers, the last layer at the end of the network is typically a fully connected (FC) layer, which acts as a classifier. This layer takes the final input and outputs an array with the same dimension as the classification categories. For two categories, such as "aneuploid exists" and "aneuploid does not exist," the last layer will output an array of length 2, representing the proportion of features in the input image that are aligned to each category, respectively. A final softmax layer is usually added, which converts the final numbers in the output array into percentages between 0 and 1, which add up to 1. Therefore, the final output can be interpreted as a confidence limit for the image to be classified into one of the categories.
[0181] A suitable DNN architecture is ResNet (and its variants; see https: / / ieeexplore.ieee.org / document / 7780459), such as ResNet152, ResNet101, ResNet50, or ResNet-18. In 2016, ResNet significantly advanced the field by using a large number of hidden layers and introducing "skipped connections" (also known as "residual connections"). Only the differences from one layer to the next are computed, which is more time-efficient, and if a small change is detected in a particular layer, that layer is skipped, thus creating a network that can very quickly adapt itself to combinations of features of varying sizes in an image.
[0182] Another suitable DNN architecture is the DenseNet variant (https: / / ieeexplore.ieee.org / document / 8099726), including DenseNet161, DenseNet201, DenseNet169, and DenseNet121. DenseNet is an extension of ResNet, now allowing each layer to skip to any other layer, with the maximum number of skip connections. This architecture requires more memory and is therefore less efficient, but can exhibit better performance than ResNet. Due to the large number of model parameters, it is also prone to overtraining / overfitting. All model architectures are typically combined with methods to control this.
[0183] Another suitable DNN architecture is Inception (-ResNet) (https: / / www.aaai.org / ocs / index.php / AAAI / AAAI17 / paper / viewPaper / 14806), such as InceptionV4 and InceptionResNetV2. Inception represents a more complex convolutional unit; therefore, instead of simply using the fixed-size filters described in Section 3.2 (e.g., 3×3 pixels), it computes several filters of different sizes in parallel: (5×5, 3×3, 1×1 pixels), with weights that are free parameters. Thus, the neural network can preferentially select the most suitable filter at each layer of the DNN. An extension would be to combine this with skip connections in the same way as ResNet to create Inception-ResNet.
[0184] As mentioned above, computer vision and deep learning methods are trained on preprocessed data using multiple training and validation cycles. The training and validation cycles follow this framework:
[0185] The training data is preprocessed and divided into multiple batches (the number of data points in each batch is a free model parameter, but it controls the speed and stability of the algorithm's learning). Augmentation can be performed before batching or during training.
[0186] After each batch, the network weights are adjusted, and the overall accuracy of the run so far is evaluated. In some embodiments, gradient accumulation is used, for example, to update the weights during batch processing. One epoch is performed when all images have been evaluated, the training set is shuffled (i.e., a new randomized result is obtained for the dataset), and training restarts from the top for the next epoch.
[0187] Depending on the size and complexity of the dataset and the model being trained, multiple epochs may be run during training. The optimal number of epochs is typically between 2 and 100, but can be more depending on the specific circumstances. After each epoch, the model is run on a validation set without any further training to provide a measure of progress in model accuracy and to guide the user on whether more epochs should be run or whether more epochs would lead to overtraining.
[0188] The validation set guides the selection of ensemble model parameters or hyperparameters, and is therefore not a true blind set. However, it is important that the image distribution of the validation set is very similar to that of the final blind test set run after training.
[0189] When reporting validation set results, each image can include augmentations (all) or no augmentations (noaug). Furthermore, augmentations for each image can be combined to provide a more robust final result. Several combination / voting strategies can be used, including: average confidence (average of the inferences from all augmented models), median confidence, majority average confidence (average confidence of the majority viability assessments, providing only the average confidence of those who agree, or the average if there is no majority), maximum confidence, weighted average, majority maximum confidence, etc.
[0190] Another approach used in machine learning is transfer learning, where a previously trained model is used as a starting point for training a new model. This is also known as pre-training. Pre-training is widely used, allowing for the rapid construction of new models. There are two types of pre-training. One example of pre-training is ImageNet pre-training. Most model architectures have a set of weights pre-trained using the standard image database ImageNet. While it is not specific to medical images and contains a thousand different types of objects, it provides the model with a way of recognizing shapes. The classifier for those thousand objects is then completely removed, and a new viability classifier is replaced. This type of pre-training is superior to other initialization strategies. Another example of pre-training is custom pre-training, which uses a previously trained embryo model derived from studies with different result sets, or studies of different images (PGS, not viability, or randomly assigned results). These models provide only a small benefit for classification.
[0191] For non-pre-trained models, or new layers added after pre-training such as classifiers, weights need to be initialized. The initialization method can affect the success of training. For example, all weights set to 0 or 1 will perform poorly. Uniform permutations of random numbers, or Gaussian distributions of random numbers, are also common options. They are often used in conjunction with normalization methods such as the Xavier or Kaiming algorithms. This addresses the problem that nodes in a neural network may be "stuck" in a certain state, i.e., saturated (close to 1) or dead (close to 0), making it difficult to measure which direction the weights associated with that particular neuron will adjust. This is especially common when introducing hyperbolic tangent or sigmoid functions, and Xavier initialization solves this problem.
[0192] In the Xavier initialization protocol, the randomization of neural network weights should ensure that the inputs to each layer of the activation function do not approach extremes such as saturation or death. However, ReLU performs better, and the benefits offered by different initializations, such as Kaiming initialization, are smaller. Kaiming initialization is more suitable for situations where ReLU is used as the non-linear activation distribution of neurons. This effectively achieves the same process as Xavier initialization.
[0193] In deep learning, a set of free parameters are used to optimize model training on a validation set. One key parameter is the learning rate, which determines how much the weights of the underlying neurons are adjusted after each batch. When training a chosen model, overtraining or overfitting the data should be avoided. This occurs when a model contains too many parameters that it cannot fit and essentially "memorizes" the data, trading generalization ability for accuracy on the training or validation set. This should be avoided because generalization ability is a true measure of whether a model correctly identifies the real underlying parameters indicative of embryonic health amidst data noise, and it should not be sacrificed for the sake of perfectly fitting the training set.
[0194] During the validation and testing phases, the success rate can sometimes drop abruptly due to overfitting during training. This can be improved through various strategies, including slowing down or decaying the learning rate (e.g., halving the learning rate every n epochs) or using cosine annealing, combined with the tensor initialization or pre-training methods mentioned above, and adding noise, such as dropout layers or batch normalization. Batch normalization is used to counteract vanishing or exploding gradients, thereby improving the stability of training large models and thus improving generalization. Dropout regularization effectively simplifies the network by introducing a random chance to set all input weights within the corrector's receiving range to zero. By introducing noise, it effectively ensures that the remaining corrector correctly fits the representation of the data without relying on over-specialization. This allows DNNs to generalize more effectively and become less sensitive to specific values of network weights. Similarly, batch normalization improves the training stability of very deep neural networks by shifting input weights to zero mean and unit variance as a precursor to the correction phase, enabling faster learning and better generalization.
[0195] When performing deep learning, methods for altering neuron weights to achieve acceptable classification include specifying an optimization protocol. That is, for a given definition of "accuracy" or "loss" (discussed below), many techniques need to be specified regarding how much weight should be adjusted and how the learning rate should be used. Suitable optimization techniques include: stochastic gradient descent (SGD) with momentum (and / or Nesterov accelerated gradients), adaptive gradient with delta (Adadelta), adaptive moment estimation (Adam), root mean square propagation (RMSProp), and the limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) algorithm. Among these, SGD-based techniques generally outperform other optimization techniques. For example, the learning rate for training an AI model on phase-contrast microscopic images of human embryos is between 0.01 and 0.0001. However, this is just an example; the learning rate will depend on the batch size, which in turn depends on hardware capacity. For example, the larger the GPU, the larger the batch size, and the higher the learning rate.
[0196] Stochastic gradient descent (SGD) with momentum (and / or Nesterov accelerated gradient) represents the simplest and most commonly used optimizer. Gradient descent algorithms typically compute the gradient (slope) of the effect of a given weight on accuracy. This is slow if the gradient needs to be computed across the entire dataset to perform weight updates, whereas stochastic gradient descent performs updates one at a time on each training image. While this can lead to fluctuations in overall target accuracy or loss, it generalizes more easily than other methods because it can jump into new regions of the loss parameter landscape and find new minimum loss functions. SGD performs well in noisy loss landscapes such as embryo selection. SGD can struggle when navigating asymmetric loss function surfaces, as one side of the surface is steeper than the other, which can be compensated for by adding a parameter called momentum. This helps to accelerate SGD in direction and suppress high fluctuations in accuracy by adding extra scoring to the weight updates derived from previous states. An extension of the method also includes the estimated position of the weights in the next state, called the Nesterov accelerated gradient.
[0197] Incremental adaptive gradient (Adadelta) is an algorithm that adapts the learning rate to the weights themselves, performing smaller updates for frequently occurring parameters and larger updates for infrequent features, making it well-suited for sparse data. While this can cause the learning rate to suddenly decrease after several epochs across the entire dataset, adding an incremental parameter limits the window allowed by accumulated past gradients to a fixed size. However, this process renders the default learning rate redundant, and the degrees of freedom of the added parameters provide some control in finding the optimal overall selection model.
[0198] Adaptive Moment Estimation (Adam) preserves the exponentially decaying averages of past squared and non-squared gradients and incorporates them into weight updates. This provides a "friction" effect for the direction of weight updates and is suitable for problems with relatively shallow or flat loss minimization and no strong fluctuations. In embryo selection models, training using Adam tends to perform well within the training set but is often overtrained and less suitable than SGD with momentum.
[0199] Root mean square propagation (RMSProp) is related to the adaptive gradient optimizer mentioned above and is almost identical to Adelta, except that the weight update term divides the learning rate by the exponentially decaying average of the squared gradient.
[0200] The limited-memory Broyden-Fletcher Goldfarb-Shanno (L-BFGS) algorithm. Although computationally intensive, the L-BFGS algorithm (rather than other methods) for practically estimating the curvature of the lost landscape does not attempt to compensate for the estimation deficiencies with additional terms. It tends to outperform Adam when the dataset is small, but it is not necessarily superior to SGD in terms of speed and accuracy.
[0201] In addition to the methods described above, non-uniform learning rates can also be included. That is, the learning rate of the convolutional layer can be specified to be much larger or smaller than that of the classifier. This is useful in the case of pre-trained models, where changes to the filters below the classifier should remain more "frozen," and the classifier should be retrained so that the pre-training is not canceled out by the additional retraining.
[0202] When the optimizer specifies how to update the weights for a given particular loss or accuracy metric, in some implementations, the loss function is modified to incorporate distribution effects. These may include cross-entropy (CE) loss, weighted CE, residual CE, inferred distribution, or a custom loss function.
[0203] Cross-entropy loss is a commonly used loss function that tends to outperform the simple mean squared error between the true and predicted values. If the network's output passes through a softmax layer, as is the case here, the distribution of cross-entropy leads to better accuracy. This is because it naturally maximizes the probability of correctly classifying the input data by passing through distant outliers without overweighting them. For an input array `batch` representing a batch of images and `class` representing live or inactive images, cross-entropy loss is defined as:
[0204]
[0205] Where C is the number of classes. In the binary case, it can be simplified to:
[0206] loss(p,C)=-(ylog(p))+(1-y)log(1-p) (4)
[0207] An optimized version is:
[0208]
[0209] If the data contains class bias, i.e., more active examples than inactive examples (or vice versa), then the loss function should be weighted proportionally so that misclassification of elements in the smaller class is penalized more severely. This is achieved by pre-multiplying the right-hand side of equation (2) by a coefficient:
[0210]
[0211] Where N[class] is the total number of images in each class, N is the total number of samples in the dataset, and C is the number of classes. If necessary, weights can be manually biased towards viable embryos to reduce the number of false negatives relative to false positives.
[0212] In some embodiments, an inferred distribution may be used. While seeking a high level of accuracy in embryo classification is important, seeking a high level of transferability in the model is also important. That is, understanding the distribution of scores is often beneficial; while seeking high accuracy is an important goal, confidently separating viable and non-viable embryos is an indicator that the model generalizes well to the test set. Since accuracy on the test set is frequently cited for comparison with important clinical benchmarks (e.g., the accuracy of embryologists in classifying the same embryos), ensuring generalization ability should also be incorporated into batch-wise evaluations of model success at each epoch.
[0213] In some embodiments, a custom loss function is used. In one embodiment, we customize how the loss function is defined to alter the optimization surface to make the global minimum more apparent, thereby improving the robustness of the model. To achieve this, a new term called the residual term is added to the differentiable loss function, which is defined based on the network weights. It encodes the collective difference between the model's predictions and the target results for each image and contributes to the normal cross-entropy loss function as an additional contribution. For N images, the residual term is formulated as follows:
[0214]
[0215] For this custom loss function, clusters with appropriately spaced scores for viable and non-viable embryos are therefore considered consistent with higher loss ratings. It's important to note that this custom loss function is not specific to embryo detection applications and can be used with other deep learning models.
[0216] In some embodiments, a loss function based on a custom confidence level is used. This is a weighted loss function with two variations: linear and nonlinear. For both cases, the goal is to encode the separation of scores as a contribution to the loss function, but in a different way than described above, by integrating the differences between classes in the predicted scores as a weighting function of the loss. The greater the difference, the more the loss is reduced. This loss function helps drive the predictive model to amplify the differences between the two classes and increase the model's confidence in the outcome. For the confidence weights: the binary target label of the i-th input sample, denoted as y∈{±1}, specifies the true class. Assume the result of the predictive model is y p =[y p0,y p1 ], y p0 ,y p1 ∈[0,1] represents the probability output of the model's estimation results, corresponding to the inputs of inactive and active results, respectively.
[0217] For a linear setup, define d = |y p0 -y p1 |; For non-linear settings, define The parameter d represents the probability difference between the model's predictions for class 0 and class 1. For the standard log softmax function, we define p... t The following (log(pt) will be included in the standard cross-entropy loss function):
[0218]
[0219] For class weights: For class 1, the weight factor ∝ ∈ [0, 1], and for class -1, the weight factor 1 - ∝. We define p with respect to... t Defined in a similar way ∝ t :
[0220]
[0221] The focus parameter γ smoothly adjusts the rate at which the difference in outcome scores affects the loss function. Finally, we propose a loss function that incorporates all three different weighting strategies:
[0222] LF=-∝ t (1-exp(d)) γ log(p t (10)
[0223] In some embodiments, a soft loss function is used, employing a technique called label smoothing. For each type of outcome or category (e.g., in a binary classification problem: live, inactive), any or all categories can exhibit label smoothing. Label smoothing is introduced to create a soft loss function, such as a weighted cross-entropy loss. Then, when computing the loss function, if any category includes label smoothing, the Kullback-Leibler (KL)-Divergence loss is computed between the inputs of the loss function; that is, the score distribution of the current batch, and a modified version of the score distribution where each category showing label smoothing has had its score changed by the number of points e / (number of categories - 1) from its actual value (e.g., 0 or 1). Therefore, this parameter e is a free parameter that controls the amount of label smoothing introduced. This KL Divergence loss is then returned as the loss function.
[0224] In some embodiments, models are combined to generate a more robust final AI model 100. That is, deep learning and / or computer vision models are combined to help predict aneuploidy in general.
[0225] In one embodiment, an ensemble approach is used. First, a well-performing model is selected. Then, each model “votes” for one of the images (using augmentations or other methods), and the voting strategy that leads to the best result is selected. Examples of voting strategies include maximum confidence, mean, majority mean, median, mean confidence, median confidence, majority mean confidence, weighted mean, majority maximum confidence, etc. Once a voting strategy is selected, an evaluation method for the combination of augmentations must also be chosen, which describes how the ensemble should handle each rotation, as previously described. In this embodiment, the final AI model 100 can therefore be defined as a collection of trained AI models using deep learning and / or computer vision models, along with a pattern that encodes the voting strategy defining how the results of the individual AI models are combined, and an evaluation pattern that defines how augmentations (if present) are combined.
[0226] Model selection should be conducted in such a way that their results are mutually comparative, i.e., their results are as independent as possible and their scores are evenly distributed. This selection process is performed by examining which images are correctly identified within the test set for each model. If the sets of correctly identified images are very similar when comparing two models, or if each model provides similar scores for a given image, these models are not considered comparative models. However, if there is little overlap between the two sets of correctly identified images, or if the scores provided for each image are significantly different, these models are considered comparative models. This process effectively assesses whether the distributions of embryo scores on the test set are similar for two different models. Due to differences in input images or segmentation, the comparison criteria drive model selection with different distributions of predicted results. This method avoids selecting models that perform well only on specific clinical datasets, thus preventing overfitting and ensuring translatability. Moreover, model selection can also use diversity criteria. Diversity criteria encourage model selection to include different hyperparameters and configurations of models. The reason is that, in practice, similar model settings lead to similar predictive results and may therefore be useless for the final ensemble model.
[0227] In one embodiment, this can be achieved by using a counting method and specifying a similarity threshold (e.g., 50%, 75%, or 90% overlap of images in two sets). In other embodiments, scores in an image set (e.g., a live set) can be added together and the two sets (totals) compared, and if the two totals are less than a threshold amount, the sets are ranked as similar. Statistical comparisons can also be used, such as considering the number of images in the sets, or otherwise comparing the distribution of images in each set.
[0228] Another approach in AI and machine learning is called "knowledge distillation" (or simply distillation) or "student-teacher" models, where the distribution of weight parameters obtained from one (or more) models (teachers) is used, via the loss function of the student model, to inform the weight updates of another model (student). We will use the term distillation to describe the process of training a student model using a teacher model. The idea behind this process is to train the student model to mimic a set of teacher models. The intuition behind this process is that the teacher models contain subtle but important relationships between predicted output probabilities (soft labels) that are not present in the raw predicted probabilities (hard labels) obtained directly from the model results without the distribution from the teacher models.
[0229] First, a set of teacher models are trained on the dataset of interest. Teacher models can be any neural network or model architecture, even architectures completely different from each other or completely different from the student models. They can share the exact same dataset, have no overlap, or have overlapping subsets of the original dataset. Once these teacher models are trained, the student will use a distillation loss function to model the outputs of these teacher models. During distillation, the teacher models are first applied to a dataset available to both the teacher and student models (called the transfer dataset). The transfer dataset can be a preserved blind dataset extracted from the original dataset or the original dataset itself. Furthermore, the transfer dataset does not need to be fully labeled; that is, some parts of the data are irrelevant to the known results. This removal of label restrictions allows for an artificial increase in the size of the dataset. The student models are then applied to the transfer dataset. The output probabilities (soft labels) of the teacher models are compared to the output probabilities of the student models calculated from the distribution using a divergence metric function (e.g., KL-Divergence) or a "relative entropy" function. The divergence metric is a recognized mathematical method for measuring the "distance" between two probability distributions. The divergence metric is then added to the standard cross-entropy classification loss function, thereby effectively minimizing both the classification loss and the discrepancy between the student and teacher models, thus improving model performance. Typically, the soft label matching loss (the divergent component of the new loss) and the hard label classification loss (the original component of the loss) are weighted together (introducing an additional adjustable parameter into the training process) to control the contribution of each of these two components to the new loss function.
[0230] A model can be defined by its network weights. This may involve using appropriate functions of the machine learning code / API to export or save checkpoint files or model files. Checkpoint files can be files generated by the machine learning code / library with a defined format, which can be exported and then read back (reloaded) as part of the machine learning code / API using provided standard functions (e.g., ModelCheckpoint() and load_weights()). The file format can be sent or copied directly (e.g., FTP or similar protocols), or it can be serialized and sent using JSON, YAML, or similar data transfer protocols. In some embodiments, additional model metadata (such as model accuracy, epochs, etc.) can be exported / saved and sent along with the network weights, which can further characterize the model or otherwise help build another model on another node / server (e.g., a student model).
[0231] Embodiments of this method can be used to generate AI models that obtain estimates regarding the presence of one or more aneuploidies in embryonic images. These AI models can be implemented in a cloud-based computing system that computationally generates an aneuploidy screening AI model. Once the model is generated, it can be deployed in the cloud-based computing system to computationally generate estimates regarding the presence of one or more aneuploidies in the embryonic image. In this system, the cloud-based computing system includes the previously generated (trained) aneuploidy screening AI model and is configured to receive images provided to the aneuploidy screening AI model from a user via a user interface to obtain estimates regarding the presence of one or more aneuploidies in the image. A report regarding the presence of one or more aneuploidies in the image is sent to the user via the user interface. Similarly, a computing system that generates estimates regarding the presence of one or more aneuploidies in the embryonic image can be provided at the clinic or similar location where the image was obtained. In this embodiment, the computing system includes at least one processor and at least one memory, the memory including instructions for configuring the processor to perform the following operations: receiving images captured during a predetermined time window following in vitro fertilization (IVF), and uploading the images captured during the predetermined time window following IVF to a cloud-based artificial intelligence (AI) model via a user interface, the AI model being used to generate an estimate regarding the presence of one or more aneuploidies in the embryo images. The estimate regarding the presence of one or more aneuploidies in the embryo images is received via the user interface and displayed by the user interface.
[0232] result
[0233] The results demonstrating that the AI model can isolate morphological features corresponding to specific chromosomes or chromosome sets purely from phase-contrast microscopy images are shown below. According to Table 1, this includes a series of example studies for several of the most serious chromosomal defects (i.e., high risk of adverse post-implantation outcomes). In the first three cases, a simple example is constructed to illustrate the existence of morphological features corresponding to specific chromosomal abnormalities. This is done by including only the affected chromosomes and viable euploid embryos. These simplified examples provide evidence that it is feasible to generate a general model based on combining individual models, where each individual model focuses on a different chromosomal defect / genetic defect. Another example using chromosome sets containing aneuploid chromosomes listed in Table 1 is also generated.
[0234] The first study was conducted to evaluate whether the AI model could detect differences between euploid, viable embryos and embryos containing any abnormalities of chromosome 21 associated with Down syndrome, including mosaic embryos. The model, trained on a blind dataset of 214 images, achieved an overall accuracy of 71.0%.
[0235] Before training the AI model, a blind test set containing characteristics of all chromosomes considered to pose a serious health risk if present in images of aneuploidy (from Table 1) and viable euploidy was retained so that it could be used as a general test set for training the model. The total number of images involved in the study is shown in Table 2.
[0236] Table 2 Breakdown of the dataset used for chromosome 21 research (1322 images)
[0237] Dataset Total number of images Number of images of aneuploid chromosome 21 Number of euploid vitality images Training (80%) 887 417 470 Verification (10%) 221 105 116 Test (10%) 214 68 146
[0238] The accuracy results on the test set are as follows:
[0239] • Embryos with any abnormality on chromosome 21: 76.47% (52 / 68 were correctly identified);
[0240] • Viable euploid embryos: 68.49% (100 / 146 were correctly identified).
[0241] The distributions of viable aneuploid and viable euploid embryos are as follows: Figure 6A and 6B As shown, the aneuploidy distribution 600 displays a small set 610 (left-hand diagonal slant bar) of chromosome 21 abnormal embryos that were incorrectly identified (missed) by the AI model, and a large set 620 (right-hand diagonal backslant bar) of chromosome 21 abnormal embryos that were correctly identified by the AI model. Similarly, the euploidy distribution 630 displays a small set 640 (left-hand diagonal slant bar) of chromosome 21 normal embryos that were incorrectly identified (missed) by the AI model, and a large set 650 (right-hand diagonal backslant bar) of chromosome 21 normal embryos that were correctly identified by the AI model. In both figures, the aneuploid and euploid embryo images are well separated, and the ploidy status scores provided by the AI model clearly cluster them.
[0242] The same method was used on chromosome 16, whose modifications are associated with autism, as in the study of chromosome 21. The total number of images involved in the study is shown in Table 3.
[0243] Table 3 Breakdown of the dataset used for chromosome 16 research (1058 images).
[0244] Dataset Total number of images Number of images of aneuploid chromosome 16 Euploid vitality image number Training (80%) 692 339 353 Verification (10%) 173 85 88 Test (10%) 193 47 146
[0245] The accuracy results on the test set are as follows:
[0246] • Embryos with any abnormality on chromosome 16: 70.21% (33 / 47 were correctly identified);
[0247] • Viable euploid embryos: 73.97% (108 / 146 were correctly identified).
[0248] The distributions of viable aneuploid and viable euploid embryos are as follows: Figure 7A and 7B As shown. The aneuploidy distribution 700 shows a small set 710 (left-hand diagonal slash filled bar) of chromosome 16 abnormal embryos that were incorrectly identified (omitted) by the AI model, and a large set 750 (right-hand diagonal backslash filled bar) of chromosome 16 abnormal embryos that were correctly identified by the AI model. Similarly, the euploidy distribution 730 shows a small set 740 (left-hand diagonal slash filled bar) of chromosome 16 normal embryos that were incorrectly identified (omitted) by the AI model, and a large set 750 (right-hand diagonal backslash filled bar) of chromosome 16 normal embryos that were correctly identified by the AI model.
[0249] As a third case study, this method was repeated on chromosome 13, which is associated with Patau syndrome. The total number of images involved in the study is shown in Table 4.
[0250] Table 4 shows the breakdown of the dataset used for the study of chromosome 13 (794 images).
[0251] Dataset Total number of images Number of images of aneuploid chromosome 13 Number of euploid vitality images Training (80%) 624 282 342 Verification (10%) 170 71 99 Test (10%) 193 44 149
[0252] The accuracy results are as follows:
[0253] • Embryos with any abnormality on chromosome 13: 54.55% (24 / 44 were correctly identified);
[0254] • Viable euploid embryos: 69.13% (103 / 149 were correctly identified).
[0255] While the accuracy for this specific chromosome is lower than for chromosomes 21 and 16, for a given dataset size, different chromosomes are expected to have different levels of confidence at which images corresponding to their specific associated aneuploidies can be identified. That is, each genetic abnormality will exhibit different visible features, so some abnormalities are expected to be easier to detect than others. However, as with most machine learning systems, increasing the size and diversity of the training dataset is expected to maximize the model's ability to detect the presence of specific chromosomal abnormalities. Therefore, combined approaches that can simultaneously and separately evaluate multiple aneuploidies can provide a useful picture of genetic abnormalities associated with the embryo, with different levels of confidence depending on the rarity of the cases included in the training dataset.
[0256] As the fourth case study, this method was used for chromosome analysis, which included viable euploid embryos and chromosome sets considered “severe,” including chromosomes 13, 14, 16, 18, 21, and 45,X (according to Table 1). For the purposes of this example, mosaicism and non-mosaicism were included together, and all types of chromosomal alterations were included together. The total number of images involved in the study is shown in Table 5.
[0257] Table 5 Breakdown of the dataset used for “severe” chromosome studies (853 images).
[0258] Dataset Total number of images Number of images of severely aneuploid chromosomes Euploid vitality image number Training (80%) 563 343 220 Verification (10%) 140 86 54 Test (10%) 150 91 59
[0259] The accuracy results are as follows:
[0260] • Severe chromosomal abnormalities were found in 54.95% of embryos (50 / 91 were correctly identified);
[0261] • Viable euploid embryos: 64.41% (38 / 59 were correctly identified).
[0262] The distributions of viable aneuploid and viable euploid embryos are as follows: Figure 8A and 8B As shown. The aneuploidy distribution 800 shows a small set 810 of aneuploid / abnormal embryos belonging to the severe chromosomal group that were incorrectly identified (omitted) by the AI model (the diagonal slant-filled bar on the left), and a large set 820 of aneuploid / abnormal embryos belonging to the severe chromosomal group that were correctly identified by the AI model (the diagonal backslant-filled bar on the right). Similarly, the euploidy distribution 830 shows a small set 840 of normal euploid embryos that were incorrectly identified (omitted) by the AI model on the left, and a large set 850 of normal euploid embryos that were correctly identified by the AI model (the diagonal backslant-filled bar on the right).
[0263] While the accuracy for this set of chromosomes is lower than for a single chromosome, it is expected that, for a given dataset size, it will be possible to identify grouped chromosomes or specific combinations based on morphology that have similar severity corresponding to their specific associated aneuploidy. That is, each genetic abnormality will exhibit different visible features, so some abnormalities are expected to be easier to detect than others. However, as with most machine learning systems, increasing the size and diversity of the training dataset is expected to maximize the model's ability to detect the presence of specific chromosomal abnormalities. Therefore, methods capable of simultaneously and separately evaluating combinations of multiple aneuploidies can provide a useful picture of genetic abnormalities associated with the embryo, with varying levels of confidence depending on the rarity of the cases included in the training dataset.
[0264] These four studies demonstrate that AI / machine learning and computer vision technologies can identify morphological features associated with abnormalities in chromosomes 21, 16, and 13, as well as combined chromosome sets.
[0265] Each AI model was able to detect morphological features associated with certain severe chromosomal abnormalities with a certain level of confidence. The score histograms associated with ploidy status provided by the selection model showed a reasonable separation between images of euploid and aneuploid embryos.
[0266] The morphological features associated with chromosomal abnormalities can be subtle and complex, making it challenging to effectively detect these patterns by training on small datasets. While this study does demonstrate a strong correlation between embryonic morphology in images and chromosomal abnormalities, higher accuracy is expected when training AI models on larger, more diverse datasets.
[0267] These studies demonstrate the feasibility of constructing a general aneuploidy assessment model by combining different models focusing on various chromosomal abnormalities. This more general aneuploidy assessment model can encompass a wider range of chromosomal abnormalities, including both severe and mild ones (as described in Table 1, or as determined by clinical practice). In other words, unlike previous systems that typically simply grouped all aneuploidies (and mosaics) together to give a presence / absence result, this system improves performance by decomposing the problem into independent sets of chromosomes and training individual models on each set separately, then combining these models to detect a diverse range of chromosomal abnormalities. Decomposing the problem into smaller sets of chromosomes and then training multiple different models, each trained in a different manner or with different configurations or architectures (e.g., hierarchical, binary, multi-class, multi-group), produces a diversity of models, each effectively solving a different optimization problem and thus generating different results for the input image. This diversity then allows for the selection of the optimal model. Furthermore, this approach aims to identify mosaicisms that are currently undetectable by invasive screening methods. In an in vitro fertilization cycle, embryos are a precious and limited resource. Current success rates (for viable pregnancies) are low, and the financial and emotional costs of performing multiple cycles are high. Therefore, an improved, non-invasive aneuploidy assessment tool, based on defined chromosome sets (e.g., those based on the severity of adverse outcomes), provides clinicians and patients with more detailed and informative results. This allows for more informed decision-making, especially in challenging situations where all available embryos (for the current cycle) exhibit aneuploidy or mosaicism, enabling clinicians and patients to balance potential risks and make more informed choices regarding which embryo to implant.
[0268] Several embodiments are discussed, including hierarchical and binary models, as well as single-set or multi-set models. Specifically, hierarchical and binary models can be used to train AI models by assigning quality labels to embryo images. In this embodiment, a hierarchical sequence of hierarchical models is generated, and a separate hierarchical and binary model can be generated for each chromosome set. In each layer, images are divided according to quality, with the highest-quality image used to train the model at that layer. That is, at each layer, the training set is divided into the highest-quality image and other images. The model at that layer is trained on the highest-quality image, and other images are passed to the next layer and the process is repeated (thus, the remaining images are divided into the next highest-quality image and other images). The models in the hierarchical and binary models can be binary models, multi-class models, or a combination of binary and multi-class models at all layers. Furthermore, this hierarchical training method can also be used to train multi-set models. The basic principle behind the hierarchical and binary model approach is that embryo images considered high-quality are likely to have the highest quality morphological features (i.e., "look like the best embryo") among the images with the fewest abnormalities, and therefore have the greatest morphological difference (i.e., "look bad or have abnormal features") compared to embryo images containing chromosomal defects. Therefore, this enables AI algorithms to better detect and predict morphological features between these two (extreme) image classifications. This process can be repeated multiple times with different numbers of layers / quality labels to generate a set of hierarchical models. Multiple independent hierarchical models are generated for each chromosome set, and the best hierarchical model can be selected from this set. This can be based on quality metrics, or it can use ensemble or distillation techniques.
[0269] In some embodiments, a set of binary models can be generated for each chromosome set, or one or more multi-set models classifying all chromosome sets (or at least multiple chromosome sets). Multiple binary models with different sets, multiple multi-class models, and multiple-set models (including multiple hierarchical multi-set models) can be generated. These provide greater diversity to the AI model. Once a set of candidate models is generated, they can be used to generate the final AI model to identify each chromosome set in an image. This can be further refined or generated using ensembles, distillation, or other similar methods to train the final single model based on multiple models. Once the final model is selected, it can be used to classify new images during IVF, thereby helping to select embryos (or multiple embryos) for implantation, for example, by identifying and excluding high-risk embryos or by identifying embryos with the lowest risk of aneuploidy.
[0270] Therefore, methods developed in conjunction with chromosomal abnormality research can be used to characterize embryo images prior to preimplantation genetic diagnosis (PGD), serving as a pre-screening tool or providing a suite of advanced genetic analyses to support clinics that cannot utilize readily available PGD technologies. For example, if images indicate a high probability / confidence of an unfavorable chromosomal abnormality, the embryo can be discarded so that only those deemed low-risk can be implanted, or further examination can be conducted using invasive (and also higher-risk) PGD techniques.
[0271] Those skilled in the art will understand that information and signals can be represented using any of a variety of technologies. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be mentioned throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or particles, light fields or particles, or any combination thereof.
[0272] Those skilled in the art will further understand that the various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software or instructions, middleware, a platform, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally according to their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in various ways for each specific application, but these determined implementations should not be construed as departing from the scope of the invention.
[0273] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be directly embodied in hardware, software modules executed by a processor, or a combination of both, including cloud-based systems. For hardware implementation, processing can be implemented within one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, other electronic units designed to perform the functions described herein, or combinations thereof. Various middleware and computing platforms can be used.
[0274] In some embodiments, the processor module includes one or more central processing units (CPUs) or graphics processing units (GPUs) for performing some steps of the method. Similarly, a computing device may include one or more CPUs and / or GPUs. The CPU may include an input / output interface, an arithmetic and logic unit (ALU), and control units and a program counter element that communicate with input and output devices via the input / output interface. The input / output interface may include a network interface and / or a communication module for communicating with an equivalent communication module in another device using predefined communication protocols (e.g., IEEE 802.11, IEEE 802.15, TCP / IP, UDP, etc.). The computing device may include a single CPU (core), multiple CPUs (multiple cores), or multiple processors. The computing device is typically a cloud-based computing device using GPU clusters, but may be a parallel processor, vector processor, or distributed computing device. Memory is operatively connected to the processor and may include RAM and ROM components, and may be located internally or externally to the device or processor module. Memory may be used to store the operating system and additional software modules or instructions. The processor may be used to load and execute the software modules or instructions stored in the memory.
[0275] A software module, also known as a computer program, computer code, or instructions, may contain multiple source code or object code segments or instructions and may reside in any computer-readable medium, such as RAM memory, flash memory, ROM memory, EPROM memory, registers, hard disk, removable disk, CD-ROM, DVD-ROM, Blu-ray disc, or any other form of computer-readable medium. In some aspects, a computer-readable medium may include non-transitory computer-readable media (e.g., tangible media). In other aspects, a computer-readable medium may include transient computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media. In another aspect, a computer-readable medium may be integrated into a processor. The processor and the computer-readable medium may reside in an ASIC or related device. Software code may be stored in memory cells, and the processor may use it to execute them. Memory cells may be implemented inside or outside the processor, in which case they may be communicatively connected to the processor by various means known in the art.
[0276] Furthermore, it should be understood that modules and / or other suitable means for performing the methods and techniques described herein can be downloaded and / or otherwise obtained by a computing device. For example, such a device can be connected to a server to facilitate the transfer of means for performing the methods described herein. Alternatively, the various methods described herein can be provided via storage devices (e.g., RAM, ROM, physical storage media such as optical discs (CDs) or floppy disks, etc.) so that the computing device can access the various methods when the storage device is connected or provided to it. Moreover, any other suitable techniques for providing the methods and techniques described herein to the device may be used.
[0277] The methods disclosed herein include one or more steps or actions for implementing the described methods. The method steps and / or actions may be interchanged with each other without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of a particular step and / or action may be modified without departing from the scope of the claims.
[0278] Throughout this specification and the appended claims, unless the context otherwise requires, the terms "comprising," "including," and variations thereof shall be construed as implying the inclusion of the stated integers or sets of integers, but not excluding any other integers or sets of integers. Any reference to prior art in this specification is not, and should not be construed as, an acceptance of any form of prior art being part of common common sense.
[0279] Those skilled in the art will understand that the purpose of this invention is not limited to the one or more specific applications described herein. The invention is also not limited to its preferred embodiments regarding the specific elements and / or features described or depicted herein. It should be understood that the invention is not limited to the one or more disclosed embodiments, but is capable of various rearrangements, modifications, and substitutions without departing from the scope set forth and defined by the appended claims.
Claims
1. A method for computationally generating an artificial intelligence (AI) model for aneuploidy screening, the AI model being used to screen for the presence of aneuploidy in embryo images, the method comprising: Define multiple chromosome set tags, where each set includes one or more different aneuploids, which contain different gene alterations or chromosomal abnormalities; A training dataset is generated from a first set of images, wherein each image includes an image of an embryo captured after in vitro fertilization and is labeled with one or more chromosome set tags, each tag indicating the presence of at least one aneuploidy associated with the corresponding chromosome set in at least one cell of the embryo, and the training dataset includes images labeled with each chromosome set; A test dataset is generated from a second set of images, wherein each image includes an image of an embryo captured after in vitro fertilization and is labeled with one or more chromosome set tags, each tag indicating the presence of at least one aneuploidy associated with the corresponding chromosome set, and the test dataset includes images labeled with each chromosome set; At least one chromosome group AI model is trained for each chromosome group using the training dataset used to train all models, wherein each chromosome group AI model is trained to identify morphological features in images labeled with the relevant chromosome group, and / or at least one multi-group AI model is trained on the training data, wherein each multi-group AI model is trained to independently identify morphological features in images labeled with each relevant chromosome group, to generate a multi-group output for the input image to indicate whether at least one aneuploidy associated with each chromosome group exists in the image; Use the test dataset to select the best chromosome set AI model for each chromosome set, or a best multi-set AI model; and The selected AI model is deployed to screen for the presence of one or more aneuploidies in embryo images. The steps of training at least one chromosome set AI model and / or training at least one multi-chromosome AI model for each chromosome set include: training a hierarchical model, wherein training the hierarchical model includes: The hierarchical sequence of training hierarchical models is provided, wherein in each layer, images associated with chromosome sets are assigned a first label and trained on a second set of images, wherein the second set of images is grouped based on the highest quality level, and in each sequential layer the second set of images is a subset of the second set of images from the previous layer whose quality is lower than the highest quality of the second set of images in the previous layer.
2. The method as described in claim 1, wherein, Training a hierarchical model includes: A quality label is assigned to each image in the training dataset, wherein the quality label set comprising the quality labels includes a graded quality label set, which includes at least "viable euploid embryos", "non-viable euploid embryos", "mild aneuploid embryos" and "severe aneuploid embryos". The top-level model was trained by dividing the training dataset into a first quality dataset labeled "viable euploid embryos" and another dataset containing all other images, and the model was trained on images labeled with chromosome sets and on images within the first quality dataset. One or more intermediate layer models are trained sequentially, wherein in each intermediate layer, images with the highest quality labels are selected from another dataset to generate the next quality-level dataset, and the model is trained on images labeled with the chromosome set and images within the next quality-level dataset; and The base layer model is trained on images labeled with the chromosome set and images from another dataset from the previous layer.
3. The method as described in claim 2, wherein, After training a first baseline model for the first chromosome set, training a hierarchical model for each other chromosome set includes training the other chromosome sets on other datasets used to train the first baseline model.
4. The method of claim 1, wherein, The step of training at least one chromosome set AI model for each chromosome set further includes: Train one or more binary models for each chromosome set, including: Images in the training dataset that have labels matching the chromosome set are labeled "existence," and all other images in the training dataset are labeled "absence." A binary model is trained using the "existence" and "absence" labels to generate a binary output about the input image, indicating whether a chromosomal abnormality associated with the chromosome set exists in the image.
5. The method of claim 1, wherein, The hierarchical and layered models mentioned above are all binary models.
6. The method of claim 1, wherein, Each chromosome set also includes multiple mutually exclusive aneuploid categories, wherein the sum of the probabilities of the aneuploid categories within the chromosome set is 1, and one or more of the AI models are multi-class AI models trained to estimate the probability of each aneuploid category within the chromosome set.
7. The method of claim 6, wherein, The aneuploid categories include: "missing", "inserted", "duplication", "deleted", and "normal".
8. The method of claim 1, further comprising: Receive multiple images, each image including an embryo image taken after in vitro fertilization and one or more aneuploid results; The plurality of images are divided into a first group of images and a second group of images, and one or more chromosome group labels are assigned to each image based on one or more associated aneuploid results, wherein the first group of images and the second group of images have each chromosome group label in a similar proportion.
9. The method of claim 1, wherein, Each group includes multiple different aneuploids with similar risks of adverse outcomes.
10. The method of claim 9, wherein, The multiple chromosome set tags include at least low-risk and high-risk groups.
11. The method of claim 10, wherein, The low-risk group includes at least chromosomes 1, 3, 4, 5, 17, 19, 20 and "47,XYY", and the high-risk group includes at least chromosomes 13, 16, 21, "45,X", "47,XXY" and "47,XXX".
12. The method of claim 1, wherein, The images were captured 3 to 5 days after fertilization.
13. The method of claim 1, wherein, The relative proportion of each chromosome group in the test dataset is similar to the relative proportion of each chromosome group in the training dataset.
14. A method for computationally generating an artificial intelligence (AI) model for aneuploid screening, comprising: Define multiple chromosome set tags, where each set includes one or more different aneuploids, which contain different gene alterations or chromosomal abnormalities; A training dataset is generated from a first set of images, wherein each image includes an image of an embryo captured after in vitro fertilization and is labeled with one or more chromosome set tags, each tag indicating the presence of at least one aneuploidy associated with the corresponding chromosome set in at least one cell of the embryo, and the training dataset includes images labeled with each chromosome set; A test dataset is generated from a second set of images, wherein each image includes an image of an embryo captured after in vitro fertilization and is labeled with one or more chromosome set tags, each tag indicating the presence of at least one aneuploidy associated with the corresponding chromosome set, and the test dataset includes images labeled with each chromosome set; At least one chromosome group AI model is trained for each chromosome group using the training dataset used to train all models, wherein each chromosome group AI model is trained to identify morphological features in images labeled with the relevant chromosome group, and / or at least one multi-group AI model is trained on the training data, wherein each multi-group AI model is trained to independently identify morphological features in images labeled with each relevant chromosome group, to generate a multi-group output for the input image to indicate whether at least one aneuploidy associated with each chromosome group exists in the image; Use the test dataset to select the best chromosome set AI model for each chromosome set, or a best multi-chromosome AI model; The selected AI model is deployed to screen for the presence of one or more aneuploidies in embryo images; as well as Generate an ensemble model for each chromosome set. The steps for generating an ensemble model for each chromosome set include: Train multiple final models, each of which is based on the best chromosome set AI model for its corresponding group, and each of the multiple final models is trained on a training dataset with different initial condition sets and image rankings; and Multiple trained final models are combined according to the ensemble voting strategy.
15. A method for computationally generating an artificial intelligence (AI) model for aneuploid screening, comprising: Define multiple chromosome set tags, where each set includes one or more different aneuploids, which contain different gene alterations or chromosomal abnormalities; A training dataset is generated from a first set of images, wherein each image includes an image of an embryo captured after in vitro fertilization and is labeled with one or more chromosome set tags, each tag indicating the presence of at least one aneuploidy associated with the corresponding chromosome set in at least one cell of the embryo, and the training dataset includes images labeled with each chromosome set; A test dataset is generated from a second set of images, wherein each image includes an image of an embryo captured after in vitro fertilization and is labeled with one or more chromosome set tags, each tag indicating the presence of at least one aneuploidy associated with the corresponding chromosome set, and the test dataset includes images labeled with each chromosome set; At least one chromosome group AI model is trained for each chromosome group using the training dataset used to train all models, wherein each chromosome group AI model is trained to identify morphological features in images labeled with the relevant chromosome group, and / or at least one multi-group AI model is trained on the training data, wherein each multi-group AI model is trained to independently identify morphological features in images labeled with each relevant chromosome group, to generate a multi-group output for the input image to indicate whether at least one aneuploidy associated with each chromosome group exists in the image; Use the test dataset to select the best chromosome set AI model for each chromosome set, or a best multi-chromosome AI model; The selected AI model is deployed to screen for the presence of one or more aneuploidies in embryo images; as well as Generate a distillation model for each chromosome set. The steps for generating a distillation model for each chromosome set include: Train multiple teacher models, each of which is based on the best chromosome set AI model for its corresponding group, and each of the multiple teacher models is trained on at least a portion of a training dataset with different initial condition sets and image rankings; and The student model is trained on the training dataset using multiple trained teacher models with the distillation loss function.
16. A method for computationally generating an estimate of the presence of one or more aneuploids in an embryo image, the method comprising: In the computing system, an aneuploid screening AI model is generated according to any one of claims 1 to 15; Images of embryos captured after in vitro fertilization are received from the user via the user interface of the computing system; The image is provided to the aneuploid screening AI model to obtain an estimate of whether one or more aneuploids exist in the image. as well as The system sends a report to the user via the user interface regarding the presence of one or more aneuploids in the image.
17. A method for obtaining an estimate of the presence of one or more aneuploids in an embryo image, the method comprising: Images captured during a predetermined time window after in vitro fertilization (IVF) are uploaded via a user interface to a cloud-based artificial intelligence (AI) model, which generates an estimate of the presence of one or more aneuploids in the images, wherein the AI model is generated by the method according to any one of claims 1 to 15. The user interface receives an estimate of the presence of one or more aneuploids in the embryo image.
18. A cloud-based computing system comprising one or more computing devices, said one or more computing devices including one or more processors and one or more memories, wherein, The cloud-based computing system is used to computationally generate an artificial intelligence (AI) model for aneuploid screening, which is configured by the method according to any one of claims 1 to 15.
19. A cloud-based computing system for computationally generating estimates of the presence of one or more aneuploids in an embryo image, wherein the computing system comprises: One or more computing servers, including one or more processors and one or more memories, the memories being used to store an aneuploidy screening artificial intelligence (AI) model, which is used to generate estimates about the presence of one or more aneuploids in an embryo image, wherein the aneuploidy screening AI model is generated by the method according to any one of claims 1 to 15, and the one or more computing servers are used for: Images are received from the user via the user interface of the computing system; The image is provided to the aneuploid screening AI model to obtain an estimate of whether one or more aneuploids exist in the image. as well as The system sends a report to the user via the user interface regarding the presence of one or more aneuploids in the image.
20. A computing system for generating estimates of the presence of one or more aneuploidies in an embryo image, wherein the computing system includes at least one processor and at least one memory, the memory including instructions for causing the at least one processor to perform the following operations: Receive images captured within a predetermined time window following in vitro fertilization (IVF); The images captured within a predetermined time window after IVF are uploaded via a user interface to a cloud-based artificial intelligence (AI) model for aneuploidy screening. The AI model is used to generate an estimate of the presence of one or more aneuploids in the embryo images, wherein the AI model is generated by the method according to any one of claims 1 to 15. The user interface receives an estimate of the presence of one or more aneuploids in the embryo image. as well as The estimation results regarding the presence of one or more aneuploids in the embryo image are displayed via the user interface.
Citation Information
Patent Citations
Apparatus, method, and system for image-based human embryo cell classification
CN105408746A
Chromosome abnormality detection model, chromosome abnormality detection system, and chromosome abnormality detection method
CN110265087A