Methods and systems for determining embryo ploidy
Patent Information
- Application Number
- PCT/US2026/020804
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2026-03-25
- Publication Date
- 2026-10-01
Smart Images

Figure US2026020804_01102026_PF_FP_ABST
Abstract
Description
96PR-701001-WO PATENT METHODS AND SYSTEMS FOR DETERMINING EMBRYO PLOIDYCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit under 35 U.S.C. §119(e) to U.S. Provisional Patent Application No. 63 / 778,342, filed March 26, 2025. The content of this related application is incorporated herein by reference in its entirety for all purposes.BACKGROUNDField
[0002] The present application generally relates to the field of assisted reproductive technologies, and particularly non-invasive embryo viability screening.Description of the Related Art
[0003] In vitro fertilization (IVF) as part of assisted reproduction technology has become an effective solution for infertility. Embryo selection based on ploidy status has been proven to be effective in reducing the risk of miscarriage while increasing the implantation rate in IVF. However, current methodologies largely rely on costly, invasive and time-consuming approaches such as PGT-A screening which hold unknown risks on the safety of embryo development before and after implantation.
[0004] There exists a need for non-invasive, cost-effective and accurate approaches for identifying embryos with chromosomal abnormalities, ideally in a high-throughput and automated manner.SUMMARY
[0005] Disclosed herein include methods for determining metabolic health of an embryo. In some embodiments, a method (or one or more actions of the method) for determining metabolic health of an embryo is under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: obtaining a plurality of training mass spectra using a mass spectrometry from a plurality of spent embryo media where embryos had been cultured in. Each of the embryos can be associated with a metabolic health classification. The metabolic health classification can comprise a healthy classification and an unhealthy classification. The method can comprise: identifying a plurality of training mass spectrometric peaks less than 2000 Daltons from each of the plurality of training mass spectra. The method can comprise: merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct atraining peak matrix. The method can comprise: training a classifier for differentiating mass spectra corresponding to embryos with the healthy classification from mass spectra corresponding to embryos with the unhealthy classification using the training peak matrix as input and the corresponding metabolic health classifications of the embryos as output. The method can comprise: obtaining a sample mass spectrum using the mass spectrometry from a sample of spent embryo medium where a sample embryo had been cultured in. The method can comprise: determining a metabolic health classification of the sample embryo using the classifier with the sample mass spectrum as input.
[0006] In some embodiments, the embryos comprise one or more euploidy embryos each with the healthy classification. In some embodiments, the embryos comprise one or more aneuploidy embryos with the unhealthy classification. In some embodiments, a euploidy embryo of the embryos is more likely to be associated with the healthy classification. In some embodiments, an aneuploidy embryo of the embryos is more likely to be associated with the unhealthy classification. In some embodiments, an embryo with the healthy classification is more likely to be a euploidy embryo. In some embodiments, an embryo with the unhealthy classification is more likely to be an aneuploidy embryo. In some embodiments, the metabolic health classification of the sample embryo is the healthy classification and the sample embryo is a euploidy embryo. In some embodiments, the metabolic health classification of the sample embryo is the unhealthy classification and the sample embryo is an aneuploidy.
[0007] Disclosed herein include methods for determining metabolic health of an embryo. In some embodiments, a method (or one or more actions of the method) for determining metabolic health of an embryo is under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: obtaining a plurality of training mass spectra using a mass spectrometry from a plurality of spent embryo media where embryos had been cultured in. Each of the embryos can be associated with a metabolic health score. The method can comprise: identifying a plurality of training mass spectrometric peaks less than 2000 Daltons from each of the plurality of training mass spectra. The method can comprise: merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix. The method can comprise: training a classifier for determining metabolic health scores of embryos using the training peak matrix as input and the corresponding metabolic health scores of the embryos as output. The method can comprise: obtaining a sample mass spectrum using the mass spectrometry from a sample of spent embryo medium where a sample embryo had been cultured in. The method can comprise: determining a metabolic health score of the sample embryo using the classifier with the sample mass spectrum as input.
[0008] In some embodiments, the embryos comprise one or more euploidy embryos each with a metabolic health score above a threshold. In some embodiments, the embryos comprise one or more aneuploidy embryos with a metabolic health score below a threshold. In some embodiments, the threshold is 50%, 60%, 70%, 80%, or 90%. In some embodiments, a euploidy embryo of the embryos is more likely to be associated with a higher metabolic health score. In some embodiments, an aneuploidy embryo of the embryos is more likely to be associated with a lower metabolic health score. In some embodiments, an embryo with a higher metabolic heath score is more likely to be a euploidy embryo. In some embodiments, an embryo with a lower metabolic health score is more likely to be an aneuploidy embryo. In some embodiments, the metabolic health score of the sample embryo is above a threshold and the sample embryo is a euploidy embryo. In some embodiments, metabolic health score of the sample embryo is below a threshold and the sample embryo is an aneuploidy. In some embodiments, the threshold is 50%, 60%, 70%, 80%, or 90%.
[0009] Disclosed herein include methods for determining metabolic health of an embryo. In some embodiments, a method (or one or more actions of the method) for determining metabolic health of an embryo is under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: obtaining a plurality of training mass spectra using a mass spectrometry from a plurality of spent embryo media where embryos had been cultured in. Each of the embryos can be associated with metabolic health information. The method can comprise: identifying a plurality of training mass spectrometric peaks less than 2000 Daltons from each of the plurality of training mass spectra. The method can comprise: merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix. The method can comprise: training a classifier for determining metabolic health information of embryos using the training peak matrix as input and the corresponding metabolic health information of the embryos as output. The method can comprise: obtaining a sample mass spectrum using the mass spectrometry from a sample of spent embryo medium where a sample embryo had been cultured in. The method can comprise: determining metabolic health information of the sample embryo using the classifier with the sample mass spectrum as input.
[0010] Disclosed herein include methods for determining metabolic health of an embryo. In some embodiments, a method (or one or more actions of the method) for determining metabolic health of an embryo is under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: obtaining a plurality of training mass spectra using a mass spectrometry from a plurality of spent embryo media where embryos had been cultured in. Each of the embryos can be associated with metabolic health information. The method can comprise: merging training mass spectrometric peaks from the plurality of training mass spectrato construct a training peak matrix. The training mass spectrometric peaks can be less than 2000 Daltons. The method can comprise: training a classifier for determining metabolic health information of embryos using the training peak matrix as input and the corresponding metabolic health information of the embryos as output. The method can comprise: obtaining a sample mass spectrum using the mass spectrometry from a sample of spent embryo medium where a sample embryo had been cultured in. The method can comprise: determining metabolic health information of the sample embryo using the classifier with the sample mass spectrum as input.
[0011] In some embodiments, the metabolic health information comprises a metabolic health classification. In some embodiments, the metabolic health classification comprises a healthy classification and an unhealthy classification. In some embodiments, the metabolic health information comprises a metabolic health score. In some embodiments, the metabolic health information comprises with a classification of ploidy comprising euploidy and aneuploidy.
[0012] Disclosed herein include methods for determining embryo ploidy. In some embodiments, a method (or one or more actions of the method) for determining embryo ploidy is under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: obtaining a plurality of training mass spectra using Matrix- Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry from a plurality of spent embryo media where embryos had been cultured in. The embryos can comprise euploidy embryos and aneuploidy embryos. Each of the embryos can be associated with a classification of ploidy comprising euploidy and aneuploidy. The method can comprise: identifying a plurality of training mass spectrometric peaks from each of the plurality of training mass spectra. The method can comprise: merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix. The method can comprise: training a classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos using the training peak matrix as input and the corresponding classifications of the embryos as output. The method can comprise: obtaining a sample mass spectrum using MALD-TOF mass spectrometry from a sample of spent embryo medium where a sample embryo had been cultured in. The method can comprise: determining a classification of the sample embryo using the classifier with the sample mass spectrum as input.
[0013] Disclosed herein include methods for determining embryo ploidy. In some embodiments, a method (or one or more actions of the method) for determining embryo ploidy is under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: obtaining a plurality of first training mass spectra using Matrix- Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry from a plurality of first spent embryo media where first embryos had been cultured in. The first embryos were cultured ina first medium. The method can comprise: obtaining a plurality of second training mass spectra using MALDI-TOF mass spectrometry from a plurality of second spent embryo media where second embryos had been cultured in. The first medium and the second medium were different. The first embryos and the second embryos can both comprise euploidy embryos and aneuploidy embryos. Each of the first embryos and each of the second embryos can be associated with a classification of ploidy comprising euploidy and aneuploidy. The method can comprise: for the plurality of first training mass spectra and for the plurality of second training mass spectra: identifying a plurality of first or second training mass spectrometric peaks from each of the plurality of first or second training mass spectra. The method can comprise: merging first or second training mass spectrometric peaks identified from the plurality of first or second training mass spectra to construct a first or second training peak matrix. The method can comprise: training a first or second classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos using the first or second training peak matrix as input and the corresponding classifications of the first or second embryos as output. The method can comprise: obtaining a sample mass spectrum using MALD-TOF mass spectrometry from a sample of spent embryo medium where a sample embryo had been cultured in. The sample embryo was cultured in the first medium or the second medium. The method can comprise: if the sample embryo was cultured in the first medium, determining a classification of the sample embryo using the first classifier with the sample mass spectrum as input, or if the sample embryo was cultured in the second medium, determining a classification of the sample embryo using the second classifier with the sample mass spectrum as input.
[0014] Disclosed herein include methods for determining embryo ploidy. In some embodiments, a method (or one or more actions of the method) for determining embryo ploidy is under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: receiving a plurality of training mass spectra obtained by subjecting a plurality of spent embryo media where embryos had been cultured into Matrix-Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry. The embryos can comprise euploidy embryos and aneuploidy embryos. Each of the embryos can be associated with a classification of ploidy comprising euploidy and aneuploidy. The method can comprise: identifying a plurality of training mass spectrometric peaks from each of the plurality of training mass spectra. The method can comprise: merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix. The method can comprise: training a classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos using the training peak matrix as input and the corresponding classifications of the embryos as output. The method can comprise:receiving a sample mass spectrum obtained by subjecting a sample spent embryo medium where a sample embryo had been cultured in to MALDI-TOF mass spectrometry. The method can comprise: determining a classification of the sample embryo using the classifier with the sample mass spectrum as input.
[0015] In some embodiments, the method further comprises providing the plurality of spent embryo media, providing the test spent embryo media, or both. The method can further comprise receiving the plurality of spent embryo media, receiving the test spent embryo media, or both. The method can further comprise culturing one or more of the embryos to generate one or more of the plurality of spent embryo media, and / or culturing the sample embryo to obtain the sample spent embryo medium.
[0016] In some embodiments, an embryo of the embryos and the sample embryo were cultured in media that were identical. In some embodiments, an embryo of the embryos and the sample embryo were cultured in media that were different. In some embodiments, the embryos were cultured in media that were identical. In some embodiments, the embryos were cultured in media that were different.
[0017] In some embodiments, the classification of the sample embryo is used to determine embryo viability of the sample embryo. In some embodiments, the method further comprises determining the embryo viability of the sample embryo. In some embodiments, determining the embryo viability of the sample embryo comprises determining the embryo viability of the sample embryo using the classification of the sample embryo. In some embodiments, determining the embryo viability of the sample embryo comprises determining whether the sample mass spectrum corresponds to that of a euploidy embryo or an aneuploidy embryo.
[0018] In some embodiments, the plurality of training mass spectra comprises at least 100 training mass spectra corresponding to euploidy embryos and / or at least 100 training mass spectra corresponding to aneuploidy embryos. In some embodiments, the plurality of training mass spectra comprises at most 10 training mass spectra corresponding to mosaic embryos. In some embodiments, the plurality of training mass spectrometric peaks comprises metabolite peaks. In some embodiments, identifying the plurality of training mass spectrometric peaks from each training mass spectrum comprise selecting training mass spectrometric peaks in the range 50-1500 Daltons. In some embodiments, the method further comprises aligning the training mass spectrometric peaks identified from the plurality of training mass spectra. In some embodiments, the method further comprises normalizing each of the plurality of training mass spectra. In some embodiments, identifying the plurality of training mass spectrometric peaks from each training mass spectrum comprises selecting peaks with a height equal to or greater than 1% of that of ahighest peak. In some embodiments, the method further comprises discarding peaks having a width of less than 1 Dalton and a prominence of less than 0.01.
[0019] In some embodiments, the classifier is selected from the group consisting of: Partial-Least Squares Discriminant Analysis (PLS-DA), Support Vector Machine (SVM), Random Forest (RF), K-Nearest Neighbour (KNN), Light Gradient-Booster Machine (LightGBM), or a combination thereof. In some embodiments, the classifier is LightGBM. In some embodiments, the classifier is RF.
[0020] In some embodiments, the method further comprises removing one or more outlier training mass spectra prior to constructing the training peak matrix. The method can further comprise generating a second plurality of training mass spectrometric peaks from a plurality of training mass spectrometric peak from one or more of the plurality of training mass spectra using SMOTE algorithm. In some embodiments, merging the training mass spectrometric peaks comprises: merging the training mass spectrometric peaks identified from the plurality of training mass spectra and the training mass spectrometric peaks generated using SMOTE algorithm to construct a training peak matrix. The method can further comprise identifying mass-to-charge values representative of a euploidy embryo or an aneuploidy embryo. In some embodiments, the embryos are from mammalian subjects.
[0021] Disclosed herein include systems for determining embryo ploidy. In some embodiments, a system for determining embryo ploidy comprises: non -transitory memory configured to store executable instructions. The non-transitory memory can be configured to store a classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos. The classifier can be generated by: receiving a plurality of training mass spectra obtained by subjecting a plurality of spent embryo media where embryos had been cultured into Matrix- Assisted Laser Desorption / Ionization Time-of-Flight (MALDL TOF) mass spectrometry. The embryos can comprise euploidy embryos and aneuploidy embryos. Each of the embryos can be associated with a classification of ploidy comprising euploidy and aneuploidy. The classifier can be generated by: identifying a plurality of training mass spectrometric peaks from each of the plurality of training mass spectra. The classifier can be generated by: merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix. The classifier can be generated by: training the classifier using the training peak matrix as input and the corresponding classifications of the embryos as output. The system can comprise: a processor (e.g., a hardware processor or a virtual processor) in communication with the non-transitory memory. The processor can be programmed by the executable instructions to perform: receiving a sample mass spectrum obtained by subjecting a sample spent embryo medium where a sample embryo had been cultured into MALDI-TOF massspectrometry. The processor can be programmed by the executable instructions to perform: determining a classification of the sample embryo using the classifier with the sample mass spectrum as input.
[0022] In some embodiments, an embryo of the embryos and the sample embryo were cultured in media that were identical. In some embodiments, an embryo of the embryos and the sample embryo were cultured in media that were different. In some embodiments, the embryos were cultured in media that were identical. In some embodiments, the embryos were cultured in media that were different.
[0023] In some embodiments, the classification of the sample embryo is used to determine embryo viability of the sample embryo. In some embodiments, the processor is further programmed by the executable instructions to perform: determining the embryo viability of the sample embryo. In some embodiments, determining the embryo viability of the sample embryo comprises determining the embryo viability of the sample embryo using the classification of the sample embryo. In some embodiments, determining the embryo viability of the sample embryo comprises determining whether the sample mass spectrum corresponds to that of a euploidy embryo or an aneuploidy embryo.
[0024] In some embodiments, the plurality of training mass spectra comprises at least 100 training mass spectra corresponding to euploidy embryos and / or at least 100 training mass spectra corresponding to aneuploidy embryos. In some embodiments, the plurality of training mass spectra comprises at most 10 training mass spectra corresponding to mosaic embryos. In some embodiments, the plurality of training mass spectrometric peaks comprises metabolite peaks. In some embodiments, identifying the plurality of training mass spectrometric peaks from each training mass spectrum comprise selecting training mass spectrometric peaks in the range 50-1500 Daltons. In some embodiments, the classifier is further generated by aligning the training mass spectrometric peaks identified from the plurality of training mass spectra. In some embodiments, the classifier is further generated by normalizing each of the plurality of training mass spectra. In some embodiments, identifying the plurality of training mass spectrometric peaks from each training mass spectrum comprises selecting peaks with a height equal to or greater than 1% of that of a highest peak. In some embodiments, the system further comprises discarding peaks having a width of less than 1 Dalton and a prominence of less than 0.01.
[0025] In some embodiments, the classifier is selected from the group consisting of: Partial-Least Squares Discriminant Analysis (PLS-DA), Support Vector Machine (SVM), Random Forest (RF), K-Nearest Neighbour (KNN), Light Gradient-Booster Machine (LightGBM), or a combination thereof. In some embodiments, the classifier is LightGBM. In some embodiments, the classifier is RF.
[0026] In some embodiments, the processor is programmed by the executable instructions to perform: removing one or more outlier training mass spectra prior to constructing the training peak matrix. In some embodiments, the classifier is further generated by generating a second plurality of training mass spectrometric peaks from a plurality of training mass spectrometric peak from one or more of the plurality of training mass spectra using SMOTE algorithm. In some embodiments, merging the training mass spectrometric peaks comprises: merging the training mass spectrometric peaks identified from the plurality of training mass spectra and the training mass spectrometric peaks generated using SMOTE algorithm to construct a training peak matrix. In some embodiments, the classifier is further generated by identifying mass-to-charge values representative of a euploidy embryo or an aneuploidy embryo. In some embodiments, the embryos are from mammalian subjects.
[0027] Also disclosed herein include a system comprising non-transitory memory configured to store: executable instructions; and a processor (e.g., a hardware processor or a virtual processor) in communication with the non-transitory memory, the processor programmed by the executable instructions to perform: any method or one or more actions of any method disclosed herein.
[0028] Also disclosed herein include a non-transitory computer-readable medium storing executable instructions, when executed by a system (e.g., a computing system), causes the system to perform any method or one or more steps of any method disclosed herein.
[0029] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Neither this summary nor the following detailed description purports to define or limit the scope of the inventive subject matter.BRIEF DESCRIPTION OF THE DRAWINGS
[0030] FIG. 1 illustrates a non-limiting workflow of the embryo ploidy calling method. Step 1: sample collection, characterization and acquisition of MALDI-TOF MS spectra. Step 2: pre-processing of each individual spectra to prepare the training set. Step 3: adjacent analyses to study the reproducibility of the data. Step 4: the preparation of a peak matrix to be used as input to ML algorithms, and Step 5: the battery of ML algorithms applied and the process of k-fold Cross Validation and obtaining metrics to evaluate the results.
[0031] FIG. 2 is a Violin plot result of the correlation analysis within spectra in each category.
[0032] FIG. 3 is a PCA plot comparing the first and second principal components ofthe training dataset. The fertility center of origin of each sample is overlaid. Green dots correspond to samples originated at Atlantic Reproductive Medicine Centre; blue crosses come from Kindbody-West Loop, yellow squares from Reproductive Care Centre and red dots from Rinehart Fertility Centre. 95% confidence ellipses are drawn over each group.’
[0033] FIG. 4A is a plot showing Receiver Operating Characteristic (ROC) curves for the optimized training of the LightGBM classifier for the Euploidy (green), Aneuploidy (yellow) and Mosaic (blue) categories. The plot uses Youden indices to show the performance of the dichotomous test and Area Under the Curve (AUC) values for each curve. FIG. 4B is a plot showing Corresponding Precision-Recall (PR) curves, including AUC and Averaged Precision (AP).
[0034] FIG. 5 is a feature importance plot for the results of the LightGBM algorithm. A higher intensity loading means more discriminatory power. Labeled are some of the more relevant peaks for discrimination among the categories.
[0035] FIG. 6 is a distance plot of the hyperparameter-optimized Random Forest (RF) training applied to the three categories Euploidy (green), aneuploidy (yellow), and mosaic (blue).
[0036] FIG. 7 is a plot showing Shapley values for the hyperparameter-optimized RF classification result. Top plot corresponds to the Euploidy category, middle plot to Aneuploidy, and bottom one to Mosaic. Each row corresponds to an m / z value representative of a detected peak. Dots in each row correspond to the area detected in each spectrum in the dataset for that m / z value. Colors towards blue indicate low area value (i.e. low intensity) in that particular peak, while colors towards red indicate high area value (i.e. high intensity). Blue horizontal bars indicate the mean of the absolute value of all Shapley values for that m / z. A large bar indicates that the feature (m / z value) is highly important to describe the category.
[0037] FIG. 8 is a plot showing heatmap matrix of the distance between all pairs of samples in the RF classification. Red means less distance between the pair, while blue means more distance.
[0038] FIG. 9 is a block diagram of an illustrative computing system that can be used in some embodiments to execute the processes and implement the features described herein.
[0039] Throughout the drawings, reference numbers may be re-used to indicate correspondence between referenced elements. The drawings are provided to illustrate example embodiments described herein and are not intended to limit the scope of the disclosure.DETAILED DESCRIPTION
[0040] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similarcomponents, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the Figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein and made part of the disclosure herein.
[0041] All patents, published patent applications, other publications, and sequences from GenBank, and other databases referred to herein are incorporated by reference in their entirety with respect to the related technology.
[0042] Disclosed herein includes a method for determining embryo ploidy. In some embodiments, the method can comprise subjecting a plurality of spent embryo media where embryos had been cultured in to Matrix- Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry to obtain a plurality of training mass spectra, wherein the embryos comprise euploidy embryos and aneuploidy embryos, and wherein each of the embryos is associated with a classification of ploidy comprising euploidy and aneuploidy, identifying a plurality of training mass spectrometric peaks from each of the plurality of training mass spectra, merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix, training a classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos using the training peak matrix as input and the corresponding classifications of the embryos as output, subjecting a sample spent embryo medium where a sample embryo had been cultured in to MALDI-TOF mass spectrometry to obtain a sample mass spectrum, and determining a classification of the sample embryo using the classifier with the sample mass spectrum as input.
[0043] In some embodiments, the method can comprise under control of a hardware processor: receiving a plurality of training mass spectra obtained by subjecting a plurality of spent embryo media where embryos had been cultured in to Matrix-Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry, wherein the embryos comprise euploidy embryos and aneuploidy embryos, and wherein each of the embryos is associated with a classification of ploidy comprising euploidy and aneuploidy, identifying a plurality of training mass spectrometric peaks from each of the plurality of training mass spectra, merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix, training a classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos using the training peak matrix as input and the corresponding classifications of the embryos asoutput, receiving a sample mass spectrum obtained by subjecting a sample spent embryo medium where a sample embryo had been cultured in to MALDI-TOF mass spectrometry, and determining a classification of the sample embryo using the classifier with the sample mass spectrum as input.
[0044] Disclosed herein also includes a system for determining embryo ploidy. In some embodiments, the system can comprise non-transitory memory configured to store executable instructions and a classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos, wherein the classifier is generated by: receiving a plurality of training mass spectra obtained by subjecting a plurality of spent embryo media where embryos had been cultured in to Matrix- Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry, wherein the embryos comprise euploidy embryos and aneuploidy embryos, and wherein each of the embryos is associated with a classification of ploidy comprising euploidy and aneuploidy, identifying a plurality of training mass spectrometric peaks from each of the plurality of training mass spectra, merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix, training the classifier using the training peak matrix as input and the corresponding classifications of the embryos as output, and a hardware processor in communication with the non-transitory memory, the hardware processor programmed by the executable instructions to perform: receiving a sample mass spectrum obtained by subjecting a sample spent embryo medium where a sample embryo had been cultured in to MALDI-TOF mass spectrometry, and determining a classification of the sample embryo using the classifier with the sample mass spectrum as input.Overview
[0045] Predicting embryo viability is a decisive aspect of in vitro fertilization (IVF) given the inherent complexities involved in culturing a human embryo. Currently, embryo selection is predicated on morphological assessment and preimplantation genetic testing (PGT). These methodologies are effective at assessing blastulation and chromosomal abnormality, respectively. However, both techniques possess limitations which can affect clinical outcomes. With morphological assessment, highly graded embryos do not always yield successful implantations; likewise, embryos determined to be chromosomally abnormal by PGT can lead to a successful pregnancy. Regardless of the efficacy of these techniques, both approaches are not sufficiently accurate. Thus, the exploration of alternative approaches is merited.
[0046] Morphological assessment has been a mainstay of IVF since the 1980s, and it has been a vital metric for predicting implantation and pregnancy outcomes. The screening process entails observing the shape, size, and developmental progress of the embryo. By prioritizing thehighest quality embryos, it reduces the risks of implantation failure and miscarriages. In addition to the extensive library of morphological data, there are established criteria and quality control steps utilized to minimize misdiagnosis. Despite the apparent advantages of morphological screening, there are limitations that bring into question its efficacy and accuracy. While morphological assessment can aid in the determination of ploidy, it is not sufficiently accurate as a sole means for diagnosis. Moreover, the approach is subjective, relying heavily on the expertise of embryologists which can lead to variability and inconsistent results. Given these limitations morphological screening is not a suitable stand-alone technique capable of providing reliable results. Thus, morphological assessment is currently used in conjunction with PGT to assess embryo viability.
[0047] PGT has become a common part of IVF treatments. Its ability to detect chromosomal abnormality reduces the risk of miscarriages, reduces the time to pregnancy, and birth defects. Moreover, being able to select euploid embryos for implantation has lowered the rate of implantation failure. The Society for Assisted Reproductive Technology (SART) data also shows that PGT reduces the maternal age effect on live birth rate. In addition to that, there are benefits to identifying euploid embryos for patients with advanced maternal age since they have a higher percentage of aneuploid embryos. In Europe there is a retrospective study that clearly demonstrates the benefits of PGT-A in terms of live births per embryo transferred, as well as per cycle started. In spite of the obvious benefits PGT offers, there is a growing concern over the accuracy and effectiveness of the technique. There is a sentiment that PGT should be used in only an experimental capacity. One particular metric, cumulative live birth rate (CLBR), is believed to be a better measure of live birth during a treatment as opposed to per cycle or per embryo transfer. With that said, a decreased CLBR in women <35 who utilized PGT-A compared to those who did not use PGT-A. No difference in LBR in women 35-37 has been reported. Further reinforcing those findings, a work in the Chinese Medical Journal showed there is no evidence that PGT-A improves the cumulative live birth rate in recurrent pregnancy loss (RPL) couples regardless of maternal age, which could be because a large number of viable embryos were not utilized. Taken together, all these findings make it difficult to discern if PGT should be used as a diagnostic tool or not. Knowing the limitations of morphological assessment and PGT, it is apparent that better testing solutions need to be investigated.
[0048] This has led to the evaluation of a number of non-invasive approaches. Initially, non-invasive PGT (niPGT) introduced in 2016 was acclaimed as a promising replacement for invasive PGT, although current data shows that niPGT-A has an 11.8% false-negative rate and a 16.0% false-positive rate, which were attributed to maternal cumulus cell contamination and embryo mosaicism. Subsequently, alternative approaches are being explored in areas other thanchromosome copy number (CNV), which include secretome analysis, gene expression profiling, metabolic images (FAD and NAD detection), and metabolomic screening. These strategies focus on in-depth analysis of spent embryo culture media (SECM) to uncover biomarkers linked to embryo viability.
[0049] Since the early 2000s, a lot of emphasis has been placed on surveying SECM. SECM contains a variety of different molecules, including small molecules, lipids, peptides, and proteins, which could serve as viable biomarkers. However, identification of SECM biomarkers is a non-trivial endeavor that has been attempted using multiple analytical techniques. Some of the earliest instances employed Raman spectroscopy, near-infrared spectroscopy, and nuclear magnetic resonance. The work demonstrates the utility of metabolic profiling to predict the reproductive potential of an embryo. Moreover, differences in metabolism also have potential as indicators for embryo viability and implantation potential. Despite having moderate success in euploidy prediction, these techniques had mixed results predicting implantation / pregnancy. Shortly thereafter, a number of biochemical techniques were employed to evaluate multiple metabolites in SECM. Although more recently, mass spectrometry has reached the forefront providing in-depth analysis of the SECM metabolome, lipidome, and proteome.
[0050] Liquid chromatography tandem mass spectrometry (LC-MS / MS) is the preeminent technique used for analyzing SECM, with some exceptions. When categorizing the previous mass spectrometry studies in the field of non-invasive embryo screening, there is a bifurcation in approaches, proteomics versus metabolomics. At the outset, a proteomic and genomic study was conducted to define the peptides and proteins in the human embryonic secretome. That work provided insight into the cellular processes that relate to embryonic development. Shortly thereafter, profiling the embryonic secretome was used to determine if embryos express and secrete unique proteins. Despite the small number of samples analyzed, this quantitative proteomics approach produced an effective way to analyze the day-3 embryo secretome. Further analysis of the day 2-3 embryonic secretome yielded more viable biomarkers. A thorough investigation of peptide and protein extraction methods was performed on SECM and identified Apolipoprotein A-l as marker for embryo competency. Another study identified Haptoglobin a-l is a viable marker for distinguishing between non-viable embryos and viable for transfer. After a fair amount of work dedicated to the embryonic secretome, protein expression in human cumulus cells (CC) was also evaluated to predict pregnancy success. Based on the previous research, proteomic analysis of the secretome and CC do hold promise for the future of non-invasive embryo screening. Alternatively, there has been a significant amount of work dedicated to profiling metabolites in SECM. Given that embryos utilize the media compounds for metabolism, both measuring the known media components and identifying byproducts ofmetabolism should reveal detectable differences. An excellent example of this is early fingerprinting work done on metabolites in embryo culture media; the study was successful at identifying 100% of samples from embryos that were implanted and 70% of some samples from the non-implanted group. However, not all studies were successful at identifying profiles or biomarkers that are significant, even though differences were observed. While individual markers can help us to understand the underlying biological mechanisms, a more comprehensive approach is needed, such as one that utilizes artificial intelligence to aid in identifying subtle patterns that are not so obvious.
[0051] More recently, there has been a litany of studies that demonstrate the effectiveness of using metabolite biomarkers. The most intriguing study to date has an implantation prediction algorithm with an 85.29% accuracy, with a PPV of 88% and a NPV of 77.78%, which is comparable if not better than current genetic testing. While both proteomic and metabolomic analysis aptly demonstrate that patterns and biomarkers can be used to aid in embryo viability, implantation success, and pregnancy, further examination, validation, and scaling must be completed before these assays can be introduced into a clinical laboratory. Based on previous literatures, the effectiveness of LC-MS / MS is apparent; its ability to detect and survey a multitude of ions and identify them makes it rather ideal. Nevertheless, there are some limitations that reduce its clinical utility. In general, LC-MS / MS can be costly, time consuming, and data analysis is complex. Alternatively, matrix assisted laser desorption ionization-time of flight mass spectrometry (MALDI-TOF MS) is a high-throughput, cost-effective, and easily automatable option. The suitability of MALDI-TOF MS in the surveying and identifying distinct spectral profiles in spent medium has been demonstrated. Similarly, other screening approaches like surface enhanced laser desorption ionization-time of flight mass spectrometry (SELDLTOF MS) have also been shown to be effective as well. The value of MALDI-TOF MS in clinical research has been aptly demonstrated over the past few decades. The technique’s attractiveness lies in its simplicity and ease of use; it doesn’t require a comprehensive scientific background for routine operation. Although, with a high-throughput technique like MALDI-TOF MS, some aspects require careful consideration. For instance, a single mass spectrum takes mere seconds to generate, which is ideal with small data sets. On the other hand, large clinical data sets generate spectra in hundreds to thousands. In these cases, data analysis cannot be handled manually by an individual. Thus, advanced artificial intelligence or machine learning algorithms are necessary to adequately examine the data.
[0052] Supervised Machine Learning (ML) is used to optimize performance metrics using past examples which are called training data. A ML model is built using training data and further learning is done to optimize the model’s performance metrics. Models can be predictive,which is used to predict output for test data or can be descriptive, which is used to describe the data to acquire knowledge like finding patterns or association rules from the data or can be both.
[0053] Disclosed herein includes methods and systems combining machine learning with non-invasive metabolomics assays, which are capable of identifying metabolic patterns indicative of embryos well-suited for implantation. This technology can process and analyze complex metabolomic data, detect subtle patterns and accurately predict results. The coupling of MALDI-TOF MS and machine learning can improve the accuracy and efficacy of predicting embryo viability, thus significantly improving embryo selection in IVF and implantation success and increasing live birth rate. The findings presented herein can also serve as the foundation for the development of a MALDI based automated platform capable of high throughput routine analysis of spent media in clinical laboratories.Methods and Systems for Determining Embryo Ploidy
[0054] Disclosed herein include methods and systems for determining embryo ploidy state. In some embodiments, the method comprises subjecting a plurality of spent embryo media where embryos had been cultured in to Matrix- Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry to obtain a plurality of training mass spectra, wherein the embryos comprise euploidy embryos and aneuploidy embryos, and wherein each of the embryos is associated with a classification of ploidy comprising euploidy and aneuploidy, identifying a plurality of training mass spectrometric peaks from each of the plurality of training mass spectra, merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix, training a classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos using the training peak matrix as input and the corresponding classifications of the embryos as output, subjecting a sample spent embryo medium where a sample embryo had been cultured in to MALDI-TOF mass spectrometry to obtain a sample mass spectrum, and determining a classification of the sample embryo using the classifier with the sample mass spectrum as input.Embryo Ploidy
[0055] As used herein, the term “ploidy” in general refers to the quantity and chromosomal identity of one or more chromosomes in a cell. The term “ploidy state” in the context of a cell (e.g., embryo) can refer to a categorization of an embryo being euploidy or aneuploidy. Accordingly, in some embodiments, a ploidy state can be euploidy or aneuploidy. Embryos in euploidy state and / or classified as being in euploidy state are considered euploidy embryos and embryos in aneuploidy state and / or classified as being in aneuploidy state are aneuploidy embryos.
[0056] In some embodiments, the embryos described herein comprise euploidy embryos and aneuploidy embryos, and each of the embryos is associated with a classification of ploidy comprising euploidy and aneuploidy. For example, a euploidy embryo is associated with euploidy classification and an aneuploidy is associated with aneuploidy classification. A euploidy embryo refers to an embryo having the correct number (i.e., normal number) of chromosomes. For example, a euploidy human embryo can have 46 chromosomes. Euploid embryos are more likely to implant, less likely to result in miscarriage, and less likely to result in a baby with birth defects caused by chromosomal abnormalities. An aneuploid embryo, also known as “abnormal” embryo, refers to an embryo having an abnormal number of chromosomes, e.g., either missing or extra chromosomes. An exemplary aneuploidy embryo disorder is Trisomy 21 or Down syndrome in which the embryo has an extra chromosome 21. In some embodiments, an embryo can be a mosaic embryo. A mosaic embryo contains both cells with normal chromosomes and cells with abnormal chromosomes. In some embodiments, the aneuploidy embryos described herein include mosaic embryos. In some other embodiments, the aneuploidy embryos exclude mosaic embryos. In some embodiments, the embryos comprise euploidy embryos, aneuploidy embryos, and mosaic embryos, and the classification of an embryo of the embryos comprises euploidy, aneuploidy, or mosaic.
[0057] The methods and systems described herein provide a non-invasive approach for determining embryo ploidy status using spent embryo media where embryo had been cultured in. The term “spent embryo media” refers to an embryo culture media where one or more embryos had been cultured in. Spent embryo media are formed by cell growth and remain after culturing embryos. During cell growth both nutrient depletion and metabolite accumulation occur in the medium. Therefore, spent embryo media generally contain lower nutrient levels and higher metabolite levels compared to fresh cell culture media.
[0058] In some embodiments, the method described herein comprises culturing one or more embryos to generate one or more spent embryo media. The method can comprise culturing one or more of the embryos described herein to generate one or more of a plurality of spent embryo media, and / or culturing a sample embryo to obtain a sample spent embryo medium.
[0059] Any cell culture media permitting in vitro culture of embryos can be used herein. The compositions of embryo culture media can vary in different embodiments. The embryo culture media used herein can be single-step solutions used throughout the culture period or sequential systems with different media compositions at different developmental stages. In general, an embryo culture media comprises a salt solution and other components that provide nutrients, energy, proper osmolality, and pH balance for embryo growth. In some embodiments, embryo culture media contains salts, buffer, energy substrates (e.g., glucose, sodium lactate,sodium pyruvate), essential amino acids, non-essential amino acids, glutamine dipeptide, chelator (e.g., EDTA), macromolecules (e.g., serum albumin), fatty acid, vitamins, antibiotics, growth factors, and / or other essential components allowing embryos to grow and divide in vitro. There are also commercially available media systems for human preimplantation embryo culture including single-step medium and sequential medium systems from Life Global, Vitrolife, Origo, Cooper Surgical, and others identifiable by a person skilled in the art. Additional information on embryo culture medium composition can be found in various published literature reviews and commercial websites.
[0060] The embryos described herein can be cultured in identical media or in different media. For example, a sample embryo and embryos used for obtaining the training mass spectra (i.e., training embryos) can be cultured in media that are identical. Alternatively, the sample embryo and the training embryos can be cultured in media that were different. In some embodiments, two or more of the training embryos can be cultured in media that were identical. Alternatively, two or more of the training embryos can be cultured in media that were different.
[0061] The spent embryo media is collected from embryo culture prior to embryo implantation. The embryos can be cultured for longer or shorter periods depending on the embryo’s development and IVF lab setting. For example, the spent embryo media can be collected after the embryos have been cultured for about, at least, at least about, at most, or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or a number or a range between any two of these values, days. In some embodiments, the embryos had been cultured for a maximum period of 14 days prior to the spent media collection.
[0062] In some embodiments, the method described herein further comprises providing a plurality of training spent embryo media, providing a sample spent embryo media, or both. In some embodiments, the method can comprise receiving the plurality of training embryo media, receiving the sample spent embryo media, or both. In some embodiments, the method described herein comprises collecting a plurality of training spent embryo media, collecting a sample spent embryo media, or both from the cell culture media wherein the embryos or embryo had been cultured in.Mass spectrometry data processins
[0063] The method described herein comprises subjecting a plurality of spent embryo media where embryos had been cultured in to mass spectrometry such as Matrix- Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry to obtain a plurality of training mass spectra.
[0064] As used herein, the term “mass spectrometry” refers to an analytical tool able to volatilize / ionize analytes to form gas-phase ion and determine their absolute or relativemolecular masses. Typically, mass spectrometers can be used to identify unknown compounds via molecular weight determination, to quantify known compounds, and to determine structure and chemical properties of molecules. Suitable methods of volatilization / ionization are matrix-assisted laser desorption ionization (MALDI), electrospray, laser / light, thermal, electrical, atomized / sprayed and the like, or combinations thereof. Suitable forms of mass spectrometry include, but are not limited to, ion trap instruments, quadrupole instruments, electrostatic and magnetic sector instruments, time of flight instruments, time of flight tandem mass spectrometer (TOF MS / MS), Fourier-transform mass spectrometers, Orbitraps and hybrid instruments composed of various combinations of these types of mass analyzers. These instruments can, in turn, be interfaced with a variety of other instruments that fractionate the samples (for example, liquid chromatography or solid-phase adsorption techniques based on chemical, or biological properties) and that ionize the samples for introduction into the mass spectrometer, including matrix-assisted laser desorption (MALDI), electrospray, or nanospray ionization (ESI) or combinations thereof.
[0065] In some embodiments, matrix assisted laser desorption ionization-time of flight mass spectrometry (MALDI-TOF MS) is used herein for metabolomics profiling of embryo spent media (e.g., human embryo spent media). In brief, a sample for analysis by MALDI MS (e.g., an embryo spent media) is prepared by mixing or coating with solution of an energy-absorbent, organic compound called matrix. When the matrix crystallizes on drying, the sample entrapped within the matrix also co-crystallizes. The sample within the matrix is ionized in an automated mode with a laser beam. Desorption and ionization with the laser beam generates singly protonated ions from analytes in the sample. The protonated ions are then accelerated at a fixed potential, where these separate from each other on the basis of their mass-to-charge ratio (m / z). The charged analytes are then detected and measured using time of flight (TOF) analyzer. During MALDI- TOF analysis, the m / z ratio of an ion is measured by determining the time required for it to travel the length of the flight tube. Based on the TOF information, a characteristic spectrum is generated for analytes in the sample.
[0066] The training mass spectra comprise mass spectra corresponding to euploidy embryos and aneuploidy embryos. The number of training mass spectra employed can vary in different embodiments. In some embodiments, the number of training mass spectra can be, be about, be at least, be at least about, be at most, or be at most about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 300, 350, 400, 500, 600, 700, 800, 900, 1000, 2000, 2500, 5000, 7500, 10000, 25000, 50000, 100000 or a number or a range between any two of these values. The number of euploidy embryos and aneuploidy embryos used for generating the training mass spectra can vary in different embodiments. The number of euploidy embryos and aneuploidyembryos can be the same or different. In some embodiments, the plurality of euploidy embryos used for generating the training mass spectra comprises, comprises about, comprises at least, comprises at least about, comprises at most, or comprises at most about, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, or a number or a range between any two of these values, of the total number of training embryos. In some embodiments, the plurality of aneuploidy embryos used for generating the training mass spectra comprises, comprises about, comprises at least, comprises at least about, comprises at most, or comprises at most about, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, or a number or a range between any two of these values, of the total number of training embryos. In some embodiments, the plurality of embryos used for generating the training mass spectra comprise more euploidy embryos than aneuploidy embryos. Alternatively, in some embodiments, the plurality of embryos used for generating the training mass spectra comprise more aneuploidy embryos than euploidy embryos.
[0067] In some embodiments, the training mass spectra can comprise mass spectra data generated from a plurality of spent embryo media corresponding to embryos collected from different subjects. In some embodiments, the subjects are mammal subjects, or human subjects. In some embodiments, a plurality training mass spectra can comprise at least 100 training mass spectra corresponding to euploidy embryos, at least 100 training mass spectra corresponding to aneuploidy embryos, or both. In some embodiments, the plurality of training mass spectra comprises mass spectra corresponding to mosaic embryos. The number of training mass spectra corresponding to mosaic embryos can vary in different embodiments. In some embodiments, the number of training mass spectra corresponding to mosaic embryos can be, be about, be at least, be at least about, be at most, or be at most about 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, 30, 40, 50 or a number or a range between any two of these values. In some embodiments, the plurality of training mass spectra used herein can comprise at most 10 training mass spectra corresponding to mosaic embryos.
[0068] The method described herein also comprises identifying a plurality of training mass spectrometric peaks from each of the plurality of training mass spectra and merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix. In some embodiments, only peaks having an intensity exceeding a certain threshold are selected. In some embodiments, the intensity threshold used for filtering out low-intensity noise peaks is about 40, 50, 60, 70, 80, 90, or 100 or a number or a range between any two of these values, Dalton. For example, in some embodiments, identifying the plurality of training mass spectrometric peaks from each training mass spectrum comprise selecting training mass spectrometric peaks in the range 50-1500 Dalton. In some embodiments, the method can further comprise identifying a highest intensity peak (i.e., the peak with the greatest signalintensity) in each training mass spectrum, and identifying the plurality of training mass spectrometric peaks from each training mass spectrum comprises selecting peaks with a height equal to or greater than 1% of that of the highest intensity peak. A peak detection algorithm can be used to identify and extract desired peaks from the mass spectrum, typically by analyzing the signal intensity across the mass-to-charge ratio axis and considering factors such as peak width and shape, signal-to-noise ratio, and local maxima as will be understood by a person skilled in the art. In some embodiments, the method can further comprise discarding peaks having a width of less than 1 Dalton and / or a prominence of less than 0.01. Peak prominence is defined as the intensity difference (i.e., vertical distance) between the peak’s height and its adjacent local minima.
[0069] In some embodiments, the method can further comprise preprocessing the mass spectra, e.g., prior to identifying peaks. Preprocessing steps can comprise variance stabilization, applying baseline correction algorithm to remove background noise, applying a smoothing function to reduce noise and improve peak identification (e.g., Savitzky-Golay smoothing), and / or spectra normalization as will be understood by a person skilled in the art.
[0070] In some embodiments, the method comprises normalizing each of the plurality of training mass spectra. Normalization can be used to remove systematic artifacts that affect mass spectral intensity. A number of different normalization approaches can be used herein including, for example, total ion current (TIC) and vector norm. In some embodiments, each training mass spectrum is normalized by its TIC value.
[0071] In some embodiments, the method can further comprise aligning the training mass spectrometric peaks identified from the plurality of training mass spectra. Peak alignment methods typically involve identifying common peaks across multiple spectra within a dataset, then using those common peaks to adjust the m / z axis of each spectrum to align them. Exemplary techniques include, for example, peak picking, cross-correlation, clustering, re-binning, or a reference spectrum-based alignment where a reference spectrum is chosen as a standard for comparison with others in the dataset. In some embodiments, the training mass spectra are aligned to each other with a certain linear mass tolerance. The linear mass tolerance can vary in different embodiments. For example, the linear mass tolerance can be about, at least, at least about, at most, or at most about 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, 500, or a number or a range between any two of these values. In some embodiments, the linear mass tolerance is selected as about 300 ppm.
[0072] In some embodiments, the method further comprises a correlation study between each training mass spectrum with the average of the plurality of training mass spectra in the same category (e.g., euploidy or aneuploidy). In some embodiments, the method can furthercomprise removing one or more outlier training mass spectra from the plurality of training mass spectra, for example, prior to constructing the training peak matrix.
[0073] The method described herein can generate a training peak matrix that can be used as an input to train a classifier for differentiating sample mass spectra corresponding to euploidy embryos from sample mass spectra corresponding to aneuploidy embryos. In some embodiments, a peak matrix can be formed by a number of rows each representing a mass spectrum and a number of columns each representing detected peaks (e.g., aligned peaks). Each cell contains the area of the corresponding peak for each spectrum, with a value of 0 if a peak was not detected in that spectrum.
[0074] In some embodiments, the method can further comprise analysis steps to remove artificial effects, such as artificial differences created by parameters such as fertility center of origin, date of MALDI-TOF analysis run, gender, maternal age, embryo grade, and other batch effects. For example, in some embodiments, prior to euploidy status analysis, principle component analysis can be performed with the training mass spectra.Training classifiers to determine embryo ploidy
[0075] The methods and systems described herein combine machine learning with mass spectral data for rapid metabolomics profiling of embryo spent media. In some embodiments, the machine learning model used herein is a supervised machine learning classifier trained based on the mass spectra data collected from the spent embryo media where embryos with known ploidy had been cultured in. Supervised machine learning classifiers are algorithms used to categorize new data points into predefined classes based on a labeled training dataset, where the model learns the relationship between input features and known output categories. Exemplary supervised machine learning classifiers include, but are not limited to, support vector machines (SVM), Partial-Least Squares Discriminant Analysis (PLS-DA), Random Forest (RF), K-Nearest Neighbour (KNN), Light Gradient-Booster Machine (LightGBM), or a combination thereof.
[0076] In some embodiments, the supervised machine learning classifier is a likelihood model classifier. A computing system can compute a likelihood of a sample embryo being euploidy or aneuploidy by applying the trained machine learning model to the mass spectral data corresponding to the sample embryo. The likelihood model classifier can output a probability score representing the likelihood of the sample data belonging to a predicted class (e.g., euploidy or aneuploidy). In some embodiments, the supervised machine learning model is LightGBM.
[0077] The methods and systems described herein comprise training a classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos using the training peak matrix as input and the correspondingclassifications of the embryos as output. In some embodiments, the training mass spectra further comprise mass spectra corresponding to mosaic embryos. Accordingly, in some embodiments, the classification of an embryo of the embryos comprises euploidy, aneuploidy, or mosaic. Training a classifier comprises training the classifier to further differentiate mass spectra corresponding to mosaic embryos from mass spectra corresponding to euploidy embryos and / or from mass spectra corresponding to aneuploidy embryos.
[0078] A computing system can receive the training peak matrix as an input. The training peak matrix can be constructed from about, at least, at least about, at most, or at most about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 300, 350, 400, 500, 600, 700, 800, 900, 1000, 2000, 2500, 5000, 7500, 10000, 25000, 50000, 100000 or a number or a range between any two of these values of training mass spectra. In some embodiments, the training peak matrix is constructed from a plurality of training mass spectra comprising at least 100 training mass spectra corresponding to euploidy embryos and / or aneuploidy embryos. In some embodiments, the plurality of training mass spectra can also comprise mass spectra corresponding to mosaic embryos. The plurality of training mass spectra can be generated from spent culture media corresponding to a number of embryos extracted from different subjects (e.g., different human subjects). The computing system can train the classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos. In some embodiments, the classifier is LightGBM. In some embodiments, the classifier is RF.
[0079] To determine a sample embryo as being euploidy or aneuploidy, a sample spent embryo medium wherein a sample embryo had been cultured is subjected to MALDI-TOF mass spectrometry to obtain a sample mass spectrum. The computing system can receive the sample mass spectrum as an input and determine the classification of the sample embryo corresponding to the sample mass spectrum using the classifier trained with the training mass spectra. The computing system can determine the ploidy state of the sample embryo by outputting a probability score between 0 and 1, indicating the predicted probability of the sample embryo belonging to a specific classification, i.e., euploidy or aneuploidy or mosaic. For example, the computing system can determine a sample embryo as being euploidy or aneuploidy using the euploidy probability score and / or aneuploidy probability score.
[0080] In some embodiments, to determine a sample embryo as being euploidy or aneuploidy, the computing system can determine the euploidy probability is greater than the aneuploidy probability, and determine the sample embryo as being euploidy. To determine a sample embryo as being euploidy or aneuploidy, the computing system can determine the euploidy probability is smaller than the aneuploidy probability and determine the sample embryo as being aneuploidy. In some embodiments, the computing system can determine the euploidy probability(and / or the aneuploidy probability) is within a range, with a lower bound of 0.4, 0.41, 0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, or 0.49 and a lower bound of 0.51, 0.52, 0.53, 0.54, 0.55, 0.56, 0.57, 0.58, 0.59, or 0.6. The computing system can determine the ploidy of the sample embryo as being undetermined. In some embodiments, a threshold on the probability score can be set to make the classification decision. Sample data points (e.g., a euploidy probability score and / or an aneuploidy probability score) exceeding the threshold are classified as belonging to a specific class. For example, a sample embryo having a euploidy probability score exceeding a threshold can be classified as being euploidy. A sample having an aneuploidy probability score exceeding a threshold can be classified as being aneuploidy.
[0081] In some embodiments, the dataset may be imbalanced when the classification categories are not approximately equally represented. For example, the training mass spectra may be generated from training embryos with a majority of normal embryos (euploidy embryos) and only a small percentage of abnormal embryos (aneuploidy embryos). Therefore, in some embodiments, techniques of over-sampling the minority class (e.g., aneuploidy embryos) can be used to increase the sensitivity of a classifier to the minority class. In some embodiments, techniques of under-sampling the majority class (e.g., euploidy embryos) can be combined with over-sampling the minority class to achieve even better classifier performance. In some embodiments, synthetic minority over-sampling technique (SMOTE) algorithm is used to upsample minority class (e.g., aneuploidy). SMOTE can generate new synthetic examples close to data points belonging to the minority class in the feature space.
[0082] In some embodiments, the methods and systems further comprise generating a second plurality of training mass spectrometric peaks from a plurality of training mass spectrometric peaks from one or more of the pluralities of training mass spectra using synthetic minority over-sampling technique (SMOTE) algorithm. Merging the training mass spectrometric peaks can comprise merging the training mass spectrometric peaks identified from the plurality of training mass spectra and the training mass spectrometric peaks generated using SMOTE algorithm to construct a training peak matrix.
[0083] In some embodiments, the methods and systems further comprise identifying mass-to-charge values representative of euploidy class, an aneuploidy class, and / or a mosaic class. The method can comprise identifying mass spectrometric peaks representing a euploidy embryo, mass spectrometric peaks representing an aneuploidy embryo, and / or mass spectrometric peaks representing a mosaic embryo. In some embodiments, the method can further comprise identifying mass spectrometric peaks with specific m / z values.
[0084] Disclosed herein also includes a computer-implemented method for determining embryo ploidy. In some embodiments, the method can comprise under control of ahardware processor: receiving a plurality of training mass spectra obtained by subjecting a plurality of spent embryo media where embryos had been cultured in to Matrix- Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry, wherein the embryos comprise euploidy embryos and aneuploidy embryos, and wherein each of the embryos is associated with a classification of ploidy comprising euploidy and aneuploidy, identifying a plurality of training mass spectrometric peaks from each of the plurality of training mass spectra;, merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix, training a classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos using the training peak matrix as input and the corresponding classifications of the embryos as output, receiving a sample mass spectrum obtained by subjecting a sample spent embryo medium where a sample embryo had been cultured in to MALDI-TOF mass spectrometry, and determining a classification of the sample embryo using the classifier with the sample mass spectrum as input.
[0085] Disclosed herein also includes a system for determining embryo ploidy. The system can comprise non-transitory memory configured to store executable instructions and a classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos, wherein the classifier is generated by receiving a plurality of training mass spectra obtained by subjecting a plurality of spent embryo media where embryos had been cultured in to Matrix- Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry, wherein the embryos comprise euploidy embryos and aneuploidy embryos, and wherein each of the embryos is associated with a classification of ploidy comprising euploidy and aneuploidy, identifying a plurality of training mass spectrometric peaks from each of the plurality of training mass spectra, merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix, training the classifier using the training peak matrix as input and the corresponding classifications of the embryos as output, and a hardware processor in communication with the non-transitory memory, the hardware processor programmed by the executable instructions to perform receiving a sample mass spectrum obtained by subjecting a sample spent embryo medium where a sample embryo had been cultured in to MALDI-TOF mass spectrometry, and determining a classification of the sample embryo using the classifier with the sample mass spectrum as input.
[0086] In some embodiments, the system can comprise non-transitory memory configured to store executable instructions and a classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos, wherein the classifier is trained using a training peak matrix as input and classifications of embryos as output, wherein a plurality of training mass spectra is obtained by subjecting a plurality of spentembryo media where the embryos had been cultured in to Matrix- Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry, wherein the embryos comprise euploidy embryos and aneuploidy embryos, and wherein each of the embryos is associated with a classification of ploidy comprising euploidy and aneuploidy, wherein a plurality of training mass spectrometric peaks is identified from each of the plurality of training mass spectra, and wherein training mass spectrometric peaks identified from the plurality of training mass spectra are merged to construct the training peak matrix; and a hardware processor in communication with the non-transitory memory, the hardware processor programmed by the executable instructions to perform receiving a sample mass spectrum obtained by subjecting a sample spent embryo medium where a sample embryo had been cultured in to MALDI-TOF mass spectrometry, and determining a classification of the sample embryo using the classifier with the sample mass spectrum as input.
[0087] Metabolomics analysis of the spent embryo culture media wherein embryos had been cultured in can lead to determination of embryo ploidy and can suggest a better selection of viable embryos for in vitro fertilization. In some embodiments, the methods and systems herein described can be used for assessing embryo viability and differentiate viable embryos from non-viable embryos. Determining the embryo viability of the sample embryo can comprise determining whether the sample mass spectrum corresponds to a euploidy embryo or an aneuploidy embryo. In some embodiments, a sample embryo classified as being euploidy can be determined as being viable. A sample embryo classified as being aneuploidy can be determined as being non-viable.Advantages of combining metabolomic analysis with machine learning
[0088] The effectiveness of IVF treatments hinge on embryo viability and implantation success. The current methods for determining these factors including morphological assessment and genetic testing, while effective, do have limitations. These limitations are evident given the suboptimal pregnancy and live birth rates reported for IVF. Additionally, the invasive nature, cost, and associated risks of PGT have led to a paradigm shift towards non-invasive approaches. Most non-invasive techniques focus on qualitative (profiling) or quantitative (measuring specific biomarkers) analysis of SECM. Previous studies of the SECM targeted various molecular classes including transcriptomics, proteomics, metabolomics, and lipidomics. As disclosed herein, the utility of MALDI-TOF MS and machine learning was investigated to elucidate metabolomic patterns indicative of chromosomal abnormality. The data shown herein provide compelling evidence of the efficacy of MALDI-TOF MS and a ML model at determining embryo viability. Out of all the ‘omics’ techniques, the practicality of metabolomic analysis isoptimal for development of a non-invasive embryo screening.
[0089] Previous research regarding non-invasive embryo testing spans a variety of different ‘omics’ approaches including transcriptomics, proteomics, metabolomics, and lipidomics. Transcriptomics aims to survey the various types of RNA molecules produced by cells. The utility of RNA analysis in reproductive medicine has been linked to embryo implantation and reproductive therapeutics. Moreover, there are multiple studies looking at the relationship between RNA and endometrium. Among them, the most prominent focuses on miRNAs and embryo implantation. Despite the promising results, there are some practical hurdles that limit the utility of transcriptomics for screening embryos. Given the labile nature of RNA, low concentration, time consuming extraction, and associated costs of analysis, it is a suboptimal molecular target for embryo screening. Ideally, a molecular target that is easily isolated and abundant would work well for rapid screening.
[0090] Like transcriptomes, a significant amount of proteomics research has been conducted to this point. There have been multiple studies examining the embryonic secretome. The secretome is composed of various proteins expressed and secreted into the extracellular environment. Different protein expression cascades will occur during culture depending on the health of the embryo. Secretome analysis has yielded multiple biomarkers capable of distinguishing which embryos are genetically viable including haptoglobin alpha- 1, apolipoprotein A-l, and lipocalin-1. While these biomarkers are promising, the methodologies used to identify them are not clinically applicable. The digestion, isolation, and concentration of proteins can be quite costly, and these processes are complex and not easily automated. Moreover, serum albumin depletion and sample pooling are not practically affordable for an IVF clinical setting. The clinical utility of targeted proteomics is hindered because of its inability to translate into clinical laboratories. Given these limitations, contemporary techniques are potentially better suited for routine embryo screening.
[0091] Cellular metabolism is a vital part of embryonic development; cultured embryos are constantly utilizing compounds in the media to drive cell division and growth. That process produces numerous metabolites which can be exploited for their predictive ability. Metabolomics studies the end-products of metabolism, enabling the elucidation of downstream events. Within the past decade, there has been a significant increase in studies aimed at surveying the metabolites present in SECM. Initially, individual metabolites were targeted, including pyruvate, glucose, lactate, and a few amino acids, to evaluate blastocyst development and pregnancy. More recently, broader approaches have been implemented examining groups of metabolites to determine embryo viability and implantation potential. Alternatively, lipidomics, which is a specialized subdivision of metabolomics, studies lipid metabolites which have beenimplicated in the determination of pregnancy outcomes. Taken together, these studies highlight the value of exploring the SECM metabolome.
[0092] One of the many advantages of metabolomics is the straightforward sample preparation. Basic liquid-liquid extractions can be employed to isolate different classes of metabolites. Nevertheless, when aiming for optimal clinical practicality certain considerations need to be taken into account. That is why metabolomics profiling is the focal point of our workflow, utilizing MALDI-TOF MS and ML is shown to potentially be a sound solution for rapid metabolite screening. The sample preparation is easily automatable, there is minimal cost, and the technique has a very high throughput. The direct analysis methodology described herein aptly demonstrates the clinical utility of combining MALDI-TOF MS and ML.
[0093] The SECM has a wealth of biomarkers capable of providing predictive patterns that relate to the biological mechanisms involved in embryo development. The metabolites present in SECM can serve as indicators of potential aberrations that may occur during embryo culture. Thus, it can help to distinguish between various outcomes. The results presented herein are very promising to build from it a solution ideally suited for clinical laboratories. MALDI-TOF MS is a cost-efficient technique which has proven success in clinical applications, and its combination with ML produces a powerful tool suited for clinical testing. In particular, this study shows cross-validation accuracy metrics, sensibility and sensitivity which outperform the current state-of-the-art. Additional data also includes the high reproducibility of the test and how it is barely affected by batch effects from samples coming from different conditions. Finally, insight is provided into potential biomarkers present in the signal, whose combination in complex patterns are used by the ML algorithm to discriminate the euploidy status of the embryos. The robustness and accuracy of our predictive ML model can be further improved with a larger cohort. Subsequently, full end-to-end automation can be implemented, therefore significantly increasing sample throughput. With optimization and comprehensive validation, it is highly expected that this approach will be a practical non-invasive solution to determine embryo viability.Determining Metabolic Health
[0094] Disclosed herein includes methods for determining metabolic health of an embryo. Metabolic health of an embryo can be determined using a classifier (or a machine learning model). A plurality of training mass spectra can be obtained using a mass spectrometry from a plurality of spent embryo media where embryos had been cultured in. Each of the embryos can be associated with metabolic health information. A plurality of training mass spectrometric peaks can be identified from each of the plurality of training mass spectra. Training mass spectrometric peaks identified can be merged to construct a training peak matrix. The training massspectrometric peaks identified from the plurality of training mass spectra and / or used construct a training peak matrix can be less than 2000 Daltons (or 1500 Daltons, 1600 Daltons, 1700 Daltons, 1800 Daltons, 1900 Daltons, 2000 Daltons, 2100 Daltons, 2200 Daltons, 2300 Daltons, 2400 Daltons, 2500 Daltons or a number or a range between any of these two values). The training mass spectrometric peaks can be more than 500 Daltons (or 100 Daltons, 200 Daltons, 300 Daltons, 400 Daltons, 500 Daltons, 600 Daltons, 700 Daltons, 800 Daltons, 900 Daltons, 1000 Daltons, or a number or a range between any of these two values). A classifier for determining metabolic health information of embryos can be trained using the training peak matrix as input and the corresponding metabolic health information of the embryos as output.
[0095] Metabolic health of an embryo can be correlated with whether the embryo is a euploidy embryo or an aneuploidy embryo. Ploidy of an embryo can be a proxy of metabolic health. In some embodiments, the metabolic health information comprises with a classification of ploidy comprising euploidy and aneuploidy.
[0096] In some embodiments, the metabolic health information comprises a metabolic health classification. The metabolic health classification can comprise a healthy classification and an unhealthy classification. In some embodiments, the embryos comprise one or more euploidy embryos each with the healthy classification. The embryos can comprise one or more aneuploidy embryos with the unhealthy classification. In some embodiments, a euploidy embryo of the embryos is more likely to be associated with the healthy classification. An aneuploidy embryo of the embryos can be more likely to be associated with the unhealthy classification. In some embodiments, an embryo with the healthy classification is more likely to be a euploidy embryo. An embryo with the unhealthy classification is more likely to be an aneuploidy embryo.
[0097] In some embodiments, the metabolic health information comprises a metabolic health score. In some embodiments, a euploidy embryos can have a metabolic health score above a threshold (or a higher metabolic health score such as 60%, 70%, 80%, 90%, or higher). The threshold can be 50%, 60%, 70%, 80%, 90%, or a number or a range between any two of these values. A aneuploidy embryo can have a metabolic health score below a threshold (or a lower metabolic health score such as 70%, 60%, 50%, 40%, 30%, or lower). The threshold can be 70%, 60%, 50%, 40%, 30%, or a number or a range between any two of these values). In some embodiments, a euploidy embryo is more likely to have a metabolic health score above a threshold (or a higher metabolic health score). An aneuploidy embryo can be more likely to have a metabolic health score below a threshold (or a lower metabolic health score). In some embodiments, an embryo with a metabolic health score above a threshold (or a higher metabolic health score) is more likely to be a euploidy embryo. An embryo with a metabolic health score below a threshold (or a lower metabolic health score) is more likely to be an aneuploidy embryo.
[0098] A sample mass spectrum can be obtained using the mass spectrometry from a sample of spent embryo medium where a sample embryo had been cultured in. Metabolic health information of the sample embryo can be determined using the classifier with the sample mass spectrum as input. In some embodiments, the metabolic health classification of the sample embryo is the healthy classification and the sample embryo is a euploidy embryo. In some embodiments, the metabolic health classification of the sample embryo is the unhealthy classification and the sample embryo is an aneuploidy. In some embodiments, the metabolic health score of the sample embryo is above a threshold (or the metabolic health score of the sample embryo is higher) and the sample embryo is (or is more likely to be) a euploidy embryo. In some embodiments, the metabolic health score of the sample embryo is below a threshold (or the metabolic health score of the sample embryo is lower) and the sample embryo is (or is more likely to be) an aneuploidy.Mass Spectrometry
[0099] In some embodiments, the mass spectrometry for spent embryo media analysis can be, for example, matrix-assisted laser desorption / ionisation time-of-flight (MALDI-TOF) mass spectrometry, electrospray ionization mass spectrometry (ESI-MS), ESI quadrupole orthogonal TOF (ESI-QTOF) mass spectrometry, ESI Fourier transform mass spectrometry (ESI-FTMS), ion trap mass spectrometry, triple quadrupole mass spectrometry (TQMS), ion mobility spectrometry-mass spectrometry (IMS-MS), Fourier transform ion cyclotron resonance mass spectrometry (FT-ICR MS), surface-enhanced laser desorption / ionization time-of-flight mass spectrometry (SELDI-TOF MS), desorption / ionization on silicon mass spectrometry (DIOS-MS), secondary ion mass spectrometry (SIMS), atmospheric pressure chemical ionization mass spectrometry (APCI-MS), atmospheric pressure photoionization mass spectrometry (APPI MS), gas chromatography -mass spectrometry (GC-MS), liquid chromatograph-mass spectrometry (LC-MS), inductively couples plasma mass spectrometry (ICP-MS), and / or tandem mass spectrometry (MS / MS).EXAMPLES
[0100] Some aspects of the embodiments discussed above are disclosed in further detail in the following examples, which are not in any way intended to limit the scope of the present disclosure.Example 1Sample preparation and spectral acquisition
[0101] This example describes the process of preparing spent media samples for MALDI-TOF mass spectrometer, spectral acquisition and data pre-processing.
[0102] A number (N=120) of spent media samples were analyzed by MALDI-TOF MS. All samples were stored at -20°C prior to sample preparation. All samples had been previously characterized using PGT-A testing via Next-Generation Sequencing. Before the samples were prepared, the calibration mixture and MALDI matrix were solubilized in diluent. The MALDI matrix used for these experiments was alpha-cyano-4-hydroxycinnamic acid (CHC A) Millipore Sigma (St. Louis, MO); it was resuspended in 80% acetonitrile, 20% deionized water, and 0.1% trifluoroacetic acid to a final concentration of 5 mg / mL. The calibration mixture Spherical Aqua Neat Peptide Low (300 - 1000 Da) Polymer Factory (Stockholm, Sweden) was resuspended in 50 mL of the calibration diluent, 50% acetonitrile and 50% deionized water. Subsequently, the samples were thawed and diluted in a 1:5 ratio with the dilution diluent, 30% acetonitrile, 70% deionized water, and 0.1% trifluoroacetic acid. After dilution the samples were vortexed for 60 seconds and placed at 4°C. Once all assay samples and reagents were prepared the three calibration spots were spotted in a 1:1 ratio with the MALDI matrix on the MALDI plate. The samples were then spotted next in a 1:1 ratio with the MALDI matrix; for each sample three technical replicates were spotted on the MALDI plate. The MALDI plate was then dried in the open air for 5 to 10 minutes. After verifying the MALDI plate was completely dried, it was placed into the MALDI-TOF mass spectrometer, and the vacuum was pumped down.
[0103] The MALDI-TOF instrument used for this study was a Shimadzu M ALDI-8020, Shimadzu Scientific Instruments (Columbia, MD). This MALDI-TOF instrument operates in linear mode and detects positive ions. The instrument operation software used to acquire the data was MALDI Solutions, Kratos Analytical (Manchester, U.K.). All spent media and calibration samples were spotted on the Fleximass DS MALDI plates; 1x48 2.5 mm diameter sample well, Kratos Analytical (Manchester, U.K.). The mass range of analysis was 500 - 2,000 Da. The data processing workflow consisted of baseline subtraction (Top Hat algorithm, factor 0.02) and smoothing (Savitzky-Golay algorithm with polynomial order 3 and a window length of 11 values). All spectra were then exported in the ASCII file format to begin the preparation of the dataset for ML training.Example 2Data analysis with machine learning
[0104] In order to train a supervised machine learning algorithm to discriminate between euploid, aneuploid and mosaic (mix of euploid and aneuploid cells) embryos, the original spectra were separated into categories: Euploid (NEUP = 75), Aneuploid (NANE = 43) and Mosaic (NMOS = 2). FIG. 1 shows a graphical abstract of the processes of analyses followed and can be used as a visual reference to understand the processes followed during this study and explained in detail during this section.Initial analyses
[0105] As a first step before model training, a reproducibility analysis was performed to study the extent to which the spectra in each category were similar to one another. Peaks of all spectra within a category were aligned with a constant tolerance of 1 Da. The analysis consisted of a correlation study between each spectra with the average of all spectra in its corresponding category. FIG. 2 shows the violin plot resulting from this analysis, with blue used to show correlation in the Euploidy category, yellow in the Aneuploidy, and green in the Mosaic. A medium level of internal correlation can be noticed overall in all three categories, indicated by the relatively flat appearance of the violin forms, with D-indices not higher than 100. However, the high vertical line in the Euploidy category corresponds to spectra showing a very high D-index, above 100. The top three dots are all corresponding to replicates of a single sample (accession number 70909AS1), and upon visual inspection it was confirmed the low quality of the resulting spectra for that sample. As a conclusion, that sample was taken out of all subsequent studies as an outlier.
[0106] The rest of the samples (n = 357, corresponding to 119 samples with three technical replicates each) were submitted to a preprocessing workflow in order to construct a peak matrix to be used as input to supervised machine learning algorithms. A peak matrix is formed by a row per spectra and aligned detected peaks in the columns, with each cell containing the area of the corresponding peak for each spectrum, with a value of 0 if a peak was not detected in a spectrum.
[0107] The workflow to prepare the peak matrix is the following: range of all spectra was cropped to 50 - 1500 Dalton (Da), because no significant peak was seen beyond that range. After that, technical replicates were averaged to form a single spectrum per sample (n = 119). Subsequently, all spectra are aligned to each other with a linear mass tolerance of 300 ppm. After that, each spectrum is normalized by its Total Ion Current (TIC) value. Then, a peak detection algorithm is used to find, in each spectrum, every peak with a height not lower than 1% of that of the highest peak. Peaks with a width of less than 1 Da and a prominence of less than 0.01 were discarded. After this process, all spectra peaks were merged into the peak matrix, again considering peaks within IDa tolerance as the same one.
[0108] Before the euploid status analyses, a PCA analysis was performed, and other categories were overlaid in the results. This is done to detect the presence of any batch effect, meaning artificial differences created by other parameters which can hinder the classification. This was done for the fertility center of origin, date of MALDI-TOF analysis run, gender, maternal age and embryo grade. None of them showed a very clear batch effect. The fertility center had a very minor degree of batch effect. This can be seen in FIG. 3, where two of the ellipses are almostconcentric (no batch effect whatsoever) while the other two show some skew toward orthogonal directions, but with a small eccentricity. Therefore, no batch effect removal step was added in this example.Batery of Supervised Classifiers
[0109] After these preliminary analyses, the training of supervised classifiers was performed to try and differentiate the euploidy status of each sample. The first step was to study the degree of imbalance of each category. After averaging replicates, n=74 sample spectra remained in the Euploidy category, n=43 in the Aneuploidy and only n=2 in the Abnormal Mosaic category. This is a very significant degree of imbalance, with the added obstacle that the number of samples in the Mosaic category is extremely reduced, not being even enough to run algorithms to reduce the imbalance, such as SMOTE. As a consequence, a new peak matrix was constructed with the same steps and parameters as explained above but skipping the averaging of replicates. In this way, the new number of samples per category increased to n=222 sample spectra remained in the Euploidy category, n=129 in the Aneuploidy and only n=6 in the Abnormal Mosaic. All ML algorithms applied to this new peak matrix for this study involve the use of the SMOTE algorithm to reduce imbalance. The best results of 10-fold Cross-Validation (CV) for training with the different supervised ML algorithms are shown in Table 1.Table 1. Exemplary ML algorithms used for classificationAlgorithm Number of samples correctly identified Balanced(% accuracy) accuracy Euploidy Aneuploidy Mosaic(n = 222) (n = 185 222) (n = 6 222)PLS-DA 119 (53.6%) 112 (50.45%) 222 (100%) 68.02%SVM 171 (77.03%) 181 (81.53%) 222 (100%) 86.19%RF 193 (86.94%) 180 (81.08%) 221 (99.55%) 89.19%KNN 145 (65.32%) 183 (82.43%) 222 (100%) 82.58% LightGBM 198 (89.19%) 188 (84.68%) 221 (99.55%) 91.14%
[0110] All algorithms were run using the SMOTE algorithm to reduce data imbalance.10-fold CV means that 10% of the data from each category is removed from each run and used blindly to analyze the accuracy of the algorithm, and this is repeated 10 times to obtain the final accuracy values. Hyperparameters for every algorithm were optimized using an exploratory method. In the table, PLS-DA corresponds to Partial Least Squares-Discriminant Analysis, SVM corresponds to Support Vector Machine, RF corresponds to Random Forest, KNN is for K-nearestneighbors, and LightGBM stands for Light Gradient-Booster Machine. Columns from second left to fourth correspond to: best accuracy achieved for the euploidy, aneuploidy and mosaic categories, and the last column shows the balanced accuracy for all three categories. A bold font indicates the best value achieved for each accuracy.[OHl] As seen from Table 1, a high balanced accuracy in CV was achieved, which means not only the training provides a high discrimination between categories but also a satisfactory validation with blind samples. The best performance was obtained with LightGBM (91.14%), followed by RF (89.19%) and SVM (86.19%). Therefore, LightGBM was chosen as the algorithm with higher potential to determine euploidy status using metabolomic data from MALDLTOF MS.
[0112] The better performance of LightGBM compared to other ML algorithms is no surprise, as this algorithm typically shows excellent results with sparse data like the spectra involved in this study, which have a large number of features but very few of them being actually relevant. This algorithm is also a good choice for other reasons, such as its efficient use of memory, its speed compared to other complex algorithms or its ability to handle large datasets.LightGBM results
[0113] Table 2 shows more detailed metrics of the LightGBM results when each of the categories are selected as the positive category.Table 2, Detailed metrics of the LightGBM results Metric / Category Euploidy Aneuploidy Mosaic Accuracy (acc.) 91.14% 91.29% 99.85%Balanced acc. 90.65% 86.94% 99.77%Fl score 87.03% 86.63% 99.77%Sensitivity 89.19% 84.68% 99.55%Specificity 92.12% 94.59% 100%Error rate 8.86% 8.71% 0.15%PPV (precision) 84.98% 88.68% 100%NPV 94.46% 92.51% 99.78%
[0114] In Table 2, accuracy indicates how often the classifier predicts correctly the outcome. The values of the optimized hyperparameters are 100 estimators, 0.1 learning rate and 32 leaves. Sensitivity is the proportion of samples which actually belong to the category and are classified as such and indicates how well a test can classify samples who truly have the outcome of interest. Specificity denotes the proportion of samples which do not belong to the category andare classified correctly and indicates how well a test can classify samples who truly do not have the outcome of interest. Balanced accuracy is the average between sensitivity and specificity. Positive Predictive Value (PPV), also called precision, shows the proportion of samples with a positive test result who truly have the category of interest, while Negative Predictive Value (NPV) is the proportion of samples with a negative test result who truly do not belong to the category of interest. Fl score is the harmonic mean of precision and sensitivity. Finally, the error rate is a measure of the prediction error of a model.
[0115] Traditionally, the positive category in diagnostics is chosen to be the one indicating abnormality, so the analysis of the results of Table 2 focuses on the data for aneuploidy and mosaic status. Firstly, it is relevant to mention that the mosaic prediction shows the most reliable results, but this is also, by far, the category less represented in the real data, and therefore where the SMOTE algorithm has created more artificial samples for balancing the training set. Therefore, strong claims cannot be taken until further research with more data in this category is used. However, the good results of this few data are very promising, and visual inspection of the spectra, as well as other statistical analyses (not shown) provide reasons for optimism towards the possibility of these good results being maintained in further studies. Secondly, by looking at the sensitivity and specificity, it can be seen that this test is highly sensitive (92.16% averaged between both positive categories), meaning very few false negative results therefore very few cases of aneuploidy are missed. The model also shows high specificity (97.30% averaged), meaning very few false positives. Note that specificity is particularly high, and this is a very promising result as screening methods are normally designed to maximize specificity. The results shown in Table 2 outperform the state of the art in the previous metabolomic assays, especially in the NPV and specificity metrics.
[0116] FIG. 4, top panel shows the Receiver Operating Characteristic (ROC) curves for the three categories (Euploidy, Aneuploidy and Mosaic) in the optimized LightGBM trained classifier. These curves show the trade-off between sensitivity (True Positive Rate or TPR) and specificity (False Positive Rate or FPR), measured as 1 - TPR. A random classifier would be expected to result in a diagonal line. All three categories show a high rate between true positive and false positive rates, indicated by the proximity of the Youden indices (red dots) to the top-left corner. This can also be seen in the Area Under the Curve (AUC) value, which is above 0.95 for all categories. FIG. 4, bottom panel shows the Precision-Recall (PR) curves for all three categories. AUC and Averaged Precision (AP) values are very close to 1 in all cases.
[0117] FIG. 5 shows the feature importance from the results of the LightGBM algorithm, which signals those peaks picked up by the algorithm as being important to distinguish among the categories. The higher the amplitude of an x-loading in the feature importance plot, themore important it is for discrimination.Random Forest results
[0118] According to Table 1, Random Forest (RF) shows comparable performance data with LightGBM. Detailed analysis of the RF results were performed in this section.
[0119] The hyperparameters optimized for this method are: 400 estimators, 16 maximum features, 10 maximum depth, 2 minimum split size, 1 minimum samples per leaf. FIG.6 shows the multidimensional scaling plot, or distance plot of the spectra in the study. The distance plot shows the distance between spectra after running the analysis. It is noticeable how the majority of samples of both euploidy and aneuploidy status have a shorter distance to spectra in their own category. However, some misclassifications are also visible in both groups. Future studies can include correlating these misclassifications with the actual outcome of pregnancies to obtain further insights in the real correlation between euploidy status and successful pregnancies.
[0120] FIG. 7 shows the Shapley values for the Euploidy (top), Aneuploidy (center) and Mosaic (bottom) categories. These values, originated from game theory, indicate the average marginal contribution of a feature considering all possible combinations. In other words, it reveals the m / z values of the peaks which are more relevant to discriminate between the different categories. These m / z values can potentially be used as biomarker candidates if shown in one category and not in the others. The data suggests that the m / z value 152.34 is very important for all three categories. However, Shapley values for this peak show that the area is predominantly low in both Euploidy and Aneuploidy categories, while it is predominantly high for the Mosaic category. This means that it is very likely that a high area at the m / z 152.34 peak is important for the identification of the Mosaic category. A similar analysis reveals the combination of high and low areas which the algorithm figured out in order to differentiate between the three categories. It is clear that there are more significant differences between the Shapley values obtained for the Mosaic category compared to the other two, which explains why Mosaic can be so well discriminated against. Note that, comparing the peaks with high Shapley values with the feature importance of FIG. 5 reveals that some of the peaks are the same, which confirms their importance as discriminators. Examples of these are m / z 65.31, 114.36, 152.34, 159.12, or 671.70.
[0121] Finally, FIG. 8 shows the Heatmap values indicating the distance between each pair of samples after the classification with the RF algorithm. Note that the three blocks of distances within the same category are clearly emphasized, meaning they can be grouped together to a significant degree as confirmed by the results shown above.Execution Environment
[0122] FIG. 9 depicts a general architecture of an example computing device 900 that can be used in some embodiments to execute the processes and implement the features described herein. The general architecture of the computing device 900 depicted in FIG. 9 includes an arrangement of computer hardware and software components. The computing device 900 may include many more (or fewer) elements than those shown in FIG. 9. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the computing device 900 includes a processing unit 910, a network interface 920, a computer readable medium drive 930, an input / output device interface 940, a display 950, and an input device 960, all of which may communicate with one another by way of a communication bus. The network interface 920 may provide connectivity to one or more networks or computing systems. The processing unit 910 may thus receive information and instructions from other computing systems or services via a network. The processing unit 910 may also communicate to and from memory 970 and further provide output information for an optional display 950 via the input / output device interface 940. The input / output device interface 940 may also accept input from the optional input device 960, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.
[0123] The memory 970 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 910 executes in order to implement one or more embodiments. The memory 970 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer-readable media. The memory 970 may store an operating system 972 that provides computer program instructions for use by the processing unit 910 in the general administration and operation of the computing device 900. The memory 970 may further include computer program instructions and other information for implementing aspects of the present disclosure.
[0124] For example, in some embodiments, the memory 970 includes a training module 974 for training a machine learning model, such as a supervised machine learning model, for determining embryo ploidy. In some embodiments, the memory 970 includes a determination module 976 for determining metabolic health of embryos using a machine learning model. In some embodiments, the memory 970 includes a determination module 976 for determining embryo ploidy using a machine learning model.
[0125] In addition, memory 970 may include or communicate with the data store 990 and / or one or more other data stores that store the machine learning model, weights of the machinelearning model (during one or more iterations of training or when trained), a plurality of training mass spectra, mass spectrometric peaks, a training peak matrix, the mass spectrum for a sample embryo for which the embryo ploidy is being determined.Additional Considerations
[0126] In at least some of the previously described embodiments, one or more elements used in an embodiment can interchangeably be used in another embodiment unless such a replacement is not technically feasible. It will be appreciated by those skilled in the art that various other omissions, additions and modifications may be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and changes are intended to fall within the scope of the subject matter, as defined by the appended claims.
[0127] One skilled in the art will appreciate that, for this and other processes and methods disclosed herein, the functions performed in the processes and methods can be implemented in differing order. Furthermore, the outlined steps and operations are only provided as examples, and some of the steps and operations can be optional, combined into fewer steps and operations, or expanded into additional steps and operations without detracting from the essence of the disclosed embodiments.
[0128] With respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Accordingly, phrases such as “a device configured to” are intended to include one or more recited devices. Such one or more recited devices can also be collectively configured to carry out the stated recitations. For example, “a processor configured to carry out recitations A, B and C can include a first processor configured to carry out recitation A and working in conjunction with a second processor configured to carry out recitations B and C. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.
[0129] It will be understood by those within the art that, in general, terms used herein, and especially in the appended claims (e.g., bodies of the appended claims) are generally intended as “open” terms (e.g., the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” the term “includes” should be interpreted as “includes but is not limited to,” etc.). It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent willbe explicitly recited in the claim, and in the absence of such recitation no such intent is present. For example, as an aid to understanding, the following appended claims may contain usage of the introductory phrases “at least one” and “one or more” to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles “a” or “an” limits any particular claim containing such introduced claim recitation to embodiments containing only one such recitation, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an” (e.g., “a” and / or “an” should be interpreted to mean “at least one” or “one or more”); the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., the bare recitation of “two recitations,” without other modifiers, means at least two recitations, or two or more recitations). Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “ a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In those instances where a convention analogous to “at least one of A, B, or C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “ a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B.”
[0130] In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.
[0131] As will be understood by one skilled in the art, for any and all purposes, such as in terms of providing a written description, all ranges disclosed herein also encompass any and all possible sub-ranges and combinations of sub-ranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As willalso be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like include the number recited and refer to ranges which can be subsequently broken down into sub-ranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 articles refers to groups having 1, 2, or 3 articles. Similarly, a group having 1-5 articles refers to groups having 1, 2, 3, 4, or 5 articles, and so forth.
[0132] It will be appreciated that various embodiments of the present disclosure have been described herein for purposes of illustration, and that various modifications may be made without departing from the scope and spirit of the present disclosure. Accordingly, the various embodiments disclosed herein are not intended to be limiting, with the true scope and spirit being indicated by the following claims.
[0133] It is to be understood that not necessarily all objects or advantages may be achieved in accordance with any particular embodiment described herein. Thus, for example, those skilled in the art will recognize that certain embodiments may be configured to operate in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other objects or advantages as may be taught or suggested herein.
[0134] All of the processes described herein may be embodied in, and fully automated via, software code modules executed by a computing system that includes one or more computers or processors. The code modules may be stored in any type of non-transitory computer-readable medium or other computer storage device. Some or all the methods may be embodied in specialized computer hardware.
[0135] Many other variations than those described herein will be apparent from this disclosure. For example, depending on the embodiment, certain acts, events, or functions of any of the algorithms described herein can be performed in a different sequence, can be added, merged, or left out altogether (for example, not all described acts or events are necessary for the practice of the algorithms). Moreover, in certain embodiments, acts or events can be performed concurrently, for example through multi -threaded processing, interrupt processing, or multiple processors or processor cores or on other parallel architectures, rather than sequentially. In addition, different tasks or processes can be performed by different machines and / or computing systems that can function together.
[0136] The various illustrative logical blocks and modules described in connection with the embodiments disclosed herein can be implemented or performed by a machine, such as a processing unit or processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereofdesigned to perform the functions described herein. A processor can be a microprocessor, but in the alternative, the processor can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor can include electrical circuitry configured to process computer-executable instructions. In another embodiment, a processor includes an FPGA or other programmable device that performs logic operations without processing computer-executable instructions. A processor can also be implemented as a combination of computing devices, for example a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Although described herein primarily with respect to digital technology, a processor may also include primarily analog components. For example, some or all of the signal processing algorithms described herein may be implemented in analog circuitry or mixed analog and digital circuitry. A computing environment can include any type of computer system, including, but not limited to, a computer system based on a microprocessor, a mainframe computer, a digital signal processor, a portable computing device, a device controller, or a computational engine within an appliance, to name a few.
[0137] Any process descriptions, elements or blocks in the flow diagrams described herein and / or depicted in the attached figures should be understood as potentially representing modules, segments, or portions of code which include one or more executable instructions for implementing specific logical functions or elements in the process. Alternate implementations are included within the scope of the embodiments described herein in which elements or functions may be deleted, executed out of order from that shown, or discussed, including substantially concurrently or in reverse order, depending on the functionality involved as would be understood by those skilled in the art.
[0138] It should be emphasized that many variations and modifications may be made to the above-described embodiments, the elements of which are to be understood as being among other acceptable examples. All such modifications and variations are intended to be included herein within the scope of this disclosure and protected by the following claims.
Claims
WHAT IS CLAIMED IS:
1. A method for determining metabolic health of an embryo, comprising:obtaining a plurality of training mass spectra using a mass spectrometry from a plurality of spent embryo media where embryos had been cultured in, wherein each of the embryos is associated with a metabolic health classification, wherein the metabolic health classification comprises a healthy classification and an unhealthy classification;identifying a plurality of training mass spectrometric peaks less than 2000 Daltons from each of the plurality of training mass spectra;merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix;training a classifier for differentiating mass spectra corresponding to embryos with the healthy classification from mass spectra corresponding to embryos with the unhealthy classification using the training peak matrix as input and the corresponding metabolic health classifications of the embryos as output;obtaining a sample mass spectrum using the mass spectrometry from a sample of spent embryo medium where a sample embryo had been cultured in;determining a metabolic health classification of the sample embryo using the classifier with the sample mass spectrum as input.
2. The method of claim 1, wherein the embryos comprise one or more euploidy embryos each with the healthy classification, and wherein the embryos comprise one or more aneuploidy embryos with the unhealthy classification.
3. The method of claim 1, wherein a euploidy embryo of the embryos is more likely to be associated with the healthy classification, and wherein an aneuploidy embryo of the embryos is more likely to be associated with the unhealthy classification.
4. The method of claim 1, wherein an embryo with the healthy classification is more likely to be a euploidy embryo, and wherein an embryo with the unhealthy classification is more likely to be an aneuploidy embryo.
5. The method of claim 1, wherein the metabolic health classification of the sample embryo is the healthy classification and the sample embryo is a euploidy embryo, and / or wherein the metabolic health classification of the sample embryo is the unhealthy classification and the sample embryo is an aneuploidy.
6. A method for determining metabolic health of an embryo, comprising:obtaining a plurality of training mass spectra using a mass spectrometry from a plurality of spent embryo media where embryos had been cultured in, wherein each of the embryos is associated with a metabolic health score;identifying a plurality of training mass spectrometric peaks less than 2000 Daltons from each of the plurality of training mass spectra;merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix;training a classifier for determining metabolic health scores of embryos using the training peak matrix as input and the corresponding metabolic health scores of the embryos as output;obtaining a sample mass spectrum using the mass spectrometry from a sample of spent embryo medium where a sample embryo had been cultured in;determining a metabolic health score of the sample embryo using the classifier with the sample mass spectrum as input.
7. The method of claim 6, wherein the embryos comprise one or more euploidy embryos each with a metabolic health score above a threshold, and wherein the embryos comprise one or more aneuploidy embryos with a metabolic health score below a threshold, optionally wherein the threshold is 50%, 60%, 70%, 80%, or 90%.
8. The method of claim 6, wherein a euploidy embryo of the embryos is more likely to be associated with a higher metabolic health score, and wherein an aneuploidy embryo of the embryos is more likely to be associated with a lower metabolic health score.
9. The method of claim 6, wherein an embryo with a higher metabolic heath score is more likely to be a euploidy embryo, and wherein an embryo with a lower metabolic health score is more likely to be an aneuploidy embryo.
10. The method of claim 6, wherein the metabolic health score of the sample embryo is above a threshold and the sample embryo is a euploidy embryo, and / or wherein metabolic health score of the sample embryo is below a threshold and the sample embryo is an aneuploidy, optionally wherein the threshold is 50%, 60%, 70%, 80%, or 90%.
11. A method for determining metabolic health of an embryo, comprising:obtaining a plurality of training mass spectra using a mass spectrometry from a plurality of spent embryo media where embryos had been cultured in, wherein each of the embryos is associated with metabolic health information;identifying a plurality of training mass spectrometric peaks less than 2000 Daltons from each of the plurality of training mass spectra;merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix;training a classifier for determining metabolic health information of embryos using the training peak matrix as input and the corresponding metabolic health information of the embryos as output;obtaining a sample mass spectrum using the mass spectrometry from a sample of spent embryo medium where a sample embryo had been cultured in;determining metabolic health information of the sample embryo using the classifier with the sample mass spectrum as input.
12. A method for determining metabolic health of an embryo, comprising:obtaining a plurality of training mass spectra using a mass spectrometry from a plurality of spent embryo media where embryos had been cultured in, wherein each of the embryos is associated with metabolic health information;merging training mass spectrometric peaks from the plurality of training mass spectra to construct a training peak matrix, wherein the training mass spectrometric peaks are less than 2000 Daltons;training a classifier for determining metabolic health information of embryos using the training peak matrix as input and the corresponding metabolic health information of the embryos as output;obtaining a sample mass spectrum using the mass spectrometry from a sample of spent embryo medium where a sample embryo had been cultured in;determining metabolic health information of the sample embryo using the classifier with the sample mass spectrum as input.
13. The method of any one of claims 11-12, wherein the metabolic health information comprises a metabolic health classification, wherein the metabolic health classification comprises a healthy classification and an unhealthy classification.
14. The method of any one of claims 11-12, wherein the metabolic health information comprises a metabolic health score.
15. The method of any one of claims 11-12, wherein the metabolic health information comprises with a classification of ploidy comprising euploidy and aneuploidy.
16. A method for determining embryo ploidy, comprising:obtaining a plurality of training mass spectra using Matrix- Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry from a plurality of spent embryo media where embryos had been cultured in, wherein the embryos comprise euploidy embryos and aneuploidy embryos, and wherein each of the embryos is associated with a classification of ploidy comprising euploidy and aneuploidy;identifying a plurality of training mass spectrometric peaks from each of the plurality of training mass spectra;merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix;training a classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos using the training peak matrix as input and the corresponding classifications of the embryos as output;obtaining a sample mass spectrum using MALD-TOF mass spectrometry from a sample of spent embryo medium where a sample embryo had been cultured in;determining a classification of the sample embryo using the classifier with the sample mass spectrum as input.
17. A method for determining embryo ploidy, comprising:obtaining a plurality of first training mass spectra using Matrix- Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry from a plurality of first spent embryo media where first embryos had been cultured in, wherein the first embryos were cultured in a first medium;obtaining a plurality of second training mass spectra using MALDI-TOF mass spectrometry from a plurality of second spent embryo media where second embryos had been cultured in, wherein the second embryos were cultured in a second medium, wherein the first medium and the second medium were different, wherein the first embryos and the second embryos both comprise euploidy embryos and aneuploidy embryos, and wherein each of the first embryos and each of the second embryos is associated with a classification of ploidy comprising euploidy and aneuploidy;for the plurality of first training mass spectra and for the plurality of second training mass spectra:identifying a plurality of first or second training mass spectrometric peaks from each of the plurality of first or second training mass spectra;merging first or second training mass spectrometric peaks identified from the plurality of first or second training mass spectra to construct a first or second training peak matrix;training a first or second classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos using the first or second training peak matrix as input and the corresponding classifications of the first or second embryos as output;obtaining a sample mass spectrum using MALD-TOF mass spectrometry from a sample of spent embryo medium where a sample embryo had been cultured in, wherein the sample embryo was cultured in the first medium or the second medium; andif the sample embryo was cultured in the first medium, determining a classification of the sample embryo using the first classifier with the sample mass spectrum as input, or if the sample embryo was cultured in the second medium, determining a classification of the sample embryo using the second classifier with the sample mass spectrum as input.
18. A method for determining embryo ploidy, comprising:under control of a hardware processor:receiving a plurality of training mass spectra obtained by subjecting a plurality of spent embryo media where embryos had been cultured into Matrix-Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry, wherein the embryos comprise euploidy embryos and aneuploidy embryos, and wherein each of the embryos is associated with a classification of ploidy comprising euploidy and aneuploidy;identifying a plurality of training mass spectrometric peaks from each of the plurality of training mass spectra;merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix;training a classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos using the training peak matrix as input and the corresponding classifications of the embryos as output;receiving a sample mass spectrum obtained by subjecting a sample spent embryo medium where a sample embryo had been cultured in to MALDI-TOF mass spectrometry; anddetermining a classification of the sample embryo using the classifier with the sample mass spectrum as input.
19. The method of any one of claims 1-18, further comprising providing the plurality of spent embryo media, providing the test spent embryo media, or both.
20. The method of any one of claims 1-18, further comprising receiving the plurality of spent embryo media, receiving the test spent embryo media, or both.
21. The method of any one of claims 1-18, further comprising culturing one or more of the embryos to generate one or more of the plurality of spent embryo media, and / or culturing the sample embryo to obtain the sample spent embryo medium.
22. The method of any one of claims 1 and 18-21, wherein an embryo of the embryos and the sample embryo were cultured in media that were identical.
23. The method of any one of claims 1 and 18-21, wherein an embryo of the embryos and the sample embryo were cultured in media that were different.
24. The method of any one of claims 1 and 18-23, wherein the embryos were cultured in media that were identical.
25. The method of any one of claims 1 and 18-23, wherein the embryos were cultured in media that were different.
26. The method of any one of claims 1-25, wherein the classification of the sample embryo is used to determine embryo viability of the sample embryo.
27. The method of any one of claims 1-26, further comprising: determining the embryo viability of the sample embryo.
28. The method of claim 27, wherein determining the embryo viability of the sample embryo comprises determining the embryo viability of the sample embryo using the classification of the sample embryo.
29. The method of claim 27, wherein determining the embryo viability of the sample embryo comprises determining whether the sample mass spectrum corresponds to that of a euploidy embryo or an aneuploidy embryo.
30. The method of any one of claims 1-29, wherein the plurality of training mass spectra comprises at least 100 training mass spectra corresponding to euploidy embryos and / or at least 100 training mass spectra corresponding to aneuploidy embryos.
31. The method of any one of claims 1-30, wherein the plurality of training mass spectra comprises at most 10 training mass spectra corresponding to mosaic embryos.
32. The method of any one of claims 1-31, wherein the plurality of training mass spectrometric peaks comprises metabolite peaks.
33. The method of any one of claims 1-32, wherein identifying the plurality of training mass spectrometric peaks from each training mass spectrum comprise selecting training mass spectrometric peaks in the range 50-1500 Daltons.
34. The method of any one of claims 1-33, further comprising aligning the training mass spectrometric peaks identified from the plurality of training mass spectra.
35. The method of any one of claims 1-34, further comprising normalizing each of the plurality of training mass spectra.
36. The method of any one of claims 1-35, wherein identifying the plurality of training mass spectrometric peaks from each training mass spectrum comprises selecting peaks with a height equal to or greater than 1% of that of a highest peak.
37. The method of claim 36, comprising discarding peaks having a width of less than 1 Dalton and a prominence of less than 0.01.
38. The method of any one of claims 1-37, wherein the classifier is selected from the group consisting of: Partial-Least Squares Discriminant Analysis (PLS-DA), Support Vector Machine (SVM), Random Forest (RF), K-Nearest Neighbour (KNN), Light Gradient-Booster Machine (LightGBM), or a combination thereof.
39. The method of any one of claims 1-37, wherein the classifier is LightGBM.
40. The method of any one of claims 1-37, wherein the classifier is RF.
41. The method of any one of claims 1-40, further comprising removing one or more outlier training mass spectra prior to constructing the training peak matrix.
42. The method of any one of claims 1-41, further comprising generating a second plurality of training mass spectrometric peaks from a plurality of training mass spectrometric peak from one or more of the plurality of training mass spectra using SMOTE algorithm, wherein merging the training mass spectrometric peaks comprises: merging the training mass spectrometric peaks identified from the plurality of training mass spectra and the training mass spectrometric peaks generated using SMOTE algorithm to construct a training peak matrix.
43. The method of any one of claims 1-42, further comprising identifying mass-to-charge values representative of a euploidy embryo or an aneuploidy embryo.
44. The method of any one of claims 1-43, wherein the embryos are from mammalian subjects.
45. A system for determining embryo ploidy comprising:non-transitory memory configured to store executable instructions and a classifier for differentiating mass spectra corresponding to euploidy embryos from mass spectra corresponding to aneuploidy embryos, wherein the classifier is generated by:receiving a plurality of training mass spectra obtained by subjecting a plurality of spent embryo media where embryos had been cultured into Matrix- Assisted Laser Desorption / Ionization Time-of-Flight (MALDI-TOF) mass spectrometry, wherein the embryos comprise euploidy embryos and aneuploidy embryos, and wherein each of the embryos is associated with a classification of ploidy comprising euploidy and aneuploidy;identifying a plurality of training mass spectrometric peaks from each of the plurality of training mass spectra;merging training mass spectrometric peaks identified from the plurality of training mass spectra to construct a training peak matrix;training the classifier using the training peak matrix as input and the corresponding classifications of the embryos as output; anda hardware processor in communication with the non-transitory memory, the hardwareprocessor programmed by the executable instructions to perform:receiving a sample mass spectrum obtained by subjecting a sample spent embryo medium where a sample embryo had been cultured into MALDI-TOF mass spectrometry;determining a classification of the sample embryo using the classifier with the sample mass spectrum as input.
46. The system of claim 45, wherein an embryo of the embryos and the sample embryo were cultured in media that were identical.
47. The system of claim 45, wherein an embryo of the embryos and the sample embryo were cultured in media that were different.
48. The system of claim 45, wherein the embryos were cultured in media that were identical.
49. The system of claim 45, wherein the embryos were cultured in media that were different.
50. The system of any one of claims 45-49, wherein the classification of the sample embryo is used to determine embryo viability of the sample embryo.
51. The system of any one of claims 45-50, wherein the hardware processor is further programmed by the executable instructions to perform: determining the embryo viability of the sample embryo.
52. The system of claim 51, wherein determining the embryo viability of the sample embryo comprises determining the embryo viability of the sample embryo using the classification of the sample embryo.
53. The system of claim 51, wherein determining the embryo viability of the sample embryo comprises determining whether the sample mass spectrum corresponds to that of a euploidy embryo or an aneuploidy embryo.
54. The system of any one of claims 45-53, wherein the plurality of training mass spectra comprises at least 100 training mass spectra corresponding to euploidy embryos and / or at least 100 training mass spectra corresponding to aneuploidy embryos.
55. The system of any one of claims 45-54, wherein the plurality of training mass spectra comprises at most 10 training mass spectra corresponding to mosaic embryos.
56. The system of any one of claims 45-55, wherein the plurality of training mass spectrometric peaks comprises metabolite peaks.
57. The system of any one of claims 45-56, wherein identifying the plurality of training mass spectrometric peaks from each training mass spectrum comprise selecting training mass spectrometric peaks in the range 50-1500 Daltons.
58. The system of any one of claims 45-57, wherein the classifier is further generated by aligning the training mass spectrometric peaks identified from the plurality of training mass spectra.
59. The system of any one of claims 45-58, wherein the classifier is further generated by normalizing each of the plurality of training mass spectra.
60. The system of any one of claims 45-59, wherein identifying the plurality of training mass spectrometric peaks from each training mass spectrum comprises selecting peaks with a height equal to or greater than 1% of that of a highest peak.
61. The system of claim 60, wherein the hardware processor is further programmed by the executable instructions to perform: discarding peaks having a width of less than 1 Dalton and a prominence of less than 0.01.
62. The system of any one of claims 45-61, wherein the classifier is selected from the group consisting of: Partial-Least Squares Discriminant Analysis (PLS-DA), Support Vector Machine (SVM), Random Forest (RF), K-Nearest Neighbour (KNN), Light Gradient-Booster Machine (LightGBM), or a combination thereof.
63. The system of any one of claims 45-61, wherein the classifier is LightGBM.
64. The system of any one of claims 45-61, wherein the classifier is RF.
65. The system of any one of claims 45-64, further comprising removing one or more outlier training mass spectra prior to constructing the training peak matrix.
66. The system of any one of claims 45-65, wherein the classifier is further generated by generating a second plurality of training mass spectrometric peaks from a plurality of training mass spectrometric peak from one or more of the plurality of training mass spectra using SMOTE algorithm, wherein merging the training mass spectrometric peaks comprises: merging the training mass spectrometric peaks identified from the plurality of training mass spectra and the training mass spectrometric peaks generated using SMOTE algorithm to construct a training peak matrix.
67. The system of any one of claims 45-66, wherein the classifier is further generated by identifying mass-to-charge values representative of a euploidy embryo or an aneuploidy embryo.
68. The system of any one of claims 45-67, wherein the embryos are from mammalian subjects.-SO-