Ranking of reproductive cellular structures using multiple instance learning
The use of multiple instance learning with an attention mechanism for ranking reproductive cellular structures in IVF addresses the instability of current AI systems, providing a reliable and efficient method for selecting embryos, thus enhancing IVF success rates.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2026-03-26
AI Technical Summary
Current AI-based systems for ranking embryos in IVF are unstable and inconsistent, leading to unpredictable performance and low reliability, despite their potential for improving IVF success rates.
A method and system utilizing multiple instance learning with an attention mechanism to rank a cohort of reproductive cellular structures, providing a stable and reliable prediction of successful IVF outcomes by evaluating the collective performance of the cohort rather than individual embryos.
The proposed approach enhances the stability and reliability of embryo ranking, offering a more robust solution for selecting the best embryos for transfer, thereby improving IVF success rates and reducing the need for multiple cycles.
Smart Images

Figure US2025047040_26032026_PF_FP_ABST
Abstract
Description
RANKING OF REPRODUCTIVE CELLULAR STRUCTURES USING MULTIPLE INSTANCE LEARNINGRelated Applications
[0001] The present application claims priority to each of U.S. Provisional Patent Application Serial No. 63 / 696,327 filed September 18, 2024 entitled COHORT-BASED MULTI INSTANCE LEARNING FOR RELATIVE SELECTION AND CLINICAL RANK ORDERING OF EMBRYOS FOR TRANSFER, which is incorporated herein by reference in its entirety for all purposes.Statement on Government Rights
[0002] This invention was made with government support under one or more of award numbers R01 AI118502, R33AI140489, and R01 EB033866, 3R01 AH 38800- 05S1 , 4R33A1140489-05 and 5R01 EB033866-03 from by the National Institutes of Health. The government has certain rights in the invention.Technical Field
[0003] The present invention relates generally to the field of assisted fertility, and more particularly, to ranking of reproductive cellular structures using multiple instance learning.Background of the Invention
[0004] Infertility is an underestimated healthcare problem that affects over forty-eight million couples globally and is a cause of distress, depression, and discrimination.Although assisted reproductive technologies (ART) such as in-vitro fertilization (IVF) has alleviated the burden of infertility to an extent, it has been inefficient with an average success rate of approximately twenty-six percent reported in 2015 in the US. IVF remains as an expensive solution, with a cost between $7000 and $20,000 per ART cycle in the US, which is generally not covered by insurance. Further, many patients require multiple cycles of IVF to achieve pregnancy. Non-invasive selection of the topquality embryo for transfer is one of the most important factors in achieving successfulART outcomes, yet this critical step remains a significant challenge. Various artificial intelligence (Al) implementations have been proposed to address this issue, with mixed results.
[0005] Trust is a fundamental driver in the adoption of breakthrough technologies, especially artificial intelligence-based solutions. Public opinion is heavily influenced by perceived risks such as bias, data safety, and lack of transparency. While often overlooked, an equally critical and challenging aspect is model stability. Unstable or variable models can lead to inconsistent and unpredictable performance. Such variability in medicine will undermine the confidence that clinicians and patients place in Al technologies. Particularly in assisted reproduction, where much of the current focus is on the development of the best performing models capable of predicting viable embryos for successful outcomes using static or time series images of individual embryos, model stability has been poorly studied. Although Al-based in-vitro fertilization (IVF) solutions are already in prospective clinical trials, one study observed that eight different commercial Al algorithms, when rank ordering embryos, had significantly lower agreement with embryologists and among themselves. In fact, two of the Al algorithms produced orders comparable to random chance. This raises questions about the reliability of these Al models and whether they are being developed and evaluated appropriately for the clinical task.Summary of the Invention
[0006] In accordance with an aspect of the present invention, a method is provided for ranking a cohort of reproductive cellular structures. A set of images, each representing a reproductive cellular structure of the cohort of reproductive cellular structures, is provided to a predictive model utilizing an attention mechanism to provide an output representing a likelihood of a successful outcome for a cycle of in-vitro fertilization associated with the cohort of reproductive cellular structures. The predictive model is trained using multiple instance learning on a set of training samples, with each training sample including a set of images representing a set of reproductive cellular structures and a value representing an outcome of a cycle of in-vitro fertilization associated with the set of reproductive cellular structures. A total attention associated with each of the set of images at the predictive model is determined. The cohort ofreproductive cellular structures is ranked according to the total attention associated with each of the set of images.
[0007] In accordance with another aspect of the present invention, a system is provided for fully automated ranking of a cohort of reproductive cellular structures. The system includes a processor and a non-transitory computer readable medium storing instructions executable by the processor. The machine executable instructions are executable by the processor to provide a predictive model that is implemented using an attention mechanism and trained using multiple instance learning on a set of training samples, with each training sample including a set of images representing a set of reproductive cellular structures and a value representing an outcome of a cycle of in- vitro fertilization associated with the set of reproductive cellular structures. The predictive model receives a set of images representing the cohort of reproductive cellular structures and provides an output representing a likelihood of a successful outcome for a cycle of in-vitro fertilization associated with the cohort of reproductive cellular structures. Ranking logic determines a total attention associated with each of the set of images at the predictive model and ranks the cohort of reproductive cellular structures according to the total attention associated with each of the set of images.
[0008] In accordance with yet another aspect of the present invention, a method is provided for selecting and implanting an embryo of a cohort of embryos. A set of images, each representing an embryo of the cohort of embryos, is provided to a convolutional neural network utilizing an attention mechanism to provide an output representing a likelihood of a successful outcome for a cycle of in-vitro fertilization associated with the cohort of embryos. The convolutional neural network is trained using multiple instance learning on a set of training samples, with each training sample including a set of images representing a set of embryos and a value representing an outcome of a cycle of in-vitro fertilization associated with the set of embryos. A total attention associated with each of the set of images at the convolutional neural network is determined. The cohort of embryos is ranked according to the total attention associated with each of the set of images. A highest-ranked embryo of the cohort of embryos is implanted into a patient.Brief Description of the Drawings
[0009] The foregoing and other features of the present invention will become apparent to those skilled in the art to which the present invention relates upon reading the following description with reference to the accompanying drawings, in which:
[0010] FIG. 1 illustrates one example of a system for fully automated ranking of a cohort of reproductive cellular structures according to their likelihood of a successful outcome;
[0011] FIG. 2 illustrates another example of a system for fully automated ranking of a cohort of embryos according to their likelihood of a successful outcome;
[0012] FIG. 3 illustrates an example of a method for ranking a cohort of reproductive cellular structures;
[0013] FIG. 4 illustrates an example of a method for selecting and implanting an embryo of a cohort of embryos; and
[0014] FIG. 5 is a schematic block diagram illustrating an exemplary system of hardware components capable of implementing examples of the systems and methods disclosed herein.Detailed Description
[0015] The phrase “continuous parameter” is used herein to distinguish numerical values from categorical values or classes and should be read to include both truly continuous data as well as data more traditionally referred to as discrete data.
[0016] As used herein, an “average” of a set of values is any measure of central tendency. For continuous parameters, this includes any of the median, arithmetic mean, and geometric mean of the values. For categorical parameters, the average is the mode of the set of values.
[0017] A “subset” of a set as used here, refers to a set containing some or all elements of the set. A “proper subset” of a set contains less than all of the elements of the set.
[0018] An “image” as used herein, is intended to encompass a single image, a time series of images, or a video.
[0019] As used herein, a “cohort” of reproductive cellular structures is a set of reproductive cellular structurse, such as sperm cells, ooctyes, embryos, or otherfertility-related cells of interest that are harvested or fertilized in a given cycle of in-vitro fertilization (IVF) or other fertility treatment from a same individual or couple. It will be appreciated that not all reproductive cellular structures in a cohort may be selected for use during IVF or other assisted fertility procedure.
[0020] A fertility endpoint can include any clinical relevant outcome of a cycle of in- vitro fertilization, including any of a live birth, a successful pregnancy, implantation of an embryo, advancement of an embryo to a given stage of development, miscarriage, biochemical pregnancy, or a cumulative measure of one of these measures across the cohort (e.g., a cumulative live birth or cumulative successful pregnancy). As used, a fertility endpoint can represent either of a patient-level clinical endpoint or a cohort-level clinical endpoint.
[0021] Conventionally used Al approaches capitalize on single instance-learning (SIL). These approaches evaluate embryos individually, focusing on the relationship between each embryo’s features and its eventual outcome independently. These models are used to determine the order of transfer of embryos in an IVF cycle, based on the model’s estimated likelihood of implantation or live birth for each specific embryo. These models are tuned to achieve the best possible performance, often measured with the area under the Receiver Operating Characteristic curve (AUC) or accuracy, in predicting the embryo transfer outcomes. The underlying assumption is that models with the highest performance in outcome prediction should also provide an equally accurate order of transfer when evaluating a patient’s entire cohort of embryos.
[0022] However, when training multiple single instance learning (SIL) replicate models to predict live-birth outcomes by only varying their initialization weights, the inventors have determined that performance metrics such as AUC and accuracy can obscure substantial variabilities in the model decision making process. This variability is evident in substantial changes in the models’ ability to correctly predict which embryos will result in positive or negative live-birth outcomes. Additionally, this variance does not seem to be considerably affected by the training dataset size or by retraining with pre-initialized weights. Studies in the literature generally do not focus on this model variability aspect because, for deployment, the single best-performing model is ultimately selected for clinical use. However, this remains a cause for concern as performance in single-instance outcome prediction does not necessarily translate toeffective rank ordering of embryos during clinical implementation, despite using common selection methods such as validation AUC, accuracy, and loss.
[0023] Thus, while current Al approaches in IVF show promise, they are often unsuitable for clinical deployment due to stability and consistency issues. The systems and methods proposed herein evaluate entire cohorts of reproductive cellular structures in unison taking into consideration the context of the cohort. Reproductive cellular structures images are processed to estimate an outcome that represents the overall set such as cumulative pregnancy or live-birth outcome, and a rank order is generated by utilizing the attention scores for each reproductive cellular structures image relative to the other reproductive cellular structures images present in the set. This approach offers a more robust and reliable solution, demonstrating substantially better performance in reproductive cellular structures selection and rank ordering
[0024] FIG. 1 illustrates one example of a system 100 for fully automated ranking of a cohort of reproductive cellular structures according to their likelihood of a successful outcome. The system 100 includes a processor 102 and a non-transitory computer readable medium 110 that stores machine executable instructions for ranking the cohort of reproductive cellular structures. The machine executable instructions include a predictive model 112 that receives data representing a cohort of reproductive cellular structures and predicts whether a selected fertility endpoint, such as a successful pregnancy or a live birth, will occur for a cycle of in-vitro fertilization conducted with the cohort of reproductive cellular structures. It will be appreciated that the provided data can include images representing each of the reproductive cellular structures in addition to any of image features generated from external convolutional neural networks or similar models, clinical, demographic, or tabular data, text, genomic or molecular data, time series of images, audio data, video data, or combinations of two or more of these data types from one of a clinical database, an electronic health records system, a laboratory system, or a textual report.. This information can be provided separately to the predictive model 112 or fused to provide a multimodal input to the predictive model. For example, the image and non-image data can be fused into a joint embedding (e.g., via early, late, or hybrid fusion) with that normalization and alignment across modalities performed before fusion.
[0025] The predictive model 112 is trained using multiple instance learning on a set of training samples, with each training sample including a set of data samples representing a cohort of reproductive cellular structures and a fertility endpoint associated with the cycle of in-vitro fertilization associated with the cohort of reproductive cellular structures. In one example, each cohort of reproductive cellular structures is labeled with a first value if the cycle of in-vitro fertilization associated with the cohort of reproductive cellular structure ended in a live birth and a second value if the cycle of in-vitro fertilization associated with the cohort of reproductive cellular structures did not end in a live birth. In another example, each cohort of reproductive cellular structures is labeled with a first value if the cycle of in-vitro fertilization associated with the cohort of reproductive cellular structures produced a successful pregnancy and a second value if the cycle of in-vitro fertilization associated with the cohort of reproductive cellular structures did not produce a successful pregnancy. In one example, the predictive model 112 can be implemented as at least one of a convolutional neural network, a transformer, or another deep learning feature extractor.
[0026] It will be appreciated that the output of the predictive model can be continuous, in which a value representing the likelihood of a successful outcome is output, or categorical, where the output is a value representing one of a plurality of categorical classes. In one example, a categorical output can take on a first value representing an expected successful output or a second value representing an expected unsuccessful outcome. In another example, the categorical parameter can assume any of a plurality of values, each representing a range of likelihoods of the successful outcome. In the illustrated example, the predictive model 1 12 uses an attention mechanism to weigh various features or regions within the input images, based on a determined importance of the features during the training process. As a result, the degree to which each image within the input image contributed to the outcome can be determined from the weights applied to the various regions or features of the images during classification at the predictive model 112.
[0027] The predictive model 112 can utilize one or more pattern recognition algorithms, each of which analyze the provided data or features extracted from the provided data to classify the patients into one of the plurality of classes and provide this information as a second clinical parameter. Where multiple classification or regressionmodels are used, an arbitration element can be utilized to provide a coherent result from the plurality of models. The training process of a given classifier will vary with its implementation, but training generally involves a statistical aggregation of training data into one or more parameters associated with the output class. For rule-based models, such as decision trees, domain knowledge, for example, as provided by one or more human experts, can be used in place of or to supplement training data in selecting rules for classifying a patient using the extracted features. Any of a variety of techniques can be utilized for the classification algorithm and related feature extraction, including support vector machines (SVMs), regression models, transformers, multimodal encoders, RNNs, foundation models, self-organized maps, fuzzy logic systems, data fusion processes, boosting and bagging methods, rule-based systems, or artificial neural networks. The predictive model 112 can utilize feature transfer from pretrained models.
[0028] A neural network includes a plurality of nodes having a plurality of interconnections. Values from the image, for example luminance and / or chrominance values associated with the individual pixels, are provided to a plurality of input nodes. The input nodes each provide these input values to layers of one or more intermediate nodes. A given intermediate node receives one or more output values from previous nodes. The received values are weighted according to a series of weights established during the training of the classifier. An intermediate node translates its received values into a single output according to an activation function at the node. For example, the intermediate node can sum the received values and subject the sum to an identify function, a step function, a sigmoid function, a hyperbolic tangent, a rectified linear unit, a leaky rectified linear unit, a parametric rectified linear unit, a Gaussian error linear unit, the softplus function, an exponential linear unit, a scaled exponential linear unit, a Gaussian function, a sigmoid linear unit, a growing cosine unit, the Heaviside function, and the mish function. A final layer of nodes provides the confidence values for the output classes of the neural network, with each node having an associated value representing a confidence for one of the associated output classes of the classifier.
[0029] Many ANN classifiers are fully-connected and feedforward. A convolutional neural network, however, includes convolutional layers in which nodes from a previous layer are only connected to a subset of the nodes in the convolutional layer. Recurrentneural networks are a class of neural networks in which connections between nodes form a directed graph along a temporal sequence. Unlike a feedforward network, recurrent neural networks can incorporate feedback from states caused by earlier inputs, such that an output of the recurrent neural network for a given input can be a function of not only the input but one or more previous inputs. As an example, Long Short-Term Memory (LSTM) networks are a modified version of recurrent neural networks, which makes it easier to remember past data in memory. In some examples, neural networks are trained in an adversarial manner in which a generative classifier is configured to generate new instances for a discriminative classifier. Typically, the generative network learns to map from a latent space to a data distribution of interest, while the discriminative network distinguishes candidates produced by the generator from a true data distribution. The generative network’s training objective is to increase the error rate of the discriminative network. Each network learns iteratively to better model the true data distribution across the classes of interest.
[0030] For example, a support vector machine (SVM) classifier can utilize a plurality of functions, referred to as hyperplanes, to conceptually divide boundaries in the N- dimensional feature space, where each of the N dimensions represents one associated feature of the feature vector. The boundaries define a range of feature values associated with each class. Accordingly, an output class and an associated confidence value can be determined for a given input feature vector according to its position in feature space relative to the boundaries. In one implementation, the SVM can be implemented via a kernel method using a linear or non-linear kernel. The predictive model 112 can also use an artificial neural network classifier as described previously.
[0031] A rule-based classifier applies a set of logical rules to the extracted features to select an output class. Generally, the rules are applied in order, with the logical result at each step influencing the analysis at later steps. The specific rules and their sequence can be determined from any or all of training data, analogical reasoning from previous cases, or existing domain knowledge. One example of a rule-based classifier is a decision tree algorithm, in which the values of features in a feature set are compared to corresponding threshold in a hierarchical tree structure to select a class for the feature vector. A random forest classifier is a modification of the decision tree algorithm using a bootstrap aggregating, or “bagging” approach. In this approach,multiple decision trees are trained on random samples of the training set, and an average (e.g., mean, median, or mode) result across the plurality of decision trees is returned. For a classification task, the result from each tree would be categorical, and thus a modal outcome can be used.
[0032] These weights are provided to a ranking logic 114, which ranks the reproductive cellular structures within the cohort represented by the set of provided data based on the likelihood of a successful outcome based on the weights applied to the features or regions associated with the provided data. The total attention used to provide the weights can be derived from any model-generated attribution or interpretability map and can represent image features, non-image features, or fused features. In one example, the weights can be represented with an attention map, with the attention associated with each pixel or defined grouping of pixels within an image in the provided data summed to provide a total attention associated with the image. It will be appreciated that the weights can be signed, such that regions that positively contribute to the likelihood that of a successful outcome for the reproductive cell structure may have weights with a first sign (e.g., positive), while portions of the provided data that negatively contribute to the likelihood that of a successful outcome for the reproductive cell structure may have weights with a second, opposing sign (e.g., negatively).
[0033] Alternatively, the attention can be provided as a class-discriminative attention map, and the ranking logic 114 can generate the total attention for the provided data as a weighted linear combination of the weights, with each class having an associated multiplier in the linear combination. For example, where the predictive model 112 represents ranges of likelihoods, the highest and lowest range of values can each have large multipliers, but with opposing signs, and one or more intermediate ranges of values can have smaller multipliers. The reproductive cellular structures can then be ranked according to the total attention applied to the portion of the provided data associated with that cell, providing a ranking of the reproductive cellular structures according to their likelihood of a successful outcome. The determined ranking can be provided to the user at an associated output device (not shown) along with a predicted outcome for the cohort of reproductive cellular structures. Further, the system 100 allows a doctor, scientist, or user to see which features the model 112 relied on most —whether that is parts of an image (e.g., areas of an embryo picture) or non-image information (e.g., age or clinical notes). It will be appreciated that the attention data provides reusable feature representations that support multiple downstream tasks, such as grading, quality prediction, and decision support. Accordingly, either or both of the attention data and latent values from the predictive model 112 can also be provided as features to a second predictive model, for example, a single instance learning model for determining a likelihood of success for a single reproductive cellular structure, a grading system, a quality prediction system, a retrieval system for locating similar reproductive cellular structures from stored records, anomaly detection systems, and clinical decision support models.
[0034] It will be appreciated that the attention data derived by the system can be provided in any of a number of ways, such as attention maps, heatmaps, or similar representations useful for clarifying, in a human-understandable way, which features were used and which were most important in making the decision. Further, the system can analyze image-derived and non-image features to identify an ideal set of features and target ranges associated with achieving a desired outcome. For example, a minimum number of blastocysts, an ideal maternal age range, an optimal growth media type and other such features can be identified. Ideally, the system can identify (i) which features are necessary / sufficient, (ii) recommended thresholds or ranges, and (iii) an “ideal feature profile” highlighting the most influential variables. The system 100 can employ sensitivity analysis to determine how predicted outcomes shift in response in response to adjustments of these features (e.g., changing age or embryo count).
[0035] FIG. 2 illustrates another example of a system 200 for fully automated ranking of a cohort of embryos according to their likelihood of a successful outcome. The system 200 includes a processor 202 and a non-transitory computer readable medium 210 that stores machine executable instructions for ranking the cohort of embryos according to their viability. The instructions stored on the computer readable medium 210 are executable by the processor to provide various software-implemented functional modules, including an input interface 212 that receives a set of images representing a cohort of embryos from either an associated imager or a storage medium. It will be appreciated that the input interface 212 can apply one or more imaging condition techniques, such as cropping and filtering, to better prepare theimages for analysis. The input interface 212 can also receive additional data representing the cohort of embryos including one or more of image features generated from external convolutional neural networks or similar models, clinical, demographic, or tabular data, text, genomic or molecular data, time series of images, audio data, video data, or combinations of two or more of these data types.
[0036] The machine executable instructions include a deep learning model 214, such as a convolutional neural network (CNN) neural network, a transformer, or another deep learning feature extractor that receives images of a cohort of embryos and predicts whether a positive outcome, such as a successful pregnancy or a live birth, will occur for a cycle of in-vitro fertilization conducted with the cohort of embryos. In the illustrated implementation, the deep learning model 214 is trained using multiple instance learning on a set of training samples, with each training sample including a cohort of embryos and labeled with a first value if the cycle of in-vitro fertilization associated with the cohort of embryos ended in a live birth and a second value if the cycle of in-vitro fertilization associated with the cohort of embryos did not end in a live birth. The deep learning model 214 is implemented as a feature extractor 216, a series of convolutional layers 218, an attention block 220, and a classifier 222.
[0037] The feature extractor 216 comprises a plurality of convolutional layers that processes the set of images to extract an initial set of visual features as an input to the series of convolutional layers 218. In one example, the feature extractor 216 can be pretrained on unrelated image data or provided by transfer learning using weights from another deep learning model. Further, the feature extractor 216 can receive data from pretrained external networks, as well as non-image data, such as text, audio, genomic data, molecular data, and fuse this data with the image data to provide a multimodal input. For example, the image and non-image data can be fused into a joint embedding (e.g., via early, late, or hybrid fusion) with that normalization and alignment across modalities performed before fusion. Each convolutional layer generates an output from the provided input, weighted according to weights provided by the attention block 220. It will be appreciated that a given convolutional layer, as the term is used herein, can include multiple operations, such as pooling and an activation function. The attention block 220 provides weights for the various features based on a determined importance of the features during the training process. The output of each layer is provided as aninput to the next layer, until an output of the layers 218 is provided to the classifier 222. The classifier 222 can include one or more convolutional layers and fully-connected layers that receives the output of the weighted convolutional layers 218 and provides an output representing the likelihood that a cycle of in-vitro fertilization associated with the cohort of embryos would have a successful outcome. It will be appreciated that the output of the predictive model can be continuous, in which a value representing the likelihood of a successful outcome is output, or categorical, where the output is a value representing one of a plurality of categorical classes.
[0038] The weights assigned to the features at the series of convolutional layers 218 by the convolutional layer 220 are provided to a ranking logic 224 that ranks the cohort of embryos according to the expected viability of the embryos. Specifically, the weights assigned to the features at the series of convolutional layers 218 can be attributed to each of the set of images that provided the features, such that a total attention applied to each image can be determined. It will be appreciated that the weights can be signed, such that features that positively contribute to the likelihood that of a successful outcome for the embryo may have weights with a first sign (e.g., positive), while features that negatively contribute to the likelihood that of a successful outcome for the embryo may have weights with a second, opposing sign (e.g., negatively). The ranking logic 224 then ranks the embryos according to the total attention applied to each image, providing a ranking of the embryos according to their likelihood of a successful outcome. The ranking of the embryos can then be stored on the non-transitory computer readable medium 210 and / or displayed to a user, along with images of the cohort of embryos, at an associated display 230.
[0039] In view of the foregoing structural and functional features described above, a method in accordance with various aspects of the present invention will be better appreciated with reference to FIGS. 3 and 4. While, for purposes of simplicity of explanation, the methods of FIGS. 3 and 4 is shown and described as executing serially, it is to be understood and appreciated that the present invention is not limited by the illustrated order, as some aspects could, in accordance with the present invention, occur in different orders and / or concurrently with other aspects from that shown and described herein. Moreover, not all illustrated features may be required to implement a method in accordance with an aspect the present invention.
[0040] FIG. 3 illustrates an example of a method 300 for ranking a cohort of reproductive cellular structures. At 302, a set of images, each representing a reproductive cellular structure of the cohort of reproductive cellular structures, is provided to a predictive model utilizing an attention mechanism to provide an output representing a likelihood of a successful outcome for a cycle of in-vitro fertilization associated with the cohort of reproductive cellular structures. In one example, the predictive model is implemented as a deep learning model, such as a transformer or convolutional neural network. It will be appreciated that the predictive model can use non-image features in addition to the set of images, and that the images and the nonimage features can be fused to provide a multimodal input to the predictive model.
[0041] The predictive model is trained using multiple instance learning on a set of training samples, with each training sample including a set of images representing a set of reproductive cellular structures and a value representing an outcome of a cycle of in- vitro fertilization associated with the set of reproductive cellular structures. In one example, the value representing the outcome of the cycle of in-vitro fertilization for each set of reproductive cellular structures is a first value when the cycle of in-vitro fertilization resulted in a positive fertility endpoint and a second value when the cycle of in-vitro fertilization did not result in a positive fertility endpoint. For example, the value representing the outcome of the cycle of in-vitro fertilization for each set of reproductive cellular structures is a first value when the cycle of in-vitro fertilization resulted in a successful pregnancy or live birth and a second value when the cycle of in-vitro fertilization did not result in a successful pregnancy or live birth.
[0042] At 304, a total attention associated with each of the set of images at the predictive model is determined. In one example, wherein the output representing a likelihood of a successful outcome for a cycle of in-vitro fertilization is one of a plurality of output classes, the total attention is determined from a class-discriminative attention map generated at the predictive model. At 306, the cohort of reproductive cellular structures are ranked according to the total attention associated with each of the set of images. Once the reproductive cellular structures in the cohort are ranked, the ranking of the cohort of reproductive cellular structures and one or more of a predicted outcome for the cohort of reproductive cellular structures, the set of images representing the cohort of reproductive cellular structures, or overlays representing attention for eachimage or non-image feature can be displayed to a user at a display to guide clinical decision, including one or more of an implantation of an reproductive cellular structure, cryopreservation, selection of reproductive cellular structures for subsequent transfers, decisions to perform biopsy or genetic testing of reproductive cellular structures, and clinician overrides of any decision by the predictive model. In one example, a highest- ranked reproductive cellular structure in the cohort of reproductive cellular structures can be used for an assisted fertility procedure.
[0043] FIG. 4 illustrates an example of a method 400 for selecting and implanting an embryo of a cohort of embryos. At 402, a set of images, each representing an embryo of the cohort of embryos, is provided to a convolutional neural network utilizing an attention mechanism to provide an output representing a likelihood of a successful outcome for a cycle of in-vitro fertilization associated with the cohort of embryos. The convolutional neural network is trained using multiple instance learning on a set of training samples, with each training sample including a set of images representing a set of embryos and a value representing an outcome of a cycle of in-vitro fertilization associated with the set of embryos. In one example, the value representing the outcome of the cycle of in-vitro fertilization for each set of embryos is a first value when the cycle of in-vitro fertilization resulted in a live birth and a second value when the cycle of in-vitro fertilization did not result in a live birth. In another example, the value representing the outcome of the cycle of in-vitro fertilization for each set of embryos is a first value when the cycle of in-vitro fertilization resulted in a successful pregnancy and a second value when the cycle of in-vitro fertilization did not result in a successful pregnancy.
[0044] At 404, a total attention associated with each of the set of images at the convolutional neural network is determined. In one example, wherein the output representing a likelihood of a successful outcome for a cycle of in-vitro fertilization is one of a plurality of output classes, the total attention is determined from a class- discriminative attention map generated at the predictive model. At 406, the cohort of embryos are ranked according to the total attention associated with each of the set of images. At 408, a highest-ranked embryo of the cohort of embryos is implanted into a patient. In one implementation, the ranking of the cohort of embryos and the set of images representing the cohort of embryos can also be displayed to a user at a display.
[0045] FIG. 5 is a schematic block diagram illustrating an exemplary system 500 of hardware components capable of implementing examples of the systems and methods disclosed in FIGS. 1 -4, such as the automated embryo evaluation system illustrated in FIG. 1 . The system 500 can include various systems and subsystems. The system 500 can be any of personal computer, a laptop computer, a workstation, a computer system, an appliance, an application-specific integrated circuit (ASIC), a server, a server blade center, or a server farm.
[0046] The system 500 can includes a system bus 502, a processing unit 504, a system memory 506, memory devices 508 and 510, a communication interface 512 (e.g., a network interface), a communication link 514, a display 516 (e.g., a video screen), and an input device 518 (e.g., a keyboard and / or a mouse). The system bus 502 can be in communication with the processing unit 504 and the system memory 506. The additional memory devices 508 and 510, such as a hard disk drive, server, stand-alone database, or other non-volatile memory, can also be in communication with the system bus 502. The system bus 502 interconnects the processing unit 504, the memory devices 506-510, the communication interface 512, the display 516, and the input device 518. In some examples, the system bus 502 also interconnects an additional port (not shown), such as a universal serial bus (USB) port.
[0047] The system 500 could be implemented in a computing cloud. In such a situation, features of the system 500, such as the processing unit 504, the communication interface 512, and the memory devices 508 and 510 could be representative of a single instance of hardware or multiple instances of hardware with applications executing across the multiple of instances (i.e., distributed) of hardware (e.g., computers, routers, memory, processors, or a combination thereof). Alternatively, the system 500 could be implemented on a single dedicated server.
[0048] The processing unit 504 can be a computing device and can include an application-specific integrated circuit (ASIC). The processing unit 504 executes a set of instructions to implement the operations of examples disclosed herein. The processing unit can include a processing core.
[0049] The additional memory devices 506, 508, and 510 can store data, programs, instructions, database queries in text or compiled form, and any other information that can be needed to operate a computer. The memories 506, 508 and 510 can beimplemented as computer-readable media (integrated or removable) such as a memory card, disk drive, compact disk (CD), or server accessible over a network. In certain examples, the memories 506, 508 and 510 can comprise text, images, video, and / or audio, portions of which can be available in formats comprehensible to human beings.
[0050] Additionally or alternatively, the system 500 can access an external data source or query source through the communication interface 512, which can communicate with the system bus 502 and the communication link 514.
[0051] In operation, the system 500 can be used to implement one or more parts of an embryo cohort ranking system in accordance with the present invention. Computer executable logic for implementing the composite applications testing system resides on one or more of the system memory 506, and the memory devices 508, 510 in accordance with certain examples. The processing unit 504 executes one or more computer executable instructions originating from the system memory 506 and the memory devices 508 and 510. It will be appreciated that a computer readable medium can include multiple computer readable media each operatively connected to the processing unit.
[0052] Specific details are given in the above description to provide a thorough understanding of the embodiments. However, it is understood that the embodiments can be practiced without these specific details. For example, circuits can be shown in block diagrams in order not to obscure the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques can be shown without unnecessary detail in order to avoid obscuring the embodiments.
[0053] Implementation of the techniques, blocks, steps, and means described above can be done in various ways. For example, these techniques, blocks, steps, and means can be implemented in hardware, software, or a combination thereof. For a hardware implementation, the processing units can be implemented within one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, micro-controllers, microprocessors, other electronic units designed to perform the functions described above, and / or a combination thereof.
[0054] Also, it is noted that the embodiments can be described as a process which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart can describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations can be re-arranged. A process is terminated when its operations are completed, but could have additional steps not included in the figure. A process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination corresponds to a return of the function to the calling function or the main function.
[0055] Furthermore, embodiments can be implemented by hardware, software, scripting languages, firmware, middleware, microcode, hardware description languages, and / or any combination thereof. When implemented in software, firmware, middleware, scripting language, and / or microcode, the program code or code segments to perform the necessary tasks can be stored in a machine readable medium such as a storage medium. A code segment or machine-executable instruction can represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a script, a class, or any combination of instructions, data structures, and / or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, and / or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, ticket passing, network transmission, etc.
[0056] For a firmware and / or software implementation, the methodologies can be implemented with modules (e.g., procedures, functions, and so on) that perform the functions described herein. Any machine-readable medium tangibly embodying instructions can be used in implementing the methodologies described herein. For example, software codes can be stored in a memory. Memory can be implemented within the processor or external to the processor. As used herein the term “memory” refers to any type of long term, short term, volatile, nonvolatile, or other storage medium and is not to be limited to any particular type of memory or number of memories, or type of media upon which memory is stored.
[0057] Moreover, as disclosed herein, the term "storage medium" can represent one or more memories for storing data, including read only memory (ROM), random access memory (RAM), magnetic RAM, core memory, magnetic disk storage mediums, optical storage mediums, flash memory devices and / or other machine readable mediums for storing information. The terms “computer readable medium” and "machine readable medium" includes, but is not limited to portable or fixed storage devices, optical storage devices, wireless channels, and / or various other storage mediums capable of storing that contain or carry instruction(s) and / or data. It will be appreciated that a “computer readable medium” or “machine readable medium” can include multiple media each operatively connected to a processing unit.
[0058] While the principles of the disclosure have been described above in connection with specific apparatuses and methods, it is to be clearly understood that this description is made only by way of example and not as limitation on the scope of the disclosure.
Claims
Having described the invention, we claim:1 . A method for ranking a cohort of reproductive cellular structures, the method comprising: providing a set of images, each representing a reproductive cell structure of the cohort of reproductive cell structures, to a predictive model utilizing an attention mechanism to provide an output representing a likelihood of a successful outcome for a cycle of in-vitro fertilization associated with the cohort of reproductive cell structures, the predictive model being trained using multiple instance learning on a set of training samples, with each training sample including a set of images representing a set of embryos and a value representing an fertility endpoint associated with a cycle of in-vitro fertilization associated with the set of embryos; determining a total attention associated with each of the set of images at the predictive model; and ranking the cohort of reproductive cellular structures according to the total attention associated with each of the set of images.
2. The method of claim 1 , wherein the value representing the fertility endpoint of the cycle of in-vitro fertilization for each set of embryos is a first value when the cycle of in-vitro fertilization resulted in a given result for the fertility endpoint and a second value when the cycle of in-vitro fertilization did not result in the given result for the fertility endpoint.
3. The method of claim 2, wherein the value representing the fertility endpoint of the cycle of in-vitro fertilization for each set of embryos is a cumulative fertility endpoint for the cycle of in-vitro fertilization.
4. The method of claim 1 , wherein providing a set of images to the predictive model comprises providing the set of images to at least one of a convolutional neural network, a transformer, or another deep learning feature extractor.
5. The method of claim 1 , wherein the output representing a likelihood of a successful outcome for a cycle of in-vitro fertilization is one of a plurality of output classes, and the total attention is determined from one of an attribution map and an interpretability map generated at the predictive model.
6. The method of claim 1 , further comprising using the ranking of the cohort of reproductive cellular structures to guide a clinical decision, including one or more implantation of an embryo, cryopreservation, selection for subsequent transfers, biopsy, or genetic testing.
7. The method of claim 1 , wherein the cohort of reproductive cellular structures is a cohort of embryos, the method further comprising implanting a highest-ranked embryo in the cohort of embryos into a patient.
8. The method of claim 1 , wherein the predictive model is a first predictive model, the method further comprising providing latent values from the predictive model as features to a second predictive model.
9. The method of claim 1 , further comprising displaying the ranking of the cohort of ranking the cohort of reproductive cellular structures and the output representing the likelihood of a successful outcome for a cycle of in-vitro fertilization associated with the cohort of reproductive cell structures at a display, the likelihood of a successful outcome for a cycle of in-vitro fertilization associated with the cohort of reproductive cell structures being provided as one of a continuous parameter and a categorical parameter.
10. The method of claim 1 , further comprising performing a sensitivity analysis on the predictive model to determine how the likelihood of the successful outcome for a cycle of in-vitro fertilization associated with the cohort of reproductive cell structures changes in response to different inputs.
11. A system for fully automated ranking of a cohort of reproductive cellular structures, the system comprising: a processor; and a non-transitory computer readable medium storing instructions executable by the processor, the machine executable instructions being executable by the processor to provide: a predictive model, implemented using an attention mechanism and trained using multiple instance learning on a set of training samples, with each training sample including a set of images representing a set of reproductive cellular structures and a value representing an outcome of a cycle of in-vitro fertilization associated with the set of reproductive cellular structures, that receives a set of images representing the cohort of reproductive cellular structures and provides an output representing a likelihood of a successful outcome for a cycle of in-vitro fertilization associated with the cohort of reproductive cellular structures; and ranking logic that determines a total attention associated with each of the set of images at the predictive model and ranks the cohort of reproductive cellular structures according to the total attention associated with each of the set of images.
12. The system of claim 11 , wherein the predictive model includes a feature extractor comprising at least one of a convolutional neural network, a transformer, or another deep learning feature extractor.
13. The system of claim 12, wherein the predictive model comprises a feature extractor, comprising at least one convolutional layer, that generates a set of features from the set of images and a set of non-image data as a fused multimodal input and an attention block that provides weights for the set of features based on a determined importance of the features during training, the weights for the set of features being provided to the ranking logic.
14. The system of claim 11 , the machine executable instructions being executable by the processor to further provide an input interface that retrieves the set of imagesfrom one of an associated imager and a storage medium and non-image data representing the cohort of reproductive cellular structure from one of a clinical database, an electronic health records system, a laboratory system, or a textual report.
15. The system of claim 11 , wherein the value representing the outcome of the cycle of in-vitro fertilization for each set of reproductive cellular structures is a first value when the cycle of in-vitro fertilization resulted in a given fertility endpoint and a second value when the cycle of in-vitro fertilization did not result in the given fertility endpoint.
16. The system of claim 11 , wherein the output representing a likelihood of a successful outcome for a cycle of in-vitro fertilization is one of a plurality of output classes, and the ranking logic determines the total attention for each of the set of images from a class-discriminative attention map generated at the predictive model.
17. The system of claim 11 , further comprising a display that receives the ranking of the cohort of reproductive cellular structure and the set of images representing the cohort of reproductive cellular structure and displays the ranking of the cohort of reproductive cellular structure and at least one of the set of images representing the cohort of reproductive cellular structure, a set of attention heatmaps representing attention, and a summary of variable-level contributions associated with the output representing a likelihood of a successful outcome for a cycle of in-vitro fertilization to a user.
18. A method for selecting and implanting an embryo of a cohort of embryos, the method comprising: providing a set of images, each representing an embryo of the cohort of embryos, to a convolutional neural network utilizing an attention mechanism to provide an output representing a likelihood of a successful outcome for a cycle of in- vitro fertilization associated with the cohort of embryos, the convolutional neural network being trained using multiple instance learning on a set of training samples,with each training sample including a set of images representing a set of embryos and a value representing an outcome of a cycle of in-vitro fertilization associated with the set of embryos; determining a total attention associated with each of the set of images at the convolutional neural network; ranking the cohort of embryos according to the total attention associated with each of the set of images; and implanting a highest-ranked embryo of the cohort of embryos into a patient.
19. The method of claim 18, wherein the value representing the outcome of the cycle of in-vitro fertilization for each set of embryos is a first value when the cycle of in-vitro fertilization resulted in a live birth and a second value when the cycle of in- vitro fertilization did not result in a live birth.
20. The method of claim 18, wherein the value representing the outcome of the cycle of in-vitro fertilization for each set of embryos is a first value when the cycle of in-vitro fertilization resulted in a successful pregnancy and a second value when the cycle of in-vitro fertilization did not result in a successful pregnancy.
Citation Information
Patent Citations
Measuring embryo development and implantation potential with timing and first cytokinesis phenotype parameters
US20140220618A1
Method and system for selecting embryos
US20220198657A1
Latent space synchronization of machine learning models for in device metrology inference
WO2023083564A1
Method and system for screening high developmental potential embryo for in-vitro fertilization
WO2023147760A1