Adaptive Image Processing Method and System in Assisted Reproductive Technology

A deep learning system using CNNs and RNNs addresses the challenges of follicle monitoring in ART by accurately detecting and counting follicles and predicting follicle growth and optimal timing for ovulation induction, thereby improving the success rate of ART procedures.

JP7691744B2Active Publication Date: 2025-06-12CYCLE CLARITY LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021574221
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-06-14
Filing Date
2020-05-12
Publication Date
2025-06-12
Estimated Expiration
2040-05-12

AI Technical Summary

Technical Problem

Current ultrasound (US) technologies for follicle monitoring in assisted reproductive technology (ART) face challenges such as inaccurate follicle size measurements due to irregular follicle shapes, significant human variability in measurements, difficulty in identifying all follicles, and variability in follicle size measurements between observers.

Method used

The development of a deep learning (DL) system that processes ultrasound images to detect, recognize, and count follicles accurately, using convolutional neural networks (CNNs) for image segmentation and classification, and recurrent neural networks (RNNs) for predicting follicle growth and optimal timing of ovulation induction.

Benefits of technology

The DL system improves the accuracy of follicle monitoring, reduces human variability, and enhances the prediction of follicle growth and optimal timing for ovulation induction, thereby increasing the success rate of ART procedures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691744000001
    Figure 0007691744000001
  • Figure 0007691744000002
    Figure 0007691744000002
  • Figure 0007691744000003
    Figure 0007691744000003
Patent Text Reader

Abstract

Adaptive image processing, image analysis, pattern recognition, and time-to-event prediction in various imaging modalities related to assisted reproductive technologies. Reference images can be processed according to one or more adaptive processing frameworks for speckle reduction or noise processing of ultrasound images. Subject images can be processed according to various computer vision techniques for object detection, recognition, annotation, segmentation, and classification of reproductive anatomical structures, such as follicles, ovaries, and uteruses. The image processing framework can also analyze secondary data along with subject image data to analyze the progression of subject images in time to events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of digital image processing and digital data processing systems, as well as corresponding image data processing frameworks, and more particularly to an adaptive digital image processing framework for use in assisted reproductive technology and ovarian stimulation.

Background Art

[0002] The subject matter of the inventions considered in this section is merely a result of the references in this section and should not be assumed to be prior art. Similarly, the problems related to the subject matter of the inventions mentioned or provided as background in this section should not be assumed to have been previously recognized in the prior art. The subject matter of the inventions in this section merely represents a different approach and may in itself correspond to an implementation of the claimed technology.

[0003] Generally, infertility is defined as the inability to conceive after having unprotected intercourse for one year (or more). In vitro fertilization (IVF) is a significant medical treatment option for a substantial number of couples experiencing infertility. Infertility can arise from a combination of disorders including male factor causes and female causes such as tubal blockage, decreased egg numbers, reduced egg quality, ovulation disorders, endometriosis, pelvic adhesions, and unexplained causes. The most aggressive treatment for infertility is called assisted reproductive technology (ART), specifically referring to the technique of extracting eggs (i.e., oocytes) from a female's ovaries, fertilizing them outside the body, and returning the resulting embryos to the patient's uterus. The goal of ART is to identify the best embryos and return them to the patient's uterus. A fundamental part of success in ART is creating the highest quality eggs possible for an individual patient. The quality and maturity of the eggs directly predict the likelihood of the eggs fertilizing and becoming embryos and the quality of the embryos. The quality of the resulting embryos predicts the implantation rate (pregnancy per embryo implantation) and the overall pregnancy rate (when multiple embryo implantations are performed). The thickness and development pattern of the endometrium also influence the prediction of the implantation rate and pregnancy rate (the inner lining of the uterus and the site of embryo implantation).

[0004] Ovarian follicles contain oocytes surrounded by granulosa cells. Follicles at different stages of development are of four different types: primordial, primary, secondary, and tertiary (or antral). The number of primordial follicles, which are the true ovarian reserve, is determined in the fetus and decreases throughout a woman's life. Primordial follicles consist of a dormant monolayer of granulosa cells surrounding the oocyte. They are quiescent but can initiate growth in response to a delicate balance between factors that promote proliferation and apoptosis (i.e., cell death). When they change into primary follicles, the granulosa cells begin to replicate and become cuboidal. The zona pellucida, a glycoprotein polymer capsule, forms around the oocyte and separates from the granulosa cells. As follicles become secondary follicles, the stroma-like theca cells surrounding the outer layer of the follicle undergo cell differentiation to become the theca externa and theca interna separated by a network of capillaries. The formation of a cavity filled with fluid adjacent to the antrum, which is the oocyte, defines the tertiary or antral follicle. Since there is no test available to evaluate the true ovarian reserve, antral follicle count (AFC) is accepted as an excellent surrogate marker. Antral follicles can be identified and counted using transvaginal ultrasound (US). AFC is frequently evaluated in reproductive-aged women for various reasons, such as predicting the risk of menopause, suspecting ovulatory dysfunction secondary to anovulation in androgen excess, infertility, and a refined examination of reproductive assistance technology.

[0005] Ultrasound (US) imaging has become an essential tool for the evaluation and management of infertility in women undergoing ART. Decline in ovarian reserve and ovarian dysfunction are the main causes of infertility, and the ovaries are the organs most frequently ultrasonically scanned in infertile women. The first step in the evaluation of infertility is the determination of ovarian status, ovarian reserve, and subsequent follicle monitoring. Ovarian antral follicles can be identified and manually counted using transvaginal US. Antral follicles can be easily identified by US when they reach a diameter of 2 millimeters (mm), which coincides with an improvement in sensitivity to follicle-stimulating hormone (FSH). Measurement of antral follicles in the range of 2 to 10 mm is "mobilizable", while antral follicles exceeding 10 mm are usually called "dominant" follicles. The ovaries are imaged for their morphology (e.g., normal, polycystic, or multilocular), their abnormalities (e.g., cysts, dermoids, endometriomas, tumors, etc.), follicle growth in ovulation monitoring, and evidence of ovulation and corpus luteum formation and function. With an ovulation scan, the physician can accurately determine the number of mobilizable eggs, the egg maturity of individual follicles, and the appropriate timing of ovulation. Generally, during infertility treatment, two-dimensional (2D) US scans are frequently performed to visualize growing follicles, and all follicles in the ovaries (usually 10 to 15 follicles) are measured to determine the mean follicle size of each follicle. This is performed 4 to 6 times over 10 days while the patient is taking medications (e.g., gonadotropin therapy). The normal time required to perform the ultrasound, in addition to about 10 to 15 minutes per patient, is the additional time for entering data into an electronic medical record (EMR) system (about 5 minutes) or an electronic health record system (EHR). Ovaries are classified into three types based on the number and size of follicles. Cystic ovaries contain 1 to 2 follicles of a size exceeding 28 mm. Polycystic ovaries are those that contain more than 12 follicles of a size less than 10 mm. Ovaries containing 1 to 10 antral follicles in the range of 2 to 10 mm and one or more antral follicles of a size in the range of 10 to 28 mm, with a "dominant" follicle, are considered normal ovaries.

[0006] Current 2D US measurements of follicles are performed under the assumption that they are round, but frequently, follicles are irregular in shape and the measurements become inaccurate. Also, there is significant human variability in the measurement of objects in millimeters by US, which further complicates the accuracy of using this modality for follicle monitoring. Additionally, it is difficult to identify all follicles within the ovary using 2D US, and measurements are frequently missed. The final complication, though not the least, is the variability in follicle size measurements between observers of sonographers, which requires further scrutiny by the reviewing physician. With the advent of three-dimensional (3D) ultrasound, resolution has been steadily improving along with data connectivity. 3D ultrasound measurements of the ovary are performed simply by placing the probe vaginally, pointing it at the ovary, and pressing a button. 3D-US imaging has the advantage of shortening the examination time and improving inter-observer reliability because it can store the data acquired for offline analysis. However, new features such as automated volume calculation (SonoAVC, GE Medical Systems) technology can potentially misidentify adjacent follicles and extra-ovarian tissue as one follicle. Despite the improvements, there is no consensus on the best US technique for performing follicle counts. All currently available semi-automated methods have advantages and disadvantages and are subject to operator preference and skill, which can lead to inaccuracies and variability.

[0007] From the perspective of a digital data processing system, follicles are regions of interest (ROIs) in ovarian ultrasound images and can be detected using image processing techniques. The basic image processing steps, namely preprocessing, segmentation, feature extraction, and classification, can be applied to this complex task of accurate follicle recognition. However, imaging modalities that form images with coherent energy such as US are plagued by speckle noise and can potentially degrade the performance of automated operations such as computer-aided diagnosis (CAD), which is a system that can distinguish between benign and malignant lesion tissues for cancer diagnosis. In the context of ART, CAD is desirable to address the cumbersome and time-consuming nature of manual follicle segmentation, sizing, counting, and ovarian classification, which require operator skills and medical expertise for accuracy. In the image classification process, the task is to specify the presence or absence of an object, and the task of counting objects also requires inference to determine the number of instances of the objects present in the scene.

[0008] Speckle (acoustic interference) refers to the inherent granular appearance within tissue that results from the interaction of an acoustic beam with small-scale interfaces that are approximately the size of or smaller than the wavelength. These non-specular reflectors scatter the beam in all directions. The scattering from these individual small interfaces combines via an interference pattern to form the visualized granular appearance. Speckle appears as noise within the tissue, degrading spatial and contrast resolution, but gives the tissue its characteristic image texture. Speckle characteristics depend on the characteristics of the imaging system (e.g., ultrasound frequency, beam shape) and the characteristics of the tissue (e.g., size distribution of scatterers, differences in acoustic impedance). Speckle is a type of multiplicative noise that is locally correlated and can significantly degrade the performance of automated operations such as classification and segmentation, which aim to extract information valuable to the end user. To suppress speckle while maintaining the relevant image features, several approaches have been proposed. Most of these approaches rely on detailed classical statistical models of the signal and speckle, either in the original domain or the transformed domain. There is a need for alternative methods to improve US resolution in order to improve the AFC accuracy and CAD of ART.

[0009] A new field of machine learning (ML), particularly deep learning, has had a significant impact on medical imaging modalities. Deep learning (DL) is a new form of ML that has dramatically improved the performance of machine learning tasks. DL uses artificial neural networks (ANNs) composed of multiple layers of interconnected linear or non-linear mathematical transformations applied to data for the purpose of solving problems such as object classification. The level of DL performance is superior to that of conventional ML and does not require humans to identify and calculate important features. Instead, during training, DL algorithms "learn" the discriminative features that best predict the results. The human labor required for training DL systems is less because it does not require feature engineering or calculations. Regarding the medical image analysis domain, datasets are often insufficient to fully exploit the potential of DL. In the computer vision domain, transfer learning and fine-tuning are often used to address the problem of small datasets. Generally, DL algorithms recognize important features of images and adjust internal parameters to predict new data, thereby assigning appropriate weights to these features, and thus performing identification, segmentation, classification, or grading, demonstrating strong processing capabilities and complete information retention.

[0010] The superiority of deep learning-based CAD has recently been reported in a wide range of diseases, including gastric cancer, diabetic retinopathy, cardiac arrhythmia, skin cancer, and colorectal polyps. In these studies, a wide type of images, including pathological slides, electrocardiograms, and radiographic images, were investigated. Algorithms that are sufficiently trained for specific diseases can improve the diagnostic accuracy and work efficiency of physicians or medical experts, free them from repetitive work, and in particular, improve diagnostic accuracy in the presence of subtle pathological changes that cannot be detected by visual evaluation. DL algorithms can be optimized by adjusting hyperparameters such as the learning rate, network architecture, and activation function. Therefore, DL-based CAD has the potential to improve the performance of ART.

[0011] Convolutional neural networks (CNNs) or ConvNets have recently been successfully adopted for image segmentation, classification, object detection and recognition tasks, and are DL network architectures that break performance benchmarks in many difficult applications. Medical image analysis applications rely heavily on feature engineering approaches, and the algorithm pipeline is used with segmentation algorithms to explicitly depict the structure of interest, measure predefined features of these structures that are believed to be predictable, and use these features to train models that predict patient outcomes. In contrast, the feature learning paradigm of CNNs adaptively learns to transform images into highly predictable features for specific learning goals. Images and patient labels are presented to a network composed of interconnected layers of convolutional filters that emphasize important patterns within the image, and the filters and other parameters of the network are mathematically adapted to minimize prediction error. Feature learning avoids a biased prior definition of features and does not require the use of segmentation algorithms that are often confused by artifacts.

[0012] A CNN consists of multiple layers with neurons that process parts of the input image. The outputs of these neurons are arranged and displayed to form an overlap that provides a filtered representation of the original image. This process is repeated for each layer until the final output is reached, which is usually the probability of the predicted class. Training a CNN requires many iterations to optimize the network parameters. During each iteration, a batch of samples is randomly selected from the input training set and propagated forward through the network layers. To achieve optimal results, the parameters within the network are updated by backpropagation to minimize the cost function. Once trained, the network can be applied to new or unseen data to obtain predictions. The main advantage of a CNN is that it can automatically learn features from the training set without the need for expertise or hard - coding. The extracted features are relatively robust to image transformations and variations. In the field of medical imaging, CNNs have mainly been used for detection, segmentation, and classification. These tasks form part of the CAD process flow, and effective feature extraction or phenotyping of patients from EMRs is an important step in the potential further applications of this technology, such as the successful performance of ART using DL techniques, which has not been considered among experts in this field until now.

[0013] Due to the sequential nature of EMR or EHR data, there have recently been several promising studies that investigate clinical events as sequential data. Since sentences can be easily modeled as sequences of signals, many of them have been inspired by natural language modeling tasks. There is growing interest in predicting treatment prescriptions and individual patient outcomes by extracting information from these data using advanced analytical approaches. In particular, the recent success of DL in image and natural language processing has made it possible to apply these state-of-the-art techniques to the modeling of clinical data. CNNs, such as recurrent neural networks (RNNs) that have proven to be powerful in language modeling and machine translation, are frequently applied to medical event data for prediction purposes because natural language and medical records share the same sequential nature. DL, more specifically RNNs, are not considered for use in improving the performance of ART, leaving important new room for improvement in the field of ART through the application of these techniques as embodied in the disclosure of this application.

[0014] The basic components for performing ART require stimulation of the ovaries to produce multiple eggs. In a natural cycle, a typical woman makes one egg per month, alternating between the two ovaries. When using ART, administration of exogenous gonadotropin, mainly follicle-stimulating hormone (e.g., FSH), encourages each ovary to produce an average of 10 - 15 eggs that grow in fluid-filled ovarian follicles. As the follicles grow, they become increasingly dependent on gonadotropin for continued development and survival. FSH promotes the proliferation and differentiation of granulosa cells and increases the size of the follicles. The follicles grow from a resting size of 3 - 7 mm to a size of 20 mm with a 10-day dosing regimen in which the dose is adjusted based on the ovarian response. During the 10 days of dosing, US is performed 3 - 4 times to measure follicle size and monitor the response. Follicle size predicts the likelihood of an egg being present in the follicle, the quality of the egg, and the likelihood of the egg maturing.

[0015] An important component of ART success is to create the highest quality eggs possible for each individual patient. Egg quality and maturity directly predict the likelihood that an egg will be fertilized and become an embryo, and predict the quality of the embryo. The quality of the resulting embryo predicts the implantation rate, defined as (pregnancy per embryo transfer), and the overall pregnancy rate (when multiple embryo transfers are performed). The number and quality of available oocytes are important factors in the success rate of ART.

[0016] One of the challenges in patient care is that eggs do not all start the same size and do not grow at the same rate. Therefore, follicles can change size at any time during the stimulation period. Thus, the timing of a patient's egg retrieval (time to event) is based on trying to determine when the majority of the follicles are mature in size. Sometimes, in order to effectively over-stimulate some of the follicles for the purpose of obtaining the majority of the mature range of follicles, it is necessary to push the ovarian stimulation longer. This follicle monitoring technique is performed with a combination of transvaginal US and blood measurements of estradiol and progesterone. ART success can benefit from ultrasound imaging, follicle monitoring, size determination, counting, determination of hormone levels, and automated connection and adjustment of cycle days to important clinical events such as follicle maturation, egg maturation, embryo number, blastocyst embryo development, and pregnancy rate. DL may have the potential to improve ART when there is insufficient patient data for optimal timing of follicle extraction and implantation.

[0017] Survival analysis is to predict the period until an event occurs. In conventional survival modeling, it is assumed that the period follows an unknown distribution. The Cox proportional hazards model is one of the most popular models among these models. The Cox model and its extensions are constructed based on the proportional hazards hypothesis, which assumes that the hazard ratio between two instances is constant over time and the risk prediction is based on a linear combination of covariates. However, in practical clinical applications such as ART, there are too many complex interactions. A more comprehensive survival model is needed to better fit clinical data with a non-linear risk function. In addition, since the health status changes over time, a patient's EHR is essentially longitudinal. Therefore, time information is required to apply CNN to analyze a patient's EMR. DL of a patient's EMR or EHR may improve the determination of the timing of follicle extraction and implantation.

[0018] Ovulation induction (OI), the most common form of infertility treatment worldwide, includes ovulation induction with oral or injectable ovulation induction agents (e.g., clomiphene citrate, letrozole, hMH, rFSH), usually inducing the growth and maturation of a cohort of oocytes over 7 - 15 days, causing ovulation, and increasing the pregnancy rate. The most common complication, the iatrogenic multiple pregnancy rate, ranges from 5 - 30%, mainly depending on the diagnosis, aggressiveness of stimulation, and the degree of monitoring by the physician. 39 - 67% of higher-order multiple births (HOMB) are estimated to be related to OI without IVF. This dramatically increasing risk is due to the challenges of accurate follicle monitoring and appropriate dosage adjustment of OI agents. Multiple pregnancies common to the use of gonadotropins (hMH, rFSH) pose significant obstetric risks to mothers and infants, including preterm birth and low birth weight, greatly increasing neonatal, maternal, family morbidity, infant mortality, and furthermore representing a substantial economic burden to the family and society.

[0019] Doctors have long been searching for ways to predict pregnancy and multiple pregnancies during ovarian induction (OI). Follicle tracking, which is the continuous evaluation of the number and size of follicles, is commonly used to evaluate the response to ovarian stimulation. The main cause of multiple pregnancies due to OI treatment is the result of the limited ability to accurately monitor ovarian stimulation and predict the number of mature oocytes that can ovulate. The treatment goal of OI is to achieve the growth of a single dominant follicle whose size determines the maturity of the oocyte, the quality of the embryo, and the pregnancy rate. OI mainly relies on the monitoring of ovarian response after the administration of exogenous OI agents, which is mainly performed by transvaginal ultrasound (TVUS) and the measurement of plasma estradiol levels (E2) and luteinizing hormone levels (LH). Follicles of different sizes develop asynchronously during ovarian stimulation, intensifying the challenges in determining the optimal timing of ovulation induction and evaluating the risk of multiple pregnancies.

[0020] It is necessary to improve the performance of assisted reproductive technology (ART) through the efficient management of in vitro fertilization (IVF), the identification, counting, measurement, systematic improvement of differential tracking of growing follicles, and the determination of the optimal timing of OI aimed at maximizing the pregnancy rate while minimizing the risk of multiple pregnancies. The applicant has developed a solution embodied by the present disclosure, the embodiments of which are described in detail below.

Summary of the Invention

[0021] The following presents a simplified overview of some embodiments of the present invention to provide a basic understanding of the present invention. This overview is not an extensive overview of the present invention. This overview is not intended to identify important / essential elements of the present invention or to define the scope of the present invention. Its sole purpose is to present some embodiments of the present invention in a simplified form as a prelude to the more detailed description presented later.

[0022] Aspects of the present disclosure provide an ensemble of deep learning (DL) systems and methods in the provision of reproductive assistive technologies (ART) for the diagnosis, treatment, and clinical management of infertility. In various embodiments, the ensemble includes processing of at least one image including the reproductive anatomical structures of one or more patients from an imaging modality using at least one artificial neural network (ANN). In various embodiments, the ensemble includes object detection, recognition, annotation, segmentation, or classification of at least one image obtained from an imaging modality using at least one ANN. In various embodiments, the ensemble further includes at least one detection framework for detection, localization, and counting of objects using at least one ANN. In various embodiments, the ensemble includes feature extraction or phenotyping of one or more patients from electronic health or medical records using at least one ANN. In various embodiments, the ensemble further includes at least one framework for predicting the outcome over time to an event using at least one ANN. In various embodiments, the ANN includes, but is not limited to, a convolutional neural network (CNN), a recurrent neural network (RNN), a fully convolutional neural network (FCNN), a dilated residual network (DRN), a generative adversarial network (GAN), etc., or combinations thereof. The ensemble is composed of a serial or parallel combination of ANNs as an artificial intelligence computer-aided diagnosis (CAD) and prediction system for the clinical management of infertility.

[0023] Aspects of the present disclosure provide the ANN system and method for preprocessing or processing one or more imaging modalities in the provision of ART for the diagnosis, treatment, and clinical management of clinical infertility. In various embodiments, the imaging modality preferably includes, but is not limited to, ultrasound including two-dimensional (2D), three-dimensional (3D), four-dimensional (4D), Doppler, and the like. In various embodiments, the image includes, but is not limited to, reproductive anatomical structures including cells, fallopian tubes, ovaries, eggs, multiple eggs, follicles, cysts, uterus, endometrium, endometrial thickness, uterine wall, ovum, blood vessels, and the like. In various embodiments, the image includes one or more normal or abnormal forms, textures, shapes, sizes, colors, etc. of the anatomical structure. In various embodiments, the image preprocessing includes at least one speckle removal or noise removal model for improving the quality of the image to enhance image search, interpretation, diagnosis, decision-making, and the like.

[0024] Aspects of the present disclosure provide the ANN systems and methods for object detection, recognition, annotation, segmentation, or classification of at least one US image in the provision of ART for the diagnosis, treatment, and clinical management of clinical infertility. In various embodiments, the systems and methods enable the detection, recognition, annotation, segmentation, or classification of ovaries, cysts, cystic ovaries, polycystic ovaries, follicles, antral follicles, etc. In various embodiments, the ANN systems and methods include, but are not limited to, at least one of a convolutional neural network (CNN), a recurrent neural network (RNN), a fully convolutional neural network (FCNN), a dilated residual network (DRN), or a generative adversarial network (GAN) architecture. In various embodiments, the architecture includes at least one input, convolution, pooling, mapping, sampling, rectification (non-linear activation function), normalization, fully connected (FC), or output layer. In various embodiments, the convolution method includes the use of one or more patches, kernels, or filters related to the reproductive anatomical structure. In various embodiments, one or more of the ANNs are trained using one or more optimization methods. In various embodiments, the input layer of the alternative ANN includes data derived from the output layer of the ANN. In various embodiments, one or more results of detection, recognition, annotation, segmentation, or classification from the output layer of the one or more ANNs are recorded in at least one electronic health record database. In alternative embodiments, the results are transmitted and stored in a database residing on a cloud-based server.

[0025] Aspects of the present disclosure provide the ANN system and method for an object detection framework in the provision of ART for the diagnosis, treatment, and clinical management of clinical infertility. In various embodiments, the system and method enable the detection, localization, counting, and tracking of one or more reproductive anatomical structures over time from one or more US images. In various embodiments, the reproductive anatomical structures include, but are not limited to, ovaries, cysts, cystic ovaries, polycystic ovaries, follicles, oocytes, antral follicles, fallopian tubes, uterus, endometrial patterns, endometrial thickness, and the like. In various embodiments, the ANN system and method include, but are not limited to, at least one of a convolutional neural network (CNN), a recurrent neural network (RNN), a fully convolutional neural network (FCNN), a dilated residual network (DRN), or a generative adversarial network (GAN) architecture. In various embodiments, the architecture includes at least one input, convolutional, pooling, mapping, sampling, rectification (non-linear activation function), normalization, fully connected (FC), or output layer. In various embodiments, the convolution method includes the use of one or more patches, kernels, or filters related to the reproductive anatomical structure. In various embodiments, the one or more ANNs are trained using one or more optimization methods. In various embodiments, the input layer of the alternative ANN includes data derived from the output layer of the ANN. In various embodiments, one or more results of detection, localization, counting, and tracking from the output layer of the one or more ANNs are recorded in at least one electronic health record database. In alternative embodiments, the results are transmitted and stored in a database residing on a cloud-based server.

[0026] Aspects of the present disclosure include the ANN system and method for analyzing electronic medical records in the provision of ART for the diagnosis, treatment, and clinical management of clinical infertility. In various embodiments, the system and method enable the extraction of features or phenotyping of one or more patients from at least one multi-year patient's electronic medical record (EMR), electronic health record (EHR), database, etc. In various embodiments, the medical records include one or more stored patient records, preferably records of patients undergoing infertility treatment, ultrasound images, ultrasound manufacturers, ultrasound models, ultrasound probes, ultrasound frequencies, images of the reproductive anatomical structures, patient age, patient ethnic background, patient demographics, physician notes, clinical notes, physician annotations, diagnostic results, body fluid biomarkers, dosage, days of drug treatment, hormone markers, hormone levels, neohormones, endocannabinoids, genomic biomarkers, proteomic biomarkers, anti-Müllerian hormone, estradiol, estrone, progesterone, FSH, luteinizing hormone (LH), inhibin, renin, relaxin, VEGF, creatine kinase, hCG, fetoprotein, pregnancy-specific b-l-glycoprotein, pregnancy-associated plasma protein-A, placental protein-14, follistatin, IL-8, IL-6, vitellogenin, calvingin-D9k, treatment procedures, treatment schedules, implantation schedules, implantation rates, follicle sizes, follicle numbers, AFC, follicle growth rates, pregnancy rates, times of implantation (i.e., events), CPT codes, HCPCS codes, ICD codes, etc. In various embodiments, one or more of the medical record fields are converted into one or more temporal matrices, preferably having time as one dimension and specific events as another dimension. In various embodiments, the ANN architecture includes at least one input, convolution, pooling, mapping, sampling, rectification (non-linear activation function), normalization, fully connected (FC), prediction, or output layer. In various embodiments, the convolution method includes the use of one or more patches, kernels, or filters related to factors for predicting follicle extraction and time of implantation. In various embodiments, one or more of the ANNs are trained using one or more optimization methods.In various embodiments, the input layer of the alternative ANN includes data derived from the output layer of the ANN. In various embodiments, the phenotypic or predictive results of one or more identified patients from the output layer of the one or more ANNs are recorded in at least one electronic health record database. In alternative embodiments, the results are transmitted and stored in a database residing on a cloud-based server.

[0027] Aspects of the present disclosure provide the ANN system and method for predictive planning in the provision of ART for the diagnosis, treatment, and clinical management of clinical infertility. In various embodiments, the ANN system and method include at least one framework for predicting the outcome of time to an event. In various embodiments, the outcomes of time to an event include, but are not limited to, start-to-end follicle stimulation, cycle days, follicle retrieval, follicle recruitment, follicle retrieval, follicular phase, follicle maturation, fertilization rate, blastocyst embryogenesis, embryo quality, implantation, etc. In various embodiments, the ANN architecture includes at least one input, convolution, pooling, mapping, sampling, rectification (non-linear activation function), normalization, fully connected (FC), Cox model, and output layer. In various embodiments, the convolution method includes the use of one or more patches, kernels, or filters related to factors for predicting the time to an event. In various embodiments, one or more of the ANNs are trained using one or more optimization methods. In various embodiments, one or more of the predictions are compared to the patient's outcome to adaptively train the network weights of one or more interconnected layers. In various embodiments, the input layer of the alternative ANN includes data derived from the output layer of the ANN. In various embodiments, the results of time to one or more events obtained from the output layer of the one or more ANNs are recorded in at least one electronic health record database. In alternative embodiments, the results are transmitted and stored in a database residing on a cloud-based server.

[0028] Aspects of the present disclosure provide a computer program product for use in providing ART for the diagnosis, treatment, and clinical management of clinical infertility. In various embodiments, the product comprises at least one patient data, US images from a US scanner / device, retrieved US images, patient medical records from an electronic medical record database, patient records related to childbirth, patient endocrinology records, patient clinical notes, physician clinical notes, data from a database on the cloud-based server, results from one or more output layers of one or more of the ANNs, and an artificial intelligence engine. In various embodiments, the artificial intelligence engine incorporates one or more results from one or more of the ANNs to generate one or more clinical insights. In various embodiments, the cloud-based server comprises one or more user applications in combination with one or more browsers, enabling a user to access clinical information, perform further data processing or analysis, and obtain or receive one or more clinical insights. In various embodiments, the user accesses the information using a mobile computing device or a desktop computing unit. In various embodiments, a mobile application enables a user to access information from the computer product.

[0029] Aspects of the present disclosure provide a computer program product for use in providing ovulation induction (OI) treatment and clinical management of clinical infertility. In various embodiments, the product comprises at least one patient data, US images from a US scanner / device, retrieved US images, patient medical records from an electronic medical record database, patient records related to childbirth, patient endocrinology records, patient clinical notes, physician clinical notes, data from a database on the cloud-based server, results from one or more output layers of one or more of the ANNs, and an artificial intelligence engine. In various embodiments, the artificial intelligence engine incorporates one or more results from one or more of the ANNs to generate one or more clinical insights. In various embodiments, the cloud-based server comprises one or more user applications in combination with one or more browsers, enabling a user to access clinical information, perform further data processing or analysis, and obtain or receive one or more clinical insights. In various embodiments, the user accesses the information using a mobile computing device or a desktop computing unit. In various embodiments, a mobile application enables a user to access information from the computer product.

[0030] Certain embodiments of the present disclosure provide a computer-aided diagnosis and prediction system for the clinical management of infertility, the system comprising: an imaging sensor operable to perform one or more imaging modalities to collect one or more images of a subject's reproductive anatomical structures; a storage device for locally or remotely storing the one or more images of the subject's reproductive anatomical structures; at least one processor operably engaged with at least one computer-readable storage medium storing computer-executable instructions, the at least one processor performing one or more actions when the computer-executable instructions are executed, the one or more actions including: receiving the one or more images of the subject's reproductive anatomical structures; processing the one or more images of the subject's reproductive anatomical structures to detect one or more reproductive anatomical structures and annotate one or more anatomical features of the one or more reproductive anatomical structures; comparing the one or more anatomical features with at least one linear or non-linear framework (i.e., a machine learning framework) to predict the result of the time to at least one event; and generating at least one graphical user output corresponding to one or more clinical actions of the subject.

[0031] Certain further embodiments of the present disclosure provide a computer-aided diagnosis and prediction system for the clinical management of infertility, the system comprising: an imaging sensor operable to perform one or more imaging modalities to collect one or more images of a patient's reproductive anatomical structure; an artificial intelligence engine configured to receive locally or remotely one or more digital images of the patient's reproductive anatomical structure, the artificial intelligence engine being configured to process one or more images of the patient's reproductive anatomical structure and generate a result prediction of time to at least one event according to at least one linear or non-linear framework (i.e., a machine learning framework); an outcome database configured to communicate clinical outcome data to the artificial intelligence engine, the clinical outcome data being incorporated into at least one linear or non-linear framework; an application server operably coupled to the artificial intelligence engine to receive a result prediction of time to at least one event, the application server being configured to generate one or more proposed clinical actions for the clinical management of infertility in response to the result prediction of time to at least one event; and a client device communicably coupled to the application server, the client device being configured to display a graphical user interface including one or more proposed clinical actions for the clinical management of infertility.

[0032] Yet another specific embodiment of the present disclosure provides at least one computer-readable storage medium storing computer-executable instructions that, when executed, perform a method for predicting clinical outcomes related to the provision of reproductive assistance technologies, the method comprising receiving one or more digital images of a patient's reproductive anatomical structures, processing the one or more digital images of the patient's reproductive anatomical structures to detect one or more reproductive anatomical structures and annotating one or more anatomical features of the one or more reproductive anatomical structures, analyzing the one or more anatomical features according to at least one linear or non-linear framework, and predicting the outcome of the time to at least one event according to at least one linear or non-linear framework.

[0033] A further aspect of the present disclosure provides a method for processing digital images in reproductive assistance technologies, the method comprising obtaining one or more digital images of a patient's reproductive anatomical structures through one or more imaging modalities, processing the one or more digital images to detect one or more reproductive anatomical structures, processing the one or more digital images to annotate, segment, or classify one or more anatomical features of the one or more reproductive anatomical structures, analyzing the one or more anatomical features according to at least one linear or non-linear framework (i.e., a machine learning framework), and predicting the outcome of the time to at least one event of a reproductive assistance procedure according to at least one linear or non-linear framework.

[0034] A further aspect of the present disclosure provides a method of image processing for a clinical plan in reproductive assistance technology, the method comprising receiving one or more digital images of a patient's reproductive anatomical structures through one or more imaging modalities; processing the one or more digital images to detect one or more reproductive anatomical structures; processing the one or more digital images to annotate, segment, or classify one or more anatomical features of the one or more reproductive anatomical structures; analyzing the one or more anatomical features according to at least one linear or non-linear framework (i.e., a machine learning framework); predicting the result of the time to at least one event of a reproductive assistance procedure according to at least one linear or non-linear framework; and generating one or more clinical recommendations related to the reproductive assistance procedure.

[0035] Yet another aspect of the present disclosure is a method for the clinical management of infertility, the method comprising obtaining an ovarian ultrasound image of an ovarian follicle of a subject using an ultrasound device; analyzing the ovarian ultrasound image according to at least one linear or non-linear framework (i.e., a machine learning framework) to annotate, segment, or classify one or more anatomical features of the ovarian follicles of the subject to predict the result of the time to an event; and generating one or more clinical recommendations regarding a reproductive assistance procedure.

[0036] Yet another aspect of the present disclosure is a method for the clinical management of infertility, the method comprising obtaining an ovarian ultrasound image of an ovarian follicle of a subject using an ultrasound device; annotating, segmenting, or classifying one or more anatomical features of the ovarian follicle of the subject to count, measure, characterize, and monitor the growth rate of the size thereof according to at least one linear or non-linear framework (i.e., a machine learning framework); and generating one or more clinical recommendations regarding the optimal timing of OI for the purpose of maximizing the pregnancy rate while minimizing the risk of multiple pregnancies.

[0037] Yet another aspect of the present disclosure provides a method for the clinical management of infertility, the method comprising obtaining an ovarian ultrasound image of an ovarian follicle of a subject using an ultrasound device; analyzing the ovarian ultrasound image according to at least one linear or non-linear framework to annotate, segment, or classify one or more anatomical features of the ovarian follicle of the subject to predict the outcome of the time to an event; and generating one or more clinical recommendations regarding the optimal timing of OI for the purpose of maximizing the pregnancy rate while minimizing the risk of multiple pregnancies.

[0038] Certain aspects of the present disclosure provide a system for digital image processing in assisted reproductive technology, the system comprising an imaging sensor configured to collect one or more digital images of a patient's reproductive anatomical structure, a computer device communicatively coupled to the imaging sensor to receive the one or more digital images of the patient's reproductive anatomical structure, at least one processor communicatively coupled to the computer device, and at least one non-transitory computer-readable medium having instructions that, when executed, cause the at least one processor to perform one or more operations, the one or more operations comprising receiving the one or more digital images of the patient's reproductive anatomical structure, processing the one or more digital images of the patient's reproductive anatomical structure to detect one or more reproductive anatomical structures and annotate one or more anatomical features of the one or more reproductive anatomical structures, analyzing the one or more anatomical features according to at least one machine learning framework to predict a result of time to at least one event, the result of time to at least one event including an ovulation induction day within an ovulation induction cycle for the patient, and generating at least one graphical user output corresponding to one or more clinical actions related to the patient, the at least one graphical user output including a proposed timing for administration of at least one pharmaceutical to the patient, the at least one pharmaceutical including an ovulation inducer.

[0039] According to certain embodiments of a system for digital image processing in assisted reproductive technology, one or more clinical actions may include the proposed timing of sperm delivery or intrauterine insemination corresponding to an ovulation induction cycle. One or more operations of a processor may further include analyzing multiple electronic health record data of a patient, along with one or more anatomical features, to predict the result of the time to at least one event. The multiple electronic health record data may include one or more data sets selected from the group consisting of diagnostic results, body fluid biomarkers, hormone markers, hormone levels, genomic biomarkers, proteomic biomarkers, therapeutic treatments, treatment schedules, follicle size and number, follicle growth rate, pregnancy rate, and ovulation induction data. In some embodiments, one or more operations of the processor may further include analyzing multiple anonymized historical data from one or more anonymized ovulation induction patients, along with one or more anatomical features, to predict the result of the time to at least one event. The multiple anonymized historical data may include one or more data sets selected from the group consisting of diagnostic results, body fluid biomarkers, hormone markers, hormone levels, genomic biomarkers, proteomic biomarkers, treatment procedures, treatment schedules, follicle size and number, follicle growth rate, pregnancy rate, and ovulation induction data.

[0040] According to certain embodiments of a system for digital image processing in assisted reproductive technology, the machine learning framework can be selected from the group consisting of artificial neural networks, regression models, convolutional neural networks, recurrent neural networks, fully convolutional neural networks, extended residual networks, and adversarial generative networks. In some embodiments, one or more reproductive anatomical structures include one or more ovarian follicles, and one or more anatomical features include the quantity and size of one or more ovarian follicles. In some embodiments, one or more operations of the processor can include receiving the patient's reproductive physiology data and analyzing the reproductive physiology data, along with one or more anatomical features, to predict the result of the time to at least one event. One or more operations of the processor can further include analyzing one or more anatomical features according to at least one machine learning framework to evaluate the patient's risk of multiple pregnancy.

[0041] Certain aspects of the present disclosure provide a method for processing digital images in reproductive assistance technology, the method comprising obtaining, by an ultrasonic device, one or more digital images of a patient's reproductive anatomical structure; receiving, by at least one processor, the one or more digital images; processing, by at least one processor, the one or more digital images to detect one or more reproductive anatomical structures of the patient's reproductive anatomical structure; processing, by at least one processor, the one or more digital images to annotate, segment, or classify one or more anatomical features of the one or more reproductive anatomical structures; analyzing, by at least one processor, the one or more anatomical features according to at least one machine learning framework to predict a result of time to at least one event, wherein the time to at least one event includes an ovulation induction day within the patient's ovulation induction cycle; and generating, by at least one processor, at least one clinical recommendation including a proposed timing of administration of at least one pharmaceutical to the patient, wherein the at least one pharmaceutical includes an ovulation inducer.

[0042] According to certain embodiments of the method of digital image processing in reproductive assistance technology, the one or more clinical procedures may include a proposed timing of sperm delivery or intrauterine insemination corresponding to the ovulation induction cycle. The one or more reproductive anatomical structures include one or more ovarian follicles, and the one or more anatomical features include the quantity and size of the one or more ovarian follicles. In some embodiments, the method may further include analyzing, using at least one processor, the one or more anatomical features according to at least one machine learning framework to evaluate the patient's risk of multiple pregnancy. The method may further include determining, using at least one processor, the maturation rate of the patient's one or more ovarian follicles by analyzing the one or more anatomical features according to at least one machine learning framework.

[0043] According to certain embodiments of a method for digital image processing in assisted reproductive technology, the method may include using at least one processor to analyze a plurality of electronic health record data of a patient, along with one or more anatomical features, to predict the result of the time to at least one event. In some embodiments, the plurality of electronic health record data includes one or more data sets selected from the group consisting of diagnostic results, body fluid biomarkers, hormone markers, hormone levels, genomic biomarkers, proteomic biomarkers, treatment procedures, treatment schedules, follicle size and number, follicle growth rate, pregnancy rate, and ovulation induction data. The method may further include using at least one processor to analyze a plurality of anonymized historical data from one or more anonymized ovulation induction patients, along with one or more anatomical features, to predict the result of the time to at least one event. In some embodiments, the plurality of anonymized historical data includes one or more data sets selected from the group consisting of diagnostic results, body fluid biomarkers, hormone markers, hormone levels, genomic biomarkers, proteomic biomarkers, treatment procedures, treatment schedules, follicle size and number, follicle growth rate, pregnancy rate, and ovulation induction data.

[0044] A further embodiment of the present disclosure is a non-transitory computer-readable storage medium storing computer-executable instructions that, when executed, cause at least one processor to perform one or more operations of a method for digital image processing, the one or more operations including receiving one or more digital images of a patient's reproductive anatomical structure, processing the one or more digital images of the patient's reproductive anatomical structure to detect one or more reproductive anatomical structures and annotate one or more anatomical features of the one or more reproductive anatomical structures, analyzing the one or more anatomical features according to at least one machine learning framework to predict a result of time to at least one event, wherein the time to at least one event includes an ovulation induction day within the patient's ovulation induction cycle, and generating at least one graphical user output corresponding to one or more clinical actions related to the patient, wherein the one or more clinical actions include a proposed timing for administration of at least one pharmaceutical to the patient, and the at least one pharmaceutical includes an ovulation inducer.

[0045] The above fairly broadly outlines more appropriate and important features of the present invention, and as a result, a better understanding of the following detailed description of the present invention can be achieved, and a more complete understanding of the contribution of the present invention to the art can be achieved. Additional features of the present invention that form the subject matter of the claims of the present invention are described below. Those skilled in the art should understand that the concepts as well as the specific methods and structures disclosed can be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present invention. Those skilled in the art should understand that such equivalent structures do not depart from the spirit and scope of the present invention as set forth in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The above and other objects, features and other advantages of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings.

Figure 1A

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 7B

Figure 7C

Figure 7D

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 14B

Figure 15

Figure 15B

Mode for Carrying Out the Invention

[0047] It should be understood that all combinations of concepts discussed in more detail below (provided that such concepts are not mutually inconsistent) are intended to be part of the subject matter of the invention disclosed herein. It should also be understood that terms explicitly used herein that may also appear in any disclosure incorporated by reference should be given the meaning that most closely matches the particular concepts disclosed herein.

[0048] Embodiments of the present disclosure provide an ensemble of deep learning (DL) systems and methods in the provision of reproductive assistance technology (ART) for the diagnosis, treatment, and clinical management of infertility. The ensemble comprises computer vision techniques (e.g., convolutional neural networks) for the detection, recognition, annotation, segmentation, classification, counting, and tracking of objects of reproductive anatomical structures such as follicles and ovaries, and one or more artificial neural network (ANN) systems and methods for speckle removal or noise processing of ultrasound images. The ANN systems and methods are also assembled to analyze electronic medical records to identify and phenotype the characteristics of a patient's outcome up to an event to predict the optimal extraction of follicles and the timing of implantation for a patient. The methods and systems are incorporated into cloud-based computer programs and mobile applications that enable physicians and patients to access clinical insights in infertility clinical and patient management.

[0049] Since the disclosed concepts are not limited to specific implementation methods, it should be understood that the various concepts introduced above and discussed in more detail below may be implemented in any of a number of ways. Examples of specific embodiments and application forms are given primarily for purposes of illustration. The present disclosure should in no way be limited to the exemplary implementations and techniques shown in the drawings and described below.

[0050] Where a range of values is provided, each intervening value, to the tenth of the unit of the lower limit between the upper and lower limits of that range, and any other stated or intervening value in the stated range, is understood to be included in the invention, unless the context clearly dictates otherwise. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also included in the invention, subject to any specifically excluded limit within the stated range. Where the stated range includes one or both of the end point limitations, ranges excluding one or both of those included end points are also included in the scope of the invention.

[0051] As used herein, "exemplary" means serving as an example or illustration and does not necessarily indicate an ideal or best one.

[0052] As used herein, the term "includes" means including but not limited to, and the term "including" means including but not limited to. The term "based on" means based at least in part on.

[0053] Turning now to the drawings, where like reference numerals refer to like elements throughout the several views, FIG. 1A shows a computing system that can implement a particular, illustrated embodiment of the present invention.

[0054] Referring now to FIG. 1A, a processor-implemented computing device that may implement one or more aspects of the present disclosure is shown. According to one embodiment, the processing system 100a generally includes at least one processor 102a, or processing unit or multiple processors, memory 104a, at least one input device 106a, and at least one output device 108a, which are coupled together via one bus or a group of buses 110a. In certain embodiments, the input device 106a and the output device 108a may be the same device. Interface 112a may also be provided to couple the processing system 100a to one or more peripheral devices. For example, interface 112a may be a PCI card or a PC card. At least one storage device 114a that houses at least one database 116a can also be provided. Memory 104a can be any form of memory device, such as volatile or non-volatile memory, solid-state storage device, magnetic device, etc. Processor 102a may include one or more separate processing devices, for example, to process different functions within the processing system 100a. The input device 106a receives input data 118a and may include, for example, a pointer device such as a keyboard, pen device, or mouse, an audio receiving device for voice control activation such as a microphone, a data receiver or antenna such as a modem or wireless data adapter, a data acquisition card, etc. The input data 118a can be obtained from different sources, such as keyboard commands combined with data received via a network. The output device 108a creates or generates output data 120a and may include, for example, a display device or monitor if the output data 120a is visual, a printer if the output data 120a is to be printed, a port such as a USB port, a peripheral component adapter, a data transmitter or antenna such as a modem or wireless network adapter. The output data 120a is separate from and / or can be obtained from different output devices, such as a visual display on a monitor combined with data transmitted to a network.The user can display data output or the interpretation of data output, for example, on a monitor or using a printer. The storage device 114a can be any form of data or information storage means, such as volatile or non-volatile memory, a solid-state storage device, a magnetic device, etc.

[0055] During use, the processing system 100a is adapted to enable data or information to be stored in and / or retrieved from at least one storage structure (e.g., a database) 116a via wired or wireless communication means. The interface 112a may enable wired and / or wireless communication between the processing unit 102a and peripheral components that can serve special purposes. Generally, the processor 102a can receive instructions as input data 118a via the input device 106a and display the processed results or other outputs to the user by utilizing the output device 108a. One or more input devices 106a and / or output devices 108a can be provided. It should be recognized that the processing system 100a can be any form of terminal, server, dedicated hardware, etc.

[0056] It should be recognized that the processing system 100a may be part of a networked communication system. The processing system 100a can be connected to a network, such as the Internet or a WAN. The input data 118a and the output data 120a can communicate with other devices via the network. The transfer of information and / or data via the network can be achieved using wired communication means or wireless communication means. The server can facilitate the transfer of data between the network and one or more databases. The server and one or more databases provide examples of information sources.

[0057] Accordingly, the processing computing system environment 100a shown in FIG. 1A may operate in a networked environment using logical connections to one or more remote computers. In an embodiment, the remote computer can be a personal computer, a server, a router, a network PC, a peer device, or other common network node, and typically includes many or all of the above elements.

[0058] It should be further recognized that the logical connections shown in FIG. 1A include local area networks (LANs) and wide area networks (WANs), but may also include other networks such as personal area networks (PANs). Such networking environments are common in offices, enterprise-wide computer networks, intranets, and the Internet. For example, when used in a LAN networking environment, the computing system environment 100a is connected to the LAN via a network interface or adapter. When used in a WAN networking environment, the computing system environment typically includes other means for establishing communications via a WAN such as a modem or the Internet. The modem, which may be internal or external, may be connected to the system bus via a user input interface or another suitable mechanism. In a networked environment, program modules depicted in relation to the computing system environment 100a or a portion thereof may be stored in a remote memory storage device. It should be recognized that the network connections shown in FIG. 1A are exemplary, and other means of establishing communication links between multiple computers may be used.

[0059] FIG. 1A is intended to provide a concise and general description of an exemplary and / or suitable exemplary environment in which embodiments of the present invention may be implemented. That is, FIG. 1A is merely an example of a suitable environment and is not intended to imply any limitations regarding the structure, scope of use, or functionality of the embodiments of the present invention illustrated therein. A particular environment should not be construed as having any dependencies or requirements related to any one or particular combination of the components shown in the illustrated operating environment. For example, in certain instances, one or more elements of the environment may be considered unnecessary and may be omitted. In other instances, one or more other elements may be considered necessary and may be added.

[0060] In the following description, certain embodiments may be described with reference to symbolic representations of acts and operations performed by one or more computing devices, such as the computing system environment 100a of FIG. 1A. Accordingly, it will be understood that such acts and operations, sometimes referred to as computer-executed, include the manipulation by a computer's processor of electrical signals representing structured data. This manipulation transforms the data or maintains it at locations in the computer's memory system, reconfiguring or changing the operation of the computer in ways well understood by those skilled in the art. A data structure in which the data is maintained is a physical location in memory having specific characteristics defined by the format of the data. However, while certain embodiments may be described in the foregoing context, those skilled in the art will understand that the acts and operations described below may also be implemented in hardware, and thus the scope of the present disclosure is not meant to be limited thereto.

[0061] Embodiments may be implemented in many other general purpose or special purpose computing devices and computing system environments or configurations. Examples of well-known computing systems, environments, and configurations suitable for use in embodiments of the present invention include, but are not limited to, personal computers, handheld or laptop devices, personal digital assistants, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, networks, minicomputers, server computers, game server computers, web server computers, mainframe computers, and distributed computing environments including any of the above systems or devices.

[0062] Embodiments can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. One embodiment may also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including memory storage devices.

[0063] Since the exemplary computing system environment 100a of FIG. 1A has been generally shown and considered, the description is directed to the illustrated embodiments of the present disclosure.

[0064] Referring to FIG. 1 here, an architecture diagram of a convolutional neural network architecture 100 is shown. A convolutional neural network (CNN) is a special type of artificial neural network (ANN). The basic difference between the densely connected layers and convolutional layers of an ANN is that an ANN learns global patterns in the input feature space. In contrast, a convolutional layer typically learns local patterns that are small 2D windows, patches, filters, or kernels 102 of the input image (or input layer) 104. The patterns learned by a CNN are translation invariant, enabling global pattern recognition within an image or sequence. A CNN can also learn a spatial hierarchy of patterns, such that the first convolutional layer (or hidden layer) 106 can learn small local patterns such as edges, and additional or subsequent layers can learn larger patterns that include the features of the previous or first layer.

[0065] A CNN learns a highly non-linear mapping by interconnecting layers of artificial neurons arranged in many different layers with non-linear activation functions. A CNN architecture includes one or more convolutional layers 106, 110 interspersed with one or more subsampling layers 108, 112 or non-linear layers, followed by one or more fully connected layers 114, 116. Each element of a CNN receives an input from a series of functions of the previous layer. Since neurons within the same feature map (or output image) 120 have the same weights or parameters, a CNN learns simultaneously. These locally shared weights reduce the complexity of the network such that when multi-dimensional input data enters the network, the CNN reduces the complexity of feature extraction and data reconstruction in the regression or classification process.

[0066] In mathematics, a tensor is a geometric object that multilinearly maps geometric vectors, scalars, and other tensors to a resulting tensor. Convolution operates on a 3D tensor (e.g., a vector) called a feature map (e.g., 120), which has two spatial axes (height and width) as well as a depth axis (also called the channel axis). In general computer vision, a CNN is usually designed to classify color images that typically contain three image channels of red, green, and blue (RGB). In the case of an RGB image, since the image has three color channels of red, green, and blue, the dimension of the depth axis is 3. In the case of a black-and-white picture, the depth is 1 (i.e., the level of gray). The convolution operation extracts patch 122 from its input feature map and applies the same transformation to all these patches to generate an output feature map 124. This output feature map remains a 3D tensor with width and height. The output depth is a parameter of the layer, and the different channels of its depth axis do not substitute for a specific color as in the RGB input, but rather act as filters, and its depth can be arbitrary. Filters encode specific aspects of the input data at a high level. A single filter can be encoded, for example, by shape, texture, or the size of follicles.

[0067] Convolution is defined by the following two main parameters: (1) The size of the patch extracted from the input -- usually 1×1, 3×3, or 5×5, and (2) The depth of the output feature map -- the number of filters calculated by the convolution. Generally, these start at a depth of 32, continue up to a depth of 64, and end at a depth of 128 or 256.

[0068] Convolution operates by sliding these windows of size 3×3 or 5×5 over the 2D or 3D input feature map, stopping at every location and extracting a patch 122 of the surrounding features [shape(window Height, window Width, input Depth)]. Next, each such patch is transformed into an ID vector of shape (output_depth) via a tensor product with the same learned weight matrix (referred to as the convolution kernel). Then, all these vectors are spatially reconfigured to, for example, a 3D output map of shape (height, width, output depth). All spatial positions of the output feature map correspond to the same positions in the input feature map (e.g., the lower right corner of the output contains information regarding the lower right corner of the input).

[0069] During training, the CNN is adjusted or trained so that the input data leads to a specific output estimate. The CNN is adjusted using backpropagation based on a comparison of the output estimate and the ground truth (i.e., the true label) until the output estimate gradually matches or approaches the ground truth. The CNN is trained by adjusting the weights (w) or parameters between neurons based on the difference between the ground truth and the actual output. The weights between neurons are free parameters that capture the data representation of the model and are learned from input / output samples. The goal of model training is to find the parameters (w) that minimize the objective loss function L(w), which measures the fit between the prediction of the model parameterized by w and the actual observations or true labels of the samples. The most common objective loss functions are cross-entropy for classification and mean squared error for regression. In other implementations, convolutional neural networks use various loss functions such as Euclidean loss and softmax loss.

[0070] Currently, CNNs are trained using stochastic gradient descent (SGD) with mini - batches. SGD is an iterative method for optimizing a differentiable objective function (e.g., a loss function) and is a stochastic approximation of gradient - descent optimization. Many variations of SGD are used to accelerate learning. Some common heuristics such as AdaGrad, AdaDelta, RMSprop, etc. adjust the learning rate adaptively to each feature. Perhaps the most popular AdaGrad adapts the learning rate by caching the sum of the squares of the gradients with respect to each parameter at each time step. The step size for each feature is multiplied by the reciprocal of the square root of this cached value. Using AdaGrad leads to faster convergence on convex error surfaces, but since the cached sum increases monotonically, the learning rate of this method decreases monotonically, which may not be desirable on non - convex error surfaces. The momentum method is another common SGD variant used in neural network training. These methods add a decaying sum of previous updates to each update. In other implementations, the gradients are computed using only selected data pairs fed to Nesterov's accelerated gradient and adaptive gradient to inject computational efficiency. The main drawback of training using gradient descent and its variants is that a large amount of labeled data is required. One way to address this difficulty is to rely on the use of unsupervised learning. Data augmentation is essential for teaching the network properties of target invariance and robustness when there are few training samples available.

[0071] The convolutional layers of a CNN (e.g., 106, 110) function as feature extractors. The convolutional layers act as adaptive feature extractors that can learn the input data and decompose it into hierarchical features. In one implementation, a convolutional layer takes two images as input and generates a third image as output. In such an implementation, the convolution operates on two two-dimensional (2D) images, one being the input image 104 and the other being the kernel (e.g., 102), which is applied as a filter to the input image 104 to generate an output image. The convolution operation involves sliding the kernel 102 over the input image 104. For each position of the kernel 102, the overlapping values of the kernel and the input image 104 are multiplied and the results are added. The sum of the products is the value of the output image 120 at the point within the input image 104 where the kernel 102 is centered. The various outputs obtained from multiple kernels are called feature maps (e.g., 120, 124).

[0072] When the convolutional layers (e.g., 106, 110) are trained, they are applied to perform recognition tasks on new inference data. Since the convolutional layers learn from the training data, they avoid explicit feature extraction and learn implicitly from the training data. The convolutional layers use the weights of the convolutional filter kernels that are determined and updated as part of the training process. The convolutional layers extract different features of the input image 104, and these are combined in higher layers (e.g., 108, 110, 112). A CNN uses various numbers of convolutional layers, each having different convolutional parameters such as kernel size, stride, padding, number of feature maps, weights, etc.

[0073] The subsampling layer (e.g., 108, 112) reduces the resolution of the features extracted by the convolutional layer, making the extracted features or feature maps (e.g., 120, 124) robust to noise and distortion, reducing computational complexity, introducing invariance properties, and reducing the possibility of overfitting. This is a summary of the statistics of features across regions within the image. In one implementation, the subsampling layer (e.g., 108, 112) uses two types of pooling operations: average pooling and max pooling. The pooling operation divides the input into non-overlapping two-dimensional spaces. In the case of average pooling, the average of the four values within the region is calculated for pooling. The output of the pooling neuron is the average value of the input values present in the input neuron set. In the case of max pooling, the maximum of the four values is selected for pooling. Max pooling identifies the most predictive features within the sampled region and reduces the image resolution and memory requirements.

[0074] In a CNN, a non-linear layer is implemented for neuron activation in combination with convolution. The non-linear layer uses various non-linear trigger functions to signal the clear identification of potential features on each hidden layer (e.g., 106, 110). The non-linear layer implements non-linear triggers using various specific functions, including Rectified Linear Unit (ReLU), Parametric Rectified Linear Unit (PreLU), hyperbolic tangent, absolute value of hyperbolic tangent, and sigmoid and continuous trigger (non-linear) functions. In a preferred implementation, ReLU is used for activation. The advantage of using the ReLU function is that the convolutional neural network is trained many times faster. ReLU is a discontinuous and non-saturating activation function that is linear with respect to the input when the input value is greater than zero and zero otherwise. In other implementations, the non-linear layer uses the activation function of the power unit.

[0075] CNN can also implement residual connections that involve reinjecting the previous representation into the downstream flow of data by adding the past output tensor to the later output tensor, thereby preventing loss of information along the data processing flow. Residual connections address two common problems that plague large-scale deep learning models, namely vanishing gradients and representation bottlenecks. With residual connections, the output of the previous layer becomes available as input to the later layer, effectively creating shortcuts in sequential networks. Instead of being concatenated to later activations, the previous output is often summed with the later activation, assuming that both activations are of the same size. If they are of different sizes, a linear transformation can be used to reshape the previous activation into the target shape.

[0076] Residual learning in CNNs was initially proposed to solve the problem of performance degradation, in which the training accuracy begins to decline as the depth of the network increases. By assuming that the residual mapping can be learned much more easily than the original un-referenced mapping, the residual network explicitly learns the residual mapping of several stacked layers. Residual networks stack a number of residual units to mitigate the decline in training accuracy. Residual blocks utilize special additive skip connections to address the vanishing gradients in deep neural networks. At the start of the residual block, the data flow is split into two streams. The first conveys the unchanged input of the block, and the second applies weights and non-linearity. At the end of the block, the two streams are merged using element-wise summation (or subtraction). The main advantage of such a configuration is that it allows gradients to flow through the network more easily. Residual networks enable CNNs to be trained easily and improve the accuracy of applications such as image classification and object detection.

[0077] A known problem in deep learning is the covariate shift where the distribution of network activations changes across layers due to the change of network parameters during training. The change in the scale and distribution of the input at each layer means that the network has to adapt its parameters significantly at each layer, thereby slowing down the training (i.e., using a small learning rate) in order to keep the loss decreasing during training (to avoid divergence during training). The general problem of covariate shift is the difference in the distribution between the training and test sets, which may lead to sub-optimal generalization performance.

[0078] In one implementation, Batch Normalization (BN) has been proposed to reduce internal covariate shift by incorporating a normalization step, a scale step, or a shift step. BN is a method to accelerate deep network training by making the normalization of data an integral part of the network architecture. BN guarantees a more regular distribution for all inputs. BN can also adaptively normalize the data as the mean and variance change over time during training. It internally maintains the exponential moving averages of the mean and variance data for each batch. The main effect is to assist in gradient propagation similar to the remaining connections. The BN layer can be used after a convolutional layer, a dense layer, or a fully connected layer, but before the output is fed into the activation function. In the case of a convolutional layer, various elements of the same feature map (i.e., activations at different locations) are normalized in the same way to follow the convolutional characteristics. Thus, all activations within a mini-batch are normalized not activation-by-activation, but at all locations.

[0079] Typically, one or more fully connected (FC) layers 114, 116 follow one or more convolutional layers 106, 110 in which one or more subsampling layers 108, 112 are interspersed. The FC layers are used to concatenate multi-dimensional feature maps (e.g., 120, 124, etc.), turn the feature maps into fixed-size categories, and generate feature vectors for the classification output layer 118. The FC layer is usually the layer that uses the most parameters and connections intensively. In one implementation, global average pooling is used to reduce the number of parameters, and optionally, one or more FC layers are replaced for classification by obtaining the spatial average of the features of the last layer for scoring. This reduces the training load and avoids the problem of overfitting. The main idea of global average pooling is to generate an average value from the feature maps of each final layer as a confidence factor for scoring and feed it directly to, for example, a softmax layer that maps 3D input to [0, 1]. This allows one or more output layers 118 to be interpreted as probabilities and select pixels (2D input) or voxels (3D input) with the highest probability.

[0080] In one implementation, one or more autoencoders are used for dimensionality reduction. An autoencoder is a neural network trained to reconstruct input data, and dimensionality reduction is achieved using a smaller number of neurons in the hidden layers (e.g., 106, 110, etc.) than in the input layer 104. A deep autoencoder is obtained by stacking multiple layers of encoders that are each trained individually (pre-training) using unsupervised learning criteria. A classification layer can be added to the pre-trained encoder and further trained with labeled data (fine-tuning).

[0081] Figure 2 is a schematic 200 of a deep learning (DL) pipeline development process according to various embodiments. The development of a DL pipeline consists of three phases: model selection (model selection and adaptation to a training dataset), model evaluation, and model distribution. In a preferred embodiment, an infrastructure is created for training, evaluating, and distributing one or more ANN networks. In various embodiments, one or more datasets are correctly separated via a dataset flow 202 and a data splitting step 204 into a test dataset 206, validation data 208, and training data 210 to avoid a biased evaluation. In various embodiments, one or more of the datasets are processed via a data IO step 212 in various different ways depending on the phase of the pipeline and then sampled 214. In applications where data is limited, the dataset is augmented 216 to supplement a small training dataset, but the training dataset is too sparse to represent the variability of the image distribution. Data augmentation artificially increases the variability of the training dataset by introducing random perturbations during training, for example, by applying random spatial transformations or adding random image noise. In various embodiments, training and validation data samples (steps 218, 220) are introduced into model selection 222. In various embodiments, the model selection process comprises a model adaptation process 224 that includes the configuration of a network 226, the selection of a loss function 228, and an optimization 230 process. In various embodiments, the output of the model selection process 222 includes one or more hyperparameters 232 and trained parameters 234. In various embodiments, the data of the parameters flows into a model inference 236 process, whereby the data from sampling 214 is systematically included to process the entire dataset 202. The model is evaluated 238 and generates one or more results 240.In various embodiments, one or more validation models can be stored 242 in the model library 244 and the stored models can be used to initialize the model at step 245 during model selection 222 or to perform a trained model comparison 246 during model inference 236. In DL, it is a common approach to partially or fully adapt a previously trained or untrained network architecture to similar or different tasks. In various embodiments, the model library 244 enables storage of models and parameters that depend on the addressed application domain, such as denoising, object detection, counting, and prediction. One or more network architectures can be constructed using a library (e.g., TensorFlow) that defines a computational pipeline and provides tools for efficient execution on hardware resources. One or more software application drivers can be used to define a common structure for one or more components of the pipeline.

[0082] Ultrasonic images are affected by speckle, a strong multiplicative noise that generally degrades the performance of automated operations such as classification and segmentation aimed at extracting information valuable to the end user. The objective of the present disclosure is a DL approach, preferably implemented via one or more CNNs. In various implementations, given an appropriate set of images, the CNN is trained to learn an implicit model of the data, such as noise that enables effective speckle removal of new data of the same type. The noise in US images can be different in shape, size, and pattern and may be non-linear. It is assumed that a non-linear model can more accurately represent the speckle noise in the image. In various embodiments, the CNN architecture is assembled to learn a non-linear end-to-end mapping between noisy and clean US images with an extended residual network (US-DRN). In various implementations, one or more skip connections are added to the denoising model along with residual learning to reduce the vanishing gradient problem. In various preferred embodiments, the model directly obtains and updates network parameters from training data and corresponding labels instead of relying on prior knowledge of a pre-determined image or noise description model. Without being bound by theory, the context information of the image can facilitate the recovery of the degraded region. Generally, a deep convolutional network can mainly enhance context information by increasing the depth of the network or expanding the receptive field by expanding the filter (e.g., 102 in FIG. 1). However, as the depth of the network increases, the accuracy reaches a "saturated" state and then rapidly decreases. Increasing the filter size increases the convolution parameters and may significantly increase the computing power and training time. In various implementations, dilated convolution is used to expand the receptive field while maintaining the filter size. The general convolutional receptive field has a linear correlation with the depth of the layer. In contrast, the expanded convolutional receptive field has an exponential correlation with the depth of the layer.For example, when the kernel size = 3×3, the dilation coefficients of the 3×3 dilated convolution of the 7-layer architecture are set to 1, 2, 3, 4, 3, 2, 1 respectively. In one implementation, the lightweight model includes seven dilated convolutions.

[0083] Referring to FIG. 3, a schematic 300 of a US-DRN architecture according to various embodiments is shown. In various implementations, the architecture includes one or more images 302 that function as input to one or more convolutional layers 304, 306, 308, 310, 312, 314, 316, each operating similarly to layer 106 of FIG. 1. In various embodiments, the architecture is a residual network that includes one or more skip connections 318, 320, element-wise sums 322, 324, passing one or more feature information from previous layers to subsequent layers while maintaining image details and avoiding or reducing the vanishing gradient problem. In various embodiments, the network learns the estimated components resulting from the speckle image (or subtracted image) 326. The output noise-removed image 328 is generated by subtracting the image 326 from the input image 302 using an element-wise subtractor 330 via a skip connection 332. In various alternative embodiments, the element-wise subtractor 330 can include elements for each division. In this implementation, the input image 302 is divided by the estimated speckle image 326. The result then passes through a non-linear function layer, optionally a hyperbolic tangent layer, to generate the noise-removed image 328. By setting the dual goal of reproducing the noise, training becomes much more effective. This is important for the detection, recognition, segmentation, or classification of the reproductive anatomical structures of ART considering the inherently insufficient training data. The implementation of residual mapping leads to more effective learning and can quickly reduce the loss function after passing through a multi-layer network. Without being bound by theory, most pixel values of the residual image should be close to zero and the spatial distribution of the residual feature map should be very sparse, thereby making the gradient descent process a smoother loss hypersurface for filtering the parameters. The search for the optimal assignment becomes quick and easy, and layers can be added to the network to improve performance.

[0084] Figure 4 is a flowchart 400 of a deep learning network for a US image speckle removal process according to various embodiments. The three main procedures included in this method are data processing of the training dataset, CNN training, and image speckle removal. In various embodiments, data processing (step 402) can include processing of a clean image 404 and a speckled image 406, which can include normalization of the images and splitting into smaller patches. The split patches provide an input to a feed-forward network for CNN training. One or more noisy patches (e.g., kernel 102 of FIG. 1) from the speckled image 406 provide the input, and clean patches of the clean image 404 are the target outputs. In one implementation, backpropagation SDG can reduce errors and increase accuracy. In various embodiments, the patches undergo a feed-forward process and then receive backpropagation to update one or more network parameter steps 408. The patches are randomly selected for input and are used only once in the training process because SGD is employed for optimization. The network selects one input patch that includes the corresponding clean patch, calculates the error between the clean patch and the network output, and then calculates the error between different hidden layers (e.g., 204, 206, 208 of FIG. 2). In each iteration, a newly learned image 410 is created and then improved by further minimization of the loss function by calculation 412. The weights are updated in network parameter step 408 by adding the current value together with the calculated result of the partial derivative function of the error. After the update, the error and loss between the reduction of patch x and patch y should be used to remove the speckles of a new image from the learned network 414 using the parameters obtained by minimum loss convergence. The result of the trained network is the collection of weights and thresholds. The learned network 414 enables the speckled noisy image 416 to be speckle-removed to a denoised image (or clean image) 418.

[0085] In various non-limiting embodiments, the training of the US-DRN further comprises the use of 100-500 images that are resized (e.g., 256×256) and optionally obtained from a US scanner or device. In various embodiments, one or more 2D channels can be assigned to corresponding axial, coronal, or sagittal slices within a volume of interest (VOI). In various embodiments, the 3D US dataset is resampled to extract one or more VOIs at different physical scales with a fixed number of voxels. Each VOI can be transformed along a random vector in 3D space with N repetitions. Each VOI can also be transformed about a randomly oriented vector for N repetitions at one or more random angles for the expansion of the training dataset. In various embodiments, the size of the kernel or patch can be selectively set (e.g., 40×40), similar to the stride (e.g., 1-10). In various embodiments, network training includes the use of a mini-batch (e.g., 16) with an optimization method as gradient descent (e.g., ADAM optimization), a learning rate (e.g., 0.0002) over several epochs (e.g., 20, etc.). In various embodiments, the training regularization parameter is set equal to a selected value (e.g., 0.002). In various embodiments, the noise removal model training platform comprises the use of options in Matlab R2014b (Mathworks company, Natick, MA, USA), the CNN toolbox is MatConvnet (MatConvnet-1.0-beta24, Mathworks, Natick, MA), and the GPU platform is Nvidia Titan X Quadro K6000 (NVIDIA Corporation, Santa Clara, CA).In various embodiments, alternative CNN toolboxes include one or more open frameworks that include their own frameworks or, but not limited to, alternative deep learning models including Caffe, Torch, GoogleNet, and, but not limited to, VGG, LeNet, AlexNet, ResNet, U Net, etc., or combinations thereof. In various embodiments, the performance evaluation of the filter system and method includes, but is not limited to, the use of measurements of primary and secondary descriptors such as standard deviation (STD), peak signal-to-noise ratio (PNSR), equivalent appearance (ENL), and edge preservation index (EPI), structural similarity index measure (SSIM), and noise-removed image ratio (UM) without the aid of quality measurements. The higher the PSNR value, the higher the noise removal ability of the algorithm. The larger the ENL value, the better the visual effect. The EPI value reflects the boundary retention ability, and the larger the value, the better. SSIM indicates the similarity of the image structure after noise removal and should be made as large as possible. UM means that the smaller the value, the less dependent on the source image for evaluating the noise-removed image, and the stronger the ability to suppress speckle. In various embodiments, this method includes the use of 3D convolution to extract more information as compared to using multiple input channels to perform 2D convolution.

[0086] The object of the present disclosure is to provide an ANN system and method for object detection, recognition, annotation, segmentation, or classification of at least one US image in the supply of ART, preferably using denoised images processed by the US-DRN architecture, for the diagnosis, treatment, and clinical management of clinical infertility. FIG. 5 is a schematic 500 of a computer-aided diagnosis (CAD) architecture and processing method according to various embodiments. In various embodiments, the CAD architecture comprises one or more modular components including an image preprocessing module 502, an image segmentation module 506, and a feature extraction and selection module 508, and a classification module 510. In various embodiments, the image preprocessing module 502 performs one or more steps including, but not limited to, enhancement, smoothing, or reduction of speckle, resulting in, for example, the denoised image 418 of FIG. 4. In various embodiments, the image segmentation module 506 performs one or more steps including, but not limited to, dividing the image into one or more non-overlapping regions, and one or more identified regions of interest (ROIs) or VOIs are separated from the background. In various embodiments, one or more ROIs or VOIs are used for feature extraction. In various embodiments, the feature extraction and selection module 508 performs one or more steps including, but not limited to, feature extraction or removal, selection of a subset of features, to construct an optimal set of features for accurately distinguishing one or more relevant features of reproductive anatomical structures. In various embodiments, the classification module 510 performs one or more steps including, but not limited to, applying classification techniques to classify one or more reproductive anatomical structures. In various embodiments, the selected CNN architecture and method enable the detection, recognition, annotation, segmentation, or classification of ovaries, cysts, cystic ovaries, polycystic ovaries, follicles, antral follicles, etc.

[0087] The object of the present disclosure is a framework for object detection, recognition, annotation, segmentation, or classification of at least one US image in the provision of ART for the diagnosis, treatment, and clinical management of clinical infertility, using one or more software application drivers. In various embodiments, one or more application drivers define a common structure for specific module functions of the CAD system described in FIG. 5. The application not only instantiates data, application objects, and workload distribution, but also functions to combine results from various computing resources (e.g., multiple CPUs or GPUs). The application driver delegates application-specific functions to separate into application classes. In various embodiments, the application driver can be programmatically configured from the command line or using a human-readable configuration file, preferably including dataset definitions and settings that may deviate from the default. The application classes encapsulate standard analysis by connecting to and including, among others, a reader driver for loading data, a sampler driver for generating data samples for processing, a network driver for processing the input (e.g., 418 in FIG. 4), preferably a CNN, and an output handler (including, e.g., the loss function 228 driver and the optimization 230 driver in FIG. 2) and an aggregator driver during inference and evaluation. In various embodiments, the driver further includes one or more sub-components, for example, to perform data augmentation.

[0088] CAD systems for analysis include many tasks or applications of clinical workflows such as detection, registration, reconstruction, augmentation, model representation, segmentation, classification, etc. Different applications use different types of inputs and outputs, different networks, and different evaluation metrics. In a preferred embodiment, the framework platform is designed in a modular manner to support the addition of any new application type by encapsulating the workflow into application classes. The application class defines the data interfaces required for the network and loss function, facilitates the instantiation of data samplers and output objects, connects them as needed, and specifies the training regimen. In a non-limiting example, during training, a uniform sampler driver enables the generation of small image patches and corresponding labels, which are processed by the CNN to generate a segmentation, and the loss function driver is used to calculate the loss used for backpropagation using the Adam Optimizer function driver. During inference, the grid sampler driver generates a set of non-overlapping patches for converting the image into segments, the network generates the corresponding segmentation, and the grid sampler aggregator driver can aggregate the patches into the final segmentation.

[0089] The DL architecture has a complex composition of simple functions that can be simplified by repeatedly reusing conceptual blocks. In one implementation, the framework platform has conceptual blocks represented by encapsulated layer classes or represented inline using, for example, the scope system of TensorFlow. In various embodiments, one or more composite layers are constructed as simple constituent layers and TensorFlow operations. In one implementation, visualization of the network graph is automatically supported as layers at various levels of detail using the TensorBoard visualizer. In various embodiments, layer objects can be repeatedly reused to define one or more scopes upon instantiation, enabling complex weight sharing without breaking encapsulation. In various embodiments, one or more reader classes enable the loading of image files from one or more medical file formats for a particular dataset and the application of preprocessing across the entire image. In various implementations, the framework platform uses nibabel to facilitate various data formats. In a preferred embodiment, the framework platform incorporates flexibility in mapping data from the input dataset into packets of data to be processed and in mapping the processed data into useful outputs. The former is encapsulated in one or more sampler classes and the latter is encapsulated in output handlers. Instantiation of matching samplers and output handlers is delegated to the application class. Samplers generate a sequence of packets of corresponding data for processing. Each packet contains all the data for one independent computation (e.g., one step of gradient descent during training), including an image, label, classification, noise sample, or other data required for processing. During training, samples are randomly drawn from the training data, but during inference and evaluation, samples are drawn systematically and the entire dataset is processed. During training, the output handler obtains the network output, calculates the loss and the gradient of the loss with respect to the trainable variables, and repeatedly trains the model using an optimizer driver.During inference, the output handler generates useful output by aggregating one or more network outputs and performing any necessary post-processing (e.g., resizing the output to the original image size). In various embodiments, data augmentation and normalization within the platform are implemented as layer classes applied to the sampler. In a preferred embodiment, the framework platform enables support for normalization of mean, variance, and histogram intensity data, as well as inversion, rotation, and scaling for spatial data augmentation.

[0090] Figure 6 is a schematic flowchart 600 of a process for detecting "mobilizable" follicles. The detection process is configured by a cascade CNN-based method. First, one or more ROIs or VOIs of at least one US image that have been noise-reduced and pre-processed, for example, using the image processing module 502 of FIG. 5, are roughly delineated or contoured by a skilled sonographer or physician as a training data set (step 602). Second, a CNN (e.g., n convolutional and n pooling layers) is trained at 604 to segment "mobilizable" follicles at 606 and generate a corresponding segmentation probability map at 608. Third, all segmentation probability maps are divided into different connected regions using one or more operators including, but not limited to, a binarization operator, an erosion operator, or a dilation operator. Finally, a CNN 610 is used to detect "mobilizable" follicles at 612 and generate an output at 614 based on the US image patches relabeled by one or more of the segmented segmentation probability maps at 608.

[0091] Referring to FIG. 7, a schematic diagram of a CNN architecture for detecting "mobilizable" follicles according to one embodiment is shown. In this non-limiting implementation, an image (e.g., 418 of FIG. 4) is introduced as an input to the first convolutional layer (Conv) 704 of the CNN604 architecture of FIG. 6, and one or more filters (e.g., 13×13), a 2-pixel stride size, and a 6×6 pixel padding size are used to generate one or more feature maps 706, and then generate objects for max pooling. The next two convolutional layers 708 both generate 265 feature maps 710 of size 45×45 via filters of size 5×5. A 2-pixel padding size and a 2-pixel stride size are used in the second convolutional layer. The size of the feature maps is further reduced by max pooling. A 2-pixel padding size and a 1-pixel stride size are used in the subsequent layer 712. The remaining convolutional layers are composed of a 1-pixel padding size and a 1-pixel stride size. Further, each of the remaining convolutional layers 712, 714, except for the last two convolutional layers, includes a filter of size 3×3 and generates 384 feature maps of size 22×22. The second last convolutional layer 714 with a filter of size 3×3 generates 256 feature maps of size 22×22. The last convolutional layer 716 with a filter of size 3×3 generates one feature map of size 44×44. Two max pooling layers 706, 710 with a 3×3 window size follow the first convolutional layer 704 and the third convolutional layer 708, respectively. In one embodiment, a 2-pixel stride size is used in these two pooling layers. A 1-pixel padding size is optionally used only in the first pooling layer 706. Additionally, the function parametric rectified linear unit (PReLu) is used as the activation function, and its parameters can be adaptively learned using one or more of the learning methods of the present disclosure.Further, for follicle detection, the CNN610 architecture includes four convolutional layers 720, 724, 728, 732, four pooling layers 722, 726, 730, 734, and two fully connected layers 736, 738 with outputs of 64 and 1, respectively. The first convolutional layer 720 generated from the map 718 (e.g., 608 of FIG. 6) is a feature map of size 64×64 via a filter of size 5×5 having a stride size of 1 pixel and a padding size of 2 pixels. The second convolutional layer 724 generates 64 feature maps of size 32×32 via a filter of size 5×5 having a padding size of 2 pixels and a stride size of 1 pixel. The third convolutional layer 728 generates 64 feature maps of size 16×16 via a filter of size 3×3 having a padding size of 1 pixel and a stride size of 1 pixel. The last convolutional layer 732 generates 384 feature maps of size 8×8 via a filter of size 3×3 having a padding size of 1 pixel and a stride size of 1 pixel. In a preferred embodiment, the feature maps of the current layer are connected to all the feature maps of the previous layer. After both of the first two convolutional layers 720, 724, max pooling layers 722, 726 with a padding of 1, a stride of 2, and a window size of 3 follow. After the third convolutional layer 728, a max pooling layer 730 with a stride of 2 and a window size of 2 follows. After the fourth convolutional layer 732, a max pooling layer 734 with a stride of 8 and a window size of 8 follows. The activation function is the rectified linear unit (ReLU), which is a pointwise non-linearity applied to all hidden units. In various embodiments, a local response normalization scheme is also applied after each ReLU operation. After the output of the second fully connected layer 738, a softmax layer 740 is used to generate a distribution over class labels (e.g., non-mobilizable or mobilizable follicles) by minimizing the cross-entropy loss between the predicted label (i.e., mobilizable follicles) and the ground truth label (i.e., pre-segmented mobilizable follicles).

[0092] In one implementation, the training process for detecting recruitable follicles comprises three steps. First, the CNN 604 in FIG. 6 is trained with randomly initialized parameters using data including patches extracted from recruitable follicle images. In one implementation, image patches of a size randomly sampled from these recruitable follicle images (e.g., 353×353 at 702 in FIG. 7) are the input to the CNN 604 in FIG. 6. They are labeled by a probability map with pixel values within an interval [e.g., 0:1;0:9], which is determined according to the relationship between the image patch and its corresponding binary mask. The segmentation probability map 608 in FIG. 6 regarding recruitable follicles is the output of the CNN 604 in FIG. 6. In various embodiments, the improved performance comprises the use of a multi-view strategy for training the CNN 604 in FIG. 6. Second, one or more segmentation methods consisting of a successive binarization operator, an erosion operator, and a dilation operator are to divide the connected regions of the segmentation probability map generated by the CNN 604 in FIG. 6 into several isolated connected regions. For example, it can be used to binarize the segmentation probability map 608 in FIG. 6 at an interval of step 0.01 [0:15;0:7]. In various embodiments, one or more workflow applications of the framework platform described in FIG. 5 are implemented to execute one or more module functions of the process for the identification, detection segmentation, or classification of one or more recruitable follicles. In various embodiments, the model training platform comprises the option of using Matlab R2014b (Mathworks company, Natick, MA, USA), the CNN toolbox is MatConvnet (MatConvnet-1.0-beta24, Mathworks, Natick, MA), and the GPU platform is Nvidia Titan X Quadro K6000 (NVIDIA Corporation, Santa Clara, CA).In various embodiments, alternative CNN toolboxes include one or more open frameworks including, but not limited to, their own frameworks, Caffe, Torch, GoogleNet, and further, but not limited to, alternative deep learning models including VGG, LeNet, AlexNet, ResNet, U Net, etc., or combinations thereof. In various embodiments, the detection process enables the detection, identification, segmentation, or classification of alternative reproductive anatomical structures, enabling "mobilizable" follicles including, but not limited to, ovaries, cysts, cystic ovaries, polycystic ovaries, follicles, antral follicles, etc. In various embodiments, one or more results are electronically recorded in at least one electronic health record database. In alternative embodiments, the results are transmitted and stored in a database residing on a cloud-based server.

[0093] The object of the present disclosure is the ANN system and method for an object detection framework in the provision of ART for the diagnosis, treatment, and clinical management of clinical infertility. In various embodiments, the system and method enable the detection, localization, counting, and tracking (e.g., follicle growth rate) of one or more reproductive anatomical structures over time from one or more US images. In various embodiments, one or more detectors or classifiers are trained in pixel space where the positions of one or more target reproductive anatomical structures are labeled (e.g., follicles). In the case of follicle detection, the output space comprises one or more sparsely labeled pixels indicating the centers of the follicles. In various embodiments, the output space is encoded into a compressed vector of fixed dimension, preferably shorter than the original sparse pixel space (i.e., compressive sensing). In various embodiments, the CNN regresses the vector from the input pixels (e.g., US image). In various embodiments, the position of the follicle on the output pixels is recovered using normalization including, but not limited to, L 1 normalization.

[0094] Without being bound by theory, the Nyquist-Shannon sampling theorem states that a specific minimum sampling rate is required for the reconstruction of band-limited signals. Compressive sensing (CS) has the potential to reduce the sampling and computational requirements for sparse signals under linear transformation. The premise of CS is that the unknown signal of interest is observed (detected) through a limited number of linear observations. It has been proven that under the assumption that the signal is sparse and matrix-inconsistent, it is possible to obtain a stable reconstruction of the unknown signal from these observations. Signal recovery techniques generally rely on convex optimization techniques with penalties, represented by L1 regularization, such as orthogonal matching pursuit and extended Lagrangian methods.

[0095] The object of the present disclosure is an ANN system and method for an object detection and characterization framework in the provision of ART and OI for the treatment and clinical management of clinical infertility. In various embodiments, the ANN enables temporal object detection (e.g., follicle detection), localization, morphological characterization, sizing, counting, and tracking (e.g., follicle growth rate) of one or more reproductive anatomical structures from one or more US images, and includes an architecture configured, for example, for segmentation. In one embodiment, the architecture includes a set of machine learning methods such as, for example, segmentation, edge detection and enhancement, quantification, sizing, and counting of one or more reproductive anatomical structures including, but not limited to, follicles. In various embodiments, the architecture includes one or more ANN architectures, including, but not limited to, fast / faster CNN, fully convolutional network (FCN), mask region convolutional neural network (mask R-CNN), or combinations thereof. In various embodiments, edge detection includes the use of one or more methods including, but not limited to, gradient, Laplacian, etc. In one embodiment, the edge method includes one or more filters such as, but not limited to, the Sobel filter. In various embodiments, the method includes a multi-task, end-to-end, deep learning framework combined with image processing methods for morphological characterization including, but not limited to, anatomical size, length, width, diameter, volume (e.g., follicle volume).

[0096] Referring to FIG. 7B, a schematic diagram of a multi-task end-to-end deep learning framework for follicle segmentation, characterization, and tracking 700b according to one embodiment is shown. In this non-limiting implementation, an image 702b (e.g., 418 of FIG. 4) is introduced as an input to a Mark R-CNN 704b architecture to perform instance segmentation. In instance segmentation, it is necessary to correctly detect all objects in the image while accurately segmenting each instance of the detected object. Instance segmentation combines the classification of individual objects and localizes each using a bounding box with semantic segmentation that classifies each pixel into a fixed set of categories without distinguishing between them. Since the Mask R-CNN network provides a rough segmentation of the region containing the follicles, the output of the 704b network is processed by one or more edge filters 706b. In various embodiments, edge detection is used to identify regions having significant luminance gradients in the TVUS image 702b. In one embodiment, a Sobel image gradient filter is selected to minimize computational overhead. The Sobel filter is a 2D filter for edge detection of one or more follicles and describes a first-order gradient operation that depends on rotations in the vertical and horizontal directions. The edges of the follicles in the image 702b correspond to high absolute responses with respect to the direction of the filter. In one embodiment, once a follicle is detected, a follicle counter 708b including a Visual Geometry Group (VGG) CNN architecture can be used to perform the count. In various embodiments, the CNN for follicle counting consists of an 11-layer 2D convolution with 3×3 pixel filters, a VGG-11 network. The number of filters varies from 64 to 512. After each convolutional layer, a batch normalization layer, a leaky ReLU, and a max pooling layer follow in sequence. VGG-11 is used for feature extraction. In count prediction, three fully connected layers (e.g., dimensions: 1024, 512, 1) separated by batch normalization layers and leaky ReLUs are used to prevent negative number outputs.In another embodiment, the morphology of the detected follicles is characterized using a morphological characterizer 710b in terms of size, area, volume, etc., or combinations thereof. In one embodiment, the area outside the follicle mask within the bounding box is removed, and the grayscale image within the mask is binarized using Otsu's method. Since the follicle mask is distributed slightly wider than the follicle, there may be noise remaining in the mask even after binarization. In one embodiment, this noise can be removed by calculating the area and eccentricity by morphological analysis and deleting regions below a preset size and eccentricity. One or more measured dimensions (length, width, etc.) of the follicle can be calculated using the skeletonization and the edges of the follicle object by the steps of: (a) finding two adjacent pixels within the skeleton, (b) calculating the direction of the pixels using the two adjacent pixels, (c) drawing the normal of the pixels according to the calculated direction, (d) finding the pixels that intersect at both ends, and (e) calculating the pixel distance between the two intersecting pixels. Skeletonization is used to convert the follicle in pixel units into a single-pixel-width representation. In various embodiments, follicle tracking includes the use of the outputs of the follicle counter 708b and the characterizer 710b as inputs to the follicle tracking device 712b. In one embodiment, the follicle tracking device 712b constructs a growth curve using a linear assignment problem-based approach and tracks follicles from 2D TVUS images during the day over a period of time to determine an individual region (diameter) versus time growth curve. For example, given a set of follicles detected throughout a time-lapse image sequence, the algorithm first links the initially detected follicles between consecutive frames, and then links the track segments generated in the first step to close the gaps and capture the follicle movement events. In a particular embodiment, a histogram (number of follicles versus diameter / day) is constructed by calculating the number of follicles using the output of the counter 708b and calculating the change in area / diameter of the follicles using the output of the characterizer 710b from daily image acquisitions for each follicle.In one embodiment, a spatial map of follicle growth rates having a single follicle resolution is constructed using the growth rates and instances of detected follicles.

[0097] Referring to FIG. 7C, a schematic diagram of the Mark R-CNN 704b architecture for follicle instance segmentation 700c according to one embodiment is shown. The architecture includes one or more Feature Pyramid Networks (FPNs) 702c based on a ResNet architecture modified according to one or more TVUS images 702b of FIG. 7B as the backbone of Mask R-CNN. Mask R-CNN complements the Faster R-CNN architecture to enable mask prediction. Mask R-CNN replaces the Region-of-Interest (RoI) pooling layer of the Faster R-CNN network with ROI Align 704c and introduces an interpolation process to solve the alignment problems caused by direct sampling by pooling. In various embodiments, one or more Fully Connected Network layers (FCNs) 702c, 706c, 708c, 710c, 712c are configured sequentially and / or in parallel and are used to predict pixel-level instance masks of one or more follicles. In various embodiments, the Mask R-CNN 704b network of FIG. 7B is constructed in three stages for rough follicle detection and localization, namely, feature extraction, region proposal, and prediction. Mask R-CNN uses the FPN to generate candidate regions RoI, performs feature extraction by ResNet 714c, for example, non-limiting ResNet-101, and obtains a pyramid feature map 716c of the follicle through pixel-level information. This addition enables the network to perform accurate localization using the high-resolution feature maps of the lower layers. The extraction process is the same as the process of Faster R-CNN that uses the Region Proposal Network (RPN) 718c to generate bounding box proposals that perform object / non-object binary classification and bounding box regression. Each RoI region of the image 702b in FIG. 7B and the feature map 716c of each RoI are corrected using ROI Align 704c. Next, after obtaining the feature map of each RoI region, the Fully Connected FC layer 720c is used to predict each classification and bounding box.Each RoI predicts the category of each pixel in the RoI region using the designed fully connected network FCN722c framework. As a final result, Mark R-CNN704b generates one or more segmentation masks for one or more follicles.

[0098] Referring now to FIG. 7D, a process flow diagram of follicle tracking 700d is shown. According to one embodiment, follicle tracking performed by the follicle tracking device 712b of FIG. 7B links follicles detected between two or more consecutive frames of the TVUS image 702b of FIG. 7B, then closes gaps, and links track segments generated in a first step to capture the movement of the follicles. In a first process flow step, one or more follicle track segments can be constructed by linking follicles detected between consecutive frames, with the constraint that a follicle in one frame can be linked to at most one follicle in the previous or next frame. Tracks are constructed from one or more image sequences (702d) by detecting follicles (704d) in each frame and determining the follicle position (706d) for each frame. In subsequent process flow steps, follicles are linked between consecutive frames (708d) and a tracking segment (710d) is created. In subsequent process flow steps, gaps between images can be closed to capture and complete merger and split events (712d) between the first track segments (714d). In various embodiments, one or more linear assignment problems (LAPs) are solved to formulate both the follicle linking step between frames and the steps of closing gaps, merging, and splitting. In the LAP framework, all potential assignments (follicle assignment in the first step, track segment assignment in the second step) are characterized by a cost function and solved by a global or local cost minimization matrix. In certain embodiments, one or more follicles or tracks are assigned to one or more assignment candidates in the follicle linking step between frames (708d). A follicle in source frame t can be linked to a follicle in target frame t + 1 (cost function A). In an alternative embodiment, a follicle in the source frame is not linked to anything and may lead to the end of a track segment (cost function B), or a particle in the target frame is not linked to anything and may lead to the start of a track segment (cost function C).In various embodiments, in the steps of gap closing, merging and splitting (712d), the six types of assignment candidates for follicles may be in cost competition. The end of a track segment can be linked to the start of another track segment, thereby filling the gap (cost function D), the end of a track segment can be linked to a midpoint of another track segment, leading to a merge (cost function E), or the start of a track segment can be linked by a midpoint of another track segment, leading to a split (cost function F). In an alternative embodiment, the end of a track segment can be linked to null, leading to the end of the track (cost function G), the start of a track segment can be linked to null, leading to the start of the track (cost function H), or the midpoints of track segments introduced for merging and splitting can be linked to null, rejecting the merge or split (cost functions D′ and B′). In this step, all track segments of the entire sequence compete with each other. In various embodiments, the cost functions are adjusted for a particular tracking application, for example, under one or more assumptions of follicle motion (such as isotropic random motion, Brownian motion, etc.).

[0099] Referring now to FIG. 8, a flowchart of a follicle detection and localization framework 800 according to various embodiments is shown. The framework consists of a follicle localization encoding phase that uses one or more random projections, a CNN-based regression model that captures one or more relationships between the US image and the encoded signal y, and a decoding phase for recovery and detection. In various embodiments, during training, the ground truth location of the follicles is indicated by a binary annotation map 804 in pixel units. In various embodiments, one or more encoding schemes 806 convert the location of the follicles from the pixel space representation of the image 802 to a compressed signal y 808. Next, a training pair 810 consisting of a US image 812 (e.g., 418 of FIG. 4) each containing one or more follicles and the compressed signal y 808 is used to train a CNN 814 to function as a multi-label regression model. In one implementation, the Euclidean loss is used during training if suitable for the regression being performed. In various embodiments, data augmentation includes one or more image rotations of the training set for robustness to rotation. In various embodiments, during testing, the trained CNN 814 generates an output of an estimated signal y' 816 for each test image 818 provided as input to the first convolutional layer. Subsequently, using a decoding scheme 820 and one or more sensing matrices 824 determined by one or more encoding or decoding schemes, the estimated signal y' 816 is subjected to L 1 The ground truth follicle location prediction 822 is estimated by performing a minimization recovery.

[0100] Several encoding schemes may be adopted by the framework. In various embodiments, the framework uses one or more random projection-based encoding schemes. In various embodiments, the centers of all follicles are marked with dot marks, cross marks, or bounding boxes. In one embodiment, the binary annotation map 804 in pixel units includes the size of w×h indicating the positions of one or more follicles by labeling 1 at the pixels of the follicle centroid, and labeling 0 at the background pixels otherwise. In one embodiment, the annotation map 804 is vectorized by concatenating all rows of the map 804 into a vector f of length wh. Thus, the positive elements of the map 804 with {x, y} coordinates are encoded at the [x+h(y-1)]-th position of the vector f. After the generation of the vector f, random projection is applied. The vector f can be represented by the matrix 824 and one or more linear observations y proportional to the sensing of the vector f. Without being restricted by theory, sensing the matrix 824 preferably satisfies one or more conditions including but not limited to isometric properties. In one implementation, the matrix 824 is a random Gaussian matrix. In an alternative embodiment, another encoding scheme 806 is used to reduce the computational load, especially for processing particularly large images. In various embodiments, the coordinates of all follicle centroids are projected onto a plurality of observation axes. The set of observation axes is created with a total of N observations. In one implementation, the observation axes are uniformly distributed around the image 802. For the case of the n observation axis oa n , the positions of the follicles are encoded into a sparse signal of length R. The orthogonal signed distance (f n ) is calculated from the follicles to the n observation axis oa n . Thus, f n includes the measurement of the distance to the side of the follicle and the signed distance as the location of the oa n . The encoding of the follicle location below the oa n is y n obtained by random projection. Similarly, y n is the signed distance f nis a proportional matrix 824 times that of. In various embodiments, the process is repeated for all N observation axes to each y n is obtained. The articular representation of the follicle location is derived from the encoding result y after concatenation of all y n . Similarly, a decoding scheme can be employed by the framework to recover the vector f. In various embodiments, the accuracy recovery from the encoded signal y is L 1 obtained by solving a regularized convex optimization problem. The recovery of f enables the localization of all true follicles localized N times, along with the N predicted positions 822.

[0101] The follicle detection and localization framework comprises one or more CNNs 814 for constructing at least one regression model between the US image 812 and its follicle location representation or the compressed signal y 808. In one implementation, the CNN 814 includes a network consisting of, but not limited to, five convolutional layers and three fully connected layers. In another implementation, the CNN 814 includes, for example, a deep neural network having a 100-layer model. In other implementations, the CNN 814 includes one or more CNNs disclosed within the present disclosure. In various embodiments, one or more loss functions, including, but not limited to, the Euclidean loss or other loss functions of the present disclosure, can be used. In various embodiments, the dimensions of the output layer of the CNN can be changed to the length of the compressed signal y 808. In various embodiments, one or more CNN 814 models can be further optimized using additional learning methods including, but not limited to, multi-task learning (MTL) for localization and follicle counting. In various embodiments, during training, one or more labels are provided to the CNN. In one implementation, an encoded vector y carrying pixel-level location information of the follicles. In another implementation, a scalar or follicle count (c) representing the total number of follicles in the training image patch, filter, or kernel. In various embodiments, two or more of the labels can be concatenated into the final training label. Next, one or more loss functions are applied to the concatenated labels. Thus, the CNN model parameters can be optimized using the monitoring information for both follicle detection and counting. A large number of square patches can be used for training. Along with each training patch, a signal (i.e., the encoding result: y) can be used to indicate the location of the target follicles present in each patch. Data augmentation can be used by performing patch rotation on a collection of training patches that are invariant to system rotation. In various embodiments, one or more MTL frameworks can be used to handle cases when follicles are in contact and clustered.In one implementation, the appearance of one or more follicles, including but not limited to texture, morphology, boundaries, and contour information, is integrated into the MTL framework to form a deep contour recognition network. Preferably, complementary appearance and contour information enhances the discriminative ability of intermediate features, thus enabling more accurate separation of individual follicles from touching follicles or clustered follicles. In various embodiments, the CNN is trained in an end-to-end manner to improve performance. In various embodiments, the model training platform is provided for use with Matlab R2014b (Mathworks company, Natick, MA, USA), the CNN toolbox is MatConvnet (MatConvnet-1.0-beta24, Mathworks, Natick, MA), and the GPU platform is Nvidia Titan X Quadro K6000 (NVIDIA Corporation, Santa Clara, CA). In various embodiments, alternative CNN toolboxes include one or more open frameworks, including but not limited to their own frameworks, or alternative deep learning models including, but not limited to, Caffe, Torch, GoogleNet, and further, but not limited to, VGG, LeNet, AlexNet, ResNet, U Net, etc., or combinations thereof. In various embodiments, the process enables the detection, localization, and counting of alternative reproductive anatomical structures, including but not limited to follicles including ovaries, cysts, cystic ovaries, polycystic ovaries, etc. In various embodiments, one or more results are electronically recorded in at least one electronic health record database. In alternative embodiments, the results are transmitted and stored in a database residing on a cloud-based server.

[0102] Referring to FIG. 9, a flowchart 900 of a follicle tracking framework according to various embodiments is shown. The follicle tracking framework comprises one or more CNNs 902, 904 that specifically task to distinguish the size, volume, or quality of follicles and generate a tracking functional map in real time using a network. In various embodiments, the tracking framework specifically learns a correlation filter for each tracked follicle and trains on the features extracted by the CNNs 902, 904. In various implementations, deep semantics is combined with the spatial resolution of the initial filter to combine accurate tracking. In a preferred embodiment, at least one of the networks is trained to perform multiple tasks by utilizing the pre-computation of the feature maps from one or more CNN network segmentation processes, e.g., 606, 608 of FIG. 6. The follicle tracking system comprises a hierarchical tracking device for performing correlation filter tracking based on the extracted features. In various embodiments, the tracking device calculates one or more circulant kernels in the Fourier space to improve performance. For each tracked follicle, a search window is positioned over the follicle in the first frame of the input. This frame is a set of neural network functions obtained from one or more segmentation map outputs (e.g., 606, 608 of FIG. 6). Once the search window is positioned, one or more correlation filters are learned by minimizing a loss function, optionally performed in the Fourier space domain. At each tracking time step, the correlation filter is the matching feature within the search window Z overlaid on the last known location of the target follicle, obtained from one or more of the feature maps 804 or follicle position prediction 822 of FIG. 8. In various embodiments, at least one filter is trained on one or more layers of a selected CNN. In various embodiments, at each time step, one or more correlation filters are the matching features within the search window Z overlaid on the last known location of the target follicle. One or more matches are preferably calculated in the Fourier space domain.In various embodiments, one or more deep filters are propagated to higher layer levels in a weighted fashion. Argmax f of {m, n}. oBy obtaining, one or more estimated new locations are detected and the new search is reoriented to the new locations. Referring again to FIG. 9, in various embodiments, FollicleTrack 906 receives, as input, one or more locations of the tracked target follicles, preferably the first frame of a time series sequence, one or more raw images, one or more processed images, or feature map outputs (e.g., 606, 608 of FIG. 6), and one or more segmented images (e.g., 814, 822 of FIG. 8) showing one or more convolutional layers of CNN 902 and CNN 904. One or more follicles can be selected for tracking or the segmented follicles can be tracked throughout the time series. In various embodiments, the time series (e.g., seconds, minutes, hours, days, etc.) sequence includes one or more sequences including, but not limited to, real-time transvaginal US images, sequences of stored transvaginal US images, images retrieved from a US scanner / device, and images transmitted from a US scanner / device. In various embodiments, CNN 902 is a deep convolutional neural network designed to further segment follicles with respect to physical dimensions and changes in dimensions. The network is trained using one or more training sets to enable segmentation of follicles by size, shape, volume, distribution, mean, standard deviation, morphology, position, displacement, and growth rate from US images (e.g., 2D, 3D, etc.). In various implementations, one or more outputs of FollicleTrack 906 are a list including, but not limited to, one or more learned filters for each follicle, the history of follicle centroid positions at each time step, the history of follicle locations, follicle distribution (size, volume, etc.), movement, displacement, growth rate (size, volume, etc.). In various embodiments, the output data enables annotation of one or more segmented or unlabeled images or enables direct mapping of the follicle trajectories (e.g., movement, growth rate). In various embodiments, the US images are obtained from a historical database containing selected information regarding recruitable follicles, non-recruitable follicles, distinguished by size or volume, or a combination thereof.In various embodiments, the US images are annotated by a skilled sonographer or physician having expertise in recognizing reproductive anatomical structures, preferably differentiating the characteristics of optimal follicles, size, follicle distribution, volume, and growth rate for extraction and implantation. In various embodiments, CNN902 includes five convolutional layers 908, 910, 912, 914, 916, followed by a fully connected layer 918, and is fed to a final FC 920. One or more outputs 922 include, but are not limited to, follicle size / volume range, distribution, mean, standard deviation, or growth rate (e.g., 1 mm per day) of one or more follicles. In various embodiments, CNN902 generates a set of one or more masks and feature maps for at least one frame of US image input obtained from at least one patient, for at least one visit for follicle tracking. In various embodiments, CNN904 is a deep convolutional neural network designed to further segment follicles with respect to quality or change in quality. In one implementation, CNN904 includes six 3×3 convolutional layers 924, 926, 928, 930, 932, 934, has ReLu activation, max pool layer 936 and softmax layer 938, and an output 940. In various implementations, dropout is implemented at one or more levels of the network. In one implementation, the network is programmed to generate output feature maps from each convolutional layer and label classification scores. In various embodiments, the convolutional layer generates one or more feature maps for follicle discrimination. In various embodiments, the hierarchical method is combined with one or more other networks of the present disclosure to improve tracking accuracy. In one implementation, the weighted correlation filter in each search window provides a cost for linear assignment and enhances its ability to track follicles.In various embodiments, the model training platform comprised the use of Matlab R2014b (Mathworks company, Natick, MA, USA) option, CNN toolbox, MatConvnet (MatConvnet-1.0-beta24, Mathworks, Natick, MA), and GPU platform Nvidia Titan X Quadro K6000 (NVIDIA Corporation, Santa Clara, CA). In various embodiments, alternative CNN toolboxes include one or more open frameworks including, but not limited to, their own frameworks, Caffe, Torch, GoogleNet, and further, but not limited to, alternative deep learning models including VGG, LeNet, AlexNet, ResNet, U Net, etc., or combinations thereof. In various embodiments, the process enables the detection, localization, counting, and tracking of alternative reproductive anatomical structures, enabling follicles including, but not limited to, oocytes, blastomeres, ovaries, cysts, cystic ovaries, polycystic ovaries, endometrial thickness, etc. In various embodiments, one or more results are electronically recorded in at least one electronic health record database. In alternative embodiments, the results are transmitted and stored in a database residing on a cloud-based server.

[0103] The object of the present disclosure is the ANN system and method for analyzing electronic medical records in the provision of ART for the diagnosis, treatment, and clinical management of clinical infertility. In various embodiments, the system and method enable the extraction of features or phenotyping of one or more patients from at least one multi-year patient's electronic medical record (EMR), electronic health record (EHR), database, etc. The electronic phenotype refers to the problem of extracting an effective phenotype from the patient's multi-year health record. The challenges in effectively extracting features from a patient's EMR or EHR are high dimensionality due to a large number of individual medical events, the transience of EHR evolving over time, data sparsity, irregularity, systematic errors, or biases. The time series representation of data is used to address the representation of a patient's medical record as a time series where one dimension corresponds to time and the other dimensions correspond to medical events. In various embodiments, the temporal EHR information or medical record is transformed into one or more binary sparse matrices with a horizontal dimension (time) and a vertical dimension (medical events). In one implementation, the (i, j) entry of a matrix for a particular patient is equal to 1 if the i-th event is recorded or observed at the time stamp j of the patient's medical record. Referring to FIG. 10, FIG. 1000 of an EMR data convolutional network architecture according to various embodiments is shown. This architecture comprises a first layer (or matrix) 1002 including one or more patient EMR matrices and a CNN 1004. In various embodiments, the CNN 1004 further includes one or more convolutional layers 1004a, preferably one-sided convolutional layers, preferably a pooling layer 1004b for introducing sparsity into the learned features, and an FC layer 1004c. In various embodiments, the convolutional layer 1004a includes a convolutional operator on the time dimension of the patient EMR matrix 1002. In one embodiment, each event matrix 1002 of length l is represented by a vector X, and x i is represented as a d-dimensional event vector corresponding to the i-th event item. In various embodiments, x i:i+j is the item x i , x i+1 , … x i+jrepresents the concatenation. The one-sided convolutional filter operation includes a filter applied to a window of n event features to generate new features. For example, feature c i is generated from a window of events (e.g., x i:i+n-1 ) using one or more non-linear activation functions, preferably ReLU. The filter is applied to each possible window of features within one or more event matrices 1002 to generate a feature map 1004d. The average pooling layer 1004b is applied over one or more feature maps 1004d to obtain the average value of c. In a preferred embodiment, one or more important features with the highest value in each feature map are captured for feature extraction. The FC layer 1004c is a fully connected layer linked to one or more softmax classifiers for classification or prediction using one or more single frames.

[0104] EMR data varies greatly over time, and temporal connections are required for prediction. In various embodiments, temporal smoothness is incorporated into the learning process using one or more temporal fusions. In various embodiments, one or more data samples are processed as a collection of short fixed-size sub-frames of a single frame that include several consecutive intervals within the time. In one implementation, the model fuses information across the time domain by modifying the convolutional layer 1004a to extend it temporally and is executed at the beginning of the network 1004. In one implementation, proximal fusion immediately combines information across the entire time window at the basic event feature level. One or more filters of the convolutional layer 1004a are modified to extend the operation on one or more sub-frames. In another implementation, the distal fusion model performs fusion on the fully connected layer 1004c. In one embodiment, one or more separate single-frame networks or sub-frames are integrated into the fully connected layer, thereby detecting patterns present in one or more sub-frames. In another implementation, the balance between proximal and distal temporal fusions enables a slow fusion of information across the network. In various embodiments, the upper layers of the network receive more global information more gradually over time. In one implementation, connectivity is extended to all convolutional layers over time, and the fully connected layer 1004c can calculate global pattern characteristics by comparing all output layers. The framework enables the generation of an insightful patient phenotype by leveraging the predominance of higher-order temporal event relationships. In various embodiments, one or more recordings of neuron activity enable the observation of patterns indicating a healthy or medical condition. In one implementation, one or more neuron outputs receive the highest preferably normalized weights in one or more top layers for positive or negative classification of the state. One or more regions that appear in the training set that highly activate one or more corresponding neurons can be identified using one or more sliding window cuts (minimum, maximum window size) to obtain one or more top-ranked regions or patterns.In another implementation, one or more weights of the neurons are aggregated and assigned to a medical or health condition, which is an important function for the purpose of extracting and predicting the phenotype of a patient. In various embodiments, the model training platform was equipped for use with Matlab R2014b (Mathworks company, Natick, MA, USA), with the CNN toolbox, MatConvnet (MatConvnet-1.0-beta24, Mathworks, Natick, MA), and the GPU platform Nvidia Titan X Quadro K6000 (NVIDIA Corporation, Santa Clara, CA). In various embodiments, alternative CNN toolboxes include one or more open frameworks including, but not limited to, their own frameworks or alternative deep learning models including, but not limited to, Caffe, Torch, GoogleNet, and further, but not limited to, VGG, LeNet, AlexNet, ResNet, U Net, etc., or combinations thereof.

[0105] In various embodiments, the medical record includes one or more stored patient records, preferably records of patients undergoing infertility treatment, ultrasound images, images of the reproductive anatomical structures, physician notes, clinical notes, physician annotations, diagnostic results, body fluid biomarkers, hormone markers, hormone levels, neohormones, endocannabinoids, genomic biomarkers, proteomic biomarkers, anti-Müllerian hormone, progesterone, FSH, inhibin, renin, relaxin, VEGF, creatine kinase, hCG, fetoprotein, pregnancy-specific b-l-glycoprotein, pregnancy-associated plasma protein-A, placental protein-14, follistatin, IL-8, IL-6, vitellogenin, calbindin-D9k, treatment procedures, treatment schedules, implantation schedules, implantation rates, follicle sizes, follicle numbers, AFC, follicle growth rates, pregnancy rates, the date and time of implantation (i.e., the event), CPT codes, HCPCS codes, ICD codes, and the like. In various embodiments, the phenotypes of one or more patients include infertility, anovulation, oligo-ovulation, endometriosis, male factor infertility, tubal factor infertility, reduced ovarian reserve, the patient risk of ovulation, implantation, patients ready for implantation, patients having one or more biomarkers indicating ovulation, patients having US images indicating optimal for extraction, and the like. In various embodiments, one or more identified patient phenotypes or prediction results from the output layer of the one or more CNNs are recorded in at least one electronic health record database. In an alternative embodiment, the results are transmitted and stored in a database residing on a cloud-based server.

[0106] The object of the present disclosure is the ANN system and method for predictive planning in the provision of ART for the diagnosis, treatment, and clinical management of clinical infertility. The unified prediction framework includes one or more survival convolutional neural networks ("SCNN") and provides one or more predictions of the time to an event from at least one US image and one or more patient phenotypes obtained from the patient's medical record. In various embodiments, the framework includes one or more image sampling and risk filtering techniques for prediction purposes. In one implementation, one or more ROIs or VOIs of at least one US image are used to train a deep CNN seamlessly integrated with a Cox proportional hazards model to predict patient outcomes, including, but not limited to, the end of induction of hormonal therapy, mobilizable follicles, dominant follicles, mature follicles, preparation for follicle extraction, and endometrial thickness optimal for implantation. Referring to FIG. 11, a schematic diagram 1100 of a survival convolutional neural network architecture according to various embodiments is shown. This architecture comprises an n-layer CNN architecture 1102 with a Cox proportional hazards model 1118 for predicting time-to-event data from image 1106. In one non-limiting implementation, image feature extraction is achieved by four groups of convolutional layers. The first group 1108 consists of two convolutional layers with 64 3×3 kernels interleaved with local normalization layers, followed by a single max pooling layer. The second group 1110 consists of two convolutional layers (128 3×3 kernels) interleaved with two local normalization layers and followed by a single max pooling layer. The third group 1112 interleaves four convolutional layers (256 3×3 kernels) with four local normalization layers and followed by a single max pooling layer. The fourth group 1114 includes an interleaving of eight convolutional (512 3×3 kernels) and eight local normalization layers, with an intermediate pooling layer and a final max pooling layer. Following these four groups is a sequence of three fully connected layers 1116, including 1,000, 1,000, and 256 nodes, respectively.The layer that is fully connected at the end outputs a prediction of the risk associated with the input image 1106 (e.g., the likelihood of being extracted for implantation). The predicted risk is input into the Cox proportional hazards layer 1118, the negative partial log-likelihood 1120 is calculated, and an error signal for backpropagation within the CNN 1102 is provided. In various embodiments, one or more optimization methods are used to optimize the weights, biases, and convolutional kernels of the model, e.g., the Adagrad algorithm. In one implementation, the non-limiting parameters of Adagra include an initial accumulator value (e.g., 0.1), an initial learning rate (e.g., 0.001), and an exponential decay coefficient (e.g., 0.1). In one implementation, the weights of the model are initialized using, e.g., the variance scaling method and weight decay (e.g., 4e-4) applied to the fully connected layers during training. In various embodiments, mini-batches (e.g., including mobilizable follicles) are preferably used for training over a plurality of epochs (e.g., 100; one epoch is one complete cycle through all training samples). In various embodiments, each mini-batch generates an update to the model, resulting in multiple updates per epoch. In one implementation, the Cox likelihood is calculated locally within each mini-batch to perform the update. In another implementation, to improve robustness, randomization of the assignment of one or more mini-batches is used at the start of each epoch. In yet another implementation, regularization is applied during training, and optionally, to avoid overfitting, 5% of the weights of the last fully connected layer of the fully connected layer 1116 of each mini-batch are randomly dropped out during training. During training, one or more identified "mobilizable" follicle fields (e.g., a pixel area or volume sufficient to distinguish mobilizable / non-mobilizable) are sampled from the region (e.g., ROI or VOI) and treated as semi-independent training samples. In various embodiments, the identity / field of each mobilizable follicle is paired with the outcome of the time to the patient's event from the medical record database 1122.In various embodiments, patient outcome information includes, but is not limited to, the number of days of hormone therapy, demographics, age, presence or absence of one or more of the diagnostic biomarkers, therapeutic treatments, chronic biomarkers, clinical records, physician observations, follicle growth rate, follicle size, pregnancy history, start - end ovarian stimulation, cycle days, follicle retrieval, embryo mobilization, oocyte retrieval, follicular phase, follicle maturation, egg maturation, fertilization rate, blastocyst embryo development, embryo fragmentation, embryo growth rate, embryo grade, trophectoderm grade of the embryo, inner cell mass grade of the embryo, embryo size, embryo growth rate, embryo cell number, embryo metabolic parameters, embryo metabolism, embryo maturity, blastocyst development rate, euploid embryos, aneuploid embryos, embryo mosaicism, embryology database, embryo quality, implantation, etc. In various embodiments, duplicate results can be paired with one or more regions containing multiple follicles. One or more regions can be sampled at the start of each training epoch to generate a new set of ROIs or VOIs. In various implementations, randomization by transformation (e.g., translation, rotation, contrast, brightness, etc.) can be applied to the acquisition field to improve robustness to follicle orientation or image variations. In various embodiments, one or more fields are sampled from each ROI or VOI for calculating risk prediction using the SCNN1102. For example, when predicting a patient's outcome, 10 fields are sampled from each ROI, a representative collection of fields is generated, and the risk of each field is predicted. In one implementation, the median of the field risks is calculated for each region, sorted and filtered, and the second - highest value is selected as the patient's risk. Selecting the second - highest risk provides robustness to outliers and high risks caused by image quality or artifacts. In various embodiments, the filtering procedure enables the selection of fields using a conservative prognosis and ensures the accurate selection of mobilizable follicles for implantation. In various embodiments, one or more diagnostic data 1124 can be incorporated into the SCNN1102 to improve the accuracy of the prognosis. In one implementation, the SCNN 1102 simultaneously learns from diagnostic biomarkers and US images by incorporating biomarker variables that affect the patterns learned by the network via the patient's blood test information during the visit.In one embodiment, the diagnostic data 1124 is incorporated into a fully connected layer 1116. In various embodiments, TensorFlow (v0.12.0) is used on a server equipped with a dual Intel(R) Xeon(R) CPU E5-2630L v2 @ 2.40 GHz CPU, 128 GB RAM, and dual NVIDIA K80 graphics cards to optionally train one or more prediction models. In various embodiments, alternative CNN toolboxes include one or more open frameworks including, but not limited to, their own frameworks, or alternative deep learning models including, but not limited to, Caffe, Torch, GoogleNet, and further, but not limited to, VGG, LeNet, AlexNet, ResNet, U Net, etc., or combinations thereof. In various embodiments, the prediction results of the time to one or more patient events from the output layer of the one or more CNNs are recorded in at least one electronic health record database. In an alternative embodiment, the results are transmitted and stored in a database residing on a cloud-based server.

[0107] The object of the present disclosure is a computer program product for use in providing ART for the diagnosis, treatment, and clinical management of clinical infertility. Referring to FIG. 12, a schematic 1200 of a computer product architecture according to various embodiments is disclosed. The architecture includes an AI engine 1202 that receives one or more inputs from one or more data sources. In one embodiment, the AI engine 1202 preferably receives data from many sources, including but not limited to physiological data 1204, US image data 1206, and environmental data 1208. In various embodiments, the physiological data 1204 includes one or more of the diagnosis results, data, biomarkers, genomic markers, proteomic markers, body fluid analytes, chemical panels, measured hormone levels, etc. In various embodiments, the US image data 1206 includes one or more US images retrieved from a US scanner / device, US images stored outside the US scanner / device, and US images processed by one or more CNNs of the present disclosure. In various embodiments, the environmental data 1208 includes one or more of the medical record data collected over the years, cycle days, time, week, or monthly data. The AI engine 1202 can also receive inputs from one or more EMR databases 1210 and one or more result databases (or cloud-based servers) 1212. In various embodiments, the EMR database 1210 includes one or more medical records of patients over the years and the records of patients under infertility treatment or management. In a similar manner, the result database 1212 includes one or more medical records of infertile patients over the years related to the success of implantation, pregnancy rate, etc. The patient's medical record database can be located within a dedicated facility or be available from an external source (e.g., cloud source, public database, etc.) accessible via a communication network. The AI engine 1202 can also receive one or more inputs of one or more data or data sets generated by one or more CNNs disclosed within the present disclosure and stored in a cloud-based server 1214.The cloud server and services are generally referred to as "cloud computing", "on-demand computing", "software as a service (SaaS)", "platform computing", "network-accessible platform", "cloud service", "data center", etc. The term "cloud" can include a collection of hardware and software that forms a shared pool of configurable computing resources (e.g., networks, servers, storage media, applications, services, etc.), and these resources can be appropriately supplied to provide features such as on-demand self-service, network access, resource pooling, elasticity, measured service, etc. The AI engine 1202 comprises one or more computing architectures, hardware, and software known to those skilled in the art to enable it to execute one or more instructions or algorithms to process one or more external inputs. One or more processed data sets or analyses or predictions or clinical insights generated by the AI engine 1202 can be transmitted and stored in a cloud-based server 1214 for distribution. In various embodiments, the cloud-based server 1214 includes one or more software applications (or desktop clients) 1216 and enables the development of one or more software application products (or mobile clients) 1218, providing one or more functions including, but not limited to, data processing, data analysis, data display in graphical form, data annotation, etc. In various embodiments, the product comprises at least one patient data, US images from a US scanner / device, retrieved US images, patient medical records from an electronic medical record database, patient records related to childbirth, patient endocrinology records, patient clinical notes, physician clinical notes, data from a database on the cloud-based server, results from one or more output layers of one or more of the ANNs, and an artificial intelligence engine. In various embodiments, at least one application enables the search or distribution of one or more clinical insights.Clinical insights may include various reproductive endocrinology insights such as the expected number of eggs retrieved from a patient, the proposed timing of egg retrieval, the expected number of eggs per patient and per day, the expected egg maturation day, the expected number of embryos, the expected quality of embryos, the amount and scheduling of embryo biopsies, patient care planning, reduction of error rates, improvement of pregnancy rates, etc. In various embodiments, a user (e.g., a physician) can access clinical insights from the desktop client 1216. Similarly, a patient can access clinical insights from the mobile client 1218, or vice versa. In one implementation, the mobile client 1218 comprises one or mobile app products 1220 that can communicate with the cloud server 1214 via the communication network 1222 to access the information. Similarly, the desktop client 1216 can access the cloud-based server 1214 via the communication network 1222. In various embodiments, the communication network 1222 comprises, but is not limited to, one or more of a LAN, WAN, wireless network, cellular network, Internet, etc., or combinations thereof.

[0108] Referring to FIG. 13, a process flow diagram for determining the follicular maturation date of follicles according to various embodiments is shown. In step 1302, one or more digital images of reproductive anatomical structures are acquired to detect one or more reproductive anatomical structures. In step 1304, the digital image is processed to detect one or more reproductive anatomical structures. In step 1306, one or more digital images are processed to annotate, segment, or classify one or more anatomical features of one or more reproductive anatomical structures. In step 1308, one or more anatomical features are analyzed according to at least one linear or non-linear framework. In step 1310, at least one result of the time to a reproductive assistance procedure event is predicted according to at least one linear or non-linear framework. In alternative step 1312, one or more digital images are processed to measure the volume of one or more follicles of a patient, thereby providing an additional input to step 1304 as an output. In another alternative step 1314, the output of step 1306 provides an input for comparing a first digital image of a patient's reproductive anatomical structure with a second digital image of the patient's reproductive anatomical structure to provide an output. In step 1316, the output of step 1314 functions as an input for determining the follicular maturation date of one or more follicles of a patient.

[0109] Referring to FIG. 14, a process flow diagram of a process for generating clinical proposals related to assisted reproductive treatment according to various embodiments is shown. In step 1402, one or more digital images of a patient's reproductive anatomical structure are received via one or more imaging modalities. In step 1404, the one or more digital images are processed to detect one or more reproductive anatomical structures. In step 1406, the one or more digital images are processed to annotate, segment, or classify one or more anatomical features of the one or more reproductive anatomical structures. In step 1408, the one or more anatomical structures are analyzed according to at least one linear or non-linear framework. In step 1410, the result of the time to at least one event of the assisted reproductive treatment is predicted according to at least one linear or non-linear framework. In step 1412, clinical proposals related to the assisted reproductive procedure are generated by the process.

[0110] Referring to FIG. 14B, a process flow diagram of a process for generating clinical recommendations related to an OI procedure according to various embodiments is shown. At step 1402b, one or more digital images of a patient's reproductive anatomy are received via one or more imaging modalities. At step 1404b, the one or more digital images are processed to detect one or more reproductive anatomical structures. At step 1406b, the one or more digital images are processed to annotate, segment, or classify one or more anatomical features of the one or more reproductive anatomical structures. At step 1408b, the one or more anatomical structures are analyzed according to at least one linear or non-linear framework. At step 1410b, the result of the time to at least one event of OI treatment is predicted according to at least one linear or non-linear framework. At step 1412b, clinical recommendations related to the determination of the optimal timing of OI incorporating the result of the time to event are generated by a process aimed at maximizing the pregnancy rate while minimizing the risk of multiple pregnancy. In various embodiments, the clinical recommendations are made in conjunction with an assessment of one or more risk factors including, but not limited to, the patient's age, duration of infertility, number or previous treatment cycles, peak serum E2 concentration on the day of induction, and number of follicles. Risk factors for high-order multiple pregnancy include ≥7 preovulatory follicles (≥10-12 mm), E2>1,000 pg / mL, initial cycle of treatment, age<32, low BMI, use of donor sperm. The recommendations may include restricting pregnancy from OI to singleton or twin births when there are only one or two preovulatory follicles of ≥10-12 mm. The recommendations may include determining the critical size of follicles predicting multiple pregnancy (i.e., 12-15 mm). When assessing the risk of multiple pregnancy, it may be necessary to consider all follicles, especially those of intermediate size (11-15 mm), before inducing ovulation.

[0111] Referring to FIG. 15, a process flow diagram of a method for generating clinical proposals related to reproductive assistance procedures according to various embodiments is shown. At step 1502, an ultrasonic image of the subject's ovarian follicles is obtained using an ultrasonic device. At step 1504, the ovarian ultrasonic image is analyzed according to at least one linear or non-linear framework to annotate, segment, or classify one or more anatomical features of the subject's ovarian follicles. At step 1506, the time-to-event results are predicted, and subsequently, at step 1508, one or more clinical proposals related to the reproductive assistance procedure are generated by the process. In another embodiment, additional input 1510 is obtained through one or more processes. At step 1512, the subject's electronic medical record is obtained as input. At step 1514, anonymized third-party electronic medical records are obtained as input. At step 1516, the subject's reproductive physiological data is obtained as input. At step 1518, environmental data related to the subject's reproductive cycle is obtained as input. At step 1520, these additional inputs 1510 are analyzed together with the ovarian ultrasonic image according to at least one linear or non-linear framework, predicting the time-to-event results at step 1506, and subsequently generating one or more clinical proposals related to the reproductive assistance procedure at step 1508.

[0112] Referring to FIG. 15B, a process flow diagram of a method for digital image processing related to an ovulation induction cycle according to various embodiments is shown. At step 1502b, one or more digital images (e.g., ultrasound images) of the target ovarian follicles are acquired using an ultrasound device. At step 1504b, the ovarian ultrasound images are analyzed according to at least one linear or non-linear framework (i.e., a machine learning framework) to annotate, segment, or classify one or more anatomical features of the target ovarian follicles. In certain embodiments, the one or more anatomical features include the quantity and size of one or more ovarian follicles. At step 1506b, at least one processor can analyze the one or more anatomical features according to at least one machine learning framework to predict the result of the time to at least one event. In certain embodiments, the time to at least one event includes the ovulation induction day within the target's ovulation induction cycle. At step 1508b, at least one processor can generate one or more clinical recommendations related to the target's ovulation induction cycle. In certain embodiments, the one or more clinical recommendations can include the proposed timing for the administration of at least one pharmaceutical to the target. In some embodiments, the at least one pharmaceutical includes an ovulation inducer. In certain embodiments, the one or more clinical recommendations can include the proposed timing for sperm delivery or intrauterine insemination corresponding to the ovulation induction cycle. Sperm delivery can include one or more means for sperm delivery, including artificial means (e.g., sperm delivered in a clinical setting) and / or natural means (sperm delivered through sexual intercourse).

[0113] According to certain aspects of the present disclosure, the additional input 1520b can be obtained via one or more data inputs and / or data transfer interfaces. In step 1512b, the electronic medical record of the subject can be obtained as a data input. According to certain embodiments, the electronic medical record of the subject can include one or more data sets selected from the group consisting of diagnostic results, bodily fluid biomarkers, hormone markers, hormone levels, genomic biomarkers, proteomic biomarkers, treatment procedures, treatment schedules, follicle size and number, follicle growth rate, pregnancy rate, and ovulation induction data. In step 1514b, an anonymized third-party electronic medical record is obtained as a data input. According to certain embodiments, the anonymized third-party electronic medical record can include one or more data sets selected from the group consisting of diagnostic results, bodily fluid biomarkers, hormone markers, hormone levels, genomic biomarkers, proteomic biomarkers, treatment procedures, treatment schedules, follicle size and number, follicle growth rate, pregnancy rate, and ovulation induction data. In step 1516b, the reproductive physiological data of the subject is obtained as an input. According to certain embodiments, the reproductive physiological data of the patient can include diagnostic results, diagnostic biomarkers, genomic markers, proteomic markers, bodily fluid analytes, chemical panels, and measured hormone levels. In step 1518b, environmental data related to the reproductive cycle of the subject is obtained as an input. According to certain embodiments, the environmental data can include the patient's medical data over the years collected based on the day, time, and time of week of the reproductive cycle.

[0114] In various embodiments, the method provides an assessment of one or more risk factors including, but not limited to, the patient's age, duration of infertility, number or previous treatment cycles, peak serum E2 concentration on the day of stimulation, and number of follicles. Risk factors for high-order multiple pregnancy include ≥7 preovulatory follicles (≥10 - 12 mm), E2 > 1,000 pg / mL, initial cycle of treatment, age < 32, low BMI, use of donor sperm. The process may require determination of a critical follicle size (i.e., 12 - 15 mm) to predict multiple pregnancy, and consideration of all follicles prior to inducing ovulation, and in particular, evaluation of follicles of intermediate size (11 - 15 mm) may be included when assessing the risk of multiple pregnancy. In step 1520b, these additional inputs 1520b are analyzed along with ovarian ultrasound images according to at least one linear or non-linear framework (i.e., a machine learning framework), and in step 1506b, the time-to-event result is predicted, and subsequently, in step 1508b, one or more clinical recommendations related to OI treatment are generated. According to certain embodiments, the one or more clinical recommendations may include recommendations for maximizing the pregnancy rate associated with OI treatment while minimizing the risk of multiple pregnancy.

[0115] As will be appreciated by those skilled in the art, the present invention may be embodied as a method (e.g., including a computer-implemented process, a business process, and / or any other process), an apparatus (e.g., a system, a machine, a device, a computer program product, etc.), or a combination of the foregoing. Accordingly, embodiments of the present invention may take the form of an all-hardware embodiment, an all-software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software aspects and hardware aspects, which may sometimes be referred to herein generically as a "system." Further, embodiments of the present invention may take the form of a computer program product on a computer-readable medium having computer-executable program code embodied therein.

[0116] Any suitable temporary or non-temporary computer-readable medium may be utilized. The computer-readable medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device. More specific examples of the computer-readable medium include, but are not limited to, tangible storage media such as electrical connections having one or more wires, portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), compact disc read-only memory (CD-ROM), or other optical or magnetic storage devices.

[0117] In the context of this document, the computer-readable medium may be any medium that can contain, store, communicate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The program code executable by a computer for use in the embodiments of the present invention may be transmitted using any suitable medium, including but not limited to the Internet, wired, fiber optic cable, radio frequency (RF) signal, or other media.

[0118] The computer-executable program code for performing the operations of the embodiments of the present invention may be written in an object-oriented, scripting, or non-scripting programming language such as Java, Perl, Smalltalk, C++, etc. However, the computer program code for performing the operations of the embodiments of the present invention may also be written in a conventional procedural programming language such as the "C" programming language or similar programming languages.

[0119] Embodiments of the present invention have been described above with reference to flowcharts and / or block diagrams of methods, apparatuses, systems, and computer program products. It will be understood that each block of the flowcharts and / or block diagrams, and / or combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-executable program code portions. These computer-executable program code portions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the code portions executed via the processor of the computer or other programmable data processing apparatus create a mechanism for implementing the specified functions / acts in one or more blocks of the flowchart and / or block diagram. Specific machines may be fabricated.

[0120] These computer-executable program code portions (i.e., computer-executable instructions) may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the code portions stored in the computer-readable memory produce a manufactured article that includes an instruction mechanism for implementing the specified functions / operations in the blocks of the flowchart and / or block diagram. Computer-executable instructions can be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of program modules can be combined or distributed as desired in various embodiments.

[0121] Computer-executable program code may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus so that the code portion executed on the computer or other programmable apparatus provides steps for implementing specific functions / acts in the blocks of a flowchart and / or block diagram. A computer-implemented process may be generated. Alternatively, to carry out the embodiments of the present invention, computer program implementation steps or acts may be combined with steps or acts implemented by an operator or a human being.

[0122] As used herein in this clause, a processor may be "operable to" or "configured to" perform a particular function in a variety of ways, including, for example, causing a general-purpose circuit to perform functions by executing specific computer-executable program code embodied on a computer-readable medium, and / or causing one or more application-specific circuits to perform functions.

[0123] The terms "program" or "software" are used herein in a general sense to refer to any type of computer code or set of computer-executable instructions that can be utilized to program a computer or other processor to implement various aspects of the present technology as described above. Additionally, according to one aspect of this embodiment, it should be understood that one or more computer programs that execute the methods of the present technology when executed need not be present on a single computer or processor and may be distributed among a plurality of different computers or processors in a modular fashion to implement various aspects of the present technology.

[0124] All definitions, as defined and used herein, shall be understood to take precedence over dictionary definitions, definitions in incorporated documents by reference, and / or ordinary meanings of defined terms.

[0125] As used herein and in the claims, the indefinite articles “a” and “an” shall be understood to mean “at least one” unless explicitly indicated to the contrary. As used herein, the terms “right,” “left,” “top,” “bottom,” “upper,” “lower,” “inner,” and “outer” refer to directions within the referenced drawings.

[0126] The term “and / or” as used herein and in the claims shall be understood to mean “either or both” of the elements so conjoined, i.e., elements that in some cases coexist and in other cases separate. Multiple elements listed using “and / or” shall likewise be construed as “one or more” of the elements so conjoined. Other elements may optionally exist, whether or not they are related to the specifically identified elements, regardless of whether they are specifically identified by the “and / or” clause. Thus, by way of non-limiting example, a reference to “A and / or B” when used with non-limiting language such as “comprising” may refer in one embodiment to only A (optionally including elements other than B), in another embodiment to only B (optionally including elements other than A), and in yet another embodiment to both A and B (optionally including other elements), and so on.

[0127] As used herein and in the claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" is inclusive, i.e., it includes at least one of a plurality of elements or a list of elements, but also includes two or more, and optionally also includes additional unlisted items. Only terms that are explicitly indicated to the contrary, such as "only one of" or "exactly one of", or "consisting of" when used in the claims, refer to exactly one element of a plurality of elements or a list of elements. In general, the term "or" as used herein should be interpreted as indicating an exclusive alternative (i.e., "one or the other, but not both") only when preceded by an exclusive term such as "either", "one of", "only one of" or "exactly one of". "Consisting essentially of" shall have its ordinary meaning as used in the field of patent law when used in the claims.

[0128] As used herein and in the claims, the phrase "at least one" referring to a list of one or more elements should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but does not necessarily include one of every element specifically listed in the element list, nor does it exclude any combination of elements in the element list. This definition also allows for the optional presence of elements other than those specifically identified in the element list referred to by the phrase "at least one", whether or not they are related to those specifically identified elements. Thus, by way of non-limiting example, "at least one of A and B" (or equivalently "at least one of A or B", or equivalently "at least one of A and / or B") in one embodiment refers to at least one A where B is absent and optionally includes two or more As (and optionally includes elements other than B), in another embodiment refers to at least one B where A is absent and optionally includes two or more Bs (and optionally includes elements other than A), and in yet another embodiment refers to at least one, optionally two or more As, and at least one, optionally two or more Bs (and optionally includes other elements), and so on.

[0129] In the claims and in the above specification, all transitional terms such as "comprises", "includes", "carries", "has", "contains", "involves", "holds", "consists of", etc. are non-limiting, that is, they should be understood to mean including but not limited to. As described in Section 2111.03 of the USPTO Patent Examination Procedure Manual, only the transitional terms "consisting of" and "consisting essentially of" should be limiting or semi-limiting transitional terms, respectively.

[0130] This disclosure includes what is contained in the appended claims, as well as the disclosure of the foregoing description. Although the invention has been described in exemplary forms with a certain degree of particularity, it is to be understood that the present disclosure has been made by way of example only, and that numerous changes in the details of the structure and combination and arrangement of parts may be resorted to without departing from the spirit and scope of the invention.

Claims

1. A method for digital image processing in assisted reproductive technology, the method comprising: acquiring, through one or more imaging modalities, one or more digital images of a patient's reproductive anatomical structures, including three-dimensional (3D) ultrasound images; processing the one or more digital images to detect one or more reproductive anatomical structures; processing the one or more digital images to annotate, segment, or classify one or more anatomical features of the one or more reproductive anatomical structures; analyzing the one or more anatomical features according to at least one linear or non-linear framework; predicting the time to at least one event for an assisted reproductive procedure according to the at least one linear or non-linear framework, the result of the time to the at least one event including predicting the egg retrieval date.

2. The method according to claim 1, further comprising processing the one or more digital images to measure the volume of one or more ovarian follicles of the patient.

3. The method according to claim 1, further comprising comparing a first digital image of the patient's reproductive anatomical structure with a second digital image of the patient's reproductive anatomical structure.

4. The method according to claim 1, wherein the one or more digital images include ultrasound images and the one or more imaging modalities include 3D ultrasound.

5. The method according to claim 1, wherein the patient's reproductive anatomical structure is selected from the group consisting of cells, fallopian tubes, ovaries, eggs, multiple eggs, follicles, cysts, uterus, endometrium, endometrial thickness, uterine wall, ovum, and blood vessels.

6. The method according to claim 1, wherein the time to the at least one event includes a dosing schedule.

7. The method according to claim 3, further comprising determining the growth pattern of one or more ovarian follicles of the patient.

8. The method according to claim 7, further comprising determining the follicle maturation date of the one or more ovarian follicles of the patient.

9. A method for digital image processing in assisted reproductive technology, the method comprising: receiving, through one or more imaging modalities, one or more digital images of a patient's reproductive anatomical structures, including three-dimensional (3D) ultrasound images; processing the one or more digital images to detect one or more reproductive anatomical structures; processing the one or more digital images to annotate, segment, or classify one or more anatomical features of the one or more reproductive anatomical structures; analyzing the one or more anatomical features according to at least one linear or non-linear framework; predicting an ovum maturation date according to at least one linear or non-linear framework; generating one or more clinical recommendations related to reproductive assistance procedures according to the ovum maturation date, the one or more clinical recommendations including recommendations regarding the timing of egg collection; a method comprising.

10. The method according to claim 9, wherein the one or more reproductive anatomical structures are selected from the group consisting of oocytes, blastomeres, ovaries, follicles, cystic ovaries, polycystic ovaries, follicles, antral follicles, and endometrial thickness.

11. The method according to claim 9, wherein the one or more clinical recommendations include recommendations regarding the amount and timing of embryo biopsy.

12. The method according to claim 9, wherein the one or more clinical recommendations include recommendations regarding one or more clinical management tasks including clinical staff allocation or patient scheduling.

13. The method according to claim 9, wherein the one or more clinical recommendations include recommendations regarding the amount and timing of embryo fertilization including standard fertilization and intracytoplasmic sperm injection.

14. The method according to claim 9, wherein the one or more clinical recommendations include recommendations regarding the amount and timing of embryo culture.

15. The method according to claim 9, wherein the one or more clinical recommendations include recommendations regarding the amount and timing of embryo cryopreservation.

16. A system for digital image processing in reproductive assistance technology, the system comprising: an imaging sensor operable to perform three-dimensional (3D) ultrasound to collect one or more 3D images of a patient's reproductive anatomical structure; a storage device for locally or remotely storing the one or more 3D images of the patient's reproductive anatomical structure; At least one computer-readable storage medium storing computer-executable instructions and at least one processor operably coupled therewith, wherein when the computer-executable instructions are executed, the at least one processor is caused to perform one or more actions, and the one or more actions include receiving the one or more 3D images of the reproductive anatomical structure of the patient; processing the one or more 3D images of the reproductive anatomical structure of the patient to detect one or more reproductive anatomical structures and annotating one or more anatomical features of the one or more reproductive anatomical structures; analyzing the one or more anatomical features according to at least one non-linear framework to predict the time to at least one event, wherein the at least one non-linear framework includes an artificial neural network selected from a convolutional neural network and a fully convolutional neural network, and the time to the at least one event includes predicting the egg retrieval date from the patient for in vitro fertilization; generating at least one graphical user output in response to the prediction of the time to the at least one event, wherein the at least one graphical user output corresponds to one or more proposed clinical actions for the clinical management of the patient's infertility. A system comprising generating.

17. The one or more actions of the at least one processor further include receiving the patient's health record from an electronic health record database and analyzing the patient's health record together with the one or more anatomical features according to the at least one non-linear framework to predict the time to the at least one event, wherein the patient's health record includes a diagnosis result, a body fluid biomarker, a hormone biomarker, a hormone level, a genomic biomarker, a proteomic biomarker, a treatment, a treatment schedule, an implantation schedule, a follicle size and number, a follicle growth rate, a pregnancy rate, and one or more selected from the group consisting of the date of implantation. The system according to claim 16.

18. The one or more actions of the at least one processor further include receiving health records of anonymized patients undergoing infertility treatment from an electronic health record database and, in accordance with the at least one non-linear framework, analyzing the health records of the anonymized patients together with the one or more anatomical features to predict the time to the at least one event, wherein the health records of the anonymized patients undergoing infertility treatment include one or more selected from the group consisting of ultrasound images, images of the reproductive anatomical structures, diagnostic results, body fluid biomarkers, hormone markers, hormone levels, genomic biomarkers, proteomic biomarkers, treatment procedures, treatment schedules, implantation schedules, follicle size and number, follicle growth rate, follicle growth rate, pregnancy rate, and the day of implantation. The system according to claim 17.

19. The one or more actions of the at least one processor include receiving clinical outcome data related to the patient, storing the clinical outcome data related to the patient in an outcome database, and further including, in accordance with the at least one non-linear framework, analyzing the clinical outcome data related to the patient together with the one or more anatomical features to predict the time to the at least one event. The system according to claim 16.

20. The system according to claim 16, wherein the at least one non-linear framework includes a fully convolutional neural network having a plurality of convolutional layers.

21. The one or more actions of the at least one processor further include receiving reproductive physiological data of the patient and, in accordance with the at least one non-linear framework, analyzing the reproductive physiological data related to the patient together with the one or more anatomical features to predict the time to the at least one event. The system according to claim 16.

22. The one or more actions of the at least one processor further include receiving environmental data related to the patient's reproductive cycle and analyzing the environmental data along with the one or more anatomical features according to the at least one non-linear framework to predict the time to the at least one event, wherein the environmental data is collected based on one or more selected from the group consisting of the day, time, and time of week of the reproductive cycle, and includes the patient's multi-year medical record data. The system according to claim 16.

23. The reproductive anatomical structure of the patient includes the anatomical structure of the patient's follicles, and the fully convolutional neural network is configured to segment the follicles based on qualitative and / or quantitative changes. The system according to claim 20.

24. A system for digital image processing in reproductive assistive technology, the system comprising an imaging sensor operable to execute one or more imaging modalities to collect one or more digital images of a patient's reproductive anatomical structure, wherein the one or more imaging modalities include three-dimensional (3D) ultrasound, the one or more digital images include one or more 3D ultrasound images, and the reproductive anatomical structure includes the ovarian anatomical structure of the patient. An imaging sensor; an artificial intelligence engine configured to receive, locally or remotely, the one or more digital images of the patient's reproductive anatomical structure and process the one or more digital images of the patient's reproductive anatomical structure to generate a prediction of the time to at least one event according to at least one non-linear framework, wherein the at least one non-linear framework is selected from the group consisting of a convolutional neural network, a recurrent neural network, a fully convolutional neural network, an extended residual network, a generative adversarial network, and combinations thereof, and the time prediction to the at least one event includes the optimal egg collection day for in vitro fertilization. An artificial intelligence engine; An electronic medical record database containing medical record data of one or more patients undergoing infertility treatment using assisted reproductive technology, wherein the electronic medical record database is configured to communicate the medical record data of the one or more patients undergoing infertility treatment using assisted reproductive technology to the artificial intelligence engine such that the medical record data is incorporated into the at least one non-linear framework. An electronic medical record database, An outcome database containing clinical outcome data of one or more patients undergoing fertilization treatment by assisted reproductive technology, wherein the outcome database is configured to communicate the clinical outcome data of the one or more patients undergoing fertilization treatment by assisted reproductive technology to the artificial intelligence engine such that the clinical outcome data is incorporated into the at least one non-linear framework, and the clinical outcome data is related to the success rate of embryo implantation and pregnancy after in vitro fertilization. An outcome database, An application server operably linked with the artificial intelligence engine to receive one or more 3D ultrasound images of the reproductive anatomical structure of the patient, the medical record data, the clinical outcome data, and a time prediction to the at least one event, wherein the application server executes an application configured to generate one or more proposed clinical actions for the clinical management of infertility through assisted reproductive technology in response to the time prediction to the at least one event, and the one or more proposed clinical actions include a proposed time for egg retrieval from the patient. An application server, A user device communicably linked with the application server, wherein the user device is configured to display a graphical user interface including the one or more proposed clinical actions for the clinical management of infertility. A system comprising the user device,

25. The system according to claim 24, wherein the user device is a smartphone.

26. The system according to claim 24, wherein the application server is configured to communicate anonymized patient history data of one or more patients undergoing infertility treatment using assisted reproductive technology to the artificial intelligence engine, and the anonymized patient history data is incorporated into the at least one non-linear framework. Claim 27 A computer-readable storage medium encoded with computer-executable instructions for instructing one or more processors to perform operations of a method for predicting the time to at least one event based on one or more ultrasound images of a patient's reproductive anatomy, the operations comprising: Receiving one or more ultrasound images of a patient's reproductive anatomy, including three-dimensional (3D) ultrasound images; Receiving one or more patient medical records from an electronic medical record database, the one or more patient medical records including stored historical medical records of one or more patients undergoing infertility treatment through reproductive assistance technology; Processing the one or more ultrasound images of the patient's reproductive anatomy to detect one or more reproductive anatomical structures and annotating one or more anatomical features of the one or more reproductive anatomical structures; Processing the one or more patient medical records of the one or more patients undergoing infertility treatment through reproductive assistance technology to identify one or more patient phenotypes, the one or more patient phenotypes being selected from the group consisting of ovulation, anovulation, oligo-ovulation, endometriosis, male factor infertility, tubal factor infertility, and reduced ovarian reserve; Analyzing the one or more anatomical features and the one or more patient phenotypes according to at least one non-linear framework, the non-linear framework including a convolutional neural network; Predicting the time to the at least one event according to the at least one non-linear framework, the time prediction to the at least one event including the optimal egg collection day from the patient for in vitro fertilization. Claim 28 The computer-readable storage medium of claim 27, wherein the operations further comprise analyzing past outcome data of one or more patients undergoing infertility treatment using the one or more anatomical features and the one or more patient phenotypes according to the at least one non-linear framework to predict the time to the at least one event. Claim 29 The operation further includes generating one or more reproductive endocrinology insights in response to the time to the at least one event, and the one or more reproductive endocrinology insights are selected from the group consisting of the expected oocyte retrieval volume from the patient, the proposed timing of oocyte retrieval from the patient, the expected amount of eggs per day of the patient, the expected maturation day of the patient's eggs, the expected number of embryos of the patient, the expected quality of the patient's embryos, the amount and schedule of embryo biopsies, and improvement in pregnancy rate. The computer-readable storage medium according to claim 27.

30. The operation further includes receiving the reproductive physiological data of the patient and analyzing the reproductive physiological data using the one or more anatomical features and the one or more phenotypes of the patient according to the at least one non-linear framework to predict the time to the at least one event, wherein the reproductive physiological data of the patient includes one or more of a diagnosis result, a biomarker, a genomic marker, a proteomic marker, a body fluid analyte, a chemical panel, and a measured hormone level. The computer-readable storage medium according to claim 27.

31. The computer-readable storage medium according to claim 27, wherein the at least one non-linear framework includes a plurality of convolutional neural networks.

32. The one or more actions include receiving reproductive physiological data related to the patient and environmental data related to the patient's reproductive cycle, analyzing the reproductive physiological data and the environmental data according to the at least one non-linear framework together with the one or more anatomical features to predict the time to the at least one event, wherein the reproductive physiological data of the patient includes one or more selected from the group consisting of a diagnosis result, a diagnostic biomarker, a genomic marker, a proteomic marker, a body fluid analyte, a chemical panel, and a measured hormone level, wherein the environmental data includes historical medical record data collected from the patient based on one or more selected from the group consisting of the day of the reproductive cycle, the date and time, and the time of the week. The system according to claim 16.

33. The at least one non-linear framework includes a survival convolutional neural network, The one or more actions include Obtaining one or more patient phenotypes from the patient's historical medical records; Analyzing the one or more patient phenotypes together with the one or more anatomical features, the reproductive physiology data, and the environmental data according to the survival convolutional neural network to predict the time to the at least one event, the system of claim 32 further comprising.

34. The system of claim 24, wherein the at least one non-linear framework includes one or more convolutional neural networks optimized for multi-task learning.

35. A system for digital image processing in assisted reproductive technology, the system comprising: An image sensor configured to collect one or more digital images of a patient's reproductive anatomical structure, including three-dimensional (3D) ultrasound images; A computing device communicatively coupled to the image sensor for receiving the one or more digital images of the patient's reproductive anatomical structure; At least one processor communicatively coupled to the computing device and at least one non-transitory computer-readable medium storing instructions that, when executed, cause the at least one processor to perform one or more operations, the one or more operations comprising: Receiving the one or more digital images of the patient's reproductive anatomical structure; Processing the one or more digital images of the patient's reproductive anatomical structure to detect one or more reproductive anatomical structures and annotating one or more anatomical features of the one or more reproductive anatomical structures; Analyzing the one or more anatomical features according to at least one machine learning framework to predict the time to at least one event, the time to the at least one event including the ovulation induction day within the patient's ovulation induction cycle; Generating at least one graphical user output corresponding to one or more clinical actions related to the patient, the one or more clinical actions including a proposed timing for administration of at least one pharmaceutical to the patient, the at least one pharmaceutical including an ovulation inducer, and a non-transitory computer-readable medium.

36. The system of claim 35, wherein the one or more clinical actions include a proposed timing for sperm delivery or intrauterine insemination corresponding to the ovulation induction cycle.

37. The system of claim 35, wherein the one or more operations of the processor further include analyzing a plurality of electronic health record data of the patient, together with the one or more anatomical features, to predict the time to the at least one event.

38. The system of claim 37, wherein the plurality of electronic health record data includes one or more data sets selected from the group consisting of diagnostic results, body fluid biomarkers, hormone markers, hormone levels, genomic biomarkers, proteomic biomarkers, treatment procedures, treatment schedules, follicle size and number, follicle growth rate, pregnancy rate, and ovulation induction data.

39. The system of claim 35, wherein the one or more operations of the processor further include analyzing a plurality of anonymized historical data from one or more anonymized ovulation induction patients, together with the one or more anatomical features, to predict the time to the at least one event.

40. The system of claim 39, wherein the plurality of anonymized historical data includes one or more data sets selected from the group consisting of diagnostic results, body fluid biomarkers, hormone markers, hormone levels, genomic biomarkers, proteomic biomarkers, treatment procedures, treatment schedules, follicle size and number, follicle growth rate, pregnancy rate, and ovulation induction data.

41. The system of claim 35, wherein the machine learning framework is selected from the group consisting of artificial neural networks, regression models, convolutional neural networks, recurrent neural networks, fully convolutional neural networks, extended residual networks, and adversarial generative networks.

42. The system of claim 35, wherein the one or more reproductive anatomical structures include one or more ovarian follicles, and the one or more anatomical features include the amount and size of the one or more ovarian follicles.

43. The system of claim 35, wherein one or more operations of the processor further comprise receiving reproductive physiological data of the patient and analyzing the reproductive physiological data together with the one or more anatomical features to predict the time to the at least one event.

44. The system of claim 35, wherein one or more operations of the processor further comprise analyzing the one or more anatomical features according to the at least one machine learning framework to evaluate the risk of multiple pregnancy of the patient.

45. A method for processing digital images in assisted reproductive technology, the method comprising: acquiring, by an ultrasonic device including a three-dimensional (3D) ultrasonic device, one or more digital images of the reproductive anatomical structure of a patient, including three-dimensional (3D) ultrasonic images; receiving, by at least one processor, the one or more digital images; processing, by the at least one processor, the one or more digital images to detect one or more reproductive anatomical structures of the reproductive anatomical structure of the patient; processing, by the at least one processor, the one or more digital images to annotate one or more anatomical features of the one or more reproductive anatomical structures; processing, by the at least one processor, the one or more anatomical features according to at least one machine learning framework to predict the time to at least one event, wherein the time to the at least one event includes the ovulation induction day within the ovulation induction cycle of the patient; generating, by the at least one processor, at least one clinical recommendation including a proposed timing of administration of at least one pharmaceutical to the patient according to the ovulation induction day, wherein the at least one pharmaceutical includes an ovulation inducer.

46. The method of claim 45, wherein the at least one clinical recommendation includes a proposed timing for sperm delivery or intrauterine insemination corresponding to the ovulation induction cycle.

47. The method of claim 45, wherein the one or more reproductive anatomical structures include one or more ovarian follicles, and the one or more anatomical features include the quantity and size of the one or more ovarian follicles.

48. The method according to claim 45, further comprising analyzing, by the at least one processor, the one or more anatomical features according to the at least one machine learning framework to evaluate the risk of multiple pregnancy of the patient.

49. The method according to claim 47, further comprising analyzing, by the at least one processor, the one or more anatomical features according to the at least one machine learning framework to determine the maturation rate of the one or more ovarian follicles of the patient.

50. The method according to claim 45, further comprising analyzing, by the at least one processor, a plurality of electronic health record data of the patient together with the one or more anatomical features to predict the time to the at least one event.

51. The method according to claim 50, wherein the plurality of electronic health record data comprises one or more data sets selected from the group consisting of diagnostic results, body fluid biomarkers, hormone markers, hormone levels, genomic biomarkers, proteomic biomarkers, treatment procedures, treatment schedules, follicle size and number, follicle growth rate, pregnancy rate, and ovulation induction data.

52. The method according to claim 45, further comprising analyzing, by the at least one processor, a plurality of anonymized historical data from one or more anonymized ovulation induction patients together with the one or more anatomical features to predict the time to the at least one event.

53. The method according to claim 52, wherein the plurality of anonymized historical data comprises one or more data sets selected from the group consisting of diagnostic results, body fluid biomarkers, hormone markers, hormone levels, genomic biomarkers, proteomic biomarkers, treatment procedures, treatment schedules, follicle size and number, follicle growth rate, pregnancy rate, and ovulation induction data.

54. An apparatus for digital image analysis in reproductive assistive technology, the apparatus comprising means for performing the method according to any one of claims 1 to 15.

55. A system comprising at least one processor and at least one non - transitory computer - readable medium communicatively coupled to the at least one processor, wherein the at least one non - transitory computer - readable medium includes one or more computer - executable instructions stored in the at least one non - transitory computer - readable medium, and when the one or more computer - executable instructions are executed by the at least one processor, the at least one processor is caused to perform one or more operations of the method according to any one of claims 1 to 15.

56. At least one non - transitory computer - readable medium for storing one or more computer - executable instructions, wherein when the one or more computer - executable instructions are executed by a processor, the processor is caused to perform one or more operations of the method according to any one of claims 1 to 15.

57. An apparatus for processing digital images in assisted reproductive technology, the apparatus comprising means for performing the method according to any one of claims 45 to 53.

58. A system comprising at least one processor and at least one non - transitory computer - readable medium communicatively coupled to the at least one processor, wherein the at least one non - transitory computer - readable medium includes one or more computer - executable instructions stored in the at least one non - transitory computer - readable medium, and when the one or more computer - executable instructions are executed by the at least one processor, the at least one processor is caused to perform one or more operations of the method according to any one of claims 45 to 53.

59. At least one non - transitory computer - readable medium for storing one or more computer - executable instructions, wherein when the one or more computer - executable instructions are executed by a processor, the processor is caused to perform one or more operations of the method according to any one of claims 45 to 53.

Citation Information

Patent Citations

  • Ovarian follicle count and size determination

    EP3363368A1

  • Ovarian Image Processing for Diagnosis of a Subject

    US20180144471A1