Systems and methods for fetal gender determination using machine learning models

WO2024243357A3PCT designated stage expired Publication Date: 2025-05-08BABYFLIX MEDIA INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/030639
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-05-23
Filing Date
2024-05-22
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Current methods for fetal gender determination, especially in early pregnancy, are limited by the need for skilled sonographers, expensive lab tests, or invasive procedures, which are not readily available in under-resourced areas, leading to potential inaccuracies and risks.

Method used

A machine learning model processes sonographic data to predict fetal gender, implemented in ultrasound devices or as a cloud service, allowing real-time identification of fetal gender as early as 8-9 weeks without relying on human expertise or invasive procedures.

Benefits of technology

This approach provides a reliable, cost-effective, and efficient method for fetal gender identification, accessible in under-resourced areas, reducing the need for skilled sonographers and invasive procedures while improving accuracy and accessibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024030639_08052025_PF_FP_ABST
    Figure US2024030639_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides systems and methods for identifying a gender of a fetus using a machine learning model. An exemplary system includes a computer processing device comprising a processor and a non-volatile storage medium with instructions for the processor to access sonographic data that depicts a genital region of a fetus, and generate, using a trained machine learning model, as output, a probability that the fetus is male, female, or undetermined.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR FETAL GENDER DETERMINATION USINGMACHINE LEARNING MODELSCROSS REFERENCE

[0001] This application claims priority to US Provisional Application No. 63 / 503,772, filed on May 23, 2023, which is incorporated herein by reference in its entirety for all purposes.BACKGROUND

[0002] Determination of fetal gender is one of the cardinal requirements of an anatomic sonography. It provides important clinical information that could profoundly affect both fetal and neonatal care and is one of the most requested questions health care providers are asked during an ultrasound visit. Recognizing different genders requires an element of skill and training that is not always readily available, for example, in more rural or otherwise underresourced areas.

[0003] Fetal gender determination is most commonly performed in the second trimester by direct visualization of genitalia through sonography. For expectant parent(s), gender determination can be stress-reducing and can offers an early connection between the parent(s) and their growing offspring. Sonography is commonly performed at 18-20 weeks of gestation. Fetal sexual differentiation typically begins at 11 weeks of gestation. Accuracy for second trimester gender determination by a trained sonographer is 99%. As sonography technology has advanced, it is now feasible to visualize fetal sex in later stages of the first trimester, as early as 12 weeks, based on identifying the genital tubercle. Due to the early pregnancy, however, there can be a substantial chance of getting a false -negative diagnosis of fetal gender. Early sonography can also be limited by maternal body mass index, placental placement, and operator skill. For these reasons, early sonographic determination of fetal gender has not been adopted by obstetricians and is not standard of care for fetal gender determination.

[0004] Alternatively, procedural testing options such as chorionic villus sampling and amniocentesis may be performed if prenatal blood work indicates the need for investigating genetic or developmental conditions of the fetus. However, these procedural techniques can come with potential risks, for example, fetal miscarriage which can be as high as 1%.

[0005] Available options for fetal gender determination typically require a skilled sonographer, an expensive lab test, or an invasive procedure. These serve as barriers in ruralor under-resourced areas where the advancement of technology or resource availability' is not afforded.SUMMARY

[0006] There are needs for a reliable, convenient, and cost-effective approach for fetal gender identification. As such, expectant parents can be offered a connection to their growing offspring to know the gender in an early stage of pregnancy without relying on expensive lab tests or invasive procedures. Moreover, expectant parents in under-resourced areas can be offered the same level of medical diagnosis as those in large cities.

[0007] The present disclosure provides systems and methods that enable the identification of fetal gender in an efficient and reliable manner. The system and methods disclosed herein provide a machine learning model that processes sonographic data of a fetus and generate a prediction that the fetus is male, female or undetermined. The machine learning model can be implemented in ultrasound imaging devices or local devices in connection therewith, which allows real-time identification while expectant mothers are under ultrasound examination. The machine learning model can also be implemented as web or cloud computing service, which allows users to upload sonographic images and videos on their personal devices and receive prediction results in a timely manner. The use of machine learning models allows gender identification at early pregnancy weeks (e.g., 8-9 weeks), and avoids naked eye identification by sonographers and any invasive procedures.

[0008] One aspect of the present disclosure provides a system for identifying a gender of a fetus. The system comprises a computer processing device comprising a processor and a nonvolatile storage medium with instructions for the processor to access sonographic data that depicts a genital region of a fetus and generate, using a trained machine learning model, as output, a probability that the fetus is male, female, or undetermined.

[0009] In some embodiments, the sonographic data may comprise one or more sonographic images of the fetus. The one or more sonographic images may comprise one or more of two- dimensional (2D) or three-dimensional (3D) images.

[0010] In some embodiments, the sonographic data may comprise one or more sonographic videos of the fetus. In other embodiments, the sonographic data may comprise one or more sonographic cine clips of the fetus.

[0011] In some embodiments, the sonographic data may comprise sonographic images and / or videos directly transmitted from an ultrasound imaging device. In other embodiments, thesonographic data may comprise sonographic images and / or videos transmitted from an image archiving system.

[0012] In some embodiments, the probability may be generated in real time when a patient is under examination by the ultrasound imaging device.

[0013] In some embodiments, the processor may be configured to display the probability with the sonographic data on the ultrasound imaging device in real time.

[0014] In some embodiments, the processor is configured to mask the probability displayed on the ultrasound imaging device. In other embodiments, the sonographic data may comprise a first sonographic image that depicts the genital region of the fetus and a second sonographic image that depicts non-genital regions of the fetus, and the processor may be configured to filter out the first sonographic image.

[0015] In some embodiments, the trained machine learning model may be trained on training data associated with a fetus, the training data comprising a first training dataset with a ground truth label as male and a second training dataset with a ground truth label as female. In other embodiments, the training data may further comprise a third training dataset with a ground truth label as undetermined.

[0016] In some embodiments, the training data may comprise sonographic images of the fetus. The sonographic images may comprise one or more of two-dimensional (2D) or three- dimensional (3D) images. In some embodiments, the training data may comprise sonographic videos of the fetus. In other embodiments, the training data may comprise sonographic cine clips of the fetus.

[0017] In some embodiments, the fetus may be at a gestational age of between 8 and 41 weeks.

[0018] In some embodiments, each of the first, second and third training datasets may comprise at least 5,000 sonographic images.

[0019] In some embodiments, the trained machine learning model may comprise a convolutional neural network. The convolutional neural network may comprise one or more of convolution layers, pooling layers, or fully connected layers.

[0020] Another aspect of the present disclosure provides a method of identifying a gender of a fetus. The method comprises accessing sonographic data that depicts a genital region of a fetus, and generating, using a trained machine learning model, as output, a probability that the fetus is male, female, or undetermined.

[0021] In some embodiments, the sonographic data may be real-time sonographic data generated from an ultrasound imaging device. The sonographic data may comprise one ormore sonographic images of the fetus. The one or more sonographic images may comprise one or more of two-dimensional (2D) or three-dimensional (3D) images. In some embodiments, the sonographic data may comprise one or more sonographic videos of the fetus. In other embodiments, the sonographic data may comprise one or more sonographic cine clips of the fetus.

[0022] In some embodiments, the method may further comprise transmitting the sonographic data to a server or a computing device implemented with the trained machine learning model. In other embodiments, the method may further comprise displaying the sonographic data and the probability on a display.

[0023] Another aspect of the present disclosure provides a method of identifying a gender of a fetus. The method comprises accessing sonographic data that depicts a genital region of a fetus generated from an ultrasound imaging device in real time, transmitting the sonographic data to a server or a computing device implemented with a trained machine learning model, receiving, from the trained machine learning model, a prediction that the fetus is male, female, or undetermined based on the sonographic data, and displaying the sonographic data and the prediction on a display coupled to the ultrasound imaging device.

[0024] In some embodiments, the sonographic data may comprise one or more sonographic images of the fetus. The one or more sonographic images comprise one or more of two- dimensional (2D) or three-dimensional (3D) images. In other embodiments, the sonographic data may comprise one or more sonographic videos of the fetus.

[0025] In some embodiments, the sonographic data and the prediction may be displayed one or more of concurrently or in real-time with the generation of the sonographic data.

[0026] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable codes that, upon execution by one or more computer processors, implement a method of identifying a gender of a fetus. The method comprises accessing sonographic data that depicts a genital region of a fetus and generating, using a trained machine learning model, as output, a probability that the fetus is male, female, or undetermined.

[0027] Another aspect of the present disclosure provides a computer-implemented method of constructing a machine learning model for identifying a gender of a fetus. The method comprises accessing training data comprising a first training dataset with a ground truth label as male and a second training dataset with a ground truth label as female, wherein at least part of the training data depicts a genital region of a fetus, and training the machine learning model that runs on one or more processors coupled to memory on the training data, byprogressively matching output of the machine learning model with corresponding ground truth labels.

[0028] In some embodiments, the training data may depict the fetus at a gestational age of between 8 and 41 weeks.

[0029] In some embodiments, the training data may further comprise a third training dataset with a ground truth label as undetermined. Each of the first, second, and third training datasets may comprise at least 5,000 sonographic images.

[0030] In some embodiments, the training data may comprise sonographic images of the fetus. The sonographic images may comprise one or more of two-dimensional (2D) or three- dimensional (3D) images. In some embodiments, the training data may comprise sonographic videos of the fetus. In other embodiments, the sonographic data may comprise sonographic cine clips of the fetus.

[0031] In some embodiments, the machine learning model may be trained using a backpropagation-based gradient update technique. In some embodiments, the machine learning model may comprise a convolutional neural network. The convolutional neural network may comprise one or more convolution layers, pooling layers, or fully connected layers.

[0032] Another aspect of the present disclosure provides a system comprising a computer processing device comprising a processor and a non-volatile storage medium with instructions to construct a machine learning model for identifying a gender of a fetus, the instructions, when executed on the processor, implement actions comprising accessing training data comprising a first training dataset with a ground truth label as male and a second training dataset with a ground truth label as female, wherein at least part of the training data depicts a genital region of a fetus, and training the machine learning model on the training data by progressively matching output of the machine learning model with corresponding ground truth labels.

[0033] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable codes that, upon execution by one or more computer processors, implement a method of constructing a machine learning model for identifying a gender of a fetus. The method comprises accessing training data comprising a first training dataset with a ground truth label as male and a second training dataset with a ground truth label as female, wherein at least part of the training data depicts a genital region of a fetus, and training the machine learning model on the training data, by progressively matching output of the machine learning model with corresponding ground truth labels.

[0034] Another aspect of the present disclosure provides a method of generating images of a fetus, the method comprising accessing sonographic data that depicts a genital region of a fetus generated from an ultrasound imaging device, the sonographic data comprising a plurality of ultrasound images, transmitting the sonographic data to a server or a computing device implemented with a trained machine learning model, determining, with the trained machine learning model, whether a sex of the fetus can be determined from each ultrasound image of the plurality of ultrasound images, and generating a set of the ultrasound images in which the sex of the fetus cannot be determined from each ultrasound image of the set.

[0035] In some embodiments, generating the set of ultrasound images in which the sex of the fetus cannot be determined from each ultrasound image of the set may comprise masking one or more features in an ultrasound image in which the sex of the fetus can be determined.

[0036] In some embodiments, the method may further comprise displaying the set of the ultrasound images in which the sex of the fetus cannot be determined from each ultrasound image of the set.

[0037] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE

[0038] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The novel features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forthillustrative embodiments, in which the principles of the present disclosure are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:

[0040] FIG. 1 is a block diagram of a non-limiting example of machine learning model that processes sonographic data of a fetus as input and identify fetal gender, according to some embodiments of the present disclosure;

[0041] FIG. 2 is a block diagram of a convolutional neural network for identifying fetal gender, according to some embodiments of the present disclosure;

[0042] FIG. 3 is a block diagram of training a machine learning model, according to some embodiments of the present disclosure;

[0043] FIG. 4 is a workflow diagram of identifying fetal gender using a machine learning model, according to some embodiments of the present disclosure;

[0044] FIG. 5 is a dataflow diagram of identifying fetal gender using a machine learning model, according to some embodiments of the present disclosure;

[0045] FIG. 6A is a workflow diagram of identifying fetal gender using a machine learning model, according to some embodiments of the present disclosure;

[0046] FIG. 6B illustrates real-time prediction of fetal gender when expectant mothers are under ultrasound examination, according to some embodiments of the present disclosure; and

[0047] FIG. 7 is a computer system that is programmed or otherwise configured to implement methods provided herein.DETAILED DESCRIPTION

[0048] While various embodiments of the present disclosure have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be employed.

[0049] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1 , greater than or equal to 2, or greater than or equal to 3.

[0050] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” or “less than or equal to” applies to each of the numerical values in thatseries of numerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0051] The term “sonography”, “sonographic”, and “ultrasound” are used interchangeably in the present disclosure.

[0052] The present disclosure provides systems, methods and non-transitory computer readable media for identifying a gender of a fetus using a machine learning model. In some embodiments, the system may comprise a computer processing device comprising a processor and a non-volatile storage medium with instructions for the processor to access sonographic data that depicts a genital region of a fetus and generate, using a trained machine learning model, as output, a probability that the fetus is male, female, or undetermined.

[0053] FIG. 1 is a block diagram of a non-limiting example of machine learning model that processes sonographic data of a fetus as input and predicts the gender of the fetus, in accordance with some embodiments. As illustrated, a trained machine learning model 120 accesses sonographic data of a fetus 110, and generates a prediction of a gender of the fetus. In some embodiments, the output may be a multi-class classification, e.g., a probability of the gender of the fetus as male, female or undetermined. As illustrated in FIG. 1, for example, the output may be a probability of 0.7 that the gender is male (see 132), a probability of 0.1 that the fetus is female (see 134), and a probability of 0.2 that the gender is undetermined (see 136). In other embodiments, the output may be a binary classification, e.g., the gender of the fetus as male or female. In some embodiments, the output includes a confidence score for the determined gender.Input to Machine Learning Model

[0054] In some embodiments, sonographic data may comprise sonographic images of the fetus. The images may be two-dimensional (2D) and / or three-dimensional (3D) images obtained from an ultrasound imaging device. The 2D images may be black-and-white images that depict the fetus, where the bones of the fetus are highlighted in white. The 3D images may be yellow / tan color images that depict specific features (e.g., face, genital tubercle) of the fetus in a well-defined formation. However, embodiments are not limited thereto, and the 2D and 3D images may include any other color schemes for identifying the bones and other features. Non-limiting examples of the size of sonographic images may comprise 1920 x 1080 pixels, 1280 x 872 pixels, 1280 x 720 pixels, 640 x 480 pixels, 852 x 480 pixels, or any other resolutions that the machine learning model may be able to process. Non-limiting examples of the format of sonographic images may comprise “.jpg”, “.jpeg”, “.gif’, “.tiff’, “ psd”, “.raw”, and others.

[0055] In other embodiments, the sonographic data may comprise sonographic videos or cine clips of the fetus obtained from a four-dimensional (4D) ultrasound imaging device. Nonlimiting examples of the resolution of videos or cine clips may comprise 1920 x 1080 pixels, 1280 x 872 pixels, 1280 x 720 pixels, 640 x 480 pixels, 852 x 480 pixels, or any other resolutions that the machine learning model may be able to process. Non-limiting examples of the format of sonographic images may comprise “ avi”, “.mp4”, “.mpeg”, “.mkv”, “.dem”, “.mov”, “,wmv”, “ flv”, “ ts”, and others. In some embodiments, videos or cine clips may be converted to a series of static images for the machine learning model to process.

[0056] The sonographic data, including sonographic images and videos, may be directly transmitted from an ultrasound imaging device. The machine learning model may be implemented in an ultrasound imaging device such that the real-time sonographic images and videos obtain from the device may be used as input to the machine learning model for identifying fetal gender. The model may generate prediction results during ultrasound examination of an expectant mother. Alternatively, sonographic data may be stored in an imaging archiving system (e.g., picture archiving and communication system, PACS), webbased or cloud-based storage system. When the machine learning model accesses sonographic data, the model may generate a prediction of fetal gender.

[0057] The sonographic data, including sonographic images and videos, may comprise a genital region of the fetus, of whom the gender is to be identified. For example, sonographic images and videos may depict one or more of vulva, clitoris, and / or labia of a female fetus, and one or more of scrotum, penis, testicles, and / or raphe of a male fetus. In other examples, the sonographic images and videos may depict internal pelvic structure of the fetus, including uterus and ovary, which may facilitate the identification of fetal gender. In some embodiments, the sonographic images and videos may be taken from different anatomical planes containing fetal genital regions. For example, the images and videos may be taken from transverse plane, sagittal plane, and / or mid-sagittal plane.

[0058] The sonographic data may depict a fetus from an early gestational age to delivery of the newborn. In some embodiments, the sonographic data may depict a fetus from pregnancy weeks 8-41, 9-40, 10-39, 11-38, 12-37, 13-36, and the like.

[0059] In some embodiments, the sonographic data may be pre-processed or augmented before being used as input to the machine learning model for prediction. For example, sections of images and videos that include personal information (e.g., information of expectant mother) and device information (e.g., device settings, scales, resolutions) may be removed or masked. The images and videos may be resized (e.g., downscaling, upscaling)preserving the original aspect of ratio or cropped. Other pre-processing steps may comprise perturbations to brightness, contrast, and color. The pre-processing may be performed by the machine learning model or other computing devices.Examples of Machine Learning Model

[0060] The machine learning model 120 may implement one or more machine learning algorithms. Machine learning may generally involve identifying and recognizing patterns in existing data in order to facilitate making predictions for subsequent data. Machine learning, whether analytical or statistical in nature, may provide deductive or abductive inference based on real or simulated data. The machine learning model 120 may be a trained model. For example, the machine learning model 120 may be a trained model that is trained through supervised learning (e.g., various parameters are determined as weights or scaling factors). Machine learning may comprise one or more of regression analysis, regularization, classification, dimensionality reduction, ensemble learning, meta learning, association rule learning, cluster analysis, anomaly detection, deep learning, or ultra-deep learning. Machine learning may comprise, but is not limited to: k-means, k-means clustering, k-nearest neighbors, learning vector quantization, linear regression, non-linear regression, least squares regression, partial least squares regression, logistic regression, stepwise regression, multivariate adaptive regression splines, ridge regression, principal component regression, least absolute shrinkage and selection operation (LASSO), least angle regression, canonical correlation analysis, factor analysis, independent component analysis, linear discriminant analysis, multidimensional scaling, non-negative matrix factorization, principal components analysis, principal coordinates analysis, projection pursuit, Sammon mapping, t-distributed stochastic neighbor embedding, AdaBoosting, boosting, gradient boosting, bootstrap aggregation, ensemble averaging, decision trees, conditional decision trees, boosted decision trees, gradient boosted decision trees, random forests, stacked generalization, Bayesian networks, Bayesian belief networks, naive Bayes, Gaussian naive Bayes, multinomial naive Bayes, hidden Markov models, hierarchical hidden Markov models, support vector machines, encoders, decoders, auto-encoders, stacked auto-encoders, perceptrons, multi-layer perceptrons, artificial neural networks, feedforward neural networks, convolutional neural networks, recurrent neural networks, long short-term memory, deep belief networks, deep Boltzmann machines, deep convolutional neural networks, deep recurrent neural networks, generative adversarial networks, vision transformers, long short-term memory networks (LSTM), masked autoencoders, etc.

[0061] The systems, the methods, and the computer-readable media disclosed herein may implement one or more computer vision techniques. Computer vision is a field of artificial intelligence that uses computers to interpret and understand the visual world at least in part by processing one or more digital images and videos. In some embodiments, computer vision may use deep learning models (e.g., convolutional neural networks). Bounding boxes and tracking techniques may be used in object detection techniques within computer vision.

[0062] The systems, the methods, the computer-readable media, and the techniques disclosed herein may implement one or more deep-learning techniques. Deep learning is an example of machine learning that may be based on a set of algorithms that attempt to model high-level abstractions in data by using multiple processing layers, with complex structures or otherwise, composed of multiple non-linear transformations. In some embodiments, a drop out method may be used to reduce overfitting. At each training stage, individual nodes are either “dropped out” of the net (e.g., ignored) with a probability 1-p or kept with probability p, so that a reduced network is left; incoming and outgoing edges to a dropped-out node may also be removed. In some embodiments, the reduced network may be trained on the data in that stage. The removed nodes may then be reinserted into the network with their original weights.

[0063] The systems, the methods, and the computer-readable media disclosed herein may implement one or more vision transformer (ViT) techniques. A ViT is a transformer-like model that handles vision processing tasks. While CNNs use convolution, a “local” operation bounded to a small neighborhood of an image, ViTs use self-attention, a “global” operation, since the ViT draws information from the whole image. This allows the ViT to capture distant semantic relevance in an image effectively. Advantageously, ViTs may be well-suited catching long-term dependencies. In some cases, ViTs may be a competitive alternative to convolutional neural networks as ViTs may outperform the current state-of-the-art CNNs by almost four times in terms of computational efficiency and accuracy. ViTs may be well- suited to object detection, image segmentation, image classification, and action recognition. Moreover, ViTs may be applied in generative modeling and multi-model tasks, including visual grounding, visual -question answering, and visual reasoning. In some embodiments, ViTs may represent images as sequences, and class labels for the image are predicted, which allows models to learn image structure independently. Input images may be treated as a sequence of patches where every patch is flattened into a single vector by concatenating the channels of all pixels in a patch and then linearly projecting it to the desired input dimension. For example, a ViT architecture may include the following operations: (A) split an image intopatches; (B) flatten the patches; (C) generate lower-dimensional linear embeddings from the flattened patches; (D) add positional embeddings; (E) provide the sequence as an input to a standard transformer encoder; (F) pretrain a model with image labels (e.g., fully supervised on a huge dataset); and (G) finetune on the downstream dataset for image classification. In some embodiments, there may be multiple blocks in a ViT encoder, with each block comprising three major processing elements: (1) Layer Norm; (2) Multi-head Attention Network; and (3) Multi-Layer Perceptrons. The Layer Norm may keep the training process on track and enable the model to adapt to the variations among the training images. The Multi-head Attention Network may be a network responsible for generating attention maps from the given embedded visual tokens. These attention maps may help the network focus on the most critical regions in the image, such as object(s). The Multi-Layer Perceptrons may be a two-layer classification network with a Gaussian Error Linear Unit at the end. The final Multi-Layer Perceptrons block may be used as an output of the transformer. An application of softmax on this output can provide classification labels (e.g., if the application is image classification).

[0064] The systems, the methods, and the computer-readable media disclosed herein may implement one or more masked autoencoder (MAE) techniques. MAEs are scalable selfsupervised learners for computer vision. The MAE leverages the success of autoencoders for various imaging and natural language processing tasks. Some computer vision models may be trained using supervised learning, such as using humans to look at images and created labels for the images, so that the model could learn the patterns of those labels (e.g., a human annotator would assign a class label to an image or draw bounding boxes around objects in the image). In contrast, self-supervised learning may not use any human-created labels. One technique for self-supervised image processing training using an MAE is for before an image is input into an encoder transformer, a certain set of masks are applied to the image. Due to the masks, pixels are removed from the image and therefore the model is provided an incomplete image. At a high level, the model’s task is to now learn what the full, original image looked like before the mask was applied.

[0065] In other words, MAE may include masking random patches of an input image and reconstructing the missing pixels. The MAE may be based on two core designs. First, an asymmetric encoder-decoder architecture, with an encoder that operates on the visible subset of patches (without mask tokens), along with a lightweight decoder that reconstructs the original image from the latent representation and mask tokens. Second, masking a high proportion of the input image, e.g., 75%, may yield a nontrivial and meaningful self-supervisory task. Coupling these two core designs enables training large models efficiently and effectively, thereby accelerating training (e.g., by 3 x or more) and improving accuracy. MAE techniques may be scalable, enabling learning of high-capacity models that generalize well, e.g., a vanilla ViT-Huge model. As mentioned, the MAE may be effective in pretraining ViTs for natural image analysis. In some cases, the MAE uses the characteristic of redundancy of image information to observe partial images to reconstruct original images as a proxy task, and the encoder of the MAE may have the capability of deducing the content of the masked image area by aggregating context information. This contextual aggregation capability may be important in the field of image processing and analysis.

[0066] FIG. 2 is a block diagram of a convolutional neural network 200 for identifying fetal gender, in accordance with some embodiments. The convolutional neural network 200 comprises a deep learning algorithm that takes in input 210 (e.g., sonographic images and / or videos of a fetus) and generates a precited output 280 as a probability of the fetus is male, female or undetermined. For example, the convolutional neural network 200 may predict the fetus has a probability of 0.2 being male, a probability of 0.7 being female, a probability of 0.1 being undetermined.

[0067] The convolutional neural network 200 may have one or more of convolution layers and pooling layers. The convolution layers may perform convolution operations between the input 210 values and convolution filters (matrix of parameters) that are learned over many gradient update iterations during the training. Convolutions operate over 3D tensors, called feature maps, with two spatial axes (height and width) as well as a depth axis (i.e., channel axis). For an RGB image, the dimension of the depth axis is 3, because the image has three color channels (red, green, and blue). For a black-and-white picture, for example, 2D sonographic images, the depth is 1 (levels of gray). The convolution operation may extract patches from its input feature map and applies the same transformation to all of these patches, producing an output feature map. This output feature map is still a 3D tensor: it has a width and a height. Its depth can be arbitrary, because the output depth is a parameter of the layer, and the different channels in that depth axis no longer stand for specific colors as in RGB input; rather, they stand for filters that encode specific aspects of the input 210.

[0068] A pooling layer may operate a pooling operation that divides input into nonoverlapping two-dimensional spaces. For example, the feature maps 220 and 240 are outputs generated from previous convolution operations and used as input to subsequent pooling layers, respectively. A filter with a size of 2 x 2 is slid over the feature maps using a stride of 2. For a receptive field with a size of 2 x 2 (e.g., the part of the feature maps 220 and 240under the filter), an average pooling operation may produce an average value of the four pixels in the receptive field, whereas a maximum pooling operation may select a maximum value of the four pixels in the receptive field. As such, pooling operations may consolidate the features learned by the convolutional neural network 200 and gradually reduce the spatial dimension of the feature maps to minimize the numbers of parameters and computations in the network.

[0069] As illustrated in FIG. 2, input 210 is processed via a convolution operation to generate feature maps 220, which in turn are processed by a pooling operation to generate pooled features maps 230. As the convolutional neural network 200 may comprise a plurality of convolution layers and pooling layers where output from a previous layer may be input to a next layer, the convolution and pooling operations may repeat, thereby generating feature maps 240 and pooled feature maps 250, respectively. In some embodiments, the pooled feature maps 250 which are 2-dimensional arrays, may be processed via a flattening operation, which generates a 1-dimensional vector 260. The vector may be processed via a fully connected layer 270, which generates predicted output 280.

[0070] In some embodiments, the convolutional neural network 200 may be scaled to construe machine learning models with better accuracy and efficiency. For example, one or more of depth, width, resolution of the convolutional neural network 200 may be scaled. Depending on the input, the convolutional neural network 200 may be scaled in a single dimension or multiple dimensions. For example, for higher resolution input sonographic data, the width and depth of the convolutional neural network 200 may be scaled, such that larger receptive fields are able to capture similar features that include more pixels in larger images / videos.

[0071] It should be noted that FIG. 2 is a non-limiting example of convolutional neural network for illustrative purposes. The number of convolution layers, pooling layers, flattening layers and fully connected layers may be adjusted without deviating from the scope of the disclosure herein.Training of Machine Learning Model

[0072] The present disclosure further provides systems, methods, and non-transitory computer readable media for constructing a machine learning model for identifying a gender of a fetus. In some embodiments, a method of constructing a machine learning model for identifying fetal gender may comprise accessing training data that depicts a genital region of a fetus, the training data comprising a first training dataset with a ground truth label as male and a second training dataset with a ground truth label as female, and training the machinelearning model that runs on numerous processors coupled to memory, on the training data by progressively matching output of the machine learning model with corresponding ground truth labels.

[0073] In some embodiments, the machine learning model 120 may be trained by way of supervised learning. A data set may be divided into a training set, a test set, and, in some cases, a validation set. In supervised learning, training data and validation data may be annotated with ground truth labels. During the training process, training data is repeatedly presented to the machine learning model 120, and for each sample presented during training, output generated by the machine learning model 120 may be compared with the corresponding ground truth label. The difference between the ground truth and the generated output may be calculated, and the machine learning model 120 may be modified to cause the output to more closely approximate or predict the ground truth. In some embodiments, a backpropagation algorithm may be utilized to cause the output to more closely approximate the ground truth. During many training iterations, the machine learning may generate outputs that progressively match the corresponding ground truth labels. Subsequently, when new and previously unseen input is presented to the machine learning model, it may generate an output classification value indicating which of the categories the new sample is most likely to fall into. In other words, the machine learning model may “generalize” from its training to new, previously unseen input.

[0074] In some embodiments, the machine learning model 120 may be validated using a validation dataset (e.g., distinct from training data set) to determine accuracy and robustness of the model. Such validation may include applying the model to the validation dataset to make predictions derived from the validation dataset. The machine learning model 120 may then be evaluated to determine whether performance is sufficient based upon the derived predictions. The sufficiency criteria applied to the machine learning model 120 may vary depending upon the size of the training data set available for training, the performance of previous iterations of trained models, or user-specified performance requirements. If the machine learning model 120 does not achieve sufficient performance, additional training may be performed. Additional training may include refinement of the model or retraining on a different training dataset, after which the newly trained model may again be validated and assessed. When the machine learning model 120 has achieved sufficient performance, in some cases, the machine learning model 120 may be stored for present or future use. The model may be stored as sets of parameter values or weights for analysis of further input (e.g., further relevant parameters to use as further predicted variables, further explanatoryvariables, further user interaction data, etc.), which may also include analysis logic or indications of model validity. In some embodiments, a plurality of machine learning models may be stored for generating predictions under different sets of input data conditions. In some embodiments, the plurality of machine learning models may be stored in a database (e.g., associated with a server).

[0075] FIG. 3 is a block diagram of training a machine learning model, in accordance with some embodiments. In the training process 300, weight parameters in each layer of the machine learning model 320 are optimized using backpropagation 350 based on comparison between the estimated / predicted output 330 and the ground truth 340 until the estimated output 330 progressively matches or approaches the ground truth 340. A single cycle of the optimization process is organized as follows. First, given a training dataset as input 310, the forward pass sequentially computes the output in each layer and propagates the function signals forward through the machine learning model 320. In the final output layer, an objective loss function measures an error between the estimated output 330 and given labels (e.g., ground truth 340) of the training data. To minimize the training error, the backward pass 360 uses the chain rule to backpropagate error signals and compute gradients with respect to all weights throughout the neural network. Finally, the weight parameters are updated using optimization algorithms based on stochastic gradient descent (SGD). Several optimization algorithms stem from stochastic gradient descent. For example, the Adagrad and Adam training algorithms perform stochastic gradient descent while adaptively modifying learning rates based on update frequency and moments of the gradients for each parameter, respectively.

[0076] In some embodiments, training data used to train the machine learning model 120 may be annotated with ground truth labels. The training data may comprise sonographic images and videos collected from a variety of ultrasound imaging devices. The annotations may be performed by skilled sonographers, where the fetus in each image / video is labeled as one of male, female and undetermined.

[0077] In some embodiments, the sonographic images may comprise 2D and 3D images. In some embodiments, the training data may comprise sonographic videos or cine clips of the fetus. The fetus depicted in the training data may be from an early gestational age to delivery of the newborn. In some embodiments, the fetus may be from pregnancy weeks 8-41, 9-40, 10-39, 11-38, 12-37, 13-36, and the like.

[0078] In some embodiments, the training data may comprise a first training dataset with a ground truth label as male and a second training dataset with a ground truth label as female.In other embodiments, the training data may further comprise a third training dataset with a ground truth label as undetermined. Each dataset may comprise at least 2,000, at least 3,000, at least 5,000, at least 8,000, at least 10,000, at least 12,000 sonographic images. At least some of the sonographic images may depict a genital region of the fetus.

[0079] Unsupervised learning may be used, in some embodiments, to train the machine learning model 120 to use input data (e.g., sonographic data) and generate a prediction of fetal gender. Unsupervised learning, in some embodiments, includes feature extraction which is performed by the machine learning model 120 on the input data. Extracted features may be used for visualization, for classification, for subsequent supervised training, and more generally for representing the input for subsequent storage or analysis.

[0080] Machine learning models that are commonly used for unsupervised training include k- means clustering, mixtures of multinomial distributions, affinity propagation, discrete factor analysis, hidden Markov models, Boltzmann machines, restricted Boltzmann machines, autoencoders, convolutional autoencoders, recurrent neural network autoencoders, and long short-term memory autoencoders. While there are many unsupervised learning models, they all have in common that, for training, they require a training set without associated labels. Inference of Machine Learning Model

[0081] Following training, the machine learning model can be used to analyze new data that the model has never encountered. Accordingly, the present disclosure provides systems, methods and non-transitory computer readable media for identifying a gender of a fetus using a trained machine learning model. In some embodiments, a method of identifying a gender of a fetus is disclosed herein. The method may comprise accessing sonographic data that depicts a genital region of a fetus, and generating, using a trained machine learning model, as output, a probability that the fetus is male, female, or undetermined.

[0082] FIG. 4 is a workflow diagram of identifying fetal gender using a trained machine learning model, in accordance with some embodiments. At step 410, the trained model accesses sonographic data that depicts a genital region of a fetus. The sonographic data may comprise sonographic images and / or videos. At step 420, the trained model processes the sonographic data and predicts that the fetus is male, female or undetermined. In some embodiments, the prediction may be a probability distribution or a confidence level of the fetus being male, female or undetermined.

[0083] FIG. 5 is a dataflow diagram of identifying fetal gender using a machine learning model, in accordance with some embodiments. A machine learning model as described herein may access sonographic data from different platforms, which provides flexibility andefficiency to users. The machine learning model may be implemented / stored at different platforms. As illustrated, the machine learning model (e.g., 120 and 200 in FIGs. 1 and 2, respectively) may be stored in an ultrasound imaging device and access sonographic data directly transmitted therefrom. By processing the sonographic data including images and / or videos, the machine learning model may generate a prediction of the fetus while the expectant mother is under examination. In other words, the prediction may occur in real-time or with minimum latency (e.g., 1-30 seconds). Alternatively, a video streaming / recording device 508 may stream or record sonographic videos on the ultrasound imaging device 506, where the videos may be stored on a server that is easily accessible to the machine learning model.

[0084] Sonographic data may be stored in an imaging archiving system (e.g., picture archiving and communication system, PACS). The sonographic data may be formatted following the Digital Imaging and Communications in Medicine (DICOM) or health level seven (HL7) standard and stored in medical information platforms or servers like Mercure and Orthanc 510 (see 516). Alternatively, the machine learning model may access sonographic data following an Application Programming Interface (API) request 502. For example, the machine learning model may access sonographic data by calling REpresentational State Transfer (REST) web service or RESTful web service (see 512). Users may also be allowed to upload sonographic data from local computing devices 504, for example, computer browser.

[0085] Sonographic data received from different platforms may be stored in cloud (e.g., Amazon Simple Storage Service S3 520). For example, sonographic images and videos may be exported from medical information platforms or servers like Mercure and Orthanc 510 and stored in the cloud (see 522). Sonographic data obtained from video streaming / recording devices 508 may be uploaded to the cloud via video service (see Amazon Interactive Video Service 518). Sonographic data uploaded by users from their local computing devices 504 and from API requests may also be stored in the cloud (see 514 and 512, respectively).

[0086] In some embodiments, the machine learning model may be implemented via webbased computing service 530 (e.g., Amazon Web Service Lambda) in connection with AWS 550. The machine learning model may access, process sonographic images and videos stored in the cloud, and generate prediction of fetus depicted in the sonograph. For sonographic data received from user’s local computing device and API request, the machine learning model may generate a log receipt of data (540). The model may send prediction results along with the log receipt to users (see 542 and 544).

[0087] The present disclosure described herein provides flexibility to users, whether or they have access to advanced medical resources. The technology herein enables users to receive real-time prediction of fetal gender during ultrasound examination, without relying on skilled sonographers. For users who have limited access to clinical resources or have sonographic images / videos without knowing the gender of the fetus, the present disclosure also allows users to upload the sonographic images from their local devices anytime and receive prediction results in a timely fashion.

[0088] FIG. 6A is a workflow diagram 600 of identifying fetal gender using a machine learning model, in accordance with some embodiments. FIG. 6B illustrates real-time prediction of fetal gender when expectant mothers are under ultrasound examination, in accordance with some embodiments. The present disclosure described herein can provide real time prediction of fetal gender during ultrasound examination, with minimum latency. An expectant mother (see 660 in FIG. 6B) is under ultrasound examination conducted by a sonographer (see 650 in FIG. 6B). The ultrasound imaging device 670 generates real-time sonographic images and / or videos depicting a fetus (see 610 in FIG. 6A). These sonographic data is transmitted to a server or other computing devices implemented with a trained machine learning model (see 620 in FIG. 6A). The machine learning model may process the images and / or videos and generate a prediction that the fetus is male, female or undetermined, and transmit the prediction to the ultrasound imaging device 670 (see 630 in FIG. 6A). The generation and transmission of the prediction may occur in real-time or with minimum latency. The sonographic images and / or videos and the prediction may be displayed one or more of concurrently or in real-time with the generation of the images and / or videos.

[0089] The ultrasound imaging device may display the sonographic images and / or videos along with the prediction (see 640 in FIG. 6A and 680 in FIG. 6B). The expectant mother may be notified of the prediction results while she is under ultrasound examination. Moreover, the machine learning model provides an efficient, effective, and non-invasive approach for gender identification, without relying on skilled sonographer, expensive lab tests, or an invasive procedure.

[0090] In some embodiments, the latency may be in a range of 1 - 60 seconds, 1 - 50 seconds, 1 - 40 seconds, 1 - 30 seconds, 1 - 20 seconds, 1 - 10 seconds, and 1 - 5 seconds.

[0091] The machine learning model may identify sonographic images that comprise genital region of fetus and filter out those images. In some embodiments, the sonographic images may comprise a first image that depicts the genital region of the fetus and a second image thatdepicts non-genital regions of the fetus. The machine learning model may fdter out the first image such that it will not be shown to expectant parents. In other embodiments, the machine learning model may filter out sonographic images that comprise a probability that the fetus is male or female. Alternatively, the machine learning model may mask the genital regions and probabilities displayed on the sonographic images. For example, the machine learning model may generate sonographic images in which the gender of the fetus cannot be determined and mask one or more features in an ultrasound image in which the gender of the fetus can be determined. In some countries where fetal gender identification is illegal or areas that have significant gender-selective abortion, the machine learning model may filter out or mask those images where fetal gender is determinable (e.g., depicting genital region of fetus and / or probability of fetal gender), such that those images will not be shown to expectant parents. Particular Implementation

[0092] The following example describes the training process of a machine learning model and performance evaluation. In particular, a dataset of 25,000 still ultrasound image frames of fetuses were collected on a variety of GE Voluson HD and Samsung premium HD Ultrasound imaging devices and used as training data to train a machine learning model. Patients were recruited from a network of obstetric practices. All images were de-identified prior to labeling and model creation. An initial dataset was labeled by an ultrasound fellowship trained physicians at Centaur Labs.

[0093] Out of 25,000 ultrasound images, 5,788 images were discarded for having textual noise present in the images. Annotations were determined by classifying each image as one of male, female, or unable to assess (i.e., undetermined). A total of 19,212 images were labeled, in particular, 6301 were labeled as male, 6738 as female, 6173 as unable to assess. Labeling was performed using proprietary technology from Centaur Labs. Labels were gathered from a diverse crowd of users on DiagnosUs, an app on iOS devices in which users are incentivized to annotate medical images at scale.

[0094] Out of 19,212 labeled images, 13,440 images were used as training dataset to train a computer vision model. The other 5,764 labeled images were used as validation dataset to evaluate the performance of the model. The computer vision model was trained using a transfer learning approach with EfficientNetB4 architecture as base. EfficientNetB4 was a convolutional neural network, which is described in “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” Mingxing Tan, and Quoc V. Le, accessible at https: / / doi.org / 10.48550 / arXiv. 1905. 11946, and incorporated herein by reference.

[0095] Accuracy, Cohen’s Kappa, and multiclass Receiver Operating Characteristic (ROC) Area Under the Curve (AUC) were used to evaluate the performance of the model. The model achieved an accuracy of 88.27% on the validation set and a Quadratic Cohen’s Kappa score of 0.843 for identifying fetal gender. The model was able to determine whether a fetus on an ultrasound still clip was male, female or unable to assess. The multiclass ROC AUC scores for identifying the fetus as male, female or unable to assess were calculated to be 0.89, 0.897, and 0.916, respectively. The accuracy, Cohen’s Kappa and multiclass ROC-AUC scores demonstrated the trained machine learning model showed viability for confidently predicting fetal gender using ultrasound images as input.Computer Systems

[0096] The present disclosure provides computer systems that are programmed to implement methods of the disclosure. FIG. 7 shows a computer system 701 that is programmed or otherwise configured to predict fetal gender using sonographic data, in accordance with some embodiments. The computer system 701 can regulate various aspects of the present disclosure, for example, implementing machine learning algorithms. The computer system 701 can be an electronic device of a user or a computer system that is remotely located with respect to the electronic device.

[0097] The computer system 701 includes a central processing unit (CPU, also “processor” and “computer processor” herein) 705, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 701 also includes memory or memory location 710 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 715 (e.g., hard disk), communication interface 720 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 725, such as cache, other memory, data storage and / or electronic display adapters. The memory 710, storage unit 715, interface 720 and peripheral devices 725 are in communication with the CPU 705 through a communication bus (solid lines), such as a motherboard. The storage unit 715 can be a data storage unit (or data repository) for storing data. The computer system 701 can be operatively coupled to a computer network (“network”) 730 with the aid of the communication interface 720. The network 730 can be the Internet, an intranet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 730 in some cases is a telecommunication and / or data network. The network 730 can include one or more computer servers, which can enable distributed computing, such as cloud computing. The network 730, in some cases withthe aid of the computer system 701, can implement a peer-to-peer network, which may enable devices coupled to the computer system 701 to behave as a client or a server.

[0098] The CPU 705 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 710. The instructions can be directed to the CPU 705, which can subsequently program or otherwise configure the CPU 705 to implement methods of the present disclosure. Examples of operations performed by the CPU 705 can include fetch, decode, execute, and writeback.

[0099] The CPU 705 can be part of a circuit, such as an integrated circuit. One or more other components of the system 701 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0100] The storage unit 715 can store files, such as drivers, libraries and saved programs. The storage unit 715 can store user data, e.g., user preferences and user programs. The computer system 701 in some cases can include one or more additional data storage units that are external to the computer system 701, such as located on a remote server that is in communication with the computer system 701 through an intranet or the Internet.

[0101] The computer system 701 can communicate with one or more remote computer systems through the network 730. For instance, the computer system 701 can communicate with a remote computer system of a user (e.g., a mobile device). Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC’s (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 701 via the network 730.

[0102] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 701, such as, for example, on the memory 710 or electronic storage unit 715. The machine executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor 705. In some cases, the code can be retrieved from the storage unit 715 and stored on the memory 710 for ready access by the processor 705. In some situations, the electronic storage unit 715 can be precluded, and machine-executable instructions are stored on memory 710.

[0103] The code can be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or can be compiled during runtime. The code can besupplied in a programming language that can be selected to enable the code to execute in a pre-compiled or as-compiled fashion.

[0104] Aspects of the systems and methods provided herein, such as the computer system 701, can be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.

[0105] Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, harddisk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0106] The computer system 701 can include or be in communication with an electronic display 735 that comprises a user interface (UI) 740 for providing, for example, a dashboard. Examples of UI include, without limitation, a graphical user interface (GUI) and web-based user interface.

[0107] Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit 705.

[0108] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

CLAIMSWHAT IS CLAIMED IS:

1. A system for identifying a gender of a fetus, the system comprising: a computer processing device comprising a processor and a non-volatile storage medium with instructions for the processor to: access sonographic data that depicts a genital region of a fetus; and generate, using a trained machine learning model, as output, a probability that the fetus is male, female, or undetermined.

2. The system according to claim 1, wherein the sonographic data comprises one or more sonographic images of the fetus.

3. The system according to claim 2, wherein the one or more sonographic images comprise one or more of two-dimensional (2D) or three-dimensional (3D) images.

4. The system according to any one of claims 1 -3 , wherein the sonographic data comprises sonographic images directly transmitted from an ultrasound imaging device.

5. The system according to any one of claims 1 -4, wherein the sonographic data comprises one or more sonographic videos of the fetus.

6. The system according to any one of claims 1-5, wherein the probability is generated in real time when a patient is under examination by the ultrasound imaging device.

7. The system according to claim 6, wherein the processor is configured to display the probability with the sonographic data on the ultrasound imaging device in real time.

8. The system according to claim 7, wherein the processor is configured to mask the probability displayed on the ultrasound imaging device.

9. The system according to any one of claims 1-8, wherein the sonographic data comprises a first sonographic image that depicts the genital region of the fetus and a second sonographic image that depicts non-genital regions of the fetus, the processor is configured to filter out the first sonographic image.

10. The system according to any one of claims 1 -9, wherein the sonographic data comprises sonographic images transmitted from an image archiving system.

11. The system according to any one of claims 1-10, wherein the trained machine learning model is trained on training data associated with a fetus, the training data comprising a first training dataset with a ground truth label as male, and a second training dataset with a ground truth label as female.

12. The system according to claim 11, wherein the training data further comprises a third training dataset with a ground truth label as undetermined.

13. The system according to claim 11 or 12, wherein the training data comprises sonographic images of the fetus.

14. The system according to claim 13, wherein the sonographic images comprise one or more of two-dimensional (2D) or three-dimensional (3D) images.

15. The system according to any one of claims 11-14, wherein the fetus is at a gestational age of between 8 and 41 weeks.

16. The system according to any one of claims 12-15, wherein each of the first, second, and third training datasets comprises at least 5,000 sonographic images.

17. The system according to any one of claims 1-16, wherein the trained machine learning model comprises a convolutional neural network.

18. The system according to claim 17, wherein the convolutional neural network comprises one or more of convolution layers, pooling layers, or fully connected layers.

19. A method of identifying a gender of a fetus, the method comprising: accessing sonographic data that depicts a genital region of a fetus; and generating, using a trained machine learning model, as output, a probability that the fetus is male, female, or undetermined.

20. The method according to claim 19, wherein the sonographic data is real-time sonographic data generated from an ultrasound imaging device.

21. The method according to claim 19 or 20, wherein the sonographic data comprises one or more sonographic images of the fetus.

22. The method according to claim 21, wherein the one or more sonographic images comprise one or more of two-dimensional (2D) or three-dimensional (3D) images.

23. The method according to any one of claims 20-22, wherein the sonographic data comprises one or more sonographic videos of the fetus.

24. The method according to any one of claims 20-23, wherein the probability is generated in real time when a patient is under examination by the ultrasound imaging device.

25. The method according to claim 24, wherein the processor is configured to display the probability with the sonographic data on the ultrasound imaging device in real time.

26. The method according to claim 25, wherein the processor is configured to mask the probability displayed on the ultrasound imaging device.

27. The method according to any one of claims 19-26, wherein the sonographic data comprises a first sonographic image that depicts the genital region of the fetus and a second sonographic image that depicts non-genital regions of the fetus, the processor is configured to filter out the first sonographic image.

28. The method according to any one of claims 19-27, further comprising: transmitting the sonographic data to a server or a computing device implemented with the trained machine learning model.

29. A method of identifying a gender of a fetus, the method comprising: accessing sonographic data that depicts a genital region of a fetus generated from an ultrasound imaging device in real time; transmitting the sonographic data to a server or a computing device implemented with a trained machine learning model; receiving, from the trained machine learning model, a prediction that the fetus is male, female, or undetermined based on the sonographic data; and displaying the sonographic data and the prediction on a display coupled to the ultrasound imaging device.

30. The method according to claim 29, wherein the sonographic data comprises one or more sonographic images of the fetus.

31. The method according to claim 30, wherein the one or more sonographic images comprise one or more of two-dimensional (2D) or three-dimensional (3D) images.

32. The method according to any one of claims 29-31, wherein the sonographic data comprises one or more sonographic videos of the fetus.

33. The method according to any one of claims 29-31, wherein the sonographic data and the prediction are displayed one or more of concurrently or in real-time with the generation of the sonographic data.

34. A non-transitory computer readable medium comprising machine executable codes that, upon execution by one or more computer processors, implement a method of identifying a gender of a fetus, the method comprising: accessing sonographic data that depicts a genital region of a fetus; and generating, using a trained machine learning model, as output, a probability that the fetus is male, female, or undetermined.

35. A computer-implemented method of constructing a machine learning model for identifying a gender of a fetus, the method comprising: accessing training data comprising a first training dataset with a ground truth label as male and a second training dataset with a ground truth label as female, wherein at least part of the training data depicts a genital region of a fetus; and training the machine learning model that runs on one or more processors coupled to memory on the training data, by progressively matching output of the machine learning model with corresponding ground truth labels.

36. The computer-implemented method according to claim 34, wherein the training data depicts the fetus at a gestational age of between 8 and 41 weeks.

37. The computer-implemented method according to claim 34 or 35, wherein the training data further comprises a third training dataset with a ground truth label as undetermined.

38. The computer-implemented method according to claim 36, wherein each of the first, second, and third training datasets comprises at least 5,000 sonographic images.

39. The computer-implemented method according to any one of claims 34-37, wherein the training data comprises sonographic images of the fetus.

40. The computer-implemented method according to claim 38, wherein the sonographic images comprise one or more of two-dimensional (2D) or three-dimensional (3D) images.

41. The computer-implemented method according to any one of claims 34-39, wherein the machine learning model is trained using a backpropagation-based gradient update technique.

42. The computer-implemented method according to any one of claims 34-40, wherein the machine learning model comprises a convolutional neural network.

43. The computer-implemented method according to claim 41, wherein the convolutional neural network comprises one or more convolution layers, pooling layers, or fully connected layers.

44. A system comprising a computer processing device comprising a processor and a nonvolatile storage medium with instructions to construct a machine learning model for identifying a gender of a fetus, the instructions, when executed on the processor, implement actions comprising: accessing training data comprising a first training dataset with a ground truth label as male and a second training dataset with a ground truth label as female, wherein at least part of the training data depicts a genital region of a fetus; andtraining the machine learning model on the training data by progressively matching output of the machine learning model with corresponding ground truth labels.

45. The system according to claim 44, wherein the training data depicts the fetus at a gestational age of between 8 and 41 weeks.

46. The system according to claim 44 or 45, wherein the training data further comprises a third training dataset with a ground truth label as undetermined.

47. The system according to claim 46, wherein each of the first, second, and third training datasets comprises at least 5,000 sonographic images.

48. The system according to any one of claims 44-47, wherein the training data comprises sonographic images of the fetus.

49. The system according to claim 48, wherein the sonographic images comprise one or more of two-dimensional (2D) or three-dimensional (3D) images.

50. The system according to any one of claims 44-49, wherein the machine learning model is trained using a backpropagation-based gradient update technique.

51. The system according to any one of claims 44-50, wherein the machine learning model comprises a convolutional neural network.

52. The system according to claim 51, wherein the convolutional neural network comprises one or more convolution layers, pooling layers, or fully connected layers.

53. A non-transitory computer readable medium comprising machine executable codes that, upon execution by one or more computer processors, implement a method of constructing a machine learning model for identifying a gender of a fetus, the method comprising: accessing training data comprising a first training dataset with a ground truth label as male and a second training dataset with a ground truth label as female, wherein at least part of the training data depicts a genital region of a fetus; and training the machine learning model on the training data, by progressively matching output of the machine learning model with corresponding ground truth labels.

54. A method of generating images of a fetus, the method comprising: accessing sonographic data that depicts a genital region of a fetus generated from an ultrasound imaging device, the sonographic data comprising a plurality of ultrasound images; transmitting the sonographic data to a server or a computing device implemented with a trained machine learning model;determining, with the trained machine learning model, whether a sex of the fetus can be determined from each ultrasound image of the plurality of ultrasound images; and generating a set of the ultrasound images in which the sex of the fetus cannot be determined from each ultrasound image of the set.

55. The method of claim 54, wherein generating the set of the ultrasound images in which the sex of the fetus cannot be determined from each ultrasound image of the set comprises masking one or more features in an ultrasound image in which the sex of the fetus can be determined.

56. The method of claim 54 or 55, further comprising displaying the set of the ultrasound images in which the sex of the fetus cannot be determined from each ultrasound image of the set.

Citation Information

Patent Citations

  • A system for gender analysis based on ultrasound imaging and method thereof

    IN202041024069A

  • Novel Algorithms for Feature Detection and Hiding from Ultrasound Images

    US20150342560A1