Detecting and classifying fetal movements using machine learning estimation of audio recordings
The system uses machine learning to analyze audio recordings and estimate fetal movements, addressing the limitations of current fetal monitoring methods by providing a more accurate and accessible assessment of fetal health.
Patent Information
- Application Number
- PCT/US2024/056414
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-16
- Filing Date
- 2024-11-18
- Publication Date
- 2025-05-22
AI Technical Summary
Current methods for fetal monitoring are invasive, costly, and do not provide a granular assessment of fetal movement over time, leading to a high rate of stillbirths due to unknown causes.
A system and method using a trained machine learning model to estimate fetal movements from audio recordings acquired from a patient's device, employing engineered features from time, frequency, and time-frequency analyses, and deep neural networks for classification.
Enables at-home fetal movement monitoring, reducing the need for specialized equipment, and providing a more accurate and granular assessment of fetal movement, potentially reducing the rate of stillbirths.
Smart Images

Figure IMGF000010_0001 
Figure IMGF000011_0001 
Figure IMGF000019_0001
Abstract
Description
DETECTING AND CLASSIFYING FETAL MOVEMENTS USING MACHINELEARNING ESTIMATION OF AUDIO RECORDINGSRELATED APPLICATION
[0001] This application claims priority to, and the benefit of, U.S. Provisional Patent Application No. 63 / 600,486, filed November 17, 2023, entitled “Detecting and classifying fetal movements using machine learning estimation of audio recordings,” and U.S. Provisional Patent Application No. 63 / 634,694, filed April 16, 2024, entitled “Detecting and classifying fetal movements using machine learning estimation of audio recordings, each of which is incorporated by reference herein in its entirety.BACKGROUND
[0002] Fetal death after 28 weeks gestation is often termed late stillbirth. Using this definition, the World Health Organization has determined that more than 2.6 million stillbirths occur worldwide on an annual basis; most of these occur in low and middle-income countries. In the United States in 2019, the rate of late stillbirth was reported to be 2.73 / 1000 births. The Stillbirth Collaborative Research Network Consortium studied 663 stillbirths in the U.S. and found that more than one-quarter of late stillbirths were related to unknown causes.
[0003] Known risk factors for stillbirth include maternal diseases such as insulindependent diabetes, chronic hypertension, pre-eclampsia, and autoimmune disorders. Additional factors include demographic factors such as Black and Native American / Pacific Islander race / ethnicity, maternal obesity, infertility treatments, a short interval of (< 12 months) between pregnancies, a prolonged interval (> 24 months) between pregnancies, and previous Cesarean section delivery. A history of previous unexplained stillbirths results in a 3-fold increased risk of recurrent stillbirth [4], After controlling for various risk factors, a history of a previous liveborn baby that subsequently died was associated with a 10-fold increased risk for stillbirth.
[0004] The most recent guidelines from the American College of Obstetricians and Gynecologists (ACOG) and the Society for Maternal -Fetal Medicine (SMFM) propose that antenatal surveillance be initiated once or twice weekly at 32 weeks gestation or starting 1-2 weeks prior to the gestational age when the last stillbirth occurred. Such testing involves the use of ultrasound (biophysical profile), fetal monitoring (non-stress test), or a combination of both (modified biophysical profile) performed once or twice weekly in an outpatient setting. These visits involve a minimum of one hour of the patient’s presence in a clinic setting.
[0005] Additionally, ACOG / SMFM have published guidelines for antenatal surveillance for more than 25 other fetal or maternal conditions associated with a relative risk for stillbirth of 2.0 or greater based on retrospective studies. The recommendations were not based on evidence from randomized clinical trials or the subject of a cost-analysis evaluation. In fact, in one study of almost 2000 pregnancies where such antenatal testing was employed for high-risk conditions, the stillbirth rate remained unchanged at 7.7 per 1000 births.
[0006] There is a benefit for improved fetal monitoring.SUMMARY
[0007] An exemplary system and method are provided for fetal monitoring using a trained Al model configured to provide an estimation of movements from an audio recording acquired from the patient’s user device. The exemplary system and method can be used to provide at- home monitoring of fetal movement to prompt the mother to seek medical attention without the need for specialized ultrasound equipment or fetal heart rate monitoring.
[0008] The exemplary system and method, in some embodiments, employs engineered ML features based on time-associated analysis, frequency-associated analysis, and time-frequency- associated analysis. The time-associated analysis, in some embodiments, includes the amplitude envelope of measured audio recording as a measure of loudness, e.g., to determine onset detection. In some embodiments, the time-associated analysis includes root-mean-square energy to measure loudness less sensitive to outliers or similarity as a measure of how similar one sound signal is to another. The machine learning can learn the time-associated analysis, frequency-associated analysis, and time-frequency-associated analysis of audio recordings in connection with an ultrasound quantified measure of fetal movement as a ground truth for the training.
[0009] The frequency-associated analysis, in some embodiments, includes band energy ratio, spectral centroid, and / or spectral flux. The time-frequency associated analysis, in some embodiments, includes spectrogram, Mel-Spectrogram, and / or constant-Q transform.
[0010] The exemplary system and method can employ a deep neural network classifier that can read the audio recording to generate an estimate of the presence or non-presence of fetal movement. The Al deep neural network classifier may be trained using an audio recording as input and an ultrasound quantified measure of fetal movement as a ground truth for the training.
[0011] In an aspect, provided is a method comprising: receiving, by a processor (e.g., at a cloud infrastructure or at an edge device), an acoustic file of a fetal sound recording; determining, by the processor, via a trained machine learning classifier, presence of each fetalmovement from the fetal sound recording; and determining, by the processor, a number of estimated presence of each fetal movement. The determined number of estimated presence of each fetal movement can be outputted via a graphical user interface or report to a user.
[0012] In some aspects, the trained machine learning classifier employs ultrasound data as ground truth in a training operation performed in conjunction with the fetal sound recording and recorded button press events registering the mother’s perception of movement.
[0013] In some aspects, the trained machine learning classifier employed labeled (e.g., manually or algorithmically labeled) for fetal movement derived from acoustic or ultrasound training data to determine the quantity of fetal movements (e.g., a count of fetal movements within a timeframe.
[0014] In some aspects, the method can further include determining, by the processor, via a second trained machine learning classifier, a quality score of each fetal movement and detection of protective fetal movements (e.g., breathing and hiccups) from the fetal sound recording. The determined quality score can be outputted via the graphical user interface or report to the user.
[0015] In some aspects, the method further includes transmitting the determined number of estimated presence of each fetal movement and / or the determined quality score to a predefined clinician.
[0016] In some aspects, the method further includes transmitting the determined number of estimated presence of each fetal movement and / or the determined quality score to a predefined clinician based on a trigger event associated with the determined number of estimated presence of each fetal movement and / or the determined quality score.
[0017] In some aspects, the processing is performed at an edge device.
[0018] In some aspects, the processing is performed at a cloud infrastructure.
[0019] In some aspects, the trained machine learning classifier is trained by simultaneously collecting a reference acoustic file and reference ultrasound data. The trained machine learning classifier can use the reference ultrasound data as a ground truth.
[0020] In some aspects, the method further includes outputting, via the graphical user interface or report, the determined number of estimated presence of each fetal movement to the user.
[0021] In another aspect, provided is a system (e.g., analysis system) comprising: a processor; and a memory having instruction stored thereon. Execution of the instructions by the processor causes the processor to: receive, an acoustic file of a fetal sound recording; determine, via a trained machine learning classifier and / or deep learning neural network, thepresence of each fetal movement from the fetal sound recording; and determine a number of estimated presence of each fetal movement. The determined number of estimated presence of each fetal movement can be outputted via a graphical user interface or report to a user.
[0022] In some aspects, the execution of the instructions by the processor causes the processor to perform any one of the disclosed methods.
[0023] In another aspect, provided is a mobile device comprising: an acoustic sensor; a network interface; a processor; and a memory having instruction stored thereon, wherein execution of the instructions by the processor causes the processor to: generate an acoustic file of a fetal sound recording from the acoustic sensor; transmit the acoustic file to a cloud infrastructure configured to: (i) receive the acoustic file of the fetal sound recording, (ii) determine, via a trained machine learning classifier and / or deep learning neural network, presence of each fetal movement from the fetal sound recording; and (iii) determine a number of estimated presence of each fetal movement; receive the determined number of estimated presence of each fetal movement; output, via the graphical user report of the device, the determined number of estimated presence of each fetal movement.
[0024] In another aspect, provided is a mobile device comprising: an acoustic sensor; a network interface; a processor; and a memory having instruction stored thereon, wherein execution of the instructions by the processor causes the processor to: generate an acoustic file of a fetal sound recording from the acoustic sensor; determine, via a trained machine learning classifier and / or deep learning neural network, presence of each fetal movement and the quality of fetal movements from the fetal sound recording; and determine a number of estimated presence of each fetal movement; output, via a graphical user report of the device, the determined number of estimated presence of each fetal movement.
[0025] In some aspects, the device is configured to perform any one of the disclosed methods.
[0026] In another aspect, provided is a harness comprising: an adjustable band releasably encircling an abdomen of a user to retain a recording device on the abdomen of a user; and a recording device holder (e.g., pocket or assembly) coupled to the adjustable band; wherein the recording device holder is configured to position a recording device adj acent to and sufficiently perpendicular to a surface of the abdomen of the user; wherein the recording device is configured to acquire a sound recording (e.g., fetal sound recording) from the abdomen of the user to determine fetal movement according to any one of the disclosed methods.
[0027] Other systems, methods, features and / or advantages will be or may become apparent to one with skill in the art upon examination of the following drawings and detaileddescription. It is intended that all such additional systems, methods, features, and / or advantages be included within this description and be protected by the accompanying claims.BRIEF DESCRIPTION OF DRAWINGS
[0028] Figs. 1 A, IB, and 1C each show an example system that employs, during its runtime operation, a trained Al model configured to provide an estimation of movements from an audio recording acquired from a patient’s user device having an acoustic sensor.
[0029] Fig. 2 shows a training system configured to receive the training data set and perform machine learning analysis for a set of engineered features.
[0030] Figs. 3 A, 3B, and 3C show examples of Time Domain Features, Frequency Domain Features, and Time-Frequency Domain Features that can be employed in the training system of Fig. 2.
[0031] Figs. 4A - 4C shows methods that may be performed for the systems described herein, including those described in relation to Fig. 1 A and IB.
[0032] FIGS. 5A-5F depict preliminary results for analysis of a single patient at approximately 38 weeks across two different visits. Specifically, FIG. 5 A shows a full normalized audio recording taken at a first visit, and FIG. 5B shows the full normalized audio recording of FIG. 5A with FMs overlaid and colored by the duration of movement. FIGS. 5C- 5D show the correlation of perceived movement with the full normalized audio recording. FIGS. 5E-5K show additional correlations. FIG. 5L shows a full normalized audio recording taken at a second visit, and FIG. 5M shows the full normalized audio recording of FIG. 5L with FMs overlaid and colored by the duration of movement. FIGS. 5N-5S show additional correlations.
[0033] FIG. 6A shows annotated ultrasounds from multiple patients.
[0034] FIGS. 6B-6C show the impact of vertical vs. horizontal phone placement on the audio recordings.
[0035] FIG. 6D shows the impact of vertical vs. horizontal phone placement on mel spectrograms.
[0036] FIG. 6E shows the impact of vertical vs. horizontal phone placement on MFCC.
[0037] FIG. 7 shows an example position of a smartphone sufficiently perpendicular to a surface of the abdomen of the user for an audio recording.DETAILED DESCRIPTION
[0038] It is appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate aspects, can also be provided in combination with a single aspect. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single aspect, can also be provided separately or in any suitable subcombination. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure.
[0039] Definitions
[0040] In this specification and in the claims that follow, reference will be made to a number of terms, which shall be defined to have the following meanings:
[0041] Throughout the description and claims of this specification, the word “comprise” and other forms of the word, such as “comprising” and “comprises,” means including but not limited to, and are not intended to exclude, for example, other additives, segments, integers, or steps. Furthermore, it is to be understood that the terms comprise, comprising, and comprises as they relate to various aspects, elements, and features of the disclosed invention also include the more limited aspects of “consisting essentially of’ and “consisting of.”
[0042] As used herein, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to an “electrode” includes aspects having two or more such electrodes unless the context clearly indicates otherwise.
[0043] Ranges can be expressed herein as from “about” one particular value and / or to “about” another particular value. When such a range is expressed, another aspect includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms another aspect. It should be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint.
[0044] As used herein, the terms “optional” or “optionally” mean that the subsequently described event or circumstance may or may not occur and that the description includes instances where said event or circumstance occurs and instances where it does not.
[0045] For the terms “for example” and “such as” and grammatical equivalences thereof, the phrase “and without limitation” is understood to follow unless explicitly stated otherwise.
[0046] Example Systems
[0047] Example System #1. Fig. 1A shows an example system 100 (shown as 100a) that employs, during its runtime operation 101, a trained Al model 102 configured to provide an estimation of movements 104 (shown as “Predicted Quantity / Quality of Fetal Movements” 104) from an audio recording 106 acquired from a patient’s user device having an acoustic sensor 108. In the example shown in Fig. 1A, the trained Al model 102 was configured in a training operation 110 by Al training system 112 that receives as its inputs, audio records 114, and Quantity / Quality of Fetal Movements 116, as ground truth, determined from ultrasound recording (shown acquired via ultrasound probe 118) concurrently acquired with the acoustic recording 114.
[0048] The data collected from a study described herein is the first to time-register annotated ultrasound video, patient-perceived movement, and smartphone audio recordings, thus providing ground-truth data for fetal movement.
[0049] In some embodiments, the exemplary system 100a can detect, classify, and characterize fetal movement from audio recordings. The system may be a part of software (e.g., mobile device APP) that can be utilized by expectant mothers to record fetal movement 107 using their smartphone. The app can upload the recordings to an Al-based deep learning system (e.g., in the cloud or server) that is configured to provide: 1) movement information and, if necessary, clinical recommendations to the patient in an accessible and user-friendly way and 2) clinical / diagnostic information to their physicians.
[0050] The smartphone can be positioned and secured in place by the patient at the midpoint of the pregnant uterus. In one embodiment, the phone can be oriented vertically (e.g., perpendicularly pointing or angled into the abdomen) so as to allow the microphone to be in direct contact with the patient’s skin. The patient then records an audible file while in a quiet room for at least 15 minutes. A hands-free harness can be employed to secure the phone in such orientations and others (e.g., parallel orientation).
[0051] The audio file, in some embodiments, may be uploaded via the app to a HIPAA- compliant data repository (e.g., server or cloud infrastructure) having an analysis engine configured to execute an artificial intelligence / deep learning operation to determine / quantify the number and / or quality of fetal movements.
[0052] The system can provide the output to the patient via a graphic display or report, e.g., of daily fetal movements over selected periods of time. The system can additionally direct clinical reporting to clinicians and / or feedback to the user through the app. In the case of decreased fetal movements, the need for immediate repeat monitoring for a prolonged periodof time or recommendations to seek medical attention are examples of messages that would be sent to patients.
[0053] The system could perform pre-processing of the acoustic recording and provide instructions to the user, e.g., via the graphic display, of proper placement of the phone to the user to ensure optimal acoustic recording for the phone, for example, instructing the patient for a change in phone placement due to poor signal acquisition. In some embodiments, prior to providing the acoustic recording to the trained Al system, the system could assess the recording to determine whether the recording of the acoustic signal achieved a sufficient amplitude or signal -to-noise ratio to instruct the user to move the phone’s microphone perpendicularly pointing or angled into the abdomen. In some embodiments, the system could trigger a video to be displayed to the user, the video showing a sequence for optimal acoustic recording for a pre-defined device (e.g., Apple device or Android device perpendicularly pointing or angled into the abdomen).
[0054] As a part of the pre-processing operation prior to providing the acoustic recording to the trained Al system, the system could also assess the ambient noise in the recording and generate a notification to the user, e.g., to change the location where the patient is recording due to high ambient noise in the present environment.
[0055] The system could additionally evaluate a given recording and notify the patient that additional recording time is desired to improve accuracy. In such embodiments, the system may generate an output from the trained Al system, the output including a likelihood score for the presence or non-presence of fetal movements. The output can be assessed by the system to generate a notification of additional recording time if the score is below a predefined threshold score.
[0056] The system could determine the presence of decreased fetal movement and provide a notification to the user’s clinician. In some embodiments, the system can be subscribed by the clinician to which the clinician can enroll the user as part of a prenatal care visit. The subscription allows the clinician's contact information and healthcare portal to be provided to the system to which the notification can be directed. The determination can be performed locally on the user’s phone or in a cloud infrastructure.
[0057] The system 100a allows the patient to use their personal cellular phone to record audio of fetal movement. The only other technology for fetal movement detection involves the application of an apron of sensors that is applied to the maternal abdomen.
[0058] The exemplary system and method provide evidence-based artificial intelligence methods to classify and extract fetal movement from audio recordings are the first approach touse time-registered annotated ultrasound video (ground truth fetal movement), patient- perceived movement, and smartphone audio recordings along with patient characteristics (e.g., abdominal wall thickness, anterior placental location, BMI, etc.), thus enabling characterization of movement type, persistence, and quality.
[0059] In some embodiments, the exemplary system and method employ deep learning algorithms to generate the clinical / diagnostic estimation of fetal movement. The algorithm may learn and refine estimations / predictions of fetal movement over time as more data for the user is uploaded.
[0060] In some embodiments, the smartphone app can be used by expectant mothers to 1) record audio of fetal movement, 2) upload data to the deep learning system, 3) provide accessible and user-friendly information for the expectant mother regarding fetal movement, and 4) provide any clinical / diagnostic information back to the expectant mother if necessary.
[0061] Training system. The training system 112 is configured to receive the training data set (e.g., 114, 116) and perform machine learning analysis for a set of engineered features (shown in Fig. 2) that involves Time Domain Feature analysis 120, Frequency Domain Feature analysis 122, and Time-Frequency Domain Feature analysis 124, to generate the trained Al model 102 (shown as machine-learned predictor / classifier 102’). Figs. 3A, 3B, and 3C show examples of Time Domain Features, Frequency Domain Features, and Time-Frequency Domain Features.Table 1
[0062] Time domain features. Time domain features can be determined as a peak amplitude value in a segment of the acoustic signal. Root mean square energy can be determined for a segment of acoustic signa.
[0063] Similarity. To calculate the similarity of one audio segment to another, the feature may calculate the cross-correlation as a measure of how well one signal aligns with another, e.g., by sliding one signal across the other and calculating the dot product at each position. The features may calculate similarity using dynamic time warping applied to the MFCCs.
[0064] Spectral centroid. Spectral centroid may be calculated as the mean or center of mass of a given spectrum over time. To calculate the spectral centroid, the feature may segment the audio signal into frames, e.g., using a pre-defined windowing (e.g., 2048), compute the fast Fourier transform of each of the windowed frames to provide the frequency spectrum, compute the magnitude of the fast Fourier transform to determine the amplitude at each frequency bin, compute the weighted average (e.g., by multiplying the magnitude times the frequency and normalizing by the total spectral energy), and repeat the operation for all the windowed frames (e.g., with a hop length of 512, among others). The center of mass or centroid over time and the shape of that line over time can provide information about what that audio segment’s shape looks like and whether the shape is reflected in the distribution of that line over time. The location of the center of mass is located in the frequency spectrum over time can provide information of whether the signal is low spectral or high spectral. In some embodiments, the spectral centroids can be compared to different audio signals to determine degree of similarity or dissimilarity to those signals.
[0065] Spectral flux. Spectral flux reflects a measure of change in the power spectrum over time. In some embodiments, spectral flux can be determined by dividing the audio signal into windowed frames (e.g., window size 2048 with hop length of 512), computing the magnitude spectrum of each frame (e.g., by taking the magnitude of the fast Fourier transform of the window and) and then computing the difference of the value at time (t), frequency bin (k) withthe value at time (t-1) and frequency bin (k). The feature can then sum the values across all frequency bins to determine the spectral flux for the current frame. The spectral flux can provide information about sudden differences in the audio signal and is typically a means for detecting the onset of a change in signal over time.
[0066] Constant-Q transform. The Constant-Q Transform transforms the audio into the frequency domain with a log frequency scale and divides them into bins. By looking at the magnitude of the output of the Constant-Q Transform, the feature can calculate the locations with the highest frequency values. Additionally, the feature can locate where there are shifts in those frequencies by looking at differences in frequency over time.
[0067] Band energy ratio. Band energy ratio is a measure of the proportion of energy in a frequency band relative to the total energy in the audio. Energy may be calculated as the magnitude squared for a given frequency bin. In some embodiments, the energy may be determined similarly to what is stated above by windowing the audio with overlap, calculating the fast Fourier transform, and calculating the values for all frequency bins for all windows over the entire audio segment.
[0068] Power spectral density. Power spectral density is a measure of the degree of the power of the audio signal being distributed across the frequency. The power may be calculated by taking the fast Fourier transform of the audio, squaring the magnitude at each (time, frequency) location and dividing each of those by the duration of the audio signal (T).
[0069] Example System #2. Fig. IB shows another example system 100 (shown as 100b) that employs, during its runtime operation 101, a trained Al model 102 (shown as 102’) configured to provide an estimation of movements 104 (shown as “Predicted Quantity / Quality of Fetal Movements” 104) from an audio recording 106 acquired from patient’s user device having an acoustic sensor 108. In the example shown in Fig. IB, the audio recording 114 is manually reviewed (128) to provide the Quantity / Quality of Fetal Movements 116 (shown as 116’) as ground truth to the Al training system 112 (shown as 112’). The Al training system 112’ also employs the audio recordings 114 as the input for the training.
[0070] Example System #3. Fig. 1C shows another system 100 (shown as 100c) that employs, during its runtime operation 101, a trained Al model 102 (shown as 102”) configured as a neural network to provide an estimate of movements 104 from an audio recording 106. Specifically, Fig. 1C shows a life cycle of the training and classification of fetal movement
[0071] Data Collection. The data collection stage may include the collection of ultrasound video co-registered with an audio recording of fetal movement, point in time recordings of patient’s perceived movements, and patient’s metadata. The audio is a recording fetalmovement, essentially recording displacement of amniotic fluid (rather than audio for music, human speech, environmental sounds, or the like).
[0072] Preprocessing Preprocessing may be performed to generate a database of uniformly sized and typed audio snippets with corresponding patient metadata. The first step (1) may include the creation of annotated files (e.g., three files) that denote a point in time an annotation (e.g., “clicks” on an annotation device) of general movement, breathing, and hiccups, respectively. The annotated files may be recorded and / or viewed by professional sonographers and replaying a video recording of the ultrasound video and registering the time when movements are observed. General movements may be recorded as consecutive “button clicks” on the annotation device if the movement is continuous. Breathing and hiccups may be registered as single button clicks each time they are observed. The annotated files may be stored as flat text files. Movements may be classified by their length and extracted for each patient. Consecutive movements may occur when the previous movement occurred less than or equal to 2 seconds prior. Extracted movements for each patient may include a series of registrations (e.g., begin time, end time, type = general, breathing, hiccups). The step may be updated as each time new data is collected and the ultrasounds have been fully annotated. Data for the entire cohort may be stored in a json file and combined with a cohort’s collection of audio recordings to extract and classify the corresponding audio snippets. A window length (e.g., 2048 padded with zeros) and a hop length (e.g., of 512) may be used to window the data such that there are overlapping segments. An ensemble of audio snippet length may be created to train the model, e.g., ranging from half a second to two seconds in length. Perturbations on the snippets may be added to the model that include noise reduction and amplification. Mel spectrograms and mel frequency cepstral coefficients may be computed on the audio snippets as 2D images and stored in the database. The images may also be taken through a series of perturbations and added to the database to include random adjustments for brightness and contrast. Multiple copies of the same images may be randomly added to the database.
[0073] Mel-frequency cepstrum (MFC), as employed by the neural network of Fig. 1C, or the features of Fig. 1 A and IB, is a representation of the short-term power spectrum of a sound based on a linear cosine transform of a log power spectrum on a nonlinear mel scale of frequency. The Mel-frequency cepstral coefficients (MFCCs) may be calculated by taking the Fourier transform of a windowed signal, mapping the powers of the spectrum onto the mel scale as a log of the powers at each of the mel frequencies, performing a discrete cosine transform of the list of mel log powers, and extracting the coefficients as the amplitudes of the resulting spectrum.
[0074] Both the mel spectrogram and the MFCCs may be treated as 2D images to be suitable for a convolutional neural network. While the MFCCs are sets of numerical coefficients that represent spectral characteristics, they may be represented as 2D images (e.g., bands of feature vectors placed on top of each other vertically). Patient metadata may be represented as a vector [Nxl] where N is the number of metadata values associated with a patient. Non-limiting examples of metadata include gestational age, body mass index (BMI), placental location, abdominal wall thickness, and amniotic fluid index. Other may be used.
[0075] Deep Learning Architecture. The deep learning architecture may concatenate the outputs of (i) the mel spectrogram input to a 2D convolutional neural network, (ii) the mel frequency cepstral coefficients input to a 2D convolutional neural network, and (iii) the patient metadata input, to a two-layer convolutional neural network. The concatenated outputs may be inputted to a two-layer convolutional neural network and classified. Inputs and outputs for the 2D CNN blocks can be implemented as a series of progressive inputs whose outputs feed into the next layer of the CNN model.
[0076] Machine Learning. The term “artificial intelligence” can include any technique that enables one or more computing devices or computing systems (i.e., a machine) to mimic human intelligence. Artificial intelligence (Al) includes but is not limited to knowledge bases, machine learning, representation learning, and deep learning. The term “machine learning” is defined herein to be a subset of Al that enables a machine to acquire knowledge by extracting patterns from raw data. Machine learning techniques include, but are not limited to, logistic regression, support vector machines (SVMs), decision trees, Naive Bayes classifiers, and artificial neural networks. The term “representation learning” is defined herein to be a subset of machine learning that enables a machine to automatically discover representations needed for feature detection, prediction, or classification from raw data. Representation learning techniques include, but are not limited to, autoencoders. The term “deep learning” is defined herein to be a subset of machine learning that enables a machine to automatically discover representations needed for feature detection, prediction, classification, etc., using layers of processing. Deep learning techniques include but are not limited to artificial neural networks or multilayer perceptron (MLP).
[0077] Machine learning models include supervised, semi-supervised, and unsupervised learning models. In a supervised learning model, the model learns a function that maps an input (also known as feature or features) to an output (also known as target or target) during training with a labeled data set (or dataset). In an unsupervised learning model, the model has a pattern in the data. In a semi-supervised model, the model learns a function that maps an input (alsoknown as a feature or features) to an output (also known as a target) during training with both labeled and unlabeled data.
[0078] Neural Networks. An artificial neural network (ANN) is a computing system including a plurality of interconnected neurons (e.g., also referred to as “nodes”). This disclosure contemplates that the nodes can be implemented using a computing device (e.g., a processing unit and memory as described herein). The nodes can be arranged in a plurality of layers, such as an input layer, an output layer, and optionally one or more hidden layers. An ANN having hidden layers can be referred to as a deep neural network or multilayer perceptron (MLP). Each node is connected to one or more other nodes in the ANN. For example, each layer is made of a plurality of nodes, where each node is connected to all nodes in the previous layer. The nodes in a given layer are not interconnected with one another, i.e., the nodes in a given layer function independently of one another. As used herein, nodes in the input layer receive data from outside of the ANN, nodes in the hidden layer(s) modify the data between the input and output layers, and nodes in the output layer provide the results. Each node is configured to receive an input, implement an activation function (e.g., binary step, linear, sigmoid, tanH, or rectified linear unit (ReLU) function), and provide an output in accordance with the activation function. Additionally, each node is associated with a respective weight. ANNs are trained with a dataset to maximize or minimize an objective function. In some implementations, the objective function is a cost function, which is a measure of the ANN’S performance (e.g., an error such as LI or L2 loss) during training, and the training algorithm tunes the node weights and / or bias to minimize the cost function. This disclosure contemplates that any algorithm that finds the maximum or minimum of the objective function can be used for training the ANN. Training algorithms for ANNs include but are not limited to backpropagation. It should be understood that an artificial neural network is provided only as an example machine learning model. This disclosure contemplates that the machine learning model can be any supervised learning model, semi-supervised learning model, or unsupervised learning model. Optionally, the machine learning model is a deep learning model. Machine learning models are known in the art and are therefore not described in further detail herein.
[0079] A convolutional neural network (CNN) is a type of deep neural network that has been applied, for example, to image analysis applications. Unlike traditional neural networks, each layer in a CNN has a plurality of nodes arranged in three dimensions (width, height, depth). CNNs can include different types of layers, e.g., convolutional, pooling, and fully- connected (also referred to herein as “dense”) layers. A convolutional layer includes a set of filters and performs the bulk of the computations. A pooling layer is optionally insertedbetween convolutional layers to reduce the computational power and / or control overfitting (e.g., by downsampling). A fully-connected layer includes neurons, where each neuron is connected to all of the neurons in the previous layer. The layers are stacked similar to traditional neural networks. GCNNs are CNNs that have been adapted to work on structured datasets such as graphs.
[0080] Other Supervised Learning Models. A logistic regression (LR) classifier is a supervised classification model that uses the logistic function to predict the probability of a target, which can be used for classification. LR classifiers are trained with a data set (also referred to herein as a “dataset”) to maximize or minimize an objective function, for example, a measure of the LR classifier’s performance (e.g., an error such as LI or L2 loss), during training. This disclosure contemplates that any algorithm that finds the minimum of the cost function can be used. LR classifiers are known in the art and are therefore not described in further detail herein.
[0081] An Naive Bayes’ (NB) classifier is a supervised classification model that is based on Bayes’ Theorem, which assumes independence among features (i.e., the presence of one feature in a class is unrelated to the presence of any other features). NB classifiers are trained with a data set by computing the conditional probability distribution of each feature given a label and applying Bayes’ Theorem to compute the conditional probability distribution of a label given an observation. NB classifiers are known in the art and are therefore not described in further detail herein.
[0082] A k-NN classifier is a supervised classification model that classifies new data points based on similarity measures (e.g., distance functions). The k-NN classifiers are trained with a data set (also referred to herein as a “dataset”) to maximize or minimize a measure of the k- NN classifier’s performance during training. The k-NN classifiers are known in the art and are therefore not described in further detail herein.
[0083] A majority voting ensemble is a meta-classifier that combines a plurality of machine learning classifiers for classification via majority voting. In other words, the majority voting ensemble’s final prediction (e.g., class label) is the one predicted most frequently by the member classification models. The majority voting ensembles are known in the art and are therefore not described in further detail herein.
[0084] Example Method
[0085] Figs. 4A - 4C shows methods (400a, 400b, 400c) that may be performed for the systems described herein, including those described in relation to Fig. 1A and IB. Method 400a of Fig. 4A includes receiving (402), by a processor, an acoustic file of a fetal soundrecording (e.g., 106). Method 400a then inludes determining (404), by the processor, via a trained machine learning classifier (e.g., 102), presence of each fetal movement from the acoustic file or fetal sound recording. Method 400a then includes determining (406), by the processor, a number of estimated presence of each fetal movement (e.g., 104), wherein the determined number (e.g., 104) of estimated presence of each fetal movement is outputted (408), via a graphical user interface or report, to a user.
[0086] In some embodiments, the trained machine learning classifier (e.g., 102) employs ultrasound data as a ground truth in a training operation performed in conjunction with the fetal sound recording, or the trained machine learning classifier employed ultrasound data labeled for fetal movement derived from acoustic or ultrasound training data.
[0087] In some embodiments, the method 400a includes determining, by the processor, via a second trained machine learning classifier, a quality score of each fetal movement from the fetal sound recording, wherein the determined quality score is outputted, via the graphical user interface or report, to the user.
[0088] In some embodiments, the method 400a includes transmitting the determined number of estimated presence of each fetal movement and / or the determined quality score to a pre-defined clinician.
[0089] In some embodiments, the method 400a includes transmitting the determined number of estimated presence of each fetal movement and / or the determined quality score to a pre-defined clinician based on a trigger event associated with the determined number of estimated presence of each fetal movement and / or the determined quality score.
[0090] In some embodiments, the processing is performed at an edge device.
[0091] In some embodiments, the processing is performed at a cloud infrastructure.
[0092] In some embodiments, the trained machine learning classifier (e.g., 102) is trained by: simultaneously collecting a reference acoustic file and reference ultrasound data, wherein the trained machine learning classifier uses the reference ultrasound data as a ground truth.
[0093] In some embodiments, the method 400a includes outputting, via the graphical user interface or report, the determined number of estimated presence of each fetal movement to the user.
[0094] In some embodiments, the trained machine learning classifier (e.g., 102) is (i) a neural network or (ii) an equation derived from engineered features associated with timedomain analysis, frequency domain analysis, or a time and frequency domain analysis.
[0095] generate an acoustic file of a fetal sound recording from the acoustic sensor;
[0096] transmit the acoustic file to a cloud infrastructure configured to: (i) receive the acoustic file of the fetal sound recording, (ii) determine, via a trained machine learning classifier, presence of each fetal movement from the fetal sound recording; and (iii) determine a number of estimated presence of each fetal movement;
[0097] receive the determined number of estimated presence of each fetal movement;
[0098] output, via the graphical user report of the device, the determined number of estimated presence of each fetal movement.
[0099] Fig. 4B shows another example method 400b of operation as performed on a mobile device having an acoustic sensor, a network interface, and a processor; and a memory having instruction stored thereon, wherein execution of the instructions by the processor. The method 400b includes generating (410) an acoustic file of a fetal sound recording from the acoustic sensor. Method 400b then includes determining (412), via a trained machine learning classifier, presence of each fetal movement from the fetal sound recording. Method 400b then includes determining (414) a number of estimated presence of each fetal movement. Method 400b then outputting (416), via a graphical user report of the device, the determined number of estimated presence of each fetal movement.
[0100] The mobile device may operate in similar matter to the step described in relation to Fig. 4A.
[0101] The mobile phone may operate with a harness configured to optimally position the microphone of the mobile device for acoustic measurement. The harness may include an adjustable band releasably encircling an abdomen of a user to retain a recording device on an abdomen of a user; and a recording device holder coupled to the adjustable band, wherein the recording device holder is configured to position a recording device adjacent to and sufficiently perpendicular to a surface of the abdomen of the user, and wherein the recording device is configured to acquire a sound recording from the abdomen of the user to determine fetal movement..
[0102] Fig. 4C shows another example 400c of operation for feature engineering features for use in machine learning model. Method 400c inncludes providing providing (418) acoustic files and ultrasound files (as ground truth for fetal movement) as training data. Method 400c then includes generating (420) features. Examples features are described herein and provided in relation to Table 1. Method 400c then includes training (422) the Al model.
[0103] Experimental Results and Additional Examples
[0104] A study was conducted to develop a digital phenotype of fetal movement by conducting a cohort study of at least 81 pregnant women (49 of 175 individual recordings and6 of 25 anticipated longitudinal patients having recordings with an additional 5 longitudinal patients in progress) to collect co-registered continuous, annotated ultrasound with I-Phone audio recordings of fetal movement and patient-perceived movements. The study also conducted extensive feature engineering to characterize and classify features in transient, aperiodic audio recording fetal movement in amniotic fluid. The study also developed machine learning models and deep learning models capable of ingesting diverse data, including audio, video, text, and clinical heuristics, to accurately quantify fetal movements and assess the quality of fetal movement. The study also conducted a comprehensive assessment and evaluation of the models to ensure accuracy and efficiency and provide transparency regarding efficacy and biases. Further study is ongoing, along with additional development and algorithm refinements.
[0105] Example A
[0106] In this part of the study (UT IRB approval #00001552), the study included 80 patients. Seventy patients were enrolled in a cross-sectional design: 10 patients at each gestational age of 26, 28, 30, 32-, 34-, 36-, and 38-week gestation. Ten additional patients were enrolled in a longitudinal design with recordings at 27, 30-, 33-, 36-, and 39-week gestation. All patients underwent 30 minutes of simultaneous recordings of maternal perception of movements, continuous ultrasound detection of movements, and continuous I-phone version 10 microphone recordings. Maternal demographics, including parity, BMI, placental location, abdominal wall thickness, and amniotic fluid volume, were recorded at each visit to assess whether any of these parameters could affect the sensitivity of the microphone assessment of FM. Audio characteristics were assessed to develop correlations with both FM assessed by ultrasound as well as the maternal perception of FM. Example data from this study is shown in FIGS. 5A-5S
[0107] To date, studies in three longitudinal patients have been completed, and one is in progress. The cross-sectional studies completed to date are as follows:
[0108] Discussion. More than five decades ago, Sadovsky and Yafee [7] reported seven cases of decreased daily fetal movements (FMs) in association with fetal compromise and fetal death. Heazell et al. [8] interviewed 291 women with late stillbirths and 733 controls to characterize the maternal perception of fetal movements during the two weeks prior to delivery. One episode of decreased fetal movements (DFM’s) was associated with a risk for stillbirth of 2.36 (95% CI: 1.69 - 3.3), while three or more episodes were associated with an OR of 5.11 (CI: 3.22 - 8.1). More recently, a meta-analysis of non-randomized studies evaluated 39 citations and noted that maternal perception of DFM’s was associated with an odds ratio of 3.44 (95% CI: 2.02 - 5.88) for subsequent stillbirth [9],
[0109] There are several challenges, however, to utilize FMs as a predictor for stillbirth. The first is that pregnant women do not perceive all FMs seen on simultaneous ultrasounds. In one study maternal perception of FM’s coincided with ultrasound movements in only 33% of cases
[0010] , When a piezo-electric crystal device was placed on the maternal abdomen, pregnant patients only detected 32% of all movements detected by the device
[0010] , Stronger movements of the fetal limbs were more likely to be detected by the pregnant woman. Interestingly, the accuracy of the device was not affected by parity, gestational age, placental site or thickness, maternal weight, or thickness of the maternal abdominal wall.
[0110] A second challenge has been the definition of the maternal perception of DFM’s. Early studies described a daily count for a period of 12 hours with the completed absence of fetal movement to less than 10 FMs in 2 hours as the definition of DFM’s. Sadovsky et al.
[0011] studied six different definitions of DFM’s and found that the movement alarm signal was the best predictor of stillbirth. The group defined this as fewer than three FMs or complete cessation of FMs over a 12-hour period. Later, Pearson and Weaver
[0012] proposed the Cardiff method which involved counting the duration of time required to achieve ten FM’s. A duration in excess of two hours was considered to represent DFM’s.[OHl] Several investigations have been undertaken to assess the utility of FM’s perceived by the pregnant patient in the prediction of late stillbirth. Two total population studies conducted as prospective cohorts with a control period followed by an intervention period have been reported. Westgate et al.
[0013] in New Zealand reported an overall relative risk of stillbirth with FM counting of 0.76 (95% CI: 0.55 - 1.04) and a relative risk of avoidable stillbirths of 0.56 (CI: 0.35 - 0.90). In the United States, Moore et al.
[0014] reported an overall relative risk of stillbirth with FM counting of 0.42 (95% CI: 0.23 - 0.76) and a relative risk of avoidable stillbirths of 0.25 (CI: 0.07 - 0.88). Neldam
[0015] in Denmark published a randomized trial of 2250 pregnant women reporting an overall relative risk of stillbirth with FM counting of 0.25(95% CI: 0.07- 0.88) and a relative risk of avoidable stillbirths of 0.27 (CI: 0.08-0.93). Saastad et al.
[0016] randomized 1076 Norwegian women to the standard of care or Cardiff FM daily counting starting at week 34 of gestation. Although there was no difference in their primary outcome between the groups, the authors reported that growth-restricted infants were more often reported in the intervention group (87% vs 60%; RR: 1.5 (95% CI: 1.0 - 2.1). The frequency of interventions and consultations was similar in both groups. In a subsequent publication, these same authors found that women who performed FM counting scored lower on the Cambridge Worry Scale as compared to the control group
[0017] , Often, the pregnant woman with decreased FM will delay seeking medical attention in the hope that the fetus will begin moving. In a Japanese study investigating the implementation of daily FM counts, the introduction of the Cardiff method starting at 34 weeks gestation was associated with a reduction in delayed self-referral for medical attention for DFM (RR: 0.31; 95% CI: 0.31 - 0.83)
[0018] , The authors noted a reduction in the regional incidence of stillbirth (3.06 to 2.70 / 1000 births during the study period, although this did not reach statistical significance based on the small sample size.
[0112] Despite these studies, the report by Grant et al.
[0019] is often quoted as negative evidence for the implementation of FM counting to reduce the rate of stillbirth. Sixty-eight thousand women were randomly allocated to 33 pairs of clusters either to a policy of routine FM counting or to standard care. The latter included selective use of FM counting or informal noting of FM. Ninety-nine late stillbirths occurred in the active intervention group as compared to 100 stillbirths in controls. Criticisms of the study included possible contamination between groups, as women in the same community were informed that they were included in the trial even though they were not assigned to the active intervention arm. In addition, only 60% of patients in the active arm were compliant with FM counting, and 50% reported medical attention when the alarm limit for FM was met (8.4% of cases).
[0113] In 2015, a Cochrane review concluded that there have been no clinical trials comparing FM counting to no FM counting
[0020] , Indirect evidence from a large cluster-RCT (Grant et al.
[0019] ) suggested that more babies at risk of death were identified in the routine FM counting group, but this did not translate to reduced perinatal mortality. The authors stated that robust research by means of studies comparing routine FM counting with selective FM counting is urgently needed.
[0114] In January 2007, Steve Jobs at Apple introduced the first version of the iPhone, which included a microphone. More recently, version 15 of the I-phone has been introduced to the market. Improvements in microphone technology have continued with the release of newerproducts. The widespread availability of smartphones has resulted in an explosion of applications. Democratization of health care with the smartphone has also been realized. As an example, machine learning with smartphone recordings was able to discriminate between the cough of COVID- 19-positive patients and the cough of healthy controls with the area under the ROC curve of 0.98
[0021] , Smartphone-based cough monitoring employed at the hospital bedside has been correlated to clinical markers of disease activity and laboratory markers of inflammation in patients with COVID-19 and other non-COVID pneumonia
[0022] ,
[0115] Discussion
[0116] Inability to objectively assess the quantity and quality of fetal movements. Expectant mothers see the obstetrician -10-20 times over the course of their pregnancy, depending on whether or not they are considered high-risk. Those visits represent 1% of the pregnancy. This invention provides a more granular assessment of fetal movement over time, thus enabling a digital phenotype of fetal movement.
[0117] Current fetal movement assessment relies on the maternal perception of movements. Ultrasound studies have shown that as many as two-thirds of fetal movements noted on ultrasound are not perceived by the pregnant patient. In addition, the pregnant patient is unable to assess the quality of fetal movement in an objective manner. There is no technology that has been built on time-registered data of annotated ultrasound that indicates ground truth fetal movement, patient’s perception of fetal movement, and audio recordings of fetal movement that is then input to artificial intelligence methods to extract, classify, and characterize these fetal movements. While there have been efforts to develop a digital phenotype of pregnancy itself, primarily for the purposes of investigating maternal health and post-partum depression, there are no digital phenotypes of fetal movement. This invention would provide that digital phenotype and through deep learning, would refine our understanding over time.
[0118] Such factors as a thickened abdominal wall or an anterior placental location may decrease the sensitivity of movement detection by the smartphone microphone. These factors could be taken into account to adjust the sensitivity based on a previous anatomical ultrasound examination (usually undertaken at 20 weeks gestation) or by correction for maternal BMI.
[0119] Example B
[0120] In another aspect of the study, a sub-study was conducted to research, develop and build a deep learning pipeline to build a digital phenotype of FM. One of the elements of the work is the integration of technology, audio, video, patient demographics, and heuristics into a deep learning platform. The sub-study employed the design and development of a smartdistributed data framework capable of learning and evolving as new data is introduced and knowledge is gained over time.
[0121] Key Health Problem Addressed and Fundamental Contributions. This sub-study sought to phenotype FMs from annotated ultrasound co-registered with smartphone audio recordings and patient-perceived movements. FMs will be quantified (both in number and strength) and characterized based on quality metrics (discussed previously). This innovative, high-risk, high-reward research can provide significant new insight into potential causes and conditions related to stillbirth. The sub-study proposed a research agenda that can contribute fundamental knowledge to both computer and information sciences and biomedical sciences. This close collaboration between team members with deep knowledge in maternal and fetal medicine and computer and information sciences has the potential to disrupt state of the art across both disciplines and can no doubt transform the fundamental understanding of fetal health and wellbeing..
[0122] Data Description. All data collected for the proposed cohort sub-study was stored in a secure RedCap® database. Analysis of the data was conducted on non-PHI data at the Texas Advanced Computing Center (Frontera, Lonestar6, and Stampede3), and those conducting the analysis only had access to non-PHI data. As a result of the preliminary cohort sub-study to determine efficacy, the sub-study was able to further define the variables of interest to collect for phenotyping FMs, focusing on quantity of movement and quality. In addition to variables regarding gestational age, patient weight / BMI and gestational age of the fetus, abdominal wall thickness, amniotic fluid index, maximum vertical pocket, and placenta location were collected, as these variables note the physical variation in the expectant mother and the placement of the fetus. Additionally, it was noted whether breathing and hiccups or fetal breathing movements were observed by the sonographer, along with any notes with specific details the sonographer felt were relevant to the sub-study.
[0123] Annotated Ultrasound: Co-registered ultrasound data was provided for each patient enrolled in the cohort sub-study. A spreadsheet with exact timings was provided to synchronize audio data and patient-perceived movement. These timings were given as offsets from the start of the sub-study exam rather than exact times of day. With each sub-study instrument coregistered in time, only offsets from the start were required to synchronize the analysis. The offsets in the annotated ultrasound were provided as a spreadsheet from the certified sonographer after the exam was completed to allow the sonographer to view the ultrasound without distraction. It became apparent during the efficacy sub-study that it was too burdensome for the sonographer to conduct a long ultrasound recording and note movementseen on the ultrasound at the same time. As such, in this sub-study, annotated ultrasound videos were collected as well to co-register what FMs “look” like with what they “sound” like. Providing this additional data enriched the ability to phenotype quality movements and provided additional “ground truth” data regarding FMs in general.
[0124] Smartphone Recordings of Fetal Movement: iPhone version 10 smartphones were used to record the audio of FM during the cohort sub-study. The audio was co-registered with annotated ultrasound movement, which served as ground truth for FM. Audio recordings were provided in MPEG-4 audio files and converted to waveform audio files for further analysis. In previous research, when a piezo-electric crystal device was placed on the maternal abdomen, pregnant patients only detected 32% of all movements detected by the device itself [18’]. Stronger movements of the fetal limbs were more likely to be detected by the pregnant woman. Notably, the accuracy of the device was not affected by parity, gestational age, placental site or thickness, maternal weight, or thickness of the maternal abdominal wall. This type of device was used in a clinical setting, making it a poor candidate for widespread adoption and at-home use. This sub-study aimed to research options that have the potential to democratize maternal care, thus making the smartphone the best candidate for recording audio long-term.
[0125] Patient Perceived Movement: Patient-perceived movement data was collected by recording a timestamp for every instance the patient entered the enter key on a synchronized laptop. This data was co-registered with both continuous ultrasound and audio recording of FM and was formatted as an exact timestamp of each perceived movement from which offsets can be calculated. These offsets were synchronized with both annotated ultrasound movements and audio recordings of FM for subsequent analysis.
[0126] Discussion. Constructing a digital phenotype of FM from annotated ultrasound, audio recordings of FM, and patient-perceived movement requires deep understanding across multiple areas of research. Not only does there need to be an understanding of the state of the art in clinical observation and recording of FMs and how this correlates with birth outcomes, but there must also be a deep understanding of detecting and extracting salient features from both audio and ultrasound, couple these features with observations and feed them into machine learning (ML) and deep learning (DL) models as well. To this end, this sub-study explored the fusion of multiple approaches: a feature engineering approach input to a machine learning model using both supervised and unsupervised learning, a convolutional neural network (CNN) well suited for deep learning on spatial data (ultrasound videos), a recurrent neural network (RNN) well suited for time-series data (audio), and the fusion strategy to combine multiple methods [43’]
[0127] Feature Engineering: A crucial first step for automatically identifying patterns of FM across the patient cohort is the extraction of features that can be input into the models. The sub-study aimed to characterize and automatically extract FM from both audio recordings and annotated ultrasound video and preserve discriminatory information and separate factors of variation relevant to FM more broadly [44’].
[0128] Audio recordings of Fetal Movement: FM has unique properties in the context of audio. The majority of research in acoustics has focused on speech and music [45’]. Tagging techniques have been used for this type of audio and, to some degree, for environmental sounds (the sounds that are typically heard in the background, including noise, natural sounds, and human activity) [46’]. FM audio closely parallels environmental sounds in several ways. First, environmental sounds are non-static (no rhythm or melody), and second, they do not have a particular structure (for example, phonetics) [47’]. Third, they have no periodicity or repetitive pattern (for example, a beat). Audio recordings of FM are also transient in nature, meaning there is no deterministic pattern to the sound, and they lie in the range of 20 Hz - 400 Hz. Additionally, this audio is recorded by placing the microphone on the abdomen of the pregnant woman, and the fetus is floating in amniotic fluid, which means that not only must one consider specific properties of the audio from a signal processing perspective, but one must also consider variations in the characteristics of both mother and baby. Furthermore, one must define what constitutes “fluid” movement versus “singular” movements.
[0129] This sub-study defined “fluid” movements as FMs that occur within two seconds of each other. That translates to assessing the annotated ultrasound timestamps and patient- perceived timestamps to group movements that have a difference of less than or equal to two seconds between them. The sub-study further classified “fluid” movements by the number of consecutive movements in a group. For example, singular movements were recorded as a single timestamp and were classified as Fl. Fluid movements that had two consecutive movements less than or equal to two seconds apart were classified as F2. Fluid movements that had three consecutive movements less than or equal to two seconds apart were classified as F3, and so on.
[0130] Preprocessing was needed to cancel out noise that may occur during the sub-study session and, if needed, normalize the audio. Noise cancellation was crucial for this sub-study as the patient clicked the button as she felt the fetal movement was audible during the previously conducted efficacy sub-study. Additionally, windowing was needed to help analyze non- stationary signals as quasi -stationary signals and, from there, extract features. The substudy explored the use of traditional audio feature extraction techniques that are widely usedin environmental sound classification. This included Mel Frequency Cepstral Coefficients (MFCC) [48’], Log-Mel Spectrogram [49’], Gammatone [50’], and Wavelets [51’]. Additionally, the sub-study examined the use of audio separation using principal component analysis (PC A), as this approach allowed for the inclusion of additional variables with the audio stream, such as amniotic fluid index, abdominal wall thickness, and BMI. The sub-study compared and contrasted the accuracy of each of the approaches by comparing feature outcomes to the annotated ultrasound (both timestamped and video).
[0131] This research produced fundamental contributions to the area of feature extraction / engineering for transient, aperiodic sound in a liquid medium.
[0132] Annotated Ultrasound Video: Research in ML and deep learning for medical images is largely focused on CT, MRI, and microscopy [52-55’]. Recent work has been done to highlight contributions that ML has made to solve current challenges with medical ultrasound [56’]. The sub-study’s efforts differed slightly but nonetheless contributed to state of the art in medical ultrasound as it applies to fetal health. The certified sonographer annotated the ultrasound video and feature recognition was conducted on that video, developed as part of this research. The features that were extracted were more salient in nature, given that they were previously identified as “areas of interest”. However, videos were preprocessed (de-speckling, etc.), the annotations were initially used to extract regions of the video images and image processing was applied to further segment and classify these regions.
[0133] These annotated videos provided additional robustness in ground truth data and enhanced the quality of both the ML and DL models.
[0134] Model Building: This sub-study explored two classes of models to support the classification of FM to accurately quantify FM, accurately assess the quality of FM, and characterize distinctions and differences of the input vectors (namely, the wide discrepancy between patient-perceived movement and ultrasound observation). The data that was input into the models included: 1) ground truth data provided in two forms: timestamps of observed movements with metadata if specific types of movements were observed (for example, hiccups) and regions of interest overlaid on the ultrasound video, both provided by the certified sonographer; 2) feature vectors extracted from audio files with supporting metadata describing what they are; 3) timestamps of patient-perceived movements; 4) patient-specific data including weight / BMI, abdominal wall thickness, amniotic fluid index, gestational age of the fetus, and when they last ate.
[0135] The sub-study explored ML models that have had success in the area of environmental scene classification and ultrasound image processing, such as support vectormachines (SVM) [57’], K-nearest neighbors (KNN) [58’], matrix factorization [59’] and extreme learning machines [60’] to deal with the engineered features. It is recognized that extensive work engineering features are a limitation to ML models [61’], but this step was useful to exhaustively research methods for the diverse types of data and features that may be important sources of data to inform the models.
[0136] The sub-study also explored deep learning models as these have significant longterm promise. Primarily, the sub-study explored CNNs for ultrasound video and RNNs for audio. It is recognized that it is challenging to obtain sufficient input data to ensure that the DL models have the desired accuracy, so the sub-study employed deep hybrid learning (DHL) strategies that employ the best of both ML and DL. These DHL strategies were used as a fusion strategy and the efficacy of each was assessed. Long term, given that there are diverse types of data that can be integrated into the models, research suggests that DL neural networks may outperform ML models [62’].
[0137] Model Assessment: Assessing a model’s quality requires numerous factors are considered: performance, efficiency, and interpretability. Performance-based metrics are well suited to evaluate supervised learning objectives, given that they measure how well the model was able to satisfy the objective. For this type of assessment, the sub-study used k-fold cross- validation to prevent the model from overfitting, and performance was measured using data not used to train the model. The sub-study measured reliability using cross-validation methods on this out-of-sample data as well, enabling comparative statistical tests [63’]. The sub-study evaluated regression models by measuring estimation errors, for example root mean square error and mean absolute percentage error. The sub-study also compared a suite of models with varying degrees of complexity as this helped identify suitable models for specific tasks. It is recognized that issues can arise when models become overly complex, increasing the risk of overfitting. The sub-study assessed and measured the complexity of models and assigned risk for overfitting so that models that were expected to perform similar tasks could be accurately compared. The sub-study also measured and assessed the level of interpretability and transparency.
[0138] For proposed deep learning models (CNN and RNN), the sub-study measured computational costs in the computing environment and set up experiments to fully assess the hidden costs of the models (memory, CPU, GPU, power, and waste). In the environment, resources were allocated on a per-node basis, and a fair measure that gets to computational efficiency was made. The sub-study also assessed the long-term efficacy of using deep learningmodels to function in an on-demand setting for applications like fetal movement monitoring with periodic data streaming in from an at-home smartphone device.
[0139] Discussion. Th sub-study proposed high-risk, high-reward research to develop a digital phenotype of fetal movement to inform causes of stillbirth. Fundamental research to extract features from transient aperiodic low-frequency audio recordings of fetal movement in amniotic fluid is potentially ground-breaking. Additionally, building both machine learning models and deep learning models that support a diverse array of input vectors (time, metadata, audio, video, patient perception, patient characteristics, and clinical heuristics) contributed to fundamental work in computer science and artificial intelligence. Finally, this research holds significant potential to transform knowledge surrounding factors contributing to stillbirth. This research has the potential to disrupt the state of the art across disciplines and can no doubt transform fundamental understanding of fetal health and wellbeing.
[0140] Example C
[0141] In another aspect of the study, a sub-study is being conducted for detecting fetal movement in smartphone audio recordings. The sub-study may eventually use a cohort of 10 longitudinal patients and 70 cross-sectional patients between 24-29 weeks of gestation. To classify fetal movement, each patient recorded perceived fetal movement while an ultrasound and a smartphone recording were simultaneously taken. The ultrasound was annotated and correlated with the patient's perceived movement to compute delta and count consecutive movements. Consecutive movements were aggregated in less than 2 seconds and classified by aggregated count. Finally, the audio file was encoded by type. FIG. 6A shows annotated ultrasounds from multiple patients.
[0142] Using this annotated ultrasound, the patient perceived movement, and the smartphone recording, the audio was discretized into 2-second snippets. The snippets were written to a database by type. Finally, all of the data was input into a mel spectrogram MFCC feature computation to generate data for machine learning training and validation, which was ultimately used to develop a deep learning convolutional neural network capable of fetal movement detection.
[0143] This sub-study also examined the impact of phone placement on the quality of the smartphone recording. Results of vertical vs. horizontal phone placement are shown in FIGS.6B-6E
[0144] Example Computing System
[0145] The exemplary system and method may be implemented (1) as a sequence of computer-implemented acts or program modules running on a computing system and / or (2) asinterconnected machine logic circuits or circuit modules within the computing system. The implementation is a matter of choice depending on the performance and other requirements of the computing system. Accordingly, the logical operations described herein are referred to variously as state operations, acts, or modules. These operations, acts, and / or modules can be implemented in software, in firmware, in special purpose digital logic, in hardware, and any combination thereof. It should also be appreciated that more or fewer operations can be performed than shown in the figures and described herein. These operations can also be performed in a different order than those described herein.
[0146] The computer system is capable of executing the software components described herein for the exemplary method or systems. In an embodiment, the computing device may comprise two or more computers in communication with each other that collaborate to perform a task. For example, but not by way of limitation, an application may be partitioned in such a way as to permit concurrent and / or parallel processing of the instructions of the application. Alternatively, the data processed by the application may be partitioned in such a way as to permit concurrent and / or parallel processing of different portions of a data set by the two or more computers. In an embodiment, virtualization software may be employed by the computing device to provide the functionality of a number of servers that are not directly bound to the number of computers in the computing device. For example, virtualization software may provide twenty virtual servers on four physical computers. In an embodiment, the functionality disclosed above may be provided by executing the application and / or applications in a cloud computing environment. Cloud computing may comprise providing computing services via a network connection using dynamically scalable computing resources. Cloud computing may be supported, at least in part, by virtualization software. A cloud computing environment may be established by an enterprise and / or can be hired on an as-needed basis from a third-party provider. Some cloud computing environments may comprise cloud computing resources owned and operated by the enterprise as well as cloud computing resources hired and / or leased from a third-party provider.
[0147] In its most basic configuration, a computing device includes at least one processing unit and system memory. Depending on the exact configuration and type of computing device, system memory may be volatile (such as random-access memory (RAM)), non-volatile (such as read-only memory (ROM), flash memory, etc.), or some combination of the two.
[0148] The processing unit may be a standard programmable processor that performs arithmetic and logic operations necessary for the operation of the computing device. While only one processing unit is shown, multiple processors may be present. As used herein,processing unit and processor refers to a physical hardware device that executes encoded instructions for performing functions on inputs and creating outputs, including, for example, but not limited to, microprocessors (MCUs), microcontrollers, graphical processing units (GPUs), and application-specific circuits (ASICs). Thus, while instructions may be discussed as executed by a processor, the instructions may be executed simultaneously, serially, or otherwise executed by one or multiple processors. The computing device may also include a bus or other communication mechanism for communicating information among various components of the computing device.
[0149] Computing devices may have additional features / functionality. For example, the computing device may include additional storage such as removable storage and nonremovable storage including, but not limited to, magnetic or optical disks or tapes. Computing devices may also contain network connection(s) that allow the device to communicate with other devices, such as over the communication pathways described herein. The network connection(s) may take the form of modems, modem banks, Ethernet cards, universal serial bus (USB) interface cards, serial interfaces, token ring cards, fiber distributed data interface (FDDI) cards, wireless local area network (WLAN) cards, radio transceiver cards such as code division multiple access (CDMA), global system for mobile communications (GSM), longterm evolution (LTE), worldwide interoperability for microwave access (WiMAX), and / or other air interface protocol radio transceiver cards, and other well-known network devices. Computing devices may also have input device(s) such as keyboards, keypads, switches, dials, mice, trackballs, touch screens, voice recognizers, card readers, paper tape readers, or other well-known input devices. Output device(s) such as printers, video monitors, liquid crystal displays (LCDs), touch screen displays, displays, speakers, etc., may also be included. The additional devices may be connected to the bus in order to facilitate the communication of data among the components of the computing device. All these devices are well known in the art and need not be discussed at length here.
[0150] The processing unit may be configured to execute program code encoded in tangible, computer-readable media. Tangible, computer-readable media refers to any media that is capable of providing data that causes the computing device (i.e., a machine) to operate in a particular fashion. Various computer-readable media may be utilized to provide instructions to the processing unit for execution. Example tangible, computer-readable media may include but is not limited to volatile media, non-volatile media, removable media, and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data.System memory, removable storage, and non-removable storage are all examples of tangible computer storage media.
[0151] Example tangible, computer-readable recording media include, but are not limited to, an integrated circuit (e.g., field-programmable gate array or application-specific IC), a hard disk, an optical disk, a magneto-optical disk, a floppy disk, a magnetic tape, a holographic storage medium, a solid-state device, RAM, ROM, electrically erasable program read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices.
[0152] In light of the above, it should be appreciated that many types of physical transformations take place in the computer architecture to store and execute the software components presented herein. It also should be appreciated that the computer architecture may include other types of computing devices, including hand-held computers, embedded computer systems, personal digital assistants, and other types of computing devices known to those skilled in the art.
[0153] In an example implementation, the processing unit may execute program code stored in the system memory. For example, the bus may carry data to the system memory, from which the processing unit receives and executes instructions. The data received by the system memory may optionally be stored on the removable storage or the non-removable storage before or after execution by the processing unit.
[0154] It should be understood that the various techniques described herein may be implemented in connection with hardware or software or, where appropriate, with a combination thereof. Thus, the methods and apparatuses of the presently disclosed subject matter, or certain aspects or portions thereof, may take the form of program code (i.e., instructions) embodied in tangible media, such as floppy diskettes, CD-ROMs, hard drives, or any other machine-readable storage medium wherein, when the program code is loaded into and executed by a machine, such as a computing device, the machine becomes an apparatus for practicing the presently disclosed subject matter. In the case of program code execution on programmable computers, the computing device generally includes a processor, a storage medium readable by the processor (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. One or more programs may implement or utilize the processes described in connection with the presently disclosed subject matter, e.g., through the use of an application programming interface (API), reusable controls, or the like. Such programs may be implemented in a high-level procedural or object-orientedprogramming language to communicate with a computer system. However, the program(s) can be implemented in assembly or machine language, if desired. In any case, the language may be a compiled or interpreted language, and it may be combined with hardware implementations.
[0155] Although example embodiments of the present disclosure are explained in some instances in detail herein, it is to be understood that other embodiments are contemplated. Accordingly, it is not intended that the present disclosure be limited in its scope to the details of construction and arrangement of components set forth in the following description or illustrated in the drawings. The present disclosure is capable of other embodiments and of being practiced or carried out in various ways.
[0156] The following patents, applications and publications as listed below and throughout this document are hereby incorporated by reference in their entirety herein.Reference list #1 of Example A[1] Tanner D, Murthy S, Lavista Ferres JM, Ramirez JM, Mitchell EA. Risk factors for late (28+ weeks' gestation) stillbirth in the United States, 2014-2015. PLoS One 2023; 18: e0289405.[2] Valenzuela CP, Gregory E, Martin JA. Decline in Perinatal Mortality in the United States, 2017-2019. NCHS Data Brief 2022: 1-8.[3] Stillbirth Collaborative Research Network Writing G. Association between stillbirth and risk factors known at pregnancy confirmation. JAMA 2011; 306: 2469- 2479.[4] American College of Obstetricians and Gynecologists SfM-FM. Obstetric Care Consensus: Management of Stillbirth. Ob stet Gynecol 2022; 135: el l0-el31.[5] Indications for Outpatient Antenatal Fetal Surveillance: ACOG Committee Opinion, Number 828. Obstet Gynecol 2021; 137: el77-el97.[6] de la Vega A, Verdiales M. Failure of intensive fetal monitoring and ultrasound in reducing the stillbirth rate. P R Health Sci J 2002; 21 : 123-125.[7] Sadovsky E, Yaffe H. Daily fetal movement recording and fetal prognosis. Obstet Gynecol 1973; 41 : 845-850.[8] Heazell AEP, Budd J, Li M, Cronin R, Bradford B, McCowan LME, Mitchell EA, Stacey T, Martin B, Roberts D, Thompson JMD. Alterations in maternally perceived fetal movement and their association with late stillbirth: findings from the Midland and North of England stillbirth case-control study. BMJ Open 2018; 8: e020031.[9] Carroll L, Gallagher L, Smith V. Pregnancy, birth and neonatal outcomes associated with reduced fetal movements: A systematic review and meta-analysis of non-randomised studies. Midwifery 2023; 116: 103524.
[0010] Valentin L, Marsal K, Lindstrom K. Recording of foetal movements: a comparison of three methods. J Med Eng Technol 1986; 10: 239-247.
[0011] Sadovsky E, Ohel G, Havazeleth H, Steinwell A, Penchas S. The definition and the significance of decreased fetal movements. Acta Obstet Gynecol Scand 1983; 62: 409-413.
[0012] Pearson JF, Weaver JB. Fetal activity and fetal wellbeing: an evaluation. Br Med J 1976; 1 : 1305-1307.
[0013] Westgate J, Jamieson M. Stillbirths and fetal movements. N Z Med J 1986; 99: 114-116.
[0014] Moore TR, Piacquadio K. A prospective evaluation of fetal movement screening to reduce the incidence of antepartum fetal death. Am J Obstet Gynecol 1989; 160: 1075-1080.
[0015] Neldam S. Fetal movements as an indicator of fetal well-being. Dan Med Bull 1983; 30: 274-278.
[0016] Saastad E, Winje BA, Stray Pedersen B, Froen JF. Fetal movement counting improved identification of fetal growth restriction and perinatal outcomes— a multicentre, randomized, controlled trial. PLoS One 2011; 6: e28482.
[0017] Saastad E, Winje BA, Israel P, Froen JF. Fetal movement counting— maternal concern and experiences: a multicenter, randomized, controlled trial. Birth 2012; 39: 10-20.
[0018] Koshida S, Tokoro S, Katsura D, Tsuji S, Murakami T, Takahashi K. Fetal movement counting is associated with the reduction of delayed maternal reaction after perceiving decreased fetal movements: a prospective study. Sci Rep 2021; 11 : 10818.
[0019] Grant A, Elbourne D, Valentin L, Alexander S. Routine formal fetal movement counting and risk of antepartum late death in normally formed singletons. Lancet 1989; 2: 345-349.
[0020] Mangesi L, Hofmeyr GJ, Smith V, Smyth RM. Fetal movement counting for assessment of fetal wellbeing. Cochrane Database Syst Rev 2015; 2015: CD004909.
[0021] Pahar M, Klopper M, Warren R, Niesler T. COVID-19 cough classification using machine learning and global smartphone recordings. Comput Biol Med 2021; 135: 104572.
[0022] Boesch M, Rassouli F, Baty F, Schwarzler A, Widmer S, Tinschert P, Shih I, Cleres D, Barata F, Fleisch E, Brutsche MH. Smartphone-based cough monitoring as a near real-time digital pneumonia biomarker. ERJ Open Res 2023; 9.Reference list #2 of Example B[U] Tanner, D., et al., Risk factors for late (28+ weeks' gestation) stillbirth in the United States, 2014-2015. PLoS One, 2023. 18(8): p. e0289405.[2’] Valenzuela, C.P., E. Gregory, and J. A. Martin, Decline in Perinatal Mortality in the United States, 2017-2019. NCHS Data Brief, 2022(429): p. 1-8.[3’] Stillbirth Collaborative Research Network Writing, G., Association between stillbirth and risk factors known at pregnancy confirmation. JAMA, 2011. 306(22): p. 2469-79.[4’] Gregory, E.C., C.P. Valenzuela, and D.L. Hoyert, Fetal Mortality: United States, 2020. Natl Vital Stat Rep, 2022. 71(4): p. 1-20.[5’] Pollock, D.D., et al., Breaking the silence: Determining Prevalence and Understanding Stillbirth Stigma. Midwifery, 2021. 93: p. 102884.[6’] American College of Obstetricians and Gynecologists, S.f. M.-F.M., Obstetric Care Consensus: Management of Stillbirth. Obstet Gynecol, 2022. 135(3): p. el lO- el31.[7’] Indications for Outpatient Antenatal Fetal Surveillance: ACOG Committee Opinion, Number 828. Obstet Gynecol, 2021. 137(6): p. el77-el97.[8’] de la Vega, A. and M. Verdiales, Failure of intensive fetal monitoring and ultrasound in reducing the stillbirth rate. P R Health Sci J, 2002. 21(2): p. 123-5.[9’] Turton, P., C. Evans, and P. Hughes, Long-term psychosocial sequelae of stillbirth: phase II of a nested case-control cohort study. Arch Womens Ment Health, 2009. 12(1): p. 35-41.[10’] Flenady, V., et al., Meeting the needs of parents after a stillbirth or neonatal death. BJOG, 2014. 121 Suppl 4: p. 137-40.[11’] UNICEF, What you need to know about stillbirths. 2023.[12’] Nahian, A. and K. Mahomed, Decreased fetal movements - An audit of predictors and an evaluation of management based on a locally developed flow chart. Eur J Obstet Gynecol Reprod Biol, 2023. 290: p. 67-73.[13’] Brown, R., et al., Maternal perception of fetal movements in late pregnancy is affected by type and duration of fetal movement. J Matern Fetal Neonatal Med, 2016. 29(13): p. 2145-50.[14’] Doppler Ultrasound in Obstetrics and Gynecology. 2 ed, ed. D. Maulik. 1997, New York, NY: Springer Berlin, Heidelberg.[15’] Sadovsky, E. and H. Yaffe, Daily fetal movement recording and fetal prognosis. Obstet Gynecol, 1973. 41(6): p. 845-50.[16’] Heazell, A.E.P., et al., Alterations in maternally perceived fetal movement and their association with late stillbirth: findings from the Midland and North of England stillbirth case-control study. BMJ Open, 2018. 8(7): p. e020031.[17’] Carroll, L., L. Gallagher, and V. Smith, Pregnancy, birth and neonatal outcomes associated with reduced fetal movements: A systematic review and metaanalysis of non-randomised studies. Midwifery, 2023. 116: p. 103524.[18’] Valentin, L., K. Marsal, and K. Lindstrom, Recording of foetal movements: a comparison of three methods. J Med Eng Technol, 1986. 10(5): p. 239-47.[19’] Sadovsky, E., et al., The definition and the significance of decreased fetal movements. Acta Obstet Gynecol Scand, 1983. 62(5): p. 409-13.[20’] Pearson, J.F. and J.B. Weaver, Fetal activity and fetal wellbeing: an evaluation. Br Med J, 1976. 1(6021): p. 1305-7.[21’] Westgate, J. and M. Jamieson, Stillbirths and fetal movements. N Z Med J, 1986. 99(796): p. 114-6.[22’] Moore, T.R. and K. Piacquadio, A prospective evaluation of fetal movement screening to reduce the incidence of antepartum fetal death. Am J Obstet Gynecol, 1989. 160(5 Pt 1): p. 1075-80.[23’] Neldam, S., Fetal movements as an indicator of fetal well-being. Dan Med Bull, 1983. 30(4): p. 274-278[24’] Saastad, E., et al., Fetal movement counting improved identification of fetal growth restriction and perinatal outcomes— a multi-centre, randomized, controlled trial. PLoS One, 2011. 6(12): p. e28482.[25’] Saastad, E., et al., Fetal movement counting— maternal concern and experiences: a multicenter, randomized, controlled trial. Birth, 2012. 39(1): p. 10-20.[26’] Koshida, S., et al., Fetal movement counting is associated with the reduction of delayed maternal reaction after perceiving decreased fetal movements: a prospective study. Sci Rep, 2021. 11(1): p. 10818.[27’] Grant, A., et al., Routine formal fetal movement counting and risk of antepartum late death in normally formed singletons. Lancet, 1989. 2(8659): p. 345-9.[28’] Mangesi, L., et al., Fetal movement counting for assessment of fetal wellbeing. Cochrane Database Syst Rev, 2015. 2015(10): p. CD004909.[29’] Bradford, B.F., et al., Fetal movements: A framework for antenatal conversations. Women Birth, 2023. 36(3): p. 238-246.[30’] Vries, J.I.P.d. and B.F. Fong, Normal fetal motility: an overview. Ultrasound Obstet Gynecol, 2006. 27(6): p. 701-711.[31’] Reissland, N. and B. Francis, The quality of fetal arm movements as indicators of fetal stress. Early Human Development, 2010. 86(12): p. 813-816.[32’] Kainer, F., et al., Assessment of the quality of general movements in fetuses and infants of women with type-I diabetes mellitus. Early Hum Dev, 1997. 50(1): p. 13- 25.[33’] Pillai, M. and D. James, Hiccups and breathing in human fetuses. Arch Dis Child, 1990. 65(10 Spec No): p. 1072-5.[34’] Woerden, E.E.v., et al., Fetal hiccups; characteristics and relation to fetal heart rate. Eur J Obstet Gynecol Reprod Biol, 1989. 30(3): p. 209-216.[35’] Howes, D., Hiccups: a new explanation for the mysterious reflex. Bioessays, 2012. 34(6): p. 451-3.[36’] Drews, C.D., J.F. Kraus, and S. Greenland, Recall bias in a case-control study of sudden infant death syndrome. Int J Epidemiol, 1990. 19(2): p. 405-11.[37’] Torous, J., et al., New Tools for New Research in Psychiatry: A Scalable and Customizable Platform to Empower Data Driven Smartphone Research. JMIR Ment Health, 2016. 3(2): p. el6.[38’] Center, P.R., Mobile Fact Sheet. 2021.[39’] Dimes, M.o., Nowhere to Go: Maternity Care Deserts Across the U.S. 2022.[40’] Ferreira-Cardoso, H., et al., Lung Auscultation Using the Smartphone- Feasibility Study in Real-World Clinical Practice. Sensors (Basel), 2021. 21(14).[41’] Pahar, M., et al., COVID-19 cough classification using machine learning and global smartphone recordings. Comput Biol Med, 2021. 135: p. 104572.[42’] Boesch, M., et al., Smartphone-based cough monitoring as a near real-time digital pneumonia biomarker. ERJ Open Res, 2023. 9(3).[43’] Fonseca, E., R. Gong, and X. Serra A Simple Fusion of Deep and Shallow Learning for Acoustic Scene Classification. Computer Science Sound, 2018. DOI: 10.48550 / arXiv.1806.07506.[44’] Goodfellow, I., Y. Bengio, and A. Courville, Deep Learning. 2016: The MIT Press.[45’] Bansal, A. and N.K. Garg, Environmental Sound Classification: A descriptive review of the literature. Intelligent Systems with Applications, 2022. 16.[46’] Duan, S., et al., A survey of tagging techniques for music, speech and environmental sound. Artif intell Rev, 2014. 42: p. 637-661.[47’] Mushtaq, Z., S.-F. Su, and Q.-V. Tran, Spectral images based environmental sound classification using CNN with meaningful data augmentation. Applied Acoustics, 2021. 172.[48’] Cotton, C.V. and D.P.W. Ellis, Spectral vs. spectro-temporal features for acoustic event detection, in IEEE workshop on applications of signal processing to audio and acoustics. 2011, IEEE. p. 69-72.[49’] Li, J., et al., A comparison of Deep Learning methods for environmental sound detection, in ICASSP IEEE International Conference on Acoustics, Speech and Signal Processing. 2017. p. 126-130.[50’] Valero, X. and F. Alias, Gammatone cepstral coefficients: Biologically inspired features for non-speech audio classification. IEEE Trans Multimed, 2012. 14(6): p. 1684-1689.[51’] Geiger, J.T. and K. Helwani, Improving event detection for audio surveillance using Gabor filterbank features, in 23rd Eur. Signal Process. Conf. EUSIPCO. 2015. p. 714-718.[52’] Wang, S. and R.M. Summers, Machine learning and radiology. Med Image Anal, 2012. 16(5): p. 933-51.[53’] Shen, D., G. Wu, and H. Suk, Deep learning in medical image analysis. Annu Rev Biomed Eng, 2017.[54’] Litjens, G., et al., A survey on deep learning in medical image analysis. Med Image Anal, 2017. 42: p. 60-88.[55’] Ravi, D., et al., Deep Learning for Health Informatics. IEEE J Biomed Health Inform, 2017. 21(1): p. 4-21.[56’] Brattain, L.J., et al., Machine learning for medical ultrasound: status, methods, and future opportunities. Abdom Radiol (NY), 2018. 43(4): p. 786-799.[57’] Wang, J.C., et al., Environmental sound classification using hybrid SVM / KNN classifier and MPEG-7 audio low-level descriptor, in International Joint Conference on Neural Networks. 2006, IEEE. p. 1731-1735.[58’] Ye, J., T. Kobayashi, and M. Masahiro, Urban sound event classification based on local and global features aggregation. Appl. Acoust., 2017. 117: p. 246-256.[59’] Bisot, V., et al., Feature learning with matrix factorization applied to acoustic scene classification. IEEE / ACM Trans. Audio Speech Lang. Process., 2017. 25(6): p. 1216-1229.[60’] Zhang, Y., et al., Multi-kernel extreme learning machine for EEG classification in brain-computer interfaces. Exp. Syst. Appl., 2017. 96(2).[61’] Mu, W., et al., Environmental sound classification using temporal-frequency attention based convolutional neural network. Sci Rep, 2021. 11(1): p. 21552.[62’] Janiesch, C., P. Zschech, and K. Heinrich Machine learning and deep learning. Electron Markets, 2021. 31, 685-695 DOI: 10.1007 / sl2525-021-00475-2.[63’] Garcia, S. and F. Herrera, An extension on “statistical comparisons of classifiers over multiple data sets. Journal of Machine Learning Research, 2008. 9(89): p. 2677-2694.[64’] Schulz, K.W., et al., Optimal mode of delivery in pregnancy: Individualized predictions using national vital statistics data. PLOS Digit Health, 2022. 1(12): p. e0000166.[65’] Drake, J., et al., Collecting and analyzing smartphone sensor data for health, in PEARC ’21 : Practice and Experience in Advanced Research Computing. 2021.
Claims
What is claimed is:
1. A method comprising: receiving, by a processor, an acoustic file of a fetal sound recording; determining, by the processor, via a trained machine learning classifier, presence of each fetal movement from the acoustic file or fetal sound recording; and determining, by the processor, a number of estimated presence of each fetal movement; and wherein the determined number of estimated presence of each fetal movement is outputted, via a graphical user interface or report, to a user.
2. The method of claim 1, wherein the trained machine learning classifier employs ultrasound data as a ground truth in a training operation performed in conjunction with the fetal sound recording, or, wherein the trained machine learning classifier employed ultrasound data labeled for fetal movement derived from acoustic or ultrasound training data.
3. The method of claim 1 or 2 further comprising: determining, by the processor, via a second trained machine learning classifier, a quality score of each fetal movement from the fetal sound recording, wherein the determined quality score is outputted, via the graphical user interface or report, to the user.
4. The method of any one of claims 1-3 further comprising: transmitting the determined number of estimated presence of each fetal movement and / or the determined quality score to a pre-defined clinician.
5. The method of any one of claims 1-4 further comprising: transmitting the determined number of estimated presence of each fetal movement and / or the determined quality score to a pre-defined clinician based on a trigger event associated with the determined number of estimated presence of each fetal movement and / or the determined quality score.
6. The method of any one of claims 1-5, wherein the processing is performed at an edge device.
7. The method of any one of claims 1-5, wherein the processing is performed at a cloud infrastructure.
8. The method of any one of claims 1-7, wherein the trained machine learning classifier is trained by: simultaneously collecting a reference acoustic file and reference ultrasound data, wherein the trained machine learning classifier uses the reference ultrasound data as a ground truth.
9. The method of any one of claims 1-8, further comprising: outputting, via the graphical user interface or report, the determined number of estimated presence of each fetal movement to the user.
10. The method of any one of claims 1-9, wherein the trained machine learning classifier is (i) a neural network or (ii) an equation derived from engineered features associated with time-domain analysis, frequency domain analysis, or a time and frequency domain analysis.
11. A system comprising: a processor; and a memory having instruction stored thereon, wherein execution of the instructions by the processor causes the processor to: receive, an acoustic file of a fetal sound recording; determine, via a trained machine learning classifier, presence of each fetal movement from the fetal sound recording; and determine a number of estimated presence of each fetal movement; and wherein the determined number of estimated presence of each fetal movement is outputted, via a graphical user interface or report, to a user.
12. The system of claim 11, wherein execution of the instructions by the processor causes the processor to perform any one of the method of claims 2-10.
13. A mobile device comprising: an acoustic sensor; a network interface; a processor; and a memory having instruction stored thereon, wherein execution of the instructions by the processor causes the processor to: generate an acoustic file of a fetal sound recording from the acoustic sensor; transmit the acoustic file to a cloud infrastructure configured to: (i) receive the acoustic file of the fetal sound recording, (ii) determine, via a trained machine learning classifier, presence of each fetal movement from the fetal sound recording; and (iii) determine a number of estimated presence of each fetal movement; receive the determined number of estimated presence of each fetal movement; output, via the graphical user report of the device, the determined number of estimated presence of each fetal movement.
14. A mobile device comprising: an acoustic sensor; a network interface; a processor; and a memory having instruction stored thereon, wherein execution of the instructions by the processor causes the processor to: generate an acoustic file of a fetal sound recording from the acoustic sensor; determine, via a trained machine learning classifier, presence of each fetal movement from the fetal sound recording; and determine a number of estimated presence of each fetal movement; output, via a graphical user report of the device, the determined number of estimated presence of each fetal movement.
15. The mobile device of claim 13 or 14, wherein the device is configured to perform the method of any one of claims 1-10.
16. A harness comprising: an adjustable band releasably encircling an abdomen of a user to retain a recording device on an abdomen of a user; anda recording device holder coupled to the adjustable band; wherein the recording device holder is configured to position a recording device adjacent to and sufficiently perpendicular to a surface of the abdomen of the user; wherein the recording device is configured to acquire a sound recording from the abdomen of the user to determine fetal movement according to the method of any one of claims 1-10.
Citation Information
Patent Citations
Fetal movement monitor
US20110306893A1
Prediction and monitoring of clinical episodes
US20150164433A1
Systems and methods for monitoring fetal wellbeing
US20200178880A1
Fetal health monitoring system and method for using the same
US20210378585A1
Ultrasound system acoustic output control using image data
US20220280139A1
Cited By
Monitoring report generation method and equipment
CN121260345A