Biophysical signal and sound-based feature data generation and analysis using artificial intelligence models
By converting biophysical signals to sound-based features and vice versa using machine learning models, the method offers a non-invasive and efficient diagnostic tool for identifying and managing medical conditions, enhancing diagnostic capabilities.
Patent Information
- Application Number
- PCT/US2025/035251
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-25
- Filing Date
- 2025-06-25
- Publication Date
- 2026-01-02
AI Technical Summary
Existing biophysical signal and sound-based diagnostic methods are limited in their ability to provide comprehensive and cost-effective insights into medical conditions, often requiring invasive and specialized tests.
A method utilizing machine learning models to convert biophysical signals into sound-based features and vice versa, enabling the generation of sound-based feature data from biophysical signals and vice versa, which can be analyzed to diagnose and prognosticate medical conditions using AI/ML models.
Provides a cost-effective, non-invasive, and efficient point-of-care diagnostic tool for identifying and managing medical conditions by generating and analyzing sound-based feature data, improving diagnosis, management, and treatment of various health states.
Smart Images

Figure US2025035251_02012026_PF_FP_ABST
Abstract
Description
BIOPHYSICAL SIGNAL AND SOUND-BASED FEATURE DATA GENERATION AND ANALYSIS USING ARTIFICIAL INTELLIGENCE MODELSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application SerialNo. 63 / 664,101, filed on June 25, 2024, and entitled “BIOPHYSICAL SIGNAL AND SOUND-BASED FEATURE DATA GENERATION AND ANALYSIS USING ARTIFICIAL INTELLIGENCE MODELS,” which is herein incorporated by reference in its entirety.STATEMENT OF FEDERALLY SPONSORED RESEARCH
[0002] This invention was made with government support under TR001436 awarded by the National Institutes of Health. The government has certain rights in the invention.BACKGROUND
[0003] Various sensors and measurement systems can be used to measure biophysical signals, such as electrophysiological signals of cardiac activity, electrophysiological signals of brain activity, electrophysiological signals of muscle activity, and the like. For example, voltages of the electrical activity of the heart are measured using electrodes places on the subject’s skin. In a conventional 12-lead ECG, electrodes are placed on the subject’s limbs and chest.
[0004] In recent years, digital stethoscopes have become increasingly popular as a diagnostic tool in medicine. These devices use advanced technology to amplify and record sounds with phonograms from the heart, lungs, and other internal organs, and can be used to detect a wide range of murmurs.
[0005] The use of an artificial intelligence-based deep learning network approach has the potential to provide further insights into biophysical signals and sound-based features acquired from a subject.SUMMARY OF THE DISCLOSURE
[0006] It is an aspect of the present disclosure to provide a method for generating sound-based feature data from biophysical signals. The method includes accessing biophysical signal data with a computer system, where the biophysical signal data may include biophysicalsignals acquired from a subject; accessing a machine learning model with the computer system, where the machine learning model has been trained on training data to convert biophysical signals to sound-based features; applying the biophysical signal data to the machine learning model using the computer system, generating sound-based feature data as an output; and outputting the sound-based feature data using the computer system. Other embodiments of this aspect include corresponding systems (e.g., computer systems), programs, algorithms, and / or modules, each configured to perform the steps of the methods.
[0007] It is another aspect of the present disclosure to provide a method for generating biophysical signals from sound-based feature data. The method includes accessing soundbased feature data with a computer system, where the sound-based feature data may include one or more sound-based features recorded from a subject: accessing a machine learning model with the computer system, where the machine learning model has been trained on training data to convert sound-based features to biophysical signals; applying the sound-based feature data to the machine learning model using the computer system, generating biophysical signal data as an output; and outputting the biophysical signal data using the computer system. Other embodiments of this aspect include corresponding systems (e.g., computer systems), programs, algorithms, and / or modules, each configured to perform the steps of the methods.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] FIG. 1 is a flowchart of an example method for converting biophysical signals to sound-based features using a trained machine learning model.
[0009] FIG. 2 is a flowchart of an example method for training a machine learning model to convert biophysical signals to sound-based features.
[0010] FIG. 3 is a flow chart of an example method for converting sound-based features to biophysical signals using a trained machine learning model.
[0011] FIG. 4 is a flowchart of an example method for training a machine learning model to convert sound-based features to biophysical signals.
[0012] FIG. 5 is a flow chart of an example method for analyzing sound-based features using a machine learning model to generate classified feature data indicative of a medical condition in a subject.
[0013] FIG. 6 is a flowchart of an example method for analyzing biophysical signals generated from sound-based features using a machine learning model to generate classified feature data indicative of a medical condition in a subject.
[0014] FIG. 7 illustrates an example wearable device that can be used to record biophysical signals and / or sound-based features.
[0015] FIG. 8 is a block diagram of example components that can implement the wearable device of FIG. 7.
[0016] FIG. 9 illustrates an example system for generating and analyzing sound-based feature data and / or biophysical signal data.
[0017] FIG. 10 is a block diagram of example components that can implement the system of FIG. 9.DETAILED DESCRIPTION
[0018] Described here are systems and methods for converting between biophysical signals and sound-based features using a machine learning model or other artificial intelligence (Al) model. As one example, sound-based feature data can be generated from biophysical signals, such as electrocardiography (ECG) signals, electroencephalography (EEG) signals, electromyography (EMG) signals, photoplethysmography (PPG) signals, and so on. Additionally or alternatively, sound-based feature data can be generated from transformed biophysical signal data (e.g., scalogram data, spectrogram data, or other N-dimensional (for N > 2) images, maps, matrices, or data structures generated from biophysical signal data).
[0019] The sound-based feature data may include sound waves, audio signals, and / or features generated from sound waves and / or audio signals. Features generated from sound waves and / or audio signals may include amplitude, frequency, and / or phase. Additionally or alternatively, sound-based features may include one or more characteristics or attributes generated from a sound wave and / or audio signal, such as time-domain features, frequencydomain features, or other features that can be computed or otherwise generated from a sound wave or audio signal. Time-domain features may include amplitude envelope, zero-crossing rate, temporal moments (e.g., mean, variance, skewness, kurtosis). energy, root mean square (RMS) energy, or the like. Frequency-domain features may include spectral centroid, spectral bandwidth, spectral contrast, spectral flatness, or the like.
[0020] Advantageously, sound-based feature data can be used as an input to a machine learning model or other Al model to identify, detect, differentiate, and / or prognosticate one or more disease conditions from normal. For instance, sound-based feature data may be input to a machine learning model or other Al model to generate classified feature data indicating the presence and / or likelihood of a particular disease condition or health state of the subject. As aresult, the systems and methods described in the present disclosure provide a point-of-care test for detecting disease conditions or otherwise assessing the health state of a subject without the need for specialized tests.
[0021] In some aspects of the present disclosure, biophysical signals (e.g., ECG signals, EEG signals, EMG signals, PPG signals, etc.) can be converted to sound-based feature data. By collecting sound-based feature data from a subject population, biophysical signal-based sound libraries can be generated. As an example, a sound library may include annotated soundbased features. For example, normal sinus rhythm, arrhythmias, left bundle branch block, etc., can be heard as a sound and a sound library7can be created from biophysical signal data measuring those conditions. The resulting sound-based features can be annotated or labeled based on the underlying medical condition.
[0022] These sound-based features and / or sound libraries can then utilized as an input to an Al or machine learning (AI / ML) model, to train one or more AI / ML models, or the like, to cost-effectively and expeditiously improve on the diagnosis, management, prognostics, and treatment of medical conditions, including medical disorders, disease conditions, or the like. In this way, sound libraries containing sound-based feature data generated from biophysical signal data can be used as a diagnostic tool, such as to diagnose conditions such as atrial fibrillation, left bundle branch block, and other specific medical conditions.
[0023] In some implementations, the sound-based feature data and / or sound library can be analyzed by audio analysis. In some other implementations, the sound-based feature data and / or sound library can be input to one or more AI / ML models to generate classified feature data indicative of a medical condition. Any suitable AI / ML or other computational modelling can be used for the analysis of sound-based feature data and / or a sound library, including deep learning, generative adversarial network (GAN), convolutional neural network (CNN), large language modeling (LLM), foundation models, Riffusion, anomaly detection, diffusion, etc. Advantageously, classical machine learning algorithms and models (e.g., random forest, support vector machines, naive Bayes classifiers, nearest neighbors, decision trees, AdaBoost, QDA, Gaussian process, etc.) can also be used in some instances. Additionally or alternatively, these classical machine learning models could be used for the analysis of sound-based feature data and / or a sound library, or could advantageously be used to initially determine which AI / ML algorithm or model is likely to provide the highest accuracy to develop further models for the analysis of sound-based feature data and / or a sound library7.
[0024] In some other aspects of the present disclosure, sounds or other audio signals collected from audio devices (e.g., digital stethoscopes, ultrasound, Doppler ultrasound) can be converted to biophysical signal data (e.g., ECG signals, EEG signals, EMG signals, PPG signals, etc.) and / or transformed biophysical signal data (e.g., scalogram data, spectrogram data, or other N-dimensional (for N > 2) images, maps, matrices, or data structures generated from biophysical signal data). These sound-based biophysical signal data can then utilized as an input to an AI / ML model, to train one or more AI / ML models, or the like, for the purpose of analyzing heart, lung, or other organ or vascular sounds to identify, detect, differentiate, and / or prognosticate medical conditions from normal.
[0025] As an example, sounds or other audio signals collected from an audio device can thus be analyzed to cost-effectively and expeditiously improve on the diagnosis, management, prognostics, and treatment of medical conditions through the application of AI / ML models and other data science techniques, including computational modeling, deep learning, GAN, CNN, LLM. foundation models, Riffusion, anomaly detection, diffusion, etc. The sounds or other audio signals can also be converted to other transformed sound libraries, enabling AI / ML model-based analysis to cost-effectively and expeditiously improve on the diagnosis, management, prognostics, and treatment of medical conditions. For example, the sound from a digital stethoscope can be converted to a standard raw ECG signal or transformed to a scalograms and / or spectrograms, and used to screen, detect, identify, or prognosticate medical conditions versus normal. In general, a digital stethoscope is a device that amplifies and records sounds with phonograms from the heart, lungs, and other internal organs, and can be used to detect a wide range of medical conditions.
[0026] In some embodiments, the disclosed systems and methods utilize a machine learning model to detect and characterize a medical condition from sound-based feature data generated from biophysical signal data. In these instances, the disclosed systems and methods provide a cost-effective, non-invasive, low-risk intervention to subjects in a point-of-care or home setting. Additionally or alternatively, the disclosed systems and methods can be implemented using a wearable patch or other wearable device with one or more channels and wearable elements, including shirts, watches, bands, and bracelets with conductive elements capable of recording physiologic signals, and / or from and from implanted devices such as loop recorders, pacemakers, defibrillators, and / or digital stethoscopes. Classified feature data can be generated from the generated sound-based feature data, processed by a machine learning model or algorithm to generate classified feature data indicative of a particular medicalcondition, and allow for notifying the user or clinicians that a medical condition has been detected (e.g., via an alert or message).
[0027] In some other embodiments, the disclosed systems and methods utilize a machine learning model to detect and characterize a medical condition from biophysical signal data generated from sounds or other audio signals. In these instances, the disclosed systems and methods provide a cost-effective, non-invasive, low-risk intervention to subj ects in a point- of-care or home setting. Additionally or alternatively, the disclosed systems and methods can be implemented using a wearable patch or other wearable device with one or more channels and wearable elements, including shirts, watches, bands, and bracelets with sensors or elements capable of recording audio signals, and / or from and from implanted devices such as loop recorders, pacemakers, defibrillators, and / or digital stethoscopes. Classified feature data can be generated from the generated biophysical signal data, processed by a machine learning model or algorithm to generate classified feature data indicative of a particular medical condition, and allow for notifying the user or clinicians that a medical condition has been detected (e.g., via an alert or message).
[0028] Referring now to FIG. 1, a flowchart is illustrated as setting forth the steps of an example method for generating sound-based feature data using a suitably trained machine learning model. As will be described, the machine learning model takes biophysical signal data as input data and generates sound-based feature data as output data. As an example, the soundbased feature data can include sound waves, audio signals, and / or features generated from sound waves and / or audio signals.
[0029] The method includes accessing biophysical signal data with a computer system, as indicated at step 102. Accessing the biophysical signal data may include retrieving such data from a memory or other suitable data storage device or medium. Additionally or alternatively, accessing the biophysical signal data may include acquiring such data with a sensor or other measurement device and transferring or otherwise communicating the data to the computer system, which may be a part of the sensor and / or measurement device.
[0030] In some implementations, the biophysical signal data may additionally or alternatively include transformed biophysical signal data, such as scalogram data, spectrogram data, or other N-dimensional (for N > 2) images, maps, matrices, or data structures generated from biophysical signal data.
[0031] As a non-limiting example, scalogram data may include one or more scalograms generated from biophysical signal data. In general, a scalogram includes an image, map, orother N-dimensional matrix or data structure depicting the time-frequency distribution of biophysical signal data. In some implementations, a scalogram may be generated by computing a continuous wavelet transform (CWT) of the biophysical signal data and constructing the scalogram based on the CWT coefficients. As an example, the scalogram may be depicted as a heat map or other image.
[0032] As another example, spectrogram data may include one or more spectrograms generated from biophysical signal data. In general, a spectrogram includes an image, map, or other N-dimensional matrix or data structure depicting a spectrum of frequencies of biophysical signal signals in the biophysical signal data as they vary' with time. In some implementations, a spectrogram may be generated by computing a Fourier transform of the biophysical signals in the biophysical signal data and constructing the spectrogram based on the coefficients of the Fourier transform. As an example, the spectrogram may be depicted as a heat map or other image.
[0033] Additionally or alternatively, other images, maps, or other N-dimensional (for N > 2) matrices or other data structures may be generated from the biophysical signals in the biophysical signal data.
[0034] In some instances, the biophysical signal data can include ECG signal data acquired with a wearable device or an ECG system (e.g., an ECG measurement system using a 12-lead configuration, or other lead or electrode combination) and transferring or otherwise communicating the data to the computer system, which may be a part of the wearable device or ECG system.
[0035] The ECG signal data may include ECG signals. Additionally or alternatively, the ECG data may include variables, parameters, or other measurements that are computed, extracted, or otherwise derived from ECG signals. By way of example, the ECG data may include ECG measurements such as ventricular rate in beats per minute (bpm), atrial rate in bpm, P-R interval in milliseconds (ms), QRS duration in ms, Q-T interval in ms, QTC Bazetf s algorithm, P axis, R axis, T axis, QRS count, P-wave onset in median beat, P-wave offset in median beat, Q-onset in median beat, Q-offset in median beat. T-onset in median beat, T-offset in median beat, number of QRS Complexes, QRS duration, QT interval, QT corrected, PR interval, ventricular rate, average R-R Interval, Q-onset (median complex sample point), Q- offset (median complex sample point), P-onset (median complex sample point), P-offset (median complex sample point). T-onset (median complex sample point), QT calculated with the Frederica algorithm, P-wave amplitude at P-onset. P-wave amplitude, P-wave duration, P-wave area, P-wave intrinsicoid (time from P-onset to peak of P), P-prime amplitude, P-prime duration. P-prime area, P-prime intrinsicoid (time from P-onset to peak of P-prime), Q-wave amplitude, Q-wave duration, Q-wave area, Q intrinsicoid (time from Q-onset to peak of Q), R amplitude, R duration, R wave area, R intrinsicoid (time from R-onset to peak of R), S amplitude, S duration, S-wave area, S intrinsicoid (time from Q onset to peak of S), R-prime amplitude, R-prime duration, R-prime wave area. R-prime intrinsicoid (time from Q onset to peak of R-prime), S-prime amplitude, S-prime duration, S-prime wave area, S intrinsicoid (time from Q onset to peak of S-prime), STJ point, end of QRS point amplitude, STM point, middle of the ST segment amplitude, STE point, end of ST segment amplitude, maximum of STJ amplitude, maximum of STM amplitude, maximum of STE amplitude, minimum of STJ amplitude, minimum of STM amplitude, special T-wave amplitude, total QRS area. QRS deflection, maximum R amplitude (R or R-prime), maximum S amplitude (S or S prime), T amplitude, T duration, T-wave area, T intrinsicoid (time from STE to peak of T), T-prime amplitude, T-prime duration, T-prime area, T-prime intrinsicoid (time from STE to peak of T), T amplitude at T offset, P-wave area (includes P and P-prime), QRS area, T-wave area (includes T and T-prime), QRS intriniscoid, RR interval, PP interval, and so on.
[0036] By way of example, a 12-lead ECG system can include a I Lateral lead (also referred to as a l lead), a II Inferior lead (also referred to as a II lead), a III Inferior lead (also referred to as a III lead), an aVR lead, an aVL Lateral lead (also referred to as an aVL lead), an aVF Inferior lead (also referred to as an aVF lead), a VI Septal lead (also referred to as a VI lead), a V2 Septal lead (also referred to as a V2 lead), a V3 Anterior lead (also referred to as a V3 lead), a V4 Anterior lead (also referred to as a V4 lead), a V5 Lateral lead (also referred to as aV5 lead), and a V6 Lateral lead (also referred to as aV6 lead). Additionally or alternatively, the ECG system can implement fewer than 12 leads, such as a single lead, six leads (e.g.. all limb leads: I, II, III, avF, avR, AVL), or the like.
[0037] In some examples, the ECG signal data may be obtained using a radiofrequency (RF)-based sensor device. These RF-based sensors are capable of measuring ECG signals in addition to other biophysical signals (e.g.. heart beats, respirator}’ rates) using transmitted RF waves, which in some instances may include RF waves transmitted according to Wi-Fi or other wireless network protocol. Such sensors enable non-contact measurement of ECG signal data or other biophysical data. The RF-based sensors can be implemented in a standalone device, integrated into a mobile device (e.g., a smartphone, a tablet computer), integrated into a wearable device (e.g., a smartwatch, a fitness tracker, a wearable patch, a band, a bracelet).integrated into other wearables (e.g., a shirt or other wearable garment with conductive elements capable of recording physiologic signals), integrated into an implanted device (e.g., loop recorders, pacemakers, defibrillators), integrated into other medical devices (e g., digital stethoscopes), or integrated into another device or system (e.g., an automobile or other vehicle, such as an autonomous vehicle that can transport an individual to a clinic or hospital if a condition is detected). The biophysical signals captured from radiofrequencies and / or Wi-Fi can then be used with the systems and methods described in the present disclosure to improve on the diagnosis, management, prognostics, and / or treatment of respiratory illnesses.
[0038] Furthermore, radiofrequency signals can detect respiratory rates and other signals that can then be synchronously combined with ECG signal data, phonocardiogram (PCG) data, and / or continuous arterial blood pressure waveform data and used in the neural networks or other Al models described in the present disclosure.
[0039] Additionally or alternatively, other biophysical signal data can be accessed, including other electrophysiology data (e.g., EEG, data, EMG data), PPG data, etc. In some cases, other biophysical signal data may include echocardiogram data, such as echocardiograms, echocardiogram findings, echocardiogram variables, or the like.
[0040] A trained machine learning model is then accessed with the computer system, as indicated at step 104. In general, the machine learning model is trained, or has been trained, on training data in order to convert biophysical signal data to sound-based feature data (e.g., by synthesizing sound-based feature data from the input biophysical signal data, by estimating or predicting sound-based feature data based on the input biophysical signal data, etc.). Accessing the trained machine learning model may include accessing model parameters (e.g., weights, biases, or both) that have been optimized or otherwise estimated by training the machine learning model on training data. In some instances, retrieving the machine learning model can also include retrieving, constructing, or otherwise accessing the particular model architecture to be implemented. For instance, data pertaining to the layers in a neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be retrieved, selected, constructed, or otherwise accessed.
[0041] In some implementations the machine learning model may be a GAN. In general, a GAN includes two neural networks: a discriminator network and a generator network. An artificial neural network generally includes an input layer, one or more hidden layers (or nodes), and an output layer. Typically, the input layer includes as many nodes asinputs provided to the artificial neural network. The number (and the type) of inputs provided to the artificial neural network may vary based on the particular task for the artificial neural network.
[0042] The input layer connects to one or more hidden layers. The number of hidden layers varies and may depend on the particular task for the artificial neural network. Additionally, each hidden layer may have a different number of nodes and may be connected to the next layer differently. For example, each node of the input layer may be connected to each node of the first hidden layer. The connection between each node of the input layer and each node of the first hidden layer may be assigned a weight parameter. Additionally, each node of the neural network may also be assigned a bias value. In some configurations, each node of the first hidden layer may not be connected to each node of the second hidden layer. That is, there may be some nodes of the first hidden layer that are not connected to all of the nodes of the second hidden layer. The connections between the nodes of the first hidden layers and the second hidden layers are each assigned different weight parameters. Each node of the hidden layer is generally associated with an activation function. The activation function defines how the hidden layer is to process the input received from the input layer or from a previous input or hidden layer. These activation functions may vary and be based on the type of task associated with the artificial neural network and also on the specific type of hidden layer implemented.
[0043] Each hidden layer may perform a different function. For example, some hidden layers can be convolutional hidden layers which can, in some instances, reduce the dimensionality' of the inputs. Other hidden layers can perform statistical functions such as max pooling, which may reduce a group of inputs to the maximum value; an averaging layer; batch normalization; and other such functions. In some of the hidden layers each node is connected to each node of the next hidden layer, which may be referred to then as dense layers. Some neural networks including more than, for example, three hidden layers may be considered deep neural networks.
[0044] The last hidden layer in the artificial neural network is connected to the output layer. Similar to the input layer, the output layer typically has the same number of nodes as the possible outputs.
[0045] The biophysical signal data and / or transformed biophy sical signal data are then input to the machine learning model, generating output as sound-based feature data, as indicated at step 106. As described above, the sound-based feature data may include soundwaves, audio signals, and / or features generated from sound waves and / or audio signals. Features generated from sound waves and / or audio signals may include amplitude, frequency, and / or phase. Additionally or alternatively, sound-based features may include one or more characteristics or attributes generated from a sound wave and / or audio signal, such as timedomain features, frequency -domain features, or other features that can be computed or otherwise generated from a sound wave or audio signal. Ime-domain features may include amplitude envelope, zero-crossing rate, temporal moments (e.g., mean, variance, skewness, kurtosis), energy. RMS energy, or the like. Frequency-domain features may include spectral centroid, spectral bandwidth, spectral contrast, spectral flatness, or the like.
[0046] The sound-based feature data generated by inputting the biophysical signal data and / or transformed biophysical signal data to the trained machine learning model can then be displayed to a user, stored for later use or further processing, or both, as indicated at step 108.
[0047] As one example, the sound-based feature data can be used to construct a sound library or other training data set for training a machine learning model to generate classified feature data indicative of a medical condition in a subject. The sound library can be constructed, for example, by annotating or otherwise labeling the sound-based feature data and storing the sound-based feature data and corresponding annotations or labels as part of the sound library'.
[0048] As another example, the sound-based feature data can be analyzed to identify, detect, predict, or otherwise evaluate a medical condition or health status of the subject. The sound-based feature data may be analyzed using an audio analysis technique. Additionally or alternatively, the sound-based feature data may by input to a machine learning model that has been trained to generate classified feature data from input sound-based feature data, where the classified feature data are indicative of a medical condition or health status of the subject.
[0049] Referring now to FIG. 2. a flowchart is illustrated as setting forth the steps of an example method for training a machine learning model on training data, such that the machine learning model is trained to receive biophysical signal data as input data in order to generate sound-based feature data as output data, where the sound-based feature data include sound waves, audio signals, and / or features generated from sound waves and / or audio signals.
[0050] In general, the machine learning model can implement any number of different machine learning or other Al model architectures. As one example, the machine learning model may be a GAN. As another example, the machine learning model may be a transformer network or model. In still other examples, the machine learning model may otherwise implement oneor more neural network architectures. For instance, the neural network(s) could implement a convolutional neural network, a residual neural network, or the like.
[0051] The method includes accessing training data with a computer system, as indicated at step 202. Accessing the training data may include retrieving such data from a memory' or other suitable data storage device or medium. Alternatively, accessing the training data may include acquiring such data with sensors or other measurement devices and transferring or otherwise communicating the data to the computer system.
[0052] In general, the training data can include biophysical signals and sound data (e.g., sound waves, audio signals, and / or features generated from sound waves and / or audio signals) collected from a plurality of subjects. Additionally, the training data may include other data, such as transformed biophysical signal data or other health information collected from the subjects. In some embodiments, the training data may include biophysical signals, transformed biophysical signal data, and / or sound data that have been labeled (e.g., labeled as containing patterns, features, or characteristics indicative of one or more particular medical conditions or health statuses; and the like).
[0053] The method can include assembling training data from biophysical signals, transformed biophysical signal data, sound data, and / or other relevant data using a computer system. This step may include assembling the biophysical signals, transformed biophysical signal data, sound data, and / or other relevant data into an appropriate data structure on which the neural network or other machine learning algorithm can be trained. Assembling the training data may include assembling biophysical signals, transformed biophysical signal data, sound data, and / or other relevant data. For instance, assembling the training data may include generating labeled data and including the labeled data in the training data. Labeled data may include biophysical signals, transformed biophysical signal data, sound data, and / or other relevant data that have been labeled as belonging to, or otherwise being associated with, one or more different classifications or categories. For instance, labeled data may include biophysical signals, transformed biophysical signal data, sound data, and / or other relevant data that have been labeled as being associated with one or more particular medical conditions or health statuses.
[0054] One or more machine learning models are trained on the training data, as indicated at step 204. In general, the machine learning model can be trained by optimizing model parameters (e.g., weights, biases, or both) based on minimizing a loss function. As one non-limiting example, the loss function may be a mean squared error loss function.
[0055] As an example, training a neural network may include initializing the neural network, such as by computing, estimating, or otherwise selecting initial network parameters (e.g., weights, biases, or both). During training, an artificial neural network receives the inputs for a training example and generates an output using the bias for each node, and the connections between each node and the corresponding weights. For instance, training data can be input to the initialized neural network, generating output as sound-based feature data. The artificial neural network then compares the generated output with the actual output of the training example in order to evaluate the quality of the sound-based feature data. For instance, the sound-based feature data can be passed to a loss function to compute an error. The current neural network can then be updated based on the calculated error (e.g., using backpropagation methods based on the calculated error). For instance, the current neural network can be updated by updating the network parameters (e.g., weights, biases, or both) in order to minimize the loss according to the loss function. The training continues until a training condition is met. The training condition may correspond to, for example, a predetermined number of training examples being used, a minimum accuracy threshold being reached during training and validation, a predetermined number of validation iterations being completed, and the like. When the training condition has been met (e.g., by determining whether an error threshold or other stopping criterion has been satisfied), the current neural network and its associated network parameters represent the trained neural network. Different types of training processes can be used to adjust the bias values and the weights of the node connections based on the training examples. The training processes may include, for example, gradient descent, Newton's method, conjugate gradient, quasi-Newton, Levenberg-Marquardt, among others.
[0056] The artificial neural network can be constructed or otherwise trained based on training data using one or more different learning techniques, such as supervised learning, unsupervised learning, reinforcement learning, ensemble learning, active learning, transfer learning, or other suitable learning techniques for neural networks. As an example, supervised learning involves presenting a computer system with example inputs and their actual outputs (e.g., categorizations). In these instances, the artificial neural network is configured to leam a general rule or model that maps the inputs to the outputs based on the provided example inputoutput pairs.
[0057] As another example, the machine learning model may be a GAN. Training a GAN can include initializing the GAN (e.g.. initializing the generator and / or discriminator network), such as by computing, estimating, or otherwise selecting initial network parameters(e.g., weights, biases, or both). During training, the generator and discriminator update their weights one at a time in an adversarial manner. In this process, the discriminator is trained to detect the synthetic sound-based feature data. On the other hand, the generator is trained to minimize the loss between the synthetic sound-based feature data and ground truth sound data. Training is complete when an equilibrium is reached between the generator and discriminator losses.
[0058] The one or more trained machine learning models are then stored for later use, as indicated at step 206. Storing the neural network(s) may include storing model parameters (e.g., weights, biases, or both), which have been computed or otherwise estimated by training the machine learning model(s) on the training data. Storing the trained machine learning model(s) may also include storing the particular model architecture to be implemented. For instance, data pertaining to the layers in a neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be stored.
[0059] To ensure the Al and / or machine learning models remain effective and accurate over time, sequential model updates can be implemented. In these cases, a model is continuously retrained and / or updated using new data from specific time periods. This approach allows for monitoring how well the model may adapt to new changing patterns in biophysical signal data and / or transformed biophysical signal data. Additionally or alternatively, this approach allows for assessing the predictive performance of the model on evolving datasets. As a non-limiting example, the model may be initially trained on data from a defined period using the techniques described in the present disclosure. The defined period may be a period of days, weeks, month, years, or other rime scales. For example, the defined period may be a period of a few years, such as 2019-2022.
[0060] Performance metrics, including AUC (i.e., area under the curve for a receiver operating characteristic (ROC) curve), as well as its derivatives (e.g., sensitivity, specificity, negative predictive value (NPV), and positive predictive value (PPV)) at a predefined cutoff may be calculated to assess the ability of the model to convert biophysical signal data and / or transformed biophysical signal data into sound-based feature data.
[0061] Once the model is trained, it may be tested on new, unseen biophysical signal data and / or transformed biophysical signal data from a subsequent time period (e.g., subsequent day or days, subsequent week or weeks, subsequent month or months, subsequent year or years, other subsequent time scales). For example, when the model is trained on data from a periodof years such as 2019-2022, the pretrained model may be tested on new unseen biophysical signal data and / or transformed biophysical signal data from a subsequent year, such as 2023. This testing phase allows for the evaluation of the generalizability of the model. Additionally or alternatively, this testing phase may be used to assess whether the model maintains its predictive accuracy when applied to a different time period. Performance metrics can be composed across the datasets from the different time periods (e.g., 2019-2022 training dataset and 2023 test set, in the described example) to evaluate how well the model adapts to potential shifts in biophysical signal data and / or transformed biophysical signal data patterns and to identify any degradation in model performance.
[0062] As additional data become available, the model may be retrained using an expanding dataset, for example, including data from 2019 to 2023 to forecast performance for 2024. That is, the subsequent data set used to test the model may be appended or otherwise concatenated with the original training data set and the updated model may then be trained on a newer subsequent data set associated with another subsequent time period. The model will again be tested using the newer subsequent data set (e.g., 2024 data, in the described example), and performance metrics will be recalculated to ensure continued accuracy and robustness of the model. This process may be repeated annually, monthly, weekly, or over other time scales, with the training dataset progressively growing to include more recent data, ensuring that the model stays current with emerging trends in biophysical signal data and / or transformed biophysical signal data.
[0063] Throughout this process, the aging of the model may be monitored by comparing its ROC curves over time. By visualizing how the ROC curve evolves with each successive test period, any changes in the ability of the model to distinguish between classes can be observed. A decline in the AUC, as a non-limiting example, may signal potential performance issues, while improvements in the AUC may indicate that the model is adapting well to new data. Additionally or alternatively, subgroup analysis can be conducted to assess whether any biases arise as the model encounters different patient populations or demographic shifts over time.
[0064] To maintain consistent performance and mitigate any biases, the model may be periodically fine-timed, retrained, and adjusted to incorporate the latest data. This iterative process ensures that the model remains aligned with clinical needs and continues to provide reliable conversion of biophysical signal data and / or transformed biophysical signal data to sound-based feature data, even as the data evolves.
[0065] Referring now to FIG. 3. a flowchart is illustrated as setting forth the steps of an example method for generating biophysical signal data using a suitably trained machine learning model. As will be described, the machine learning model takes sound-based feature data (sound waves, audio signals, and / or features generated from sound waves and / or audio signals) as input data and generates biophysical signal data as output data. As an example, the biophysical signal data can include ECG signals, EEG signals, EMG signals, and / or PPG signals. In some cases, other biophysical signal data may include echocardiogram data, such as echocardiograms, echocardiogram findings, echocardiogram variables, or the like. Additionally or alternatively, the biophysical signal data can include transformed biophysical signal data, including scalogram data, spectrogram data, or other N-dimensional (for N > 2) images, maps, matrices, or data structures generated from biophysical signal data.
[0066] The method includes accessing sound-based feature data with a computer system, as indicated at step 302. Accessing the sound-based feature data may include retrieving such data from a memory or other suitable data storage device or medium. Additionally or alternatively, accessing the sound-based feature data may include acquiring such data with an audio device (e.g., a microphone, a digital stethoscope, an ultrasound system) and transferring or otherwise communicating the data to the computer system.
[0067] A trained neural network (or other suitable machine learning algorithm) is then accessed with the computer system, as indicated at step 304. In general, the neural network is trained, or has been trained, on training data in order to generate biophysical signal data from sound-based feature data (e.g., by synthesizing biophysical signal data and / or transformed biophysical signal data from the input sound-based feature data, by estimating or predicting biophysical signal data and / or transformed biophysical signal data based on the input soundbased feature data. etc.). Accessing the trained machine learning model may include accessing model parameters (e g., weights, biases, or both) that have been optimized or otherwise estimated by training the machine learning model on training data. In some instances, retrieving the machine learning model can also include retrieving, constructing, or otherwise accessing the particular model architecture to be implemented. For instance, data pertaining to the layers in a neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be retrieved, selected, constructed, or otherwise accessed.
[0068] In some implementations the machine learning model may be a GAN. In general, a GAN includes two neural networks: a discriminator network and a generatornetwork. As another example, the machine learning model may be a transformer network or model. In still other examples, the machine learning model may otherwise implement one or more neural network architectures. For instance, the neural network(s) could implement a convolutional neural network, a residual neural network, or the like.
[0069] The sound-based feature data are then input to the machine learning model, generating output as biophysical signal data and / or transformed biophysical signal data, as indicated at step 306. For example, the biophysical signal data may include ECG signal data, EEG signal data, EMG signal data, PPG signal data, or the like. Transformed biophysical signal data may include scalogram data, spectrogram data, or other N-dimensional (for N > 2) images, maps, matrices, or data structures that represent a transformation of underlying biophysical signal data.
[0070] The biophysical signal data and / or transformed biophysical signal data generated by inputting the sound-based featured data to the trained machine learning model(s) can then be displayed to a user, stored for later use or further processing, or both, as indicated at step 308. As an example, the biophysical signal data and / or transformed biophysical signal data can be analyzed to identify, detect, predict, or otherwise evaluate a medical condition or health status of the subject. The biophysical signal data and / or transformed biophysical signal data may be analyzed, for example, by inputting the biophysical signal data and / or transformed biophysical signal data to a machine learning model that has been trained to generate classified feature data from input biophysical signal data and / or transformed biophysical signal data, where the classified feature data are indicative of a medical condition or health status of the subject.
[0071] Referring now to FIG. 4. a flowchart is illustrated as setting forth the steps of an example method for training a machine learning model on training data, such that the machine learning model is trained to receive sound-based feature data as input data in order to generate biophysical signal data and / or transformed biophysical signal data as output data, where the biophysical signal data can include ECG signal data, EEG signal data. EMG signal data, PPG signal data, or the like, and the transformed biophysical signal data may include scalogram data, spectrogram data, or other N-dimensional (for N > 2) images, maps, matrices, or data structures that represent a transformation of underlying biophysical signal data.
[0072] In general, the machine learning model can implement any number of different machine learning or other Al model architectures. As one example, the machine learning model may be a GAN. As another example, the machine learning model may be a transformer networkor model. In still other examples, the machine learning model may otherwise implement one or more neural network architectures. For instance, the neural network(s) could implement a convolutional neural network, a residual neural network, or the like.
[0073] The method includes accessing training data with a computer system, as indicated at step 402. Accessing the training data may include retrieving such data from a memory or other suitable data storage device or medium. Alternatively, accessing the training data may include acquiring such data with an audio device (e.g., a microphone, a digital stethoscope, an ultrasound system) and transferring or otherwise communicating the data to the computer system.
[0074] In general, the training data can include sound-based feature data (e.g., sound waves, audio signals, and / or features generated from sound waves and / or audio signals) and biophysical signals (ECG signals, EEG signals, EMG signals, PPG signals) and / or transformed biophysical signal data collected from a plurality of subjects. Additionally, the training data may include other data, such as other health information collected from the subjects. In some embodiments, the training data may include sound-based feature data, biophysical signal data, and / or transformed biophysical signal data that have been labeled (e.g., labeled as containing patterns, features, or characteristics indicative of a particular medical condition of health status; and the like).
[0075] The method can include assembling training data from sound-based feature data, biophysical signal data, and / or transformed biophysical signal data using a computer system. This step may include assembling the sound-based feature data, biophysical signal data, and / or transformed biophysical signal data into an appropriate data structure on which the neural network or other machine learning algorithm can be trained. Assembling the training data may include assembling sound-based feature data, biophysical signal data, and / or transformed biophysical signal data and other relevant data. For instance, assembling the training data may include generating labeled data and including the labeled data in the training data. Labeled data may include sound-based feature data, biophysical signal data, and / or transformed biophysical signal data other relevant data that have been labeled as belonging to, or otherwise being associated with, one or more different classifications or categories. For instance, labeled data may include sound-based feature data, biophysical signal data, and / or transformed biophysical signal data that have been labeled as being associated with a particular medical condition and / or health status.
[0076] One or more machine learning models are trained on the training data, as indicated at step 404. In general, the machine learning model can be trained by optimizing model parameters (e.g., weights, biases, or both) based on minimizing a loss function. As one non-limiting example, the loss function may be a mean squared error loss function.
[0077] As an example, training a neural network may include initializing the neural network, such as by computing, estimating, or otherwise selecting initial network parameters (e.g., weights, biases, or both). During training, an artificial neural network receives the inputs for a training example and generates an output using the bias for each node, and the connections between each node and the corresponding w eights. For instance, training data can be input to the initialized neural network, generating output as biophysical signal data. The artificial neural network then compares the generated output with the actual output of the training example in order to evaluate the quality of the biophysical signal data. For instance, the biophysical signal data can be passed to a loss function to compute an error. The current neural network can then be updated based on the calculated error (e.g., using backpropagation methods based on the calculated error). For instance, the current neural network can be updated by updating the network parameters (e.g., weights, biases, or both) in order to minimize the loss according to the loss function. The training continues until a training condition is met. The training condition may correspond to, for example, a predetermined number of training examples being used, a minimum accuracy threshold being reached during training and validation, a predetermined number of validation iterations being completed, and the like. When the training condition has been met (e g., by determining whether an error threshold or other stopping criterion has been satisfied), the current neural netw ork and its associated netw ork parameters represent the trained neural network. Different ty pes of training processes can be used to adjust the bias values and the weights of the node connections based on the training examples. The training processes may include, for example, gradient descent, Newton's method, conjugate gradient, quasi -Newton. Levenberg-Marquardt, among others.
[0078] The artificial neural netw ork can be constructed or otherwise trained based on training data using one or more different learning techniques, such as supervised learning, unsupervised learning, reinforcement learning, ensemble learning, active learning, transfer learning, or other suitable learning techniques for neural netw orks. As an example, supervised learning involves presenting a computer system with example inputs and their actual outputs (e.g., categorizations). In these instances, the artificial neural network is configured to leam ageneral rule or model that maps the inputs to the outputs based on the provided example inputoutput pairs.
[0079] As another example, the machine learning model may be a GAN. Training a GAN can include initializing the GAN (e.g., initializing the generator and / or discriminator network), such as by computing, estimating, or otherwise selecting initial network parameters (e.g., weights, biases, or both). During training, the generator and discriminator update their weights one at a time in an adversarial manner. In this process, the discriminator is trained to detect the synthetic biophysical signal data. On the other hand, the generator is trained to minimize the loss between the synthetic biophysical signal data and ground truth biophysical signal data. Training is complete when an equilibrium is reached between the generator and discriminator losses.
[0080] The one or more trained machine learning models are then stored for later use, as indicated at step 406. Storing the neural network(s) may include storing model parameters (e.g., weights, biases, or both), which have been computed or otherwise estimated by training the machine learning model(s) on the training data. Storing the trained machine learning model(s) may also include storing the particular model architecture to be implemented. For instance, data pertaining to the layers in a neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) maybe stored.
[0081] To ensure the Al and / or machine learning models remain effective and accurate over time, sequential model updates can be implemented. In these cases, a model is continuously retrained and / or updated using new data from specific time periods. This approach allows for monitoring how well the model may adapt to new changing patterns in sound-based feature data. Additionally or alternatively, this approach allows for assessing the predictive performance of the model on evolving datasets. As a non-limiting example, the model may be initially trained on data from a defined period using the techniques described in the present disclosure. The defined period may be a period of days, weeks, month, years, or other time scales. For example, the defined period may be a period of a few years, such as 2019-2022.
[0082] Performance metrics, including AUC (i.e., area under the curve for a receiver operating characteristic (ROC) curve), as well as its derivatives (e.g., sensitivity-, specificity-, negative predictive value (NPV), and positive predictive value (PPV)) at a predefined cutoffmay be calculated to assess the ability of the model to convert sound-based feature data to biophysical signal data and / or transformed biophysical signal data.
[0083] Once the model is trained, it may be tested on new, unseen sound-based feature data from a subsequent time period (e.g., subsequent day or days, subsequent week or weeks, subsequent month or months, subsequent year or years, other subsequent time scales). For example, when the model is trained on data from a period of years such as 2019-2022, the pretrained model may be tested on new unseen sound-based feature data from a subsequent year, such as 2023. This testing phase allows for the evaluation of the generalizability of the model. Additionally or alternatively, this testing phase may be used to assess whether the model maintains its predictive accuracy when applied to a different time period. Performance metrics can be composed across the datasets from the different time periods (e.g., 2019-2022 training dataset and 2023 test set, in the described example) to evaluate how well the model adapts to potential shifts in sound-based feature data patterns and to identify any degradation in model performance.
[0084] As additional data become available, the model may be retrained using an expanding dataset, for example, including data from 2019 to 2023 to forecast performance for 2024. That is, the subsequent data set used to test the model may be appended or otherwise concatenated with the original training data set and the updated model may then be trained on a newer subsequent data set associated with another subsequent time period. The model will again be tested using the newer subsequent data set (e.g., 2024 data, in the described example), and performance metrics will be recalculated to ensure continued accuracy and robustness of the model. This process may be repeated annually, monthly, weekly, or over other time scales, with the training dataset progressively growing to include more recent data, ensuring that the model stays current with emerging trends in sound-based feature data.
[0085] Throughout this process, the aging of the model may be monitored by comparing its ROC curves over time. By visualizing how the ROC curve evolves with each successive test period, any changes in the ability of the model to distinguish between classes can be observed. A decline in the AUC, as a non-limiting example, may signal potential performance issues, while improvements in the AUC may indicate that the model is adapting well to new data. Additionally or alternatively, subgroup analysis can be conducted to assess whether any biases arise as the model encounters different patient populations or demographic shifts over time.
[0086] To maintain consistent performance and mitigate any biases, the model may be periodically fine-tuned, retrained, and adjusted to incorporate the latest data. This iterative process ensures that the model remains aligned with clinical needs and continues to provide reliable conversion of sound-based feature data to biophysical signal data and / or transformed biophysical signal data, even as the data evolves.
[0087] Referring now to FIG. 5. a flowchart is illustrated as setting forth the steps of an example method for generating classified feature data using a suitably trained neural network or other machine learning algorithm (e.g., large language model, generative pretrained transformer model, or the like). As will be described, the neural network or other machine learning algorithm takes sound-based feature data generated from biophysical signals as input data and generates classified feature data as output data. As an example, the classified feature data can be indicative of the presence of a medical condition or health status.
[0088] The method includes accessing sound-based feature data with a computer system, as indicated at step 502. Accessing the sound-based feature data may include retrieving such data from a memory or other suitable data storage device or medium. Additionally or alternatively, accessing the sound-based feature data may include generating such data using the methods described herein and transferring or otherwise communicating the data to the computer system.
[0089] In still other examples, additional data may be accessed, such as patient health data. The patient health data may include data stored in, retrieved from, extracted from, or otherwise derived from the patient’s electronic medical record (EMR) and / or electronic health record (EHR). The patient health data can include unstructured text, questionnaire response data, clinical laboratory data, histopathology data, genetic sequencing, medical imaging, and other such clinical data types. Examples of clinical laboratory data and / or histopathology data can include genetic testing and laboratory information, such as performance scores, lab tests, pathology results, prognostic indicators, date of genetic testing, testing method used, and so on.
[0090] Patient health data can include a set of clinical features associated with information derived from clinical records of a patient, which can include records from family members of the patient. These clinical features and data may be abstracted from unstructured clinical documents, EMR, EHR, or other sources of patient history . Such data may include patient symptoms, diagnosis, treatments, medications, therapies, responses to treatments, laboratory testing results, medical history, geographic locations of each, demographics, or otherfeatures of the patient which may be found in the patient’s EMR and / or EHR. For example, features derived from structured, curated, and / or EMR or EHR data may include clinical features such as diagnoses; symptoms; therapies; outcomes; patient demographics, such as patient name, date of birth, gender, and / or ethnicity; diagnosis dates for cancer, illness, disease, or other physical or mental conditions; personal medical history: family medical history; clinical diagnoses, such as date of initial diagnosis; and the like. Additionally, the patient health data may also include features such as treatments and outcomes, such as line of therapy, therapy groups, clinical trials, medications prescribed or taken, surgeries, imaging, adverse effects, and associated outcomes.
[0091] The patient health data may also include measurement data collected from wearable devices (e.g., physiological measurements or other data recorded with a wearable device). Physiological measurements that may be recorded with a wearable device include heart rate, temperature, or other physical parameters.
[0092] The patient health data may also include epidemiological data on the prevalence and incidence of relevant diseases, such as weekly observed incidence of new cardiac disease cases, which may change over time. This epidemiological data can be sourced from public health records, patient registries, and other relevant databases. Integrating these disease trends into the model can help ensure that the model is learning from the shifting disease landscape, thereby improving its ability to capture evolving patterns in sound-based feature data tied to specific health conditions.
[0093] A trained neural network (or other suitable machine learning algorithm) is then accessed with the computer system, as indicated at step 504. In general, the neural network is trained, or has been trained, on training data in order to detect, identify, or otherwise characterize paterns in sound-based feature data that are indicative of a medical condition or health status in the subject from whom the biophysical signals used to generate the sound-based feature data were acquired.
[0094] Accessing the trained neural network may include accessing network parameters (e.g., weights, biases, or both) that have been optimized or otherwise estimated by training the neural network on training data. In some instances, retrieving the neural network can also include retrieving, constructing, or otherwise accessing the particular neural network architecture to be implemented. For instance, data pertaining to the layers in the neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections betweenlayers, hyperparameters for layers) may be retrieved, selected, constructed, or otherwise accessed.
[0095] The sound-based feature data are then input to the one or more trained neural networks, generating output as classified feature data, as indicated at step 506. For example, the classified feature data may include a risk score. The risk score can provide physicians or other clinicians with a recommendation to consider additional monitoring for subjects whose sound-based feature data indicate the likelihood of the subject suffering from a particular medical condition.
[0096] As another example, the classified feature data may indicate the probability for a particular classification (i.e., the probability’ that the sound-based feature data include patterns, features, or characteristics indicative of detecting, differentiating, and / or determining the severity of one or more medical conditions).
[0097] Additionally or alternatively, the classified feature data may classify the soundbased feature data as indicating a particular medical condition. The classified feature data may include a single classification output, or may include multiple classification outputs (e.g., an indication of more than one medical condition being positive and / or negative). In these instances, the classified feature data can differentiate between different medical conditions. In still other embodiments, the classified feature data may indicate a severity' of a medical condition. For example, the classified feature data may include a severity score that quantifies a severity’ of a medical condition.
[0098] The classified feature data generated by inputting the sound-based feature data to the trained neural network(s) can then be displayed to a user, stored for later use or further processing, or both, as indicated at step 508.
[0099] Referring now to FIG. 6. a flowchart is illustrated as setting forth the steps of an example method for generating classified feature data using a suitably trained neural network or other machine learning algorithm (e.g., large language model, generative pretrained transformer model, or the like). As will be described, the neural netw ork or other machine learning algorithm takes biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data as input data and generates classified feature data as output data. As an example, the classified feature data can be indicative of the presence of a medical condition or health status. In some implementations, the medical condition may include a heart murmur, a cardiac arrhythmia, a left bundle branch block, or other such cardiac-related medical conditions. Additionally or alternatively, the medical condition may include a respiratory condition, or a condition with another internal organ or organ system.
[0100] The method includes accessing sound-based feature data with a computer system, as indicated at step 602. Accessing the sound-based feature data may include retrieving such data from a memory or other suitable data storage device or medium. Additionally or alternatively, accessing the sound-based feature data may include acquiring such data with an audio device (e.g., a microphone, a digital stethoscope, an ultrasound system) and transferring or otherwise communicating the data to the computer system.
[0101] In still other examples, additional data may be accessed, such as patient health data. The patient health data may include data stored in, retrieved from, extracted from, or otherwise derived from the patient's electronic medical record (EMR) and / or electronic health record (EHR). The patient health data can include unstructured text, questionnaire response data, clinical laboratory data, histopathology data, genetic sequencing, medical imaging, and other such clinical data types. Examples of clinical laboratory data and / or histopathology data can include genetic testing and laboratory information, such as performance scores, lab tests, pathology results, prognostic indicators, date of genetic testing, testing method used, and so on.
[0102] Patient health data can include a set of clinical features associated with information derived from clinical records of a patient, which can include records from family members of the patient. These clinical features and data may be abstracted from unstructured clinical documents, EMR, EHR, or other sources of patient history. Such data may include patient symptoms, diagnosis, treatments, medications, therapies, responses to treatments, laboratory testing results, medical history , geographic locations of each, demographics, or other features of the patient which may be found in the patient’s EMR and / or EHR. For example, features derived from structured, curated, and / or EMR or EHR data may' include clinical features such as diagnoses; symptoms; therapies; outcomes; patient demographics, such as patient name, date of birth, gender, and / or ethnicity ; diagnosis dates for cancer, illness, disease, or other physical or mental conditions; personal medical history’; family medical history; clinical diagnoses, such as date of initial diagnosis; and the like. Additionally, the patient health data may also include features such as treatments and outcomes, such as line of therapy, therapy groups, clinical trials, medications prescribed or taken, surgeries, imaging, adverse effects, and associated outcomes.
[0103] The patient health data may also include measurement data collected from wearable devices (e.g., physiological measurements or other data recorded with a wearable device). Physiological measurements that may be recorded with a wearable device include heart rate, temperature, or other physical parameters.
[0104] The patient health data may also include epidemiological data on the prevalence and incidence of relevant diseases, such as weekly observed incidence of new cardiac disease cases, which may change over time. This epidemiological data can be sourced from public health records, patient registries, and other relevant databases. Integrating these disease trends into the model can help ensure that the model is learning from the shifting disease landscape, thereby improving its ability to capture evolving patterns in sound-based feature data tied to specific health conditions.
[0105] The sound-based feature data are then converted to biophysical signal data and / or transformed biophysical signal data, as indicated at step 604. For example, the soundbased feature data can be converted to biophysical signal data and / or transformed biophysical signal data using the methods described herein (e.g., the method of FIG. 4). As described above, in some implementations, the biophysical signal data may additionally or alternatively include transformed biophysical signal data, such as scalogram data, spectrogram data, or other N- dimensional (for N > 2) images, maps, matrices, or data structures generated from biophysical signal data or otherwise representative of a transformation of underlying biophysical signal data.
[0106] As a non-limiting example, scalogram data may include one or more scalograms generated from biophysical signal data. In general, a scalogram includes an image, map, or other N-dimensional matrix or data structure depicting the time-frequency distribution of biophysical signal data. In some implementations, a scalogram may be generated by computing a continuous wavelet transform (CWT) of the biophysical signal data and constructing the scalogram based on the CWT coefficients. As an example, the scalogram may be depicted as a heat map or other image.
[0107] As another example, spectrogram data may include one or more spectrograms generated from biophysical signal data. In general, a spectrogram includes an image, map, or other N-dimensional matrix or data structure depicting a spectrum of frequencies of biophysical signal signals in the biophysical signal data as they vary' with time. In some implementations, a spectrogram may be generated by computing a Fourier transform of the biophysical signals in the biophysical signal data and constructing the spectrogram based on the coefficients of theFourier transform. As an example, the spectrogram may be depicted as a heat map or other image.
[0108] Additionally or alternatively, other images, maps, or other N-dimensional (for N > 2) matrices or other data structures may be generated from the biophysical signals in the biophysical signal data.
[0109] A trained neural network (or other suitable machine learning algorithm) is then accessed with the computer system, as indicated at step 606. In general, the neural network is trained, or has been trained, on training data in order to detect, identify, or otherwise characterize patterns in biophysical signal data and / or transformed biophysical signal data that are indicative of a medical condition or health status in the subject from whom the sound-based feature data used to generate the biophysical signal data were acquired.
[0110] Accessing the trained neural network may include accessing network parameters (e.g., weights, biases, or both) that have been optimized or otherwise estimated by training the neural network on training data. In some instances, retrieving the neural network can also include retrieving, constructing, or otherwise accessing the particular neural network architecture to be implemented. For instance, data pertaining to the layers in the neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be retrieved, selected, constructed, or otherwise accessed.
[0111] The biophysical signal data and / or transformed biophysical signal data are then input to the one or more trained neural networks, generating output as classified feature data, as indicated at step 608. For example, the classified feature data may include a risk score. The risk score can provide physicians or other clinicians with a recommendation to consider additional monitoring for subjects whose biophysical signal data and / or transformed biophysical signal data indicate the likelihood of the subj ect suffering from a particular medical condition.
[0112] As another example, the classified feature data may indicate the probability for a particular classification (i.e., the probability that the biophysical signal data and / or transformed biophysical signal data include patterns, features, or characteristics indicative of detecting, differentiating, and / or determining the severity of one or more medical conditions).
[0113] Additionally or alternatively, the classified feature data may classify the biophysical signal data and / or transformed biophysical signal data as indicating a particular medical condition. The classified feature data may include a single classification output, ormay include multiple classification outputs (e.g., an indication of more than one medical condition being positive and / or negative). In these instances, the classified feature data can differentiate between different medical conditions. In still other embodiments, the classified feature data may indicate a severity of a medical condition. For example, the classified feature data may include a severity' score that quantifies a severity' of a medical condition.
[0114] The classified feature data generated by inputting the biophysical signal data and / or transformed biophysical signal data to the trained neural netyvork(s) can then be displayed to a user, stored for later use or further processing, or both, as indicated at step 610.
[0115] In some embodiments, foundation models can be used to additionally or alternatively process the sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data. Foundation models can receive various t pes of data (e.g., text data, image data, sound data, other ID, 2D, and / or 3D data types) and generate various types of clinically relevant outputs. For example, the foundation models may generate outputs as predictive scores, text outputs, classifications, and so on. As a non-limiting example, text outputs may include ansyvers to questions posed by a clinical user (e.g., medical question ansyvering), interpretive reports of sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data, or other text-based reports and / or summaries of the input data. Beyond giving vital measurements and diagnostic capabilities, generating reports across temporal and spatial domains yvith the sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data is advantageous for identifying if a patient’s vitals and / or conditions are improving, worsening, or staying the same.
[0116] As one example, a language model, such as a large language model (LLM), can be used. Large language modeling with electrocardiograms and medical images involves using machine learning models (e.g., deep learning models) to process and analyze medical images and associated clinical text data. This approach can combine natural language processing (NLP) with computer vision to understand the content of medical images or other physiologic signal data (e.g., sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data) and extract relevantinformation from them. Advantageously, large language modeling can be applied to soundbased feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data, and / or medical images to enable automatic analysis of the input data and to extract clinically relevant information, such as the presence of respiratory illness or abnormalities. This approach can help healthcare professionals make more accurate diagnoses and develop personalized treatment plans for patients.
[0117] One example method for large language modeling with sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data uses CNNs to extract features from the input data, which are then fed into a recurrent neural network (RNN) to generate text descriptions of the sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data. The RNN can also be used to generate clinical reports based on the input data. Another example approach is to use a transformer-based model, such as a BERT or GPT-4 model, to analyze both the sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data and accompanying text data. These models can be pre-trained on large datasets of sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data and associated text to improve their accuracy and ability to identify relevant features in the input data.
[0118] As a non-limiting example, the sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data can be provided as an input to the LLM. The sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data are first tokenized and converted into a numerical format. The tokenized input data may then be applied to an embedding layer to transform each tokenized input into a high-dimensional vector that captures relationships between the sound-based feature data generated from biophysical signal dataand / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data. In some instances, the embedding layer may additionally or alternatively capture semantic relationships between text data and the sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data. The resulting embedded vectors form a one-dimensional sequence that will be input to the LLM. Each element of the embedded vectors corresponds to a token in the tokenized input data.
[0119] Any suitable LLM can be used. As one non-limiting example, the LLM may be based on a recurrent layer model, such as a long short term memory (LSTM) model. As another non-limiting example, the LLM may be based on an attention mechanism (e.g., transformer). For instance, the LLM may be based on a generative pre-trained transformer (GPT) model. In still other examples, the LLM may be based on combinations of such model types.
[0120] GPT is a type of large language model that is pre-trained on a massive corpus of text data and can generate human-like language. Language models can also be trained on data from other modalities (e.g., images, audio recordings, videos, etc.) to enable more diverse capabilities, provide a stronger learning signal, and increase learning speed. The GPT model is based on the transformer architecture, which allows it to process long sequences of text efficiently. GPT can be used in a wide range of natural language processing (NLP) tasks, including language translation, text summarization, and question answering.
[0121] Sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data contain valuable information that can aid in the diagnosis and treatment of various medical conditions. By combining GPT- based language modeling with the analysis of sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data, the disclosed systems and methods can analyze the input sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data to generate textual reports summarizing the findings. For example, a GPT-based model can be trained on a large corpus of sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformedbiophysical signal data generated from sound-based feature data and medical reports and then used to generate reports for new sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data. The generated reports can include information such as the location and size of any abnormalities associated with a respiratory’ illness, as well as recommendations for further testing or treatment.
[0122] These LLMs can also be used to develop a chatbot based on a foundation model that can serve as a physician’s assistant to support more accurate diagnosis and tailored therapy selection. These capabilities can improve the accuracy and efficiency of patient care while increasing patient engagement and adherence to therapy. Once diagnosis is determined as a result, the output can be incorporated into the patient’s medical documentation through electronic medical records.
[0123] Once diagnosis is made with these models, clinical documents can be generated as described above. A foundation model can then generate tailored patient education materials and explain their care plan at the appropriate reading level based on the clinical documents. The models can also be used to draft a clinic note in real-time based on the results. As another advantage, the models can also be used to optimize clinic scheduling or to simplify generation of medical codes for billing (e.g., current procedural terminology (CPT) codes), disease surveillance, and even automated follow-up reminders.
[0124] In some instances, the code to these models can be automated and / or updated using an LLM, GPT, or the like. For example, auto-GPT can be used to write and update its own code and execute scripts. This allows the model to recursively debug, develop, and selfimprove. While input data are applied to an auto-GPT-based model, the model can update itself automatically.
[0125] As described herein, multimodal language modeling can be used to receive multiple sources of information to train a language model. In the context of sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data, this can mean combining textual information from clinical notes, laboratories, or reports with visual information from sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data, or other sources, including for example medical images such as x-rays, CT scans, or MRIs. One approach tomultimodal language modeling is to use the GPT architecture, which can be effective in a variety of NLP tasks. The GPT architecture is based on a transformer network, which can leam to model long-range dependencies between words in a sentence. To adapt GPT for multimodal language modeling with sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data, visual information can be incorporated into the GPT model through pretraining with contrastive learning. This process involves training the model to predict which in put data (e.g., sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data) and text pairs are related, while also ensuring that unrelated pairs are distinguishable from each other. The resulting model can then be fine-tuned for specific tasks, such as biophysical signal captioning, audio signal captioning, sound library captioning, medical report generation, or disease diagnosis. For example, a model trained on clinical notes and sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data could be fine-tuned to generate reports to predict the presence of certain medical conditions based on the sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data.
[0126] As another example, a visual transformer-based model can be used. In these instances, higher-dimensional inputs can be provided to the underlying foundation model. For example, the sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data may be tokenized and embedded, as described above, into a 2D or other N-dimensional (N > 1) vector, matrix, or tensor. These higher dimensional input data can then be applied to a suitable foundation model, such as a visual transformer-based model.
[0127] Vision-language processing can be improved by using paired samples sharing semantics. For instance, given an image and text pair, the text should describe the image with minimal extraneous detail. Without knowledge of this initial image, temporal information in the text modality (e.g., "‘condition is improving,” “condition is worsening,” “condition isstable’') could pertain to any image including “condition"’, creating vagueness during contrastive training. Vision-language processing implementations can, in some instances, assume alignment between single images and reports, removing temporal content from reports in training data to prevent hallucinations in downstream report generation.
[0128] In other implementations of vision-language processing, temporal information can provide complementary self-supervision by using an existing structure without requiring additional data. Rather than treating all image-report pairs in the dataset as independent, temporal correlations can be used by making previous images available for comparison to a given report. A temporal vision-language processing pre-training framework can be learned from this structure. In some implementations, a multi-image encoder that can handle the absence of previous images and potential spatial misalignment between images across time can be used in this vision-language processing. Prior images can be accounted for where available, thereby removing cross-modal ambiguity. Linking multiple images has the advantage of improving image and text models and performance on temporal image classification and report generation. Prefixing the prior report can significantly improve performance. When available during training and fine-tuning, earlier images and labels can also be accounted for.
[0129] As one non-limiting example, a convolutional neural network (CNN) or LLM visual transformer hybrid multi-image encoder can be trained jointly with a text model. The hybrid model can provide improved processing for tasks in both single-image and multi-image structures, achieving performance on disease progression classification, phrase grounding, and document generation, while offering consistent modifications on disease category and sentence-similarity7tasks. The similarity7between the sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data, or other signals and text embeddings can be computed to obtain probabilities, which can be used to classify the various categories of the sound-based feature data generated from biophysical signal data and / or transformed biophysical signal data, or the biophysical signal data and / or transformed biophysical signal data generated from sound-based feature data and then reported textually. The outputs can also include classified feature data indicating risk categories of no risk, low7risk, medium risk, and high risk. These outputs can be displayed into risk and probability categories where the model draw s more attention for action or testing if it falls into a high risk category versus normal.
[0130] As yet another example, a sound transformer-based model can be used. In these instances, the sound-based feature data (or biophysical signal data converted to an audio data format) may be input to the sound transformer-based model. As above, these audio data can be tokenized and embedded into one-dimensional embedded vectors that are applied to the sound transformer-based model. Additional audio data (e.g., stethoscope recordings) may also be input to the sound transformer-based model.
[0131] Referring now to FIGS. 7 and 8, an example of a wearable device 800 for recording ECG signals, EEG signals, EMG signals, PPG signals, other physiological signals, or other biophysical signals and / or generating classified feature data in accordance with some embodiments is shown.
[0132] As shown in FIG. 7, the wearable device 800 can include a device that is configured to be worn on a user’s wrist or limb (e.g., a smart watch, a band, a bracelet), placed on a user’s skin (e.g., a wearable patch), or worn as an article of clothing (e.g., a shirt). The wearable device 800 can be in communication with an external device 810 and / or a server 812 either directly or indirectly via a network 808.
[0133] The network 808 may be a long-range wireless network such as the Internet, a local area network (LAN), a wide area network (WAN), or a combination thereof. In other embodiments, the network 808 may be a short-range wireless communication network, and in yet other embodiments, the network 808 may be a wired network using, for example, USB cables. In some embodiments, the network 808 may include both wired and wireless devices and connections. Similarly, the server 812 may transmit information to the external device 810 to be forwarded to the wearable device 800.
[0134] In some embodiments, the wearable device 800 communicates directly with the external device 810. For example, the wearable device 800 can transmit data (e.g., physiological sensor data, other data collected or generated by the wearable device 800) to the external device 810. Similarly, the wearable device 800 can receive data (e.g., settings, machine learning algorithm parameters, firmware updates, etc.) from the external device 810.
[0135] In some other embodiments, the wearable device 800 bypasses the external device 810 to access the network 808 and communicate with the server 812 via the network 808. In some embodiments, the wearable device 800 is equipped with a long-range transceiver instead of or in addition to a short-range transceiver. In such embodiments, the wearable device 800 communicates directly with the server 812 or with the server 812 via the network 808 (in either case, bypassing the external device 810).
[0136] In some embodiments, the wearable device 800 may communicate directly with both the server 812 and the external device 810. In such embodiments, the external device 810 may, for example, generate a graphical user interface to facilitate control and programming of the wearable device 800 while the server 812 may store and analyze larger amounts of data (e.g., training data, trained machine learning models and parameters) for future programming or operation of the wearable device 800. In other embodiments, the wearable device 800 may communicate directly with the server 812 without utilizing a short-range communication protocol with the external device 810.
[0137] In the illustrated embodiment, the wearable device 800 communicates with the external device 810. The external device 810 may include, for example, a smartphone, a tablet computer, a cellular phone, a laptop computer, a smart watch, another wearable device, and the like. The wearable device 800 communicates with the external device 810, for example, to transmit at least a portion of the physiological sensor data or other data collected or generated by the wearable device 800, which in some instances may include classified feature data generated by the wearable device 800.
[0138] In some embodiments, the external device 810 may include a short-range transceiver to communicate with the wearable device 800, and a long-range transceiver to communicate with the sen' er 812. In the illustrated embodiment, the wearable device 800 can also include a transceiver to communicate with the external device 810 via, for example, a short-range communication protocol such as Bluetooth®. In some embodiments, the external device 810 bridges the communication between wearable device 800 and the server 812. That is, the wearable device 800 transmits data to the external device 810, and the external device 810 forwards the data from wearable device 800 to the server 812 over the network 808.
[0139] The server 812 includes a server electronic control assembly having a server electronic processor, a server memory, and a transceiver. The transceiver allows the server 812 to communicate with the wearable device 800, the external device 810, or both. The server electronic processor receives physiological sensor data or other data collected with or generated by the wearable device 800, and stores the received data in the server memory. The server 812 may maintain a database (e.g., on the server memory) for containing sound-based feature data, sound libraries, biophysical signal data, transformed biophysical signal data, physiological data, training data, trained machine learning controls (e.g., trained machine learning models and / or algorithms), artificial intelligence controls (e.g., rules and / or other control logic implemented in an artificial intelligence model and / or algorithm), and the like.
[0140] Although illustrated as a single device, the server 812 may be a distributed device in which the server electronic processor and server memory are distributed among two or more units that are communicatively coupled (e g., via the network 808).
[0141] The wearable device 800 includes an electronic controller 820, a power source 852, a wireless communication device 860, and one or more sensors 872, among other components. In some embodiments, the wearable device 800 may not include a wireless communication device 860.
[0142] The electronic controller 820 can include an electronic processor 830 and memory 840. The electronic processor 830 and the memory 840 can communicate over one or more control buses, data buses, etc., which can include a device communication bus 876. The control and / or data buses are shown generally in FIG. 8 for illustrative purposes. The use of one or more control and / or data buses for the interconnection between and communication among the various modules, circuits, and components would be known to a person skilled in the art.
[0143] The electronic processor 830 can be configured to communicate with the memory 840 to store data and retrieve stored data. The electronic processor 830 can be configured to receive instructions 842 and data from the memory' 840 and execute, among other things, the instructions 842. In particular, the electronic processor 830 executes instructions 842 stored in the memory 840. Thus, the electronic controller 820 coupled with the electronic processor 830 and the memory 840 can be configured to perform the methods described herein (e g., the process 100 of FIG. 1).
[0144] The memory' 840 can include read-only memory' (ROM), random access memory (RAM), other non-transitory computer-readable media, or a combination thereof. The memory 840 can include instructions 842 for the electronic processor 830 to execute. The instructions 842 can include software executable by the electronic processor 830 to enable the electronic controller 820 to, among other things, receive data and / or commands, transmit data, and the like. The software can include, for example, firmware, one or more applications, program data, filters, rules, one or more program modules, and other executable instructions.
[0145] The electronic processor 830 is configured to retrieve from memory 840 and execute, among other things, instructions related to the control processes and methods described herein. The electronic processor 830 is also configured to store data on the memory 840 including physiological sensor data, classified feature data, and the like.
[0146] The wearable device 800 receives electrical power from the power source 852, which as an example may include a battery. Additionally or alternatively, the power source 852 may include an external power source (e.g., a wall outlet when the wearable device 800 is connected to a wall outlet).
[0147] In some embodiments, the wearable device 800 may also include a wireless communication device 860. In these embodiments, the wireless communication device 860 is coupled to the electronic controller 820 (e.g., via the device communication bus 876). The wireless communication device 860 may include, for example, a radio transceiver and antenna, a memory', and an electronic processor. In some examples, the wireless communication device 860 can further include a global navigation satellite system (GNSS) receiver configured to receive signals from GNSS satellites (e.g., global positioning system (GPS) satellites), land- based transmitters, etc. The radio transceiver and antenna operate together to send and receive wireless messages to and from the external device 810, the server 812, and the like. The memory of the wireless communication device 860 stores instructions to be implemented by the electronic processor of the wireless communication device 860 and / or may store data related to communications between the wearable device 800 and the external device 810 and / or the server 812.
[0148] The electronic processor for the wireless communication device 860 controls wireless communications between the wearable device 800 and the external device 810 and / or the server 812. For example, the electronic processor of the wireless communication device 860 buffers incoming and / or outgoing data, communicates with the electronic processor 830, and determines the communication protocol and / or settings to use in wireless communications.
[0149] In some embodiments, the wireless communication device 860 is a Bluetooth® controller. The Bluetooth® controller communicates with the external device 810 and / or the server 812 employing the Bluetooth® protocol. In such embodiments, therefore, the external device 810 and / or the server 812 and the wearable device 800 are within a communication range (i.e., in proximity) of each other while they exchange data. In other embodiments, the wireless communication device 860 communicates using other protocols (e.g., Wi-Fi, cellular protocols, a proprietary protocol, etc.) over a different type of wireless network. For example, the wireless communication device 860 may be configured to communicate via Wi-Fi through a wide area network such as the Internet or a local area netw ork, or to communicate through a piconet (e.g.. using infrared or NFC communications). The communication via the wirelesscommunication device 860 may be encrypted to protect the data exchanged between the wearable device 800 and the external device 810 and / or the server 812 from third parties.
[0150] The wireless communication device 860, in some embodiments, exports physiological sensor data or other data collected with or generated by the wearable device 800 (e.g., classified feature data). The wireless communication device 860 also enables the wearable device 800 to sync or otherwise communicate data with other devices, such as an external device 810 that is configured as a smart watch or smartphone. In some instances, the wearable device 800 can send and receive additional health information for the user (e.g., by syncing with a heartrate monitor, smart watch, or other wearable device).
[0151] The wearable device 800 can include or be coupled to one or more physiological sensors 872. For example, the physiological sensors 872 can include ECG leads, electrodes, or other conductive elements capable of measuring electrophysiology signals. That is, in some embodiments, the physiological sensors 872 are capable of recording or otherwise measuring ECG data. The physiological sensors 872 can implement various lead or electrode configurations, including a 12-lead configuration, a 6-lead configuration, a 3-lead configuration, a 1-lead configuration, or the like. Additionally or alternatively, the physiological sensors 872 can also record or measure other electrophysiology signals, such as electromyography (EMG) signals.
[0152] In other embodiments, the physiological sensors 872 can include additional sensors, such PPG sensors or PPG sensing circuits, temperature sensors or temperature sensing circuits, inertial sensors or inertial sensing circuits (e g., accelerometers, gyroscopes, magnetometers), a pressure sensor or pressure sensing circuit (e.g., a barometer), or the like. In still other embodiments, the physiological sensors 872 can include one or more microphones or other audio recording devices to collect sound data (e.g.. sound wave measurements, audio signal recordings, etc.).
[0153] In some embodiments, the wearable device 800 can include one or more inputs 890 (e.g., one or more buttons, switches, touchscreen, and the like) that are coupled to the electronic controller 820 and allow a user to interact with and control the wearable device 800. In some embodiments, the input 890 includes an interactive graphical user interface (GUI) element that enables user interaction with the wearable device 800.
[0154] In some embodiments, the wearable device 800 may include one or more outputs 892 that are also coupled to the electronic controller 820. The output(s) 892 can receive control signals from the electronic controller 820 to present data or information to a user inresponse, or to generate other visual, audio, or other outputs. As one example, the output(s) 892 can generate a visual signal to convey information regarding the physiological sensor data, other health data, generated classified feature data, or the like. The output(s) 892 may include, for example, LEDs or a display screen and may generate various signals indicative of, for example, sound-based feature data, biophysical signal data, physiological sensor data, other health data, generated classified feature data, or the like. For example, the output(s) 892 may indicate the detection of a medical condition or health status in a subject based on classified feature data generated by inputting sound-based feature data and / or biophysical signal data to a trained neural network or other machine learning model.
[0155] FIG. 9 shows an example of a system 900 for generating sound-based feature data from biophysical signals (and vice versa) and for analyzing such data in accordance with some embodiments described in the present disclosure. As shown in FIG. 9, a computing device 950 can receive one or more types of data (e.g., sound-based feature data, biophysical signal data, transformed biophysical signal data, other patient health data) from data source 902. In some embodiments, computing device 950 can execute at least a portion of a sound-based feature data and biophysical signal generation and analysis system 904 to generate and / or analyze sound-based feature data and / or biophysical signal data from data received from the data source 902.
[0156] Additionally or alternatively, in some embodiments, the computing device 950 can communicate information about data received from the data source 902 to a server 952 over a communication network 954, which can execute at least a portion of the sound-based feature data and biophysical signal generation and analysis system 904. In such embodiments, the server 952 can return information to the computing device 950 (and / or any other suitable computing device) indicative of an output of the sound-based feature data and biophysical signal generation and analysis system 904.
[0157] In some embodiments, computing device 950 and / or server 952 can be any suitable computing device or combination of devices, such as a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable computer, a server computer, a virtual machine being executed by a physical computing device, and so on. The computing device 950 and / or server 952 can also reconstruct images from the data.
[0158] In some embodiments, data source 902 can be any suitable source of data (e.g., sound-based feature data, biophysical signal data, transformed biophysical signal data, etc.), such as a wearable device, an audio device, a physiological sensor or measurement device.another computing device (e.g., a server storing sound-based feature data, biophysical signal data, transformed biophysical signal data), and so on. In some embodiments, data source 902 can be local to computing device 950. For example, data source 902 can be incorporated with computing device 950 (e.g., computing device 950 can be configured as part of a device for measuring, recording, estimating, acquiring, or otherwise collecting or storing data). As another example, data source 902 can be connected to computing device 950 by a cable, a direct wireless link, and so on. Additionally or alternatively, in some embodiments, data source 902 can be located locally and / or remotely from computing device 950, and can communicate data to computing device 950 (and / or server 952) via a communication network (e.g., communication network 954).
[0159] In some embodiments, communication network 954 can be any suitable communication network or combination of communication networks. For example, communication network 954 can include a Wi-Fi network (which can include one or more wireless routers, one or more switches, etc.), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e.g., a 3G network, a 4G network, etc., complying with any suitable standard, such as CDMA, GSM, LTE, LTE Advanced, WiMAX, etc.), other types of wireless network, a wired network, and so on. In some embodiments, communication netw ork 954 can be a local area network, a wide area network, a public network (e.g., the Internet), a private or semi-private network (e.g., a corporate or university intranet), any other suitable type of netw ork, or any suitable combination of networks. Communications links shown in FIG. 9 can each be any suitable communications link or combination of communications links, such as wired links, fiber optic links, Wi-Fi links, Bluetooth links, cellular links, and so on.
[0160] Referring now to FIG. 10, an example of hardware 1000 that can be used to implement data source 902, computing device 950, and server 952 in accordance with some embodiments of the systems and methods described in the present disclosure is shown.
[0161] As shown in FIG. 10, in some embodiments, computing device 950 can include a processor 1002, a display 1004, one or more inputs 1006, one or more communication systems 1008, and / or memory 1010. In some embodiments, processor 1002 can be any suitable hardware processor or combination of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), and so on. In some embodiments, display 1004 can include any suitable display devices, such as a liquid cry stal display (LCD) screen, a light-emitting diode (LED) display, an organic LED (OLED) display, an electrophoretic display (e.g., an “e- ink" display), a computer monitor, a touchscreen, a television, and so on. In someembodiments, inputs 1006 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.
[0162] In some embodiments, communications systems 1008 can include any suitable hardware, firmware, and / or software for communicating information over communication network 954 and / or any other suitable communication networks. For example, communications systems 1008 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1008 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0163] In some embodiments, memory 1010 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1002 to present content using display 1004, to communicate with server 952 via communications system(s) 1008, and so on. Memory 1010 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1010 can include random-access memory (RAM), read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), other forms of volatile memory, other forms of non-volatile memory7, one or more forms of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1010 can have encoded thereon, or otherwise stored therein, a computer program for controlling operation of computing device 950. In such embodiments, processor 1002 can execute at least a portion of the computer program to present content (e.g., images, user interfaces, graphics, tables), receive content from server 952, transmit information to server 952. and so on. For example, the processor 1002 and the memory 1010 can be configured to perform the methods described herein (e.g., the method of FIG. 1, the method of FIG. 2, the method of FIG. 3, the method of FIG. 4, the method of FIG. 5, the method of FIG. 6).
[0164] In some embodiments, server 952 can include a processor 1012, a display 1014, one or more inputs 1016, one or more communications systems 1018, and / or memory 1020. In some embodiments, processor 1012 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, display 1014 can include any suitable display devices, such as an LCD screen, LED display, OLED display, electrophoretic display, a computer monitor, a touchscreen, a television, and so on. In someembodiments, inputs 1016 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.
[0165] In some embodiments, communications systems 1018 can include any suitable hardware, firmware, and / or software for communicating information over communication network 954 and / or any other suitable communication networks. For example, communications systems 1018 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1018 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0166] In some embodiments, memory 1020 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1012 to present content using display 1014, to communicate with one or more computing devices 950, and so on. Memory 1020 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1020 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other ty pes of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1020 can have encoded thereon a server program for controlling operation of server 952. In such embodiments, processor 1012 can execute at least a portion of the server program to transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 950, receive information and / or content from one or more computing devices 950, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone), and so on.
[0167] In some embodiments, the server 952 is configured to perform the methods described in the present disclosure. For example, the processor 1012 and memory' 1020 can be configured to perform the methods described herein (e.g., the method of FIG. 1, the method of FIG. 2. the method of FIG. 3, the method of FIG. 4. the method of FIG. 5, the method of FIG. 6).
[0168] In some embodiments, data source 902 can include a processor 1022, one or more data acquisition systems 1024, one or more communications systems 1026, and / or memory 1028. In some embodiments, processor 1022 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, theone or more data acquisition systems 1024 are generally configured to acquire data, images, or both, and can include an audio device, a physiological measurement device, or the like. Additionally or alternatively, in some embodiments, the one or more data acquisition systems 1024 can include any suitable hardware, firmware, and / or software for coupling to and / or controlling operations of an audio device, a physiological measurement device, or the like. In some embodiments, one or more portions of the data acquisition system(s) 1024 can be removable and / or replaceable.
[0169] Note that, although not shown, data source 902 can include any suitable inputs and / or outputs. For example, data source 902 can include input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, a trackpad, a trackball, and so on. As another example, data source 902 can include any suitable display devices, such as an LCD screen, an LED display, an OLED display, an electrophoretic display, a computer monitor, a touchscreen, a television, etc., one or more speakers, and so on.
[0170] In some embodiments, communications systems 1026 can include any suitable hardware, firmware, and / or software for communicating information to computing device 950 (and, in some embodiments, over communication network 954 and / or any other suitable communication networks). For example, communications systems 1026 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 1026 can include hardware, firmware, and / or software that can be used to establish a wired connection using any suitable port and / or communication standard (e.g., VGA, DVI video, USB, RS-232, etc ), Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.
[0171] In some embodiments, memory' 1028 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1022 to control the one or more data acquisition systems 1024, and / or receive data from the one or more data acquisition systems 1024; to generate images from data; present content (e.g., data, images, a user interface) using a display; communicate with one or more computing devices 950; and so on. Memory 1028 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1028 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other ty pes of non-volatile memory7, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 1028 can have encoded thereon, or otherwise storedtherein, a program for controlling operation of data source 902. In such embodiments, processor 1022 can execute at least a portion of the program to generate images, transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 950, receive information and / or content from one or more computing devices 950, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone, etc.), and so on.
[0172] In some embodiments, any suitable computer-readable media can be used for storing instructions for performing the functions and / or processes described herein. For example, in some embodiments, computer-readable media can be transitory' or non-transitory . For example, non-transitory computer-readable media can include media such as magnetic media (e.g., hard disks, floppy disks), optical media (e.g., compact discs, digital video discs. Blu-ray discs), semiconductor media (e.g., RAM, flash memory, EPROM, EEPROM), any suitable media that is not fleeting or devoid of any semblance of permanence during transmission, and / or any suitable tangible media. As another example, transitory computer- readable media can include signals on networks, in wires, conductors, optical fibers, circuits, or any suitable media that is fleeting and devoid of any semblance of permanence during transmission, and / or any suitable intangible media.
[0173] As used herein in the context of computer implementation, unless otherwise specified or limited, the terms “component.” “system,” “module,” “framework,” and the like are intended to encompass part or all of computer-related systems that include hardware, softw are, a combination of hardw are and software, or software in execution. For example, a component may be, but is not limited to being, a processor device, a process being executed (or executable) by a processor device, an object, an executable, a thread of execution, a computer program, or a computer. By way of illustration, both an application running on a computer and the computer can be a component. One or more components (or system, module, and so on) may reside within a process or thread of execution, may be localized on one computer, may be distributed between two or more computers or other processor devices, or may be included within another component (or system, module, and so on).
[0174] In some implementations, devices or systems disclosed herein can be utilized or installed using methods embodying aspects of the disclosure. Correspondingly, description herein of particular features, capabilities, or intended purposes of a device or system is generally intended to inherently include disclosure of a method of using such features for the intended purposes, a method of implementing such capabilities, and a method of installingdisclosed (or otherwise known) components to support these purposes or capabilities. Similarly, unless otherwise indicated or limited, discussion herein of any method of manufacturing or using a particular device or system, including installing the device or system, is intended to inherently include disclosure, as embodiments of the disclosure, of the utilized features and implemented capabilities of such device or system.
[0175] The present disclosure has described one or more preferred embodiments, and it should be appreciated that many equivalents, alternatives, variations, and modifications, aside from those expressly stated, are possible and within the scope of the invention.
Claims
CLAIMS1. A method for generating sound-based feature data from biophysical signals, the method comprising: accessing biophysical signal data with a computer system, wherein the biophysical signal data comprise biophysical signals acquired from a subject; accessing a machine learning model with the computer system, wherein the machine learning model has been trained on training data to convert biophysical signals to sound-based features; applying the biophysical signal data to the machine learning model using the computer system, generating sound-based feature data as an output; and outputting the sound-based feature data using the computer system.
2. The method of claim 1, wherein the biophysical signal data comprise at least one of electrocardiography (ECG) signal data, electroencephalography (EEG) data, electromyography (EMG) data, or photoplethysmography (PPG) data.
3. The method of claim 1 or 2, wherein biophysical signal data are obtained with a radiofrequency (RF)-based biophysical sensor.
4. The method of any one of claims 1-3, wherein the biophysical signal data are acquired using one or more physiological sensors coupled to a wearable device.
5. The method of claim 1. wherein the sound-based feature data comprise at least one of a sound wave, an audio signal, or features computed from at least one of sound waves or audio signals.
6. The method of claim 5, wherein the sound-based feature data comprise the features computed from at least one of sound waves or audio signals, wherein features computed from at least one of sound waves or audio signals comprise at least one of amplitude, frequency, phase, time-domain features, or frequency domain features.
7. The method of claim 6, wherein the time-domain features comprise at least one of amplitude envelope, zero-crossing rate, temporal moments (e.g., mean, variance, skewness, kurtosis), energy, or root mean square (RMS) energy'.
8. The method of claim 6, wherein the frequency-domain features comprise at least one of spectral centroid, spectral bandwidth, spectral contrast, or spectral flatness.
9. The method of claim 1, wherein the biophysical signal data comprise transformed biophysical signal data.
10. The method of clam 9, wherein the transformed biophysical signal data comprise scalogram data generated from the biophysical signals.
11. The method of clam 9, wherein the transformed biophysical signal data comprise spectrogram data generated from the biophysical signals.
12. The method of claim 1. wherein the machine learning model comprises a generative adversarial network (GAN).
13. The method of claim 1, wherein the machine learning model comprises a transformer network.
14. The method of claim 1, further comprising: accessing a second machine learning model with the computer system, wherein the second machine learning model has been trained on training data to detect a presence of a medical condition based on sound-based feature data; applying the sound-based feature data to the second machine learning model using the computer system, generating classified feature data indicative of a medical condition in the subject as an output; and outputting the classified feature data using the computer system.
15. The method of claim 14, wherein the second machine learning model comprises a neural network.
16. The method of claim 14, wherein the second machine learning model comprises a large language model (LLM).
17. The method of claim 16, wherein the LLM implements a generative pretrained transformer (GPT) architecture.
18. The method of claim 16, wherein the classified feature data include text data comprising a report.
19. The method of claim 18, wherein the report indicates whether the medical condition in the subject is improving, worsening, or staying the same.
20. The method of claim 18, wherein the report indicates recommendations for further testing or treatment of the subj ect.
21. The method of claim 18, wherein the report includes patient education materials that explain a care plan tailored to the subject.
22. The method of claim 18, wherein the report comprises clinic notes.
23. The method of claim 18, wherein the report comprises medical billing codes.
24. The method of claim 14, wherein the classified feature data comprise a risk score for the subject having the medical condition.
25. The method of claim 14, wherein the classified feature data comprise a classification of the sound-based feature data as being indicative of the medical condition.
26. A method for generating biophysical signals from sound-based feature data, the method comprising:accessing sound-based feature data with a computer system, wherein the sound-based feature data comprise one or more sound-based features recorded from a subject; accessing a machine learning model with the computer system, wherein the machine learning model has been trained on training data to convert sound-based features to biophysical signals; applying the sound-based feature data to the machine learning model using the computer system, generating biophysical signal data as an output; and outputting the biophysical signal data using the computer system.
27. The method of claim 26, wherein the sound-based feature data comprise at least one of a sound wave, an audio signal, or features computed from at least one of sound waves or audio signals.
28. The method of claim 27, wherein the sound-based feature data comprise the features computed from at least one of sound waves or audio signals, wherein features computed from at least one of sound waves or audio signals comprise at least one of amplitude, frequency, phase, time-domain features, or frequency domain features.
29. The method of claim 28, wherein the time-domain features comprise at least one of amplitude envelope, zero-crossing rate, temporal moments (e.g., mean, variance, skewness, kurtosis), energy, or root mean square (RMS) energy.
30. The method of claim 28, wherein the frequency -domain features comprise at least one of spectral centroid, spectral bandwidth, spectral contrast, or spectral flatness.
31. The method of claim 26, wherein the sound-based feature data are acquired from the subject using a digital stethoscope.
32. The method of claim 26, wherein the sound-based feature data are acquired from the subject using an audio recording device.
33. The method of claim 26, wherein the sound-based feature data are acquired from the subject using an ultrasound system.
34. The method of clam 33, wherein the sound-based feature data include Doppler ultrasound data.
35. The method of claim 26, wherein the biophysical signal data comprise at least one of electrocardiography (ECG) signal data, electroencephalography (EEG) data, electromyography (EMG) data, or photoplethysmography (PPG) data.
36. The method of claim 26, wherein the biophysical signal data comprise transformed biophysical signal data.
37. The method of clam 36, wherein the transformed biophysical signal data comprise scalogram data generated from the biophysical signals.
38. The method of clam 36, wherein the transformed biophysical signal data comprise spectrogram data generated from the biophysical signals.
39. The method of claim 26, wherein the machine learning model comprises a generative adversarial network (GAN).
40. The method of claim 26, wherein the machine learning model comprises a transformer network.
41. The method of claim 26, further comprising: accessing a second machine learning model with the computer system, wherein the second machine learning model has been trained on training data to detect a presence of a medical condition based on biophysical signal data; applying the biophysical signal data to the second machine learning model using the computer system, generating classified feature data indicative of a medical condition in the subject as an output; and outputting the classified feature data using the computer system.
42. The method of claim 41, wherein the second machine learning model comprises a neural network.
43. The method of claim 41, wherein the second machine learning model comprises a large language model (LLM).
44. The method of claim 43, wherein the LLM implements a generative pretrained transformer (GPT) architecture.
45. The method of claim 43, wherein the classified feature data include text data comprising a report.
46. The method of claim 45, wherein the report indicates whether the medical condition in the subject is improving, worsening, or staying the same.
47. The method of claim 45, wherein the report indicates recommendations for further testing or treatment of the subj ect.
48. The method of claim 45, wherein the report includes patient education materials that explain a care plan tailored to the subject.
49. The method of claim 45, wherein the report comprises clinic notes.
50. The method of claim 45, wherein the report comprises medical billing codes.
51. The method of claim 41 , wherein the classified feature data comprise a risk score for the subject having the medical condition.
52. The method of claim 41, wherein the classified feature data comprise a classification of the sound-based feature data as being indicative of the medical condition.
Citation Information
Patent Citations
Using colored probes in patient monitoring
US20100286494A1
Device and method for sleep monitoring
US20160089078A1
Integrated sport electrocardiography system
US20210212625A1
System And Method For Measuring Human Intention
US20230162719A1