Ensemble machine learning models for detecting respiratory syndromes

Ensemble machine learning models utilizing multiple channels of input data improve the accuracy and specificity of COVID-19 detection, addressing the limitations of single-channel analysis and remote diagnostic challenges.

JP7850135B2Active Publication Date: 2026-04-22THE COVID DETECTION FOUNDATION D B A VIRUFY
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
THE COVID DETECTION FOUNDATION D B A VIRUFY
Filing Date
2021-08-03
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing methods for diagnosing COVID-19 are time-consuming and economically burdensome, especially in remote areas with limited medical resources, and current computer-based speech analysis for COVID-19 is often limited to single-channel information, resulting in lower accuracy and specificity.

Method used

Implementing ensemble machine learning models that utilize multiple channels of input data, including audio and image analysis, to enhance the detection of COVID-19 using smartphones, public kiosks, or other devices, by preprocessing and feature extraction, and combining outputs from heterogeneous models to improve accuracy.

Benefits of technology

The ensemble machine learning models enhance the detection accuracy and specificity of COVID-19 infection by leveraging multiple input modalities, providing a powerful tool for predicting COVID-19 status in advance, enhancing the accuracy and reducing the time and cost associated with traditional diagnostic methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007850135000001
    Figure 0007850135000001
  • Figure 0007850135000002
    Figure 0007850135000002
  • Figure 0007850135000003
    Figure 0007850135000003
Patent Text Reader

Abstract

The present invention provides a process comprising: obtaining, by one or more processors, a dataset including a plurality of patient records; selecting a subset of the plurality of parameters for input to a machine learning system; and generating a classifier using the machine learning system based on training data and the subset of the plurality of parameters for input; receiving, by one or more processors, a patient record of a first user; and performing, by one or more processors, an analysis to identify acoustic measurements from a voice sample of the first user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - reference to Related Applications) This application claims priority based on U.S. Provisional Patent Application No. 63 / 060,297, filed on August 3, 2020, with the title "Ensemble Machine Learning Model for Detecting Respiratory Syndromes", and U.S. Provisional Patent Application No. 63 / 117,394, filed on November 23, 2020, with the title: "Cross - continental Applicability of Crowdsourcing - based Datasets and Clinical Datasets for AI Detection of COVID - 19 from Cough". The entire contents of the above applications are hereby incorporated by reference for all purposes.

[0002] The present disclosure generally relates to computer models for detecting infections, and more specifically, to machine learning models for detecting individuals infected with respiratory viruses and other pathogens.

Background Art

[0003] The novel coronavirus has spread, and more than 73 million COVID - 19 patients have been discovered worldwide. At the same time, the clinical diagnosis of COVID - 19 imposes a waste of time and an economic burden on people, especially those in remote areas with few medical institutions for COVID - 19.

Summary of the Invention

[0004] The following non - exhaustively lists some aspects of the present technology. These and other aspects are described in the subsequent disclosure.

[0005] Some aspects of the present invention provide a computer-implemented method. The method comprises the step of acquiring a dataset containing a plurality of patient records using one or more processors, each patient record containing a plurality of parameters and corresponding values ​​for the patient, the plurality of parameters and corresponding values ​​for the patient containing an audio file of the patient's voice noise such as cough, breathing or speaking, the dataset containing diagnostic information indicating whether or not the patient has been diagnosed with COVID-19, the method further comprises the step of selecting a subset of the plurality of parameters to be input to a machine learning system, the subset of the plurality of parameters containing at least two parameters and corresponding values ​​for the patient, one of the parameters of the subset of the plurality of parameters being the patient's cough The method comprises the steps of: dividing the dataset into training data and validation data; generating a classifier using the machine learning system based on the training data and a subset of the plurality of parameters for the input; receiving a patient record of a first user with one or more processors; performing an analysis with one or more processors to identify acoustic measurements from a voice sample of the first user; determining the likelihood of the first user being infected with COVID-19 using the classifier based on the identified acoustic measurements of the voice sample of the first user; and outputting the likelihood of the first user being infected with COVID-19.

[0006] Some embodiments provide a tangible, non-temporary, machine-readable medium that, when executed by a data processing device, stores instructions causing the data processing device to perform operations including the processes described above.

[0007] Some embodiments provide a system comprising one or more processors and memory for storing instructions, wherein when an instruction is executed by at least a portion of the one or more processors, the processing of the above-mentioned process is performed. [Brief explanation of the drawing]

[0008] The above-described and other aspects of this technology will be better understood by reading this application with reference to the following figures. In these figures, the same numbers indicate similar or identical elements.

[0009] [Figure 1] A logical diagram and a physical architecture diagram illustrating one embodiment of a controller configured to classify data as an indication of infection, in accordance with part of the technology of the present invention.

[0010] [Figure 2] This flowchart shows an example of a process for determining the likelihood of COVID-19 infection using a machine learning model, according to some of the techniques of the present invention.

[0011] [Figure 3] An example of a computer device in which the technology of the present invention can be implemented is shown.

[0012] The technology of the present invention can take on various modifications and alternative forms, specific embodiments of which are shown as examples in the drawings and described in detail herein. The drawings may not be to actual scale. However, it should be understood that the drawings and the detailed description therefrom are not intended to limit the technology to any particular form disclosed, but rather to cover all modifications, equivalents, and alternatives that fall within the spirit and scope of the technology as defined by the appended claims. [Modes for carrying out the invention]

[0013] To mitigate the problems described in this book, the inventors had to devise solutions, and in some cases, equally importantly, they had to recognize problems that others in the field of machine learning had overlooked (or had not yet foreseen). Indeed, the inventors want to emphasize the difficulty of recognizing problems in their early stages. These problems will become far more apparent in the future if industry trends continue as the inventors hope. Furthermore, because there are multiple problems to address, it should be understood that some embodiments are specialized for one of these problems, and not all embodiments address all of the problems of the conventional systems described herein, nor do they offer all of the advantages described herein. In other words, improvements that address various permutations of these problems are described below.

[0014] Machine learning algorithms have the potential to be a powerful tool for predicting a person's COVID-19 status in advance. In some embodiments, such models are implemented to accurately predict COVID-19 infection from audio and images acquired via smartphones. Given the high and continuously rising smartphone usage, even in economically disadvantaged areas, these devices are expected to be an ideal, versatile, and low-cost platform for collecting respiratory audio recordings and conducting voice-based COVID-19 testing. However, this technology can also be used on other platforms, such as public kiosks, desktop computers, and servers that receive similar data from remote client devices.

[0015] Some forms of computer-based speech analysis for COVID-19 are often limited to single-channel information, such as speech alone, resulting in lower accuracy and specificity than those predicted to be achievable with a broader set of features and a suitable ensemble model.

[0016] In some embodiments, the native application may operate on raw data or inputs that undergo preprocessing, filtering, and feature extraction before the final model is trained (if in training) or inferred (e.g., classifying an input set as indicating COVID-19 infection). Some embodiments ensemble multiple (potentially heterogeneous) machine learning models to train and infer classifications from inputs of many channels other than speech. Alternatively, some embodiments may operate on speech only. In some embodiments, the machine learning model may operate by fusing multiple channels of input data. For speech, for example, some embodiments use both Mel-frequency septal coefficients (MFCCs) and Mel-spectrograms to train a deep neural network. In another embodiment, the output of an ensemble machine learning model may be combined with the output of a computer vision model to infer COVID-related features from images in the ensemble machine learning model.

[0017] Some embodiments run on a smartphone (e.g., exclusively as a monolithic application or as part of a distributed application running partially on a remote server) and acquire multi-channel data about the user; such examples are described below. In other embodiments, audio or images may be acquired from other sources, for example, a call from a user making a call to a call center, or from a smart speaker or other host of a voice-based digital assistant. Such sources also constitute examples of the user's mobile computer device for this purpose. Some embodiments classify the state of COVID-19 (or other syndromes) (e.g., performed locally or by a remote server) and respond in real time (e.g., within 1 minute or 10 minutes of data acquisition). Some embodiments utilize sensor hardware present in the smartphone. Some embodiments utilize a single modality test, while others combine various modalities as an ensemble technique to improve accuracy (e.g., measured by sensitivity, specificity, type 1 error, type 2 error, or F2 score). In some embodiments, the user may be prompted to perform various actions via the user interface of a native application on the smartphone, which acquire inputs to various upstream submodels that feed into the ensemble model. For example, this includes allowing you to fill out a text questionnaire, breathe or cough into a phone microphone, speak within the microphone's audible range, take videos or photos of your fingers, other appendages, face or other bodily excretions (e.g., feces, saliva, blood, mucus, etc.), or allow data to be collected from wearable devices (such as wrist-worn pulse oximeters, inertial measuring devices (pedometers, etc.), heart rate sensors, temperature monitors, etc.).

[0018] Figure 1 is a schematic block diagram of an example of a controller 12 operating within a computer system 100 in which the present technology may be implemented. A variety of different computer architectures are possible. Therefore, the term “computer system” shall be used as a general term for both a single computer device (which could be, for example, a smartphone or a server) and a collection of computer devices (which could include both a smartphone and multiple different servers in a microservices architecture, where each device performs a different subset of tasks performed by the computer system). In some embodiments, some or all of the components of the controller 12 may be hosted by different entities, for example, in a client-server architecture, model training or inference may be performed on the server side, and data may be acquired from the smartphone, which is the client side. In some cases, the model may be trained on the server side, but inference may be performed on the client side using a trained model downloaded to a native application. In some embodiments, the controller 12 and its components may be implemented, for example, as a monolithic application, and the various components shown may be implemented as different software modules or processes that communicate with each other, for example, via function calls, and in some cases, some or all of the components may be implemented as multiple different processes running concurrently on a single computer device. In some embodiments, some or all of the illustrated components may be implemented as separate services running on different network hosts, which communicate with each other, for example, through messages exchanged via their respective network stacks according to the application programming interfaces of each service.

[0019] In some embodiments, the computer system 100 can train a model using multiple source datasets 10, and the controller 12 may cause a computer device such as a smartphone to present a user interface 18. In some embodiments, the controller 12 may include an artificial intelligence (AI) module 14 (such as one that implements a machine learning model) having multiple modality classifiers 16 (e.g., a cough classifier, deep breathing analysis, time data analysis, face video, fingertip video, and biometric image). The classifiers 16 may be capable of classifying the input according to whether or not it has been indicated as infected, or some embodiments of the classifiers 16 may extract features from the input for downstream processing by an ensemble model.

[0020] In some embodiments, the controller 12 may be configured to perform the process 200 described below with reference to Figure 2. In some embodiments, several different subsets of this process 200 may be performed by the illustrated components of the controller 12, the features of which are described concurrently herein. Embodiments of the process 200 are not limited to the implementation according to the architecture of Figure 1, and the architecture of Figure 1 may perform processes different from those described with reference to Figure 2, neither of which suggests that the rest of the description herein is limited.

[0021] In some embodiments, process 200 includes obtaining multiple datasets of training data, as shown by block 102 in Figure 2. The training data may be labeled data for supervised learning, or unlabeled data for unsupervised or semi-supervised learning. An example is a labeled dataset for the same channel input data used for inference. In some cases, each training set may include inputs for each channel and labels indicating whether the person has COVID-19, when they contracted COVID-19, the stage of their infection at the time the sample was taken, whether they were hospitalized, demographic data, comorbidities, complications from the infection, and whether they died from the infection. In some cases, the models described above may be used to infer the likelihood of hospitalization or death. In some cases, some of the input channels for information may include these fields of data entered by the user when filling out a survey presented through UI 18.

[0022] In some embodiments, a subset of multiple parameters (e.g., one or more of multiple channels) may be selected as input to an AI module (e.g., a machine learning model), as shown by block 104. In some embodiments, a text-based questionnaire may be used to increase the reliability of the COVID-19 prediction.

[0023] In some embodiments, a smartphone or medical device may be used to assess the user's likelihood of infection with the novel coronavirus (SAR-CoV-2) or other pathogens. Depending on the type of modality, the smartphone or medical device may include a camera (such as one with a high-resolution (e.g., 1 megapixel or more) complementary metal-oxide-semiconductor (CMOS) image sensor), a temperature sensor, a Global Positioning System (GPS) sensor, an accelerometer, a gyroscope, a magnetometer, an ambient light sensor, a microphone, a touchscreen interface, an oxygen concentration sensor (Apple® Watch Series 6), and the like.

[0024] In some embodiments, deep breathing (e.g., 80% or more of the maximum breathing depth) analysis for COVID-19 detection may be used. The prediction accuracy of this modality is currently considered inferior to voice because the signal strength is weak, but it still far exceeds random guessing and is expected to be useful as an additional metric for measuring confidence in an ensemble model. In some cases, different forms of voice input, such as coughing, reading a specified phrase, repeating syllables (e.g., asking the user to say "ah, ah, ah..." or "ee, ee, ee..." for 5 seconds), and deep breathing may each constitute input for different channels. The voice input may be performed using the microphone of the user's smartphone.

[0025] In some embodiments, temporal data analysis may be used. By using the user interface to record data multiple times over several days and weeks using data from the same patient, the algorithm is expected to infer the stage of the user in COVID-19 disease and predict the onset and outcome of the disease. Even after recovering from COVID-19, there are cases where the tissues of the patient's ears, nose, throat, and lungs are affected along with the presence of antibodies. The biological and physical differences caused by these changes are expected to be detectable by some embodiments, and in some embodiments, COVID-19 immunity may be inferred from such data.

[0026] Some embodiments can obtain an image (or a collection of images such as a video) and perform face image analysis, for example, from the camera of a user's smartphone. In some embodiments, distinct features in the faces of COVID-19 positive and negative patients are detected (e.g., on the client device or server side), such as the color of the lips, which tend to be bluish in COVID-19 patients due to oxygen deficiency, and changes in skin color / texture. In some embodiments, various states such as heart rate, heart rate variability, oxygen saturation, respiratory rate, etc. are inferred from a video of the face based on changes in the intensity of the redness of the face (due to blood flow around blood vessels).

[0027] In some embodiments, voice-based detection of COVID-19 infection may also be used. Further, in some embodiments, in order to further enhance the effectiveness of a system for accurately detecting COVID-19 carriers, age, gender, and ethnicity can be inferred as features from the voice of a speaker. In some cases, these features can be input by the user in a survey presented via UI18.

[0028] In some embodiments, a video (or individual images) of a fingertip used to measure and record blood oxygen saturation and heart rate may be acquired (e.g., from a mobile device camera) and processed. Patients with COVID-19 often experience respiratory system effects leading to decreased oxygen intake, which can be detected by visual features (e.g., color) indicating decreased oxygen concentration in the blood vessels of the finger. In some cases, patients may be instructed to shine a light on their finger when taking the video. Similarly, patients with COVID-19 often experience increased heart rate or arrhythmias, which are complications of COVID-19 that occur with increased difficulty in oxygen intake. In some embodiments, a smartphone can serve as a substitute for a photoelectric blood pressure monitor (PPG) by implementing a pulse oximeter, taking a video by firmly pressing one finger against the camera lens with the flash on, and analyzing the intensity of the captured red pixels (e.g., the intensity of the red channel and its temporal variation). In some further embodiments, the acquired PPG may be further analyzed for heart rate to infer various patient vital signs. For example, some embodiments implement the techniques of the following papers incorporated herein by reference. Hasan et al., SmartHeLP: Hemoglobin level prediction function using an artificial neural network on smartphones, AMIA Annu Symp Proc. December 5, 2018; 2018:535-544. eCollection 2018, PMID:30815094 PMCID:PMC6371334.

[0029] In some embodiments, bio-images can be used to identify individuals infected with COVID-19. COVID-19 can affect various biophysical systems in the body. In some embodiments, changes in various bodily secretions such as saliva, feces, urine, vomit, and mucus can be detected by analyzing images taken with a user's smartphone. Subtle differences in images of these substances associated with COVID-19 are expected to be detected in some embodiments. For example, statistical values ​​of the dimensions (and color) of multiple blobs (small clumps or bubbles) detected using a blob detection algorithm with known reference dimensions (such as a credit card) set within the field of view (or on such a surface at a specified angle) may indicate fluid viscosity, surface tension, or other attributes associated with COVID-19 infection. Changes in surface tension and / or color reported by patients can also be used as input features in some embodiments.

[0030] In some embodiments, audio / image compression on mobile devices may be adjusted to enhance inference. Some embodiments of the machine learning models described herein are expected to be able to pick up COVID-19 from signals that are indistinguishable to the human eye and ear, which are often lost by conventional lossy compression techniques. Some embodiments may adjust audio compression / decompression of data to preserve features relevant to the classification of COVID-19 by such models. For example, some embodiments may apply lossy compression techniques to some of the human audible frequency bands, while prioritizing compression with relatively less data loss for frequency bands determined to be related to COVID-19. Similar techniques may be applied to image compression (e.g., video compression) to preserve relevant features, for example, by adjusting the quantization matrix. In some cases, compression may be adjusted by applying techniques to machine learning models trained to improve their interpretability, for example, by measuring the effect of removing specific parts of the neural network in the F2 score. Removed parts of the model (e.g., perceptrons, convolutional filters, connections, etc.) that have a relatively large effect on the F2 score are considered important. In some embodiments, the effect of various compression parameters on the features output by those parts of the model may be measured, and parameter values ​​that maintain accuracy may be determined while considering acceptable trade-offs in compression.

[0031] Some embodiments may comprise multiple upstream submodels that generate multiple outputs which are combined in a downstream ensemble model that outputs the final classification. In some cases, each of the above modalities expected to have discriminative ability may have different submodels or be a combination thereof. In some cases, each submodel is trained separately and independently and optimized for accuracy in detecting COVID-19 (or other respiratory diseases as referenced with respect to COVID-19). Alternatively, in some cases, end-to-end training may be applied in a single global optimization, although this approach is considered more computationally intensive because memory is required simultaneously for multiple model parameters.

[0032] Examples of techniques include stochastic gradient descent, simulated annealing, and evolutionary optimization algorithms. In some cases, each submodel is trained before the ensemble model is trained. In some embodiments, model parameter values ​​are randomly assigned, the partial derivative of each parameter with respect to the objective function is calculated, and the model is locally optimized by adjusting the parameters in the direction indicated by the partial derivatives. This calculation and adjustment is repeated until the change in the objective function between iterations falls below a threshold indicating a local or global optimality. In some embodiments, this process may be repeated multiple times with several different randomly assigned initial parameter values, and from these iterations, a version of the trained model that yields the best result as measured by the objective function may be selected.

[0033] Ensemble models can take various architectures. Examples include deep neural networks, decision trees, random forests, regression trees, classification trees, and Bainian networks. Along with initial coupling, methods such as soft voting and hard voting may be implemented. In some cases, these approaches may also be used in submodels. In some cases, some submodels, for example those processing time-series data (e.g., video or audio), can use transformer architectures, such as multi-head attention, long short-term memory models, or other recurrent neural networks. In particular, if the training data (or positive examples within it) is sparse, techniques such as Siamese networks or triplet loss networks may be applied, and in some cases, time-contrastive networks for time-series data may be used.

[0034] In some embodiments, data augmentation (such as adding background audio noise like white noise or Gaussian noise, or blurring images) and auxiliary data (such as audio and visual datasets of various respiratory and other diseases) can also be used to enhance and improve the effectiveness of the algorithm.

[0035] In some embodiments, data collection can be carried out from multiple angles, combining global grassroots crowdsourcing efforts with clinical research and trials in various countries.

[0036] In some embodiments, the algorithm may be configured to detect and distinguish a variety of diseases, including influenza, the common cold, SARS, COVID-19, and other coronaviruses, along with respiratory illnesses such as whooping cough and asthma. In some embodiments, other potentially detectable disorders by voice (e.g., child abuse, domestic violence, depression, etc.) may be detected.

[0037] In some embodiments, the set of labeled training data may be split into several different subgroups (e.g., a training dataset and a validation dataset), as shown in block 106 of Figure 2. In some cases, the training data may be a fairly unbalanced dataset due to the relatively low number of positive cases. In some cases, data augmentation techniques may be applied to create a more balanced training dataset. The number of COVID-19 labeled samples may be increased by adding Gaussian noise or white noise (or other examples above), adjusting the volume, pitch shifting, time signal shifting, and time signal stretching. Before the augmentation stage, the data may be split into a training dataset, a validation dataset, and a test dataset, so that augmentation is applied separately to the split datasets. In some cases, each class may be represented by one-third of the number of split samples, which is considered to ensure that the data is perfectly balanced across all classes.

[0038] In some embodiments, the classifier may be generated (e.g., trained) using machine learning techniques, as shown by block 108 in Figure 2. In some embodiments, a deep neural network may be trained using a publicly available dataset of cough sounds with COVID-19 status labels such as Coswara, Coughvid, and Iatos.

[0039] In some embodiments, additional datasets with more detailed labels beyond the Coswara and Coughvid crowdsourced data may be compiled to validate the model's performance. All data may have COVID-19 PCR labels and be acquired under conditions intended to simulate real-world use. Audio files may be a mix of compressed and uncompressed files (e.g., wav, ogg, flac, webm, mp3 files) depending on the mode of data acquisition. Potential privacy risks and security threats may be addressed through regional privacy policies and patient consent forms, along with data protection impact assessments (DPIAs) and several internal information security policies. In some cases, datasets may be anonymized and encrypted both in processing and unprocessed.

[0040] In some embodiments, to mimic one potential use case of detecting COVID-19 from the voices of typical smartphone users, the samples used within the model are crowdsourced using a mobile data collection app.

[0041] In some embodiments, smartphones may be used to collect samples in hospitals to determine the performance of a COVID-19 detection algorithm in a clinical setting. Explicit patient consent forms, which are presented electronically and signed by all patients, are drafted in advance. Data is collected directly from patients under a clinical research protocol approved by the hospital's Institutional Review Board (IRB).

[0042] In some embodiments, multiple features from a crowdsourced dataset may be used to train the model. After searching for various features and architectures using a grid search, an ensemble model of three features with parameters as described below may be used. The first feature is the Mel-frequency cepstrum coefficient (MFCC), which is an audio feature obtained from the short-term power spectrum. Each audio file may be resampled to 22.5 kHz, and the first 39 MFCCs may be extracted using the librosa package with a sampling rate of 22.5 kHz, a hop length of 23 ms, a window length of 93 ms, and a Hann window type. The output is averaged over time to obtain an average of 39 MFCC features per audio file.

[0043] In some embodiments, the second feature extracted may be another audio feature, a Mel frequency spectrogram. The MFCC is derived from the spectrogram, which encodes raw power information without any transformation. The spectrogram may be extracted using the librosa package with the same parameters as the MFCC and interpolated to a predetermined size.

[0044] In some embodiments, the method of extracting speech features from audio files can affect the performance of the model. Several useful features are considered for training the network, for example, Mel-frequency cepstrum coefficients and Mel-frequency spectrograms, both of which are speech features. In some embodiments, multiple heterogeneous classifiers can be used, one of which is trained on Mel-spectrograms and another on MFCCs. Each audio file may be downsampled to half of its original frequency (22.5 kHz) and divided into 3-second chunks. The first 13 MFCCs are extracted from the pre-processed chunks using the librosa package in Python, and the Hann window type may have a hop length of 10 ms and a window length of 20 ms.

[0045] In some embodiments, mel-spectrograms may be extracted using the librosa package for the same parameters used to extract MFCCs. Each mel-spectrogram color image may be reshaped to the original input size of the ResNet-50 convolutional neural network, (224,224,3). Additionally, other useful clinical information from the COUGHVID dataset, such as a history of respiratory illness or fever symptoms, may be used to further improve the accuracy of the model predicting COVID-19 infections. This clinical information can be passed as a one-dimensional binary vector, as the presence or absence of symptoms or conditions is represented in binary.

[0046] In some embodiments, multiple different types of features extracted from vocal sound chunks may be stored in a hash table along with the key for each record. The data may be randomly (e.g., pseudorandomly) grouped into training-validation-test sets using an 80-10-10 split.

[0047] In some embodiments, slice-based analysis is performed to divide the test dataset into groups based on age and sex. The test dataset may be divided into multiple groups based on age. For example, in the case of four groups, the first group may consist of patients under 20 years of age, the second group may consist of patients between 20 and 40 years of age, the third group may consist of patients between 40 and 60 years of age, and the fourth group may consist of patients 60 years and older. Alternatively, in some embodiments, the groups may be 18-30 years of age, 30-45 years of age, 46-60 years of age, and older. For sex, the test dataset may be divided into corresponding groups.

[0048] In some embodiments, the model is a multi-branch ensemble learning architecture based on a ResNet-50 3D convolutional neural network pre-trained on the ImageNet dataset and with the top layer (e.g., classification layer) removed. The input to the CNN may be a Mel spectrogram color image of a predetermined size (224 pixels, 224 pixels, three RGB layers, or larger or smaller than any of these dimensions), and the output of the CNN may be passed to both a global mean pooling layer and a global max pooling layer in two separate parallel links. These layers are followed by a batch normalization layer and a dropout layer, respectively, which may be linked together in a single dense layer (e.g., a nonlinear layer such as a layer with a sigmoid or hyperbolic tangent activation function) to form the first branch.

[0049] In some embodiments, the second branch may be a multilayer feedforward neural network including two dense layers, each having 8 and 64 nodes, respectively. Each layer may be followed by a batch normalization layer and a dropout layer. The input to the first branch may be a binary ID vector. The binary number may encode one of the clinical features related to the patient record, such as a history of respiratory illness, type of cough, or presence or absence of fever in the patient. This branch is expected to enrich the clinical information.

[0050] In some embodiments, the third branch may be a dual parallel feedforward neural network whose input vector is a vector of Mel-frequency cepstrum coefficients of a predetermined size (13, 1, or greater or less than any of these dimensions). Each of the two parallel links may be a multilayer feedforward neural network containing two layers, each layer of which may be followed by a batch normalization layer and a dropout layer. The higher ends of both links may be connected by a single dense layer.

[0051] In some embodiments, high-level features extracted at the higher end of the third branch may be combined before being passed to a sequential feedforward neural network (SFFN) followed by a softmax layer for the multi-label classification task. The three labels in some embodiments are: COVID-19 negative (healthy), COVID-19 negative (symptomatic), and COVID-19 positive. In other embodiments, more labels may be included, such as low confidence negative, high confidence negative, low confidence positive, high confidence positive, and uncertain. Alternatively, some embodiments may output a real score, such as a value between 0 and 1, where a higher value indicates a stronger inference that the person is infected.

[0052] In some embodiments, the network architecture may use multiple heterogeneous classifiers, and may combine high-level features extracted from spectrogram images using a ResNet-50 CNN (convolutional neural network) and high-level features extracted from MFCCs using a deep neural network. The network architecture, the number of hidden layers relative to the branch, and the number of units per layer are hyperparameters that can be determined using grid search. The model may be trained using a stochastic gradient descent optimizer with categorical cross-entropy loss, learning rate le-2, and decay steps of 2500.

[0053] Beyond the audio files, each sample may contain additional rich information that could improve predictive accuracy. In some embodiments, two further features reflecting the patient's clinical presentation may be used for each audio file. Detectable changes in cough sounds have been shown to occur in diseases other than COVID-19. Therefore, binary labels regarding the presence or absence of a current respiratory illness can be integrated and fed into the algorithm as a single additional feature. COVID-19 presents with symptoms other than cough, most notably fever and muscle pain. The presence or absence of these symptoms may also influence the probability of having COVID-19. In some embodiments, a second binary label for fever or muscle pain can also be integrated from the entire dataset and fed into the model as a second additional feature.

[0054] Various architectures can be used to maximize the accuracy of detecting individuals infected with the novel coronavirus. In some embodiments, 1D CNN, 2D CNN, LSTM, and CRNN architectures may be used individually or in combination.

[0055] In some embodiments, an ensemble of three different networks may be used, and the structure and hyperparameters of the ensemble may be fine-tuned using grid grinding to minimize overfitting. The outputs from each network may be combined to predict the probability of having COVID-19.

[0056] In some embodiments, the first network is for an MFCC with an input size of (39,) and includes two hidden layers with ReLU (rectified linear activation function) activation, followed by a dropout layer. The second network may be a convolutional neural network with a Mel spectrogram image of size (64,64,1) as input. The second network may include three 2D convolutional layers, with the first convolutional layer having a kernel size of 3 and a stride size of 2, and the remaining two convolutional layers having a kernel size of 3 and a stride size of 1, followed by 2D mean pooling, batch normalization, and ReLU activation, respectively. The third network corresponds to two additional features for each sample: fever or muscle pain and respiratory status. Similar to the first network, the third network includes two hidden layers with a ReLU activation function, followed by a dropout layer, respectively. The outputs from each network may be integrated and fed into two additional hidden layers, each with a ReLU activation function, and finally combined into a sigmoid (activation function) output determination layer.

[0057] In some embodiments, the ensemble network may be trained using cross-entropy loss, an Adam optimizer, and a learning rate of 0.001. The training data may be randomly split into training-validation-test datasets using a 70-15-15 split. Each training instance may be repeated five times using a different random data split. Mean statistics and 95% confidence intervals may be reported and stored in memory.

[0058] In some embodiments, both accuracy and the area under the receiver operating characteristic (ROC) curve (AUC) can be used as evaluation metrics. Because training data can be imbalanced, AUC can better represent how the model is performing.

[0059] In some embodiments, longitudinal crowdsourced and clinical studies conducted across various countries may be carried out to train machine learning algorithms using more information about human respiratory sound features, including cough and speech (or other forms of vocalization), both pre-symptomatic and during the course of COVID-19 infection. After collecting more vocal data in relation to PCR and evolving in vitro diagnostic methods for COVID-19, demographics, and disease progression labels, sub-analyses may be performed to validate the performance of ML models across a large number of symptomatic and demographic groups.

[0060] In some embodiments, machine learning algorithms include decision tree learning, artificial neural networks, deep learning neural networks, support vector machines, rule-based machine learning, random forests, and the like. Algorithms such as linear regression or logistic regression may be used as part of the machine learning process.

[0061] In some embodiments, support vector machines (SVMs) can be used as supervised learning models to analyze data for classification and regression analysis. An SVM may plot a set of data points in an n-dimensional space (e.g., where n is the number of clinical parameters), and classification is performed by finding a hyperplane that separates the set of data points into multiple classes. In some embodiments, the hyperplane is linear, and in other embodiments, the hyperplane is nonlinear. SVMs are effective in high-dimensional spaces, when the number of dimensions is greater than the number of data points, and generally work well with datasets where the separation margins are clear.

[0062] In some embodiments, decision trees can be used as a type of supervised learning algorithm also used in classification problems. Decision trees can be used to identify the most important variables that provide the best homogeneous set of data. A decision tree can split multiple groups of data points into one or more subsets, and each subset into one or more further categories, and so on, until it forms a terminal node (e.g., an unsplit node). Various algorithms, such as entropy, Gini impurity, chi-squared, information gain, and variance reduction, can be used to determine where the splits occur. Decision trees are often useful for quickly identifying the most important variables from a large number of variables or for identifying relationships between two or more variables. Furthermore, decision trees can handle both numerical and non-numerical data. This method is generally considered a non-parametric approach, meaning, for example, that the data does not need to fit a normal distribution.

[0063] In some embodiments, random forests (or random decision forests) can be used as a suitable approach for both classification and regression. In some embodiments, the random forest method constructs a collection of decision trees with small variance. Generally, for M input variables, fewer than M variables (nvars) are used to partition groups of data points. The process is repeated until the optimal partition is selected and a terminal node is reached. Random forests are particularly well-suited for processing a large number of input variables (e.g., thousands) to identify the most important variables. Random forests are also effective for estimating missing data.

[0064] In some embodiments, another machine learning technique, deep learning neural networks, may be used. These networks may have multiple hidden layers and can be manipulated in an automated manner (e.g., feature extraction).

[0065] In some embodiments, to train the machine learning system, the dataset is randomly divided into training data and validation data. A classifier is generated using the machine learning system based on the training data, a subset of the input, and other parameters related to the machine learning system described herein. It is determined whether the classifier satisfies predetermined receiver operator characteristic (ROC) statistics that define the sensitivity and specificity required to correctly classify patients. In embodiments, the specificity and sensitivity thresholds may be optimized to conform to FDA and WHO standards for medical devices; for example, for an antigen test, a specificity of 90% or higher and a sensitivity of 80% or higher may be specified.

[0066] If a classifier does not meet a predetermined ROC statistic, the classifier may be repeatedly generated based on different subsets of training data and inputs until the classifier meets the predetermined ROC statistic. If the machine learning system meets the predetermined ROC statistic, a static configuration of the classifier may be generated. This static configuration may be deployed in a hospital or healthcare facility, or stored on a remote server accessible to the hospital or healthcare facility, for use in identifying patients at risk of contracting COVID-19. In some cases, the results may be written to a patient's file on an electronic health record system.

[0067] In some embodiments, the exact nature and duration of a cough may vary from disease to disease, but intensity (strength), frequency (number of occurrences), and duration of cough (time since onset) are variables that can help identify an infectious disease (e.g., COVID-19) and distinguish an infected individual from a non-infectious state. For example, unlike certain acute conditions (e.g., COVID-19), coughs resulting from infectious diseases usually last longer. In some diseases, such as tuberculosis, a cough may last for several weeks.

[0068] Furthermore, one marker of respiratory infection is a change in voice quality caused by factors such as laryngeal inflammation or upper airway obstruction. In some embodiments, the likelihood of COVID-19 infection can be determined by combining information about voice behavior with other bioparameters (e.g., oxygen levels). In some embodiments, a voice sample from before infection and a recent voice recording may be obtained. In some embodiments, the difference (e.g., frequency) between these two voices can be calculated and used as an input feature.

[0069] In some embodiments, the audio stream is analyzed by receiving an audio stream via a telephone / microphone (e.g., mobile phone, VoIP, internet, etc.), segmenting the audio stream into short windows, calculating acoustic measurements from each window (e.g., Mel-frequency cepstrum coefficients), comparing the acoustic measurements across multiple consecutive windows, identifying cough acoustic patterns by developing and training a machine learning pattern recognition engine, and determining the likelihood that a particular window (or set of windows) contains an instance of coughing.

[0070] When a cough (or other target speech sample) is detected in the audio stream, the frequency, intensity, or other characteristics of the cough signal can be extracted and used as model input features (or intermediate features) to distinguish between diseases (e.g., COVID-19 and seasonal cold). For example, one disease may produce a "wet" cough characterized by a rumbling voice quality, while another disease may produce a "dry" cough (e.g., associated with COVID-19 patients) characterized by a hard initial consonant (fast attack time) followed by non-periodic (noise) energy.

[0071] In some embodiments, the first user may be asked to provide their patient record (which may be fully or partially anonymized where possible, and may omit personally identifiable and health information that is not relevant to the analysis at hand) as data input 20 to the controller 12 in Figure 1, as shown in block 110 in Figure 2. In some embodiments, the user may be asked to perform various actions via the user interface of a native app, which will take input to various upstream submodels that will feed into the ensemble model. Specifically, these include filling out a text questionnaire, breathing or coughing into a phone microphone, reading a text aloud within the microphone's pickup range, recording a video of a finger or other body part, and allowing data acquisition from wearable devices (such as a wrist-worn pulse oximeter, an inertial measurement unit (such as a step counter or a smartphone configured to extract the user's gait characteristics), a heart rate sensor, temperature, etc.).

[0072] Based on the patient record, several different analyses (e.g., cough classifier, deep breathing analysis, time-based data analysis, facial video, fingertip video, and biometric images) may be performed to assess the likelihood of COVID-19 infection in the first user, as shown in block 114 of Figure 2, as shown in block 112 of Figure 2.

[0073] In some embodiments, an individual's vocal behavior may be tracked over a longer period (e.g., by repeating the described sampling process and reprocessing new data) to determine how the cough changes over time. The change and its rate may serve as features of the models described herein. A rapid change or prolonged worsening of cough (or other vocal) behavior may indicate a particular disease state.

[0074] In some embodiments, voice samples may be used to determine novel clinically relevant outcome variables: the Cough Awakening Index (CAI) and the Cough Disturbance Index (CDI). The CAI reflects the number of nocturnal coughs associated with electroencephalogram (EEG) awakenings during each hour of sleep. If nocturnal coughs do not occur with EEG awakenings, they are counted in the Cough Disturbance Index (CDI), which is defined as the number of coughs per hour of sleep without awakenings. These novel indices can be used not only for the medical management of individual patients but also for medical research, for example, to understand the antitussive effects and antitussive profiles of pharmacological compounds.

[0075] In some embodiments, the user is then notified of the possibility of COVID-19 infection via the user interface, as shown by block 116 in Figure 2, by updating the user interface 18 to present such information on the user's mobile computer device, such as a smartphone. In some embodiments, the machine learning model is expected to be able to classify individuals who are likely to be infected with COVID-19 with a sensitivity of at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99%, when the specificity is set to 80%. This is expected to be superior to linear statistical models such as threshold classification with a single variable or multivariate logistic regression with multiple variables. In some embodiments, machine learning techniques are used to achieve at least a 5%, at least a 10%, at least a 15%, at least a 20%, at least a 25%, or at least a 30% improvement compared to conventional statistical methods such as conventional logistic regression or multivariate linear regression.

[0076] Figure 3 illustrates an exemplary computer system 1000 in which embodiments of the present technology may be implemented. For example, features of system 1000 may exist in both a smartphone and a server as described above. Various parts of the systems and methods described herein may include or be executed on one or more computer systems similar to computer system 1000. Furthermore, the processes and modules described herein may be executed by one or more processing systems similar to those of cocomputer system 1000.

[0077] The computer system 1000 may include one or more processors (e.g., processors 1010a to 1010n) coupled to system memory 1020, input / output I / O device interface 1030, and network interface 1040 via input / output (I / O) interface 1050. The processor may include a single processor or multiple processors (e.g., distributed processors). The processor may be any suitable processor capable of executing instructions. The processor may include a central processing unit (CPU) and / or image processing unit (GPU) that execute program instructions for the arithmetic, logical, and input / output operations of the computer system 1000. The processor may execute code that makes up the execution environment for program instructions (e.g., processor firmware, protocol stack, database management system, operating system, or a combination thereof). The processor may include a programmable processor. The processor may include a general-purpose or special-purpose microprocessor. The processor may receive instructions and data from memory (e.g., system memory 1020). The computer system 1000 may be a uniprocessor system including one processor (e.g., processor 1010a) or a multiprocessor system including any number of suitable processors (e.g., 1010a to 1010n). Multiple processors may be employed to achieve parallel or sequential execution of one or more parts of the technology described herein. Processes such as logic flows described herein are executed by one or more programmable processors that execute one or more computer programs and perform functions by manipulating input data to produce corresponding outputs. Processes described herein may also be executed by special-purpose logic circuits such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits), and the devices described herein may also be implemented by these.The computer system 1000 may include multiple computer devices (for example, a distributed computer system) to implement various processing functions.

[0078] The I / O device interface 1030 may provide an interface for connecting one or more I / O devices 1060 to the computer system 1000. The I / O devices may include devices that receive input (e.g., from a user) or output information (e.g., to a user). The I / O devices, for example, the client device 202, may include a graphical user interface presented on a display (e.g., a CRT (Cathode Ray Tube) or LCD (Liquid Crystal Display) monitor), a pointing device (e.g., a computer mouse or trackball), a keyboard, a keypad, a touchpad, a scanning device, a voice recognition device, a gesture recognition device, a printer, an audio speaker, a microphone, a camera, and the like. The I / O devices 1060 may be connected to the computer system 1000 via wired or wireless connections. The I / O devices 1060 may be connected to the computer system 1000 from a remote location. The I / O devices 1060 located in a remote computer system may be connected to the computer system 1000 via, for example, a network and the network interface 1040.

[0079] The network interface 1040 may include a network adapter that provides connectivity for the computer system 1000 to a network. The network interface 1040 may facilitate data exchange between the computer system 1000 and other devices connected to the network. The network interface 1040 may support wired or wireless communication. The network may include electronic communication networks such as the Internet, a local area network (LAN), a wide area network (WAN), or a cellular communication network.

[0080] System memory 1020 may be configured to store program instructions 1100 or data 1110. Program instructions 1100 may be executable by a processor (e.g., one or more of processors 1010a to 1010n) to implement one or more embodiments of the present technology. Instructions 1100 may include modules of computer program instructions for implementing one or more technologies described herein with respect to various processing modules. Program instructions may include computer programs (known in certain forms as programs, software, software applications, scripts, or code). Computer programs may be written in a programming language such as a compiled language, an interpreted language, a declarative language, or a procedural language. Computer programs may include units suitable for use in a computer environment, such as standalone programs, modules, components, and subroutines. Computer programs may or may not correspond to files in a file system. A program may be stored in part of a file that stores other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple collaborative files (e.g., a file that stores one or more modules, subprograms, or parts of code). Computer programs may be located locally at one site, or they may be distributed across multiple remote sites and configured to run on one or more computer processors interconnected by a communication network.

[0081] System memory 1020 may include a tangible program carrier for storing program instructions. The tangible program carrier may include a non-temporary computer-readable storage medium. The non-temporary computer-readable storage medium may include a machine-readable storage device, a machine-readable storage board, a storage device, or any combination thereof. The non-temporary computer-readable storage medium may include non-volatile memory (e.g., flash memory, ROM, PROM, EPROM, EEPROM memory), volatile memory (e.g., random access memory (RAM), static random access memory (SRAM), synchronous dynamic RAM (SDRAM)), bulk storage memory (e.g., CD-ROM and / or DVD-ROM, hard drive), etc. System memory 1020 may include a non-temporary computer-readable storage medium for storing program instructions that can be executed by a computer processor (e.g., one or more of processors 1010a to 1010n) in order to achieve the subject matter and functional operations described herein. Memory (e.g., system memory 1020) may include a single memory device and / or multiple memory devices (e.g., distributed memory devices). Instructions or other logarithmic power code that provide the functions described herein may be stored in a tangible, non-temporary, computer-readable medium. In some cases, the entire set of instructions may be stored on the medium simultaneously, or in some cases, different parts of the instructions may be stored on the same medium at different times.

[0082] The I / O interface 1050 may be configured to coordinate I / O traffic between processors 1010a-1010n, system memory 1020, network interface 1040, I / O device 1060, and / or other peripheral devices. The I / O interface 1050 may perform protocol conversion, timing conversion, or other data conversion to convert data signals from one component (e.g., system memory 1020) into a format suitable for use by another component (e.g., processors 1010a-1010n). The I / O interface 1050 may support devices connected via various types of peripheral buses, such as variations of the PCI (Peripheral Component Interconnect) bus standard, Bluetooth, WiFi, and USB (Universal Serial Bus) standard.

[0083] In implementing embodiments of the technology described herein, a single instance of computer system 1000 may be used, or multiple computer systems 1000 configured to host different parts or instances of the embodiment may be used. The multiple computer systems 1000 may provide parallel or sequential processing / execution of one or more parts of the technology described herein.

[0084] Those skilled in the art will understand that computer system 1000 is merely illustrative and is not intended to limit the scope of the technology described herein. Computer system 1000 may include any combination of devices or software that perform or otherwise provide the performance of the technology described herein. For example, computer system 1000 may include, or be a combination thereof, a cloud computing system, a data center, a server rack, a server, a virtual server, a desktop computer, a laptop computer, a tablet computer, a server device, a client device, a mobile phone, a PDA (Personal Digital Assistant), a portable audio / video player, a game console, an in-vehicle computer, or a GPS (Global Positioning System). Furthermore, computer system 1000 may be connected to other devices not shown, or it may operate as a standalone system. In addition, the functions provided by the illustrated components may, in some embodiments, be combined into fewer components or distributed among additional components. Similarly, in some embodiments, some functions of the illustrated components may not be provided, or other additional functions may be available.

[0085] Furthermore, although various items are illustrated to be stored on memory or storage during use, those skilled in the art will understand that these items or some of them may be transferred between memory and other storage devices for the purposes of memory management and data integrity. Alternatively, in other embodiments, some or all of the software components may run in memory on another device and communicate with the illustrated computer system via intercomputer communication. Also, some or all of the system components or data structures may be stored (e.g., as instructions or structured data) on a computer-accessible medium or portable device read by a suitable drive, various examples of which are described above. In some embodiments, instructions stored on a computer-accessible medium separate from computer system 1000 may be transmitted to computer system 1000 as a transmission medium or signal such as an electrical signal, electromagnetic signal, or digital signal transmitted via a communication medium such as a network or wireless link. Various embodiments may further include receiving, transmitting, or storing instructions or data implemented on a computer-accessible medium in accordance with the above description. Thus, the art of the present invention may be implemented in configurations of other computer systems.

[0086] In the block diagrams, the illustrated components are depicted as separate functional blocks, but embodiments are not limited to systems in which the functions described herein are organized as illustrated. The functions provided by each component may also be provided by software or hardware modules organized in a manner different from that currently illustrated, for example, such software or hardware may be mixed, combined, replicated, divided, distributed (e.g., within a data center or geographically), or otherwise organized in a different manner. The functions described herein may also be provided by one or more processors of one or more computers executing code stored in tangible, non-transient, machine-readable media. In some cases, regardless of the use of the singular term “medium,” instructions may be distributed on different storage devices associated with different computer devices, in which case, for example, each computer device may have a different subset of instructions. This is an implementation consistent with the use of the singular term “medium” herein. In some cases, a third-party content distribution network may host some or all of the information transmitted over the network, in which case the information (e.g., content) may be provided by sending instructions to retrieve the information from the content distribution network, to the extent that the information can be expressed as being supplied or otherwise provided.

[0087] Readers should understand that this application describes several individually useful technologies. The applicant has combined these technologies into a single document rather than dividing them into multiple separate patent applications, because the subject matter of these technologies is related, leading to economic efficiency in the filing process. However, the distinct merits or embodiments of such technologies should not be confused. In some cases, embodiments address all of the defects pointed out herein, but the technologies are independently useful, and it should be understood that some embodiments address only a subset of such problems or offer other unmentioned merits that would be obvious to a person skilled in the art browsing this disclosure. Due to cost constraints, some technologies disclosed herein may not currently be claimed, and may be claimed in a subsequent application, such as a continuation application, or by amending the current claims. Similarly, for space reasons, the “Abstract” or “Summary of Invention” sections of this document should not be considered to comprehensively describe all of such technologies or all embodiments of such technologies.

[0088] It should be understood that the detailed description and drawings are not intended to limit the Art to any particular form disclosed, but rather to cover all modifications, equivalents, and substitutions that fall within the spirit and scope of the Art as defined by the appended claims. Further modifications and alternative embodiments of various aspects of the Art will be apparent to those skilled in the art by reading this description. Therefore, this description and drawings should be interpreted as illustrative only and are intended to teach those skilled in the art a general way of carrying out the Art. It should be understood that the forms of the Art illustrated and described herein should be considered as examples of embodiments. Various elements and materials may be used in place of those illustrated and described herein, parts and processes may be reversed or omitted, and certain features of the Art may be used independently, all of which will be apparent to those skilled in the art after benefiting from this description of the Art. Modifications to the elements described herein may be made without departing from the spirit and scope of the Art set forth in the following claims. The headings used herein are for organizational purposes only and are not intended to limit the scope of the description.

[0089] As used throughout this application, the word “may” is used in an allowable sense (i.e., it may be done) rather than an essential sense (i.e., it must be done). Words such as “include,” “including,” and “includes” mean that they include but are not limited to them. In this application, the singular forms “a,” “an,” and “the” include plural things unless the content explicitly indicates a different meaning. Thus, for example, a reference to “an element” or “a element” includes a combination of two or more elements, regardless of the use of other terms and phrases for one or more elements, such as “one or more.” The term “or” is non-exclusive unless another meaning is explicitly stated, i.e., it encompasses both “and” and “or.” Terms expressing conditional relationships, such as "in response to X,Y," "on X,Y," "if X,Y," and "when X,Y," encompass causal relationships where the antecedent is a necessary causal condition, a sufficient causal condition, or a strong causal condition for the result. For example, "State X occurs when condition Y is met" is comprehensive to both "X occurs only when Y is met" and "X occurs when Y and Z are met." Such conditional relationships are not limited to those where the result occurs immediately upon the fulfillment of the antecedent condition; some may have a delayed result. Furthermore, in conditional statements, the antecedent condition and the result are linked; for example, the antecedent condition relates to the likelihood of the result occurring.A description that maps multiple attributes or functions to multiple objects (for example, one or more processors performing steps A, B, C, and D) includes, unless otherwise indicated, both cases where all of those attributes or functions map to all of those objects, and cases where a subset of those attributes or functions maps to a subset of those attributes or functions (for example, both cases where all processors each perform steps A through D, and cases where processor 1 performs step A, processor 2 performs steps B and part of step C, and processor 3 performs part of step C and step D). Similarly, expressions such as "computer system" performing step A and "computer system" performing step B may include the same computer device within the computer system performing both steps, or different computer devices within the computer system performing steps A and B. Furthermore, a description that one value or action "based on" another condition or value includes, unless otherwise indicated, both cases where that condition or value is the sole factor, and cases where that condition or value is one of several factors. A statement that "each" instance of a collection has a certain characteristic should not be interpreted, unless otherwise indicated, as excluding cases where identical or similar members of a larger collection do not possess that characteristic. In other words, "each" does not necessarily mean all. For example, unless explicitly specified, such as "perform X, then perform Y," a claim should not be read as restricting the order of the described steps. Conversely, a statement that could be inappropriately claimed to imply an order restriction, such as "perform X on an item, then perform Y on an item that has been Xed," is used to make the claim easier to read, rather than to specify an order. Also, a statement such as "at least Z of A, B, and C" (e.g., "at least Z of A, B, or C") refers to at least Z of each enumerated category (A, B, and C), not that each category requires at least Z units.As is evident from the discussion, in this specification, discussions using terms such as “processing,” “computer,” “calculation,” and “decision” are understood to refer to the operation or process of a specific device, such as a special-purpose computer or similar special-purpose electronic processing / calculation device, unless otherwise specified. Features described in reference to geometric structures such as “parallel,” “perpendicular / orthogonal,” “square,” and “cylindrical” should be interpreted as encompassing items that substantially embody the characteristics of that geometric structure; for example, a reference to a “parallel” surface would encompass substantially parallel surfaces. The permissible deviation of these geometric structures from the Platonic concept should be determined by reference to the scope in the specification, and if no such scope is specified, to industry norms in the art of use; if no such scope is defined, to industry norms in the art of manufacture of the specified feature; and if no such scope is defined, to interpret that features that substantially embody a geometric structure include features no more than 15% of the defining attributes of that geometric structure. The terms “first,” “second,” “third,” and “predetermined” used in the claims are used for distinction or identification purposes and do not indicate a sequential or numerical limitation. As is common in the art, data structures and formats described with reference to a use prominent to humans do not need to be presented in a human-readable form to constitute the data structure or format described above. For example, to constitute text, it is not necessary to render the text or encode it in Unicode or ASCII; to constitute images, maps, and data visualizations, it is not necessary to display and decode the images, maps, and data visualizations, respectively; and to constitute speech, music, and other sounds, it is not necessary to emit or decode the speech, music, and other sounds, respectively. Instructions, commands, etc., implemented in a computer are not limited to executable code, but can also be implemented in the form of data that provides functionality, such as arguments to functions or API calls.If noun phrases (and other neologisms) created for a specific purpose are used in a claim, and to the extent that there is no obvious interpretation, the definition of such phrase may be stated in the claim itself, in which case the use of such noun phrase should not be deemed to impose additional restrictions by reference to the specification or external evidence.

[0090] This patent specification incorporates certain U.S. patents, U.S. patent applications, or other materials (e.g., papers) by reference. However, the text of such U.S. patents, U.S. patent applications, and other materials is incorporated by reference only to the extent that there is no conflict between such materials and the descriptions and drawings contained herein. In the event of such conflict, the text of this specification shall prevail, and the terms used herein should not be construed more narrowly because they are used in other materials incorporated by reference.

[0091] The technology of the present invention will be better understood by referring to the embodiments listed below. [Embodiment 1] A tangible, non-temporary, machine-readable medium for storing instructions, wherein when the instructions are executed by one or more processors, a trained machine learning model is acquired which is configured to use a computer system to infer whether the user has a respiratory disease based on both audio and images acquired by the user's mobile computer device, the trained machine learning model is trained by acquiring a training set which includes a plurality of training records, each of the plurality of training records in the training set which includes a plurality of parameters and corresponding values ​​for the person, each of the plurality of training records in the training set which includes an audio of the person's voice and an image of at least a portion of the person, and each of the plurality of training records in the training set which includes the person calling A machine-readable medium in which a process is performed, the computer system includes information indicating whether or not the user has been diagnosed with a respiratory illness, trains the machine learning model with the training set, infers from both the audio and the images whether or not the user has the respiratory illness, obtains the trained machine learning model, then receives a first user record of the first user, the first user record includes an audio file or audio stream of the first user's cough and an image of at least a portion of the first user, the computer system infers that the first user has the respiratory illness based on the audio file or audio stream of the first user's cough and the image of at least a portion of the first user, and the computer system stores in memory information indicating that the first user has the respiratory illness. [Embodiment 2] The machine-readable medium according to Embodiment 1 includes at least two of the following: text questionnaire responses, respiration data, time data, facial video, fingertip video, and bio-images. [Embodiment 3] The machine-readable medium according to Embodiment 1 includes, of course, text questionnaire responses, respiration data, time data, facial videos, fingertip videos, and bio-images. [Embodiment 4] The machine-readable medium according to any one of Embodiments 1 to 3, wherein the plurality of training records include videos of fingertips, and the machine learning model is trained using the videos of fingertips to measure blood oxygen saturation and heart rate as features to be used for inference. [Embodiment 5] The machine-readable medium according to any one of embodiments 1 to 4, wherein the processing further comprises training the machine learning model. [Embodiment 6] A machine-readable medium according to any one of Embodiments 1 to 5, wherein training the machine learning model includes calculating partial derivatives of the parameters of the machine learning model with respect to an objective function, and adjusting the parameters of the machine learning model in the direction indicated by the partial derivatives so that the machine learning model is locally optimized. [Embodiment 7] The machine-readable medium according to any one of Embodiments 1 to 6, wherein the machine learning model outputs a first output indicating infection with the novel coronavirus and a second output indicating the stage of infection with the novel coronavirus. [Embodiment 8] The machine-readable medium according to any one of Embodiments 1 to 7, wherein the machine learning model has means for combining the outputs of a plurality of submodels. [Embodiment 9] The machine-readable medium according to any one of embodiments 1 to 8, further comprising setting up lossy compression of the audio file or audio stream of the cough to store data that is not perceptible to humans and affects the accuracy of the trained machine learning model. [Embodiment 10] The machine-readable medium according to any one of embodiments 1 to 9, wherein the training of the machine learning model is performed by a set of computers in the computer system that is different from the set of computers that perform the inference that the first user has the respiratory disease. [Embodiment 11] The machine-readable medium according to any one of Embodiments 1 to 10 is configured such that the inference that the first user has the respiratory disease is performed by the first user's smartphone, which is part of the computer system. [Embodiment 12] The machine-readable medium according to any one of embodiments 1 to 11, wherein the trained machine learning model comprises an ensemble of at least three different neural networks having outputs combined by means for ensembling multiple submodels. [Embodiment 13] The machine-readable medium according to any one of embodiments 1 to 12, wherein the processing includes preprocessing the voice cough sample, cleaning the voice cough sample and selecting a segment of the voice cough sample to be input to the trained machine learning model, before inputting it to the trained machine learning model. [Embodiment 14] The machine-readable medium according to any one of Embodiments 1 to 11, wherein the processing includes extracting cepstrum coefficients from a voice cough sample. [Embodiment 15] Extracting the cepstrum coefficients from the voice cough sample comprises constructing a spectrogram from the voice cough sample, calculating logarithmic power frame by frame from the spectrogram, applying a filter to the magnitude of the logarithmic power, performing logarithmic compression to convert the output of the filter into the cepstrum region, and forming a vector of cepstrum coefficients frame by frame, according to Embodiment 14 of the machine-readable medium. [Embodiment 16] The machine-readable medium according to any one of Embodiments 1 to 15, wherein the processing includes extracting Mel-frequency cepstrum coefficients from the power spectrum of the audio of a second user's cough sample. [Embodiment 17] The machine-readable medium according to any one of Embodiments 1 to 16, wherein the trained machine learning model includes a multilayer feedforward neural network having at least two nonlinear layers. [Embodiment 18] The machine-readable medium according to any one of Embodiments 1 to 17, wherein the trained Mel machine learning model includes a dual parallel feedforward neural network that takes a vector of Mel-frequency cepstrum coefficients as input. [Embodiment 19] A method comprising the process described in any one of Embodiments 1 to 18. [Embodiment 20] A system comprising one or more processors and memory for storing instructions, wherein when an instruction is executed by one or more processors, a process including the process described in any one of embodiments 1 to 12 is executed.

Claims

1. A tangible, non-temporary, machine-readable medium for storing instructions, wherein when the instructions are executed by one or more processors, A computer system is used to obtain a trained machine learning model configured to infer whether or not a user has a respiratory illness based on both audio and images acquired by the user's mobile computer device. The aforementioned trained machine learning model is trained by obtaining a training set containing multiple training records. Each of the multiple training records in the training set includes multiple parameters for the individual corresponding to that training record, and a value corresponding to each of the multiple parameters, wherein the value represents data obtained for that individual. Each of the training records in the training set includes an audio recording of the individual's voice and an image of at least a portion of the individual. Each of the training records within the aforementioned training set includes information indicating whether or not the individual has been diagnosed with a respiratory disease. The machine learning model is trained with the training set, and from both the audio and the images, it is inferred whether or not the user has the respiratory disease. After obtaining the trained machine learning model, the computer system receives the first user record of the first user. The first user record includes an audio file or audio stream of the first user's voice and an image of at least a portion of the first user. In the computer system, based on the audio file or audio stream of the first user's voice and images of at least a portion of the first user, the system infers that the first user has the respiratory disease. A machine-readable medium in which a process is performed to store in memory information indicating that the first user has the respiratory disease in the computer system.

2. The machine-readable medium according to claim 1, wherein the plurality of training records include at least two of the following (1) to (6). (1) Responses to a text-based questionnaire (2) Data indicating respiration (3) Time data (4) Video of the face (5) Video of fingertips (6) A biological image of skin, feces, mucus, urine, or vomit.

3. The aforementioned training records include videos of fingertips, The machine-readable medium according to claim 1, wherein the machine learning model is trained using the video of the fingertip to measure blood oxygen concentration and heart rate as features to be used as the basis for inference.

4. The machine-readable medium according to claim 1, wherein training the machine learning model includes calculating partial derivatives of the parameters of the machine learning model with respect to an objective function, and adjusting the parameters of the machine learning model in the direction indicated by the partial derivatives so that the machine learning model is locally optimized.

5. The machine-readable medium according to claim 1, wherein the machine learning model includes at least two outputs: a first output indicating infection with the novel coronavirus and a second output indicating the stage of infection with the novel coronavirus.

6. The machine-readable medium according to claim 1, further comprising setting up lossy compression of the audio file or audio stream of the individual's voice in order to store data that is not perceptible to humans and affects the accuracy of the trained machine learning model.

7. The machine-readable medium according to claim 1, wherein the training of the machine learning model is performed by a set of computers in the computer system that is different from the set of computers that perform the inference that the first user has the respiratory disease.

8. The machine-readable medium according to claim 1, wherein the trained machine learning model comprises an ensemble of at least three different machine learning algorithms having outputs combined with the trained ensemble model.

9. The aforementioned process is, The machine-readable medium according to claim 1, further comprising preprocessing the audio sample, which includes cleaning the audio sample and selecting a segment of the audio sample to be input to the trained machine learning model, before inputting it to the trained machine learning model.

10. The aforementioned process is, The machine-readable medium according to claim 1, further comprising extracting cepstrum coefficients from the audio file or audio stream of the first user's voice.

11. Extracting the aforementioned cepstrum coefficients is Constructing a spectrogram from the audio file or audio stream of the first user's voice, Calculating the logarithmic power for each frame from the aforementioned spectrogram, Applying a filter to the magnitude of the logarithmic power, The output of the filter is subjected to logarithmic compression and converted to the cepstrum region, and The machine-readable medium according to claim 10, comprising forming a vector of cepstrum coefficients for each frame.

12. The aforementioned process is, The machine-readable medium according to claim 1, further comprising extracting Mel-frequency cepstrum coefficients from the audio power spectrum of a second user's voice sample.

13. The trained machine learning model includes a multilayer feedforward neural network comprising at least two nonlinear layers, The machine-readable medium according to claim 1, wherein the trained machine learning model includes a dual-parallel feedforward neural network that takes a vector of Mel-frequency cepstrum coefficients as input.

14. A method comprising the processing described in any one of claims 1 to 13.

15. It comprises one or more processors and memory for storing instructions, A system in which, when the instruction is executed by at least a portion of the one or more processors, the process described in any one of claims 1 to 13 is performed.

Citation Information

Patent Citations

  • Evaluation of pulmonary disease by voice analysis

    JP2018534026A

  • Information processing system

    JP2020035178A

  • Respiratory status management based on respiratory system sounds

    JP2021524958A

  • Intelligent Health Monitoring

    US20200146623A1

  • Managing respiratory conditions based on sounds of the respiratory system

    WO2019229543A1