Computer-Implemented Method, System, and Computer Program Product (Multimodal Spirometry for Respiratory Disease Prediction)

The computer-implemented multimodal spirometry system addresses the challenge of undetected respiratory diseases by using audio and video processing to predict lung capacity, offering a cost-effective and accessible method for early detection.

JP7759691B2Active Publication Date: 2025-10-24INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021165451
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-08
Filing Date
2021-10-07
Publication Date
2025-10-24
Estimated Expiration
2041-10-07

AI Technical Summary

Technical Problem

Respiratory diseases often go undetected until they become severe due to the slow decline in lung capacity, and current methods like X-rays and CT scans are expensive and not readily accessible.

Method used

A computer-implemented method and system for multimodal spirometry using audio and video processing to assess lung capacity through speech analysis, utilizing machine learning models to predict respiratory health.

Benefits of technology

Enables early detection of respiratory disorders with a low-cost, accessible solution that allows users to self-assess their lung capacity reliably, broadening the scope of testing in terms of population and frequency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007759691000001
    Figure 0007759691000001
  • Figure 0007759691000002
    Figure 0007759691000002
  • Figure 0007759691000003
    Figure 0007759691000003
Patent Text Reader

Abstract

To provide a computer packaging method, a system and a computer program product that enable multi-modal lung capacity measurement for respiratory illness prediction.SOLUTION: Determination of lung capacity involves capturing an audio waveform of a user performing an utterance presented to the user. It is possible to capture a video of the speaking user. A captured audio waveform and the video are analyzed for specifications. A respiratory function indicator is determined based on the audio waveform. The indicator is compared to a reference indicator to determine a health condition of the user. Machine learning models, such as neural networks, may be trained to predict indicators of respiratory function based on input features including speech spectra and temporal features of utterances. Determining indicators or respiratory function may include performing trained machine learning models.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates generally to computers and computer applications, and more particularly to multimedia and audio processing for health detection. [Background technology]

[0002] Respiratory diseases often affect the lungs slowly and are only detected once the impact becomes severe. As the disease affects the body, the lung's capacity to hold air slowly decreases. Early detection of changes in lung capacity is useful, but current methods, such as X-rays and CT scans, require special setups and experts to analyze the scans. For example, these methods are expensive and not readily available. Summary of the Invention [Problem to be solved by the invention]

[0003] A computer-implemented method, system, and computer program product for implementing multimodal spirometry for respiratory disease prediction is provided. [Means for solving the problem]

[0004] A method and system for multimodal spirometry for respiratory disease prediction may be provided. In one aspect, the method may include presenting a user with a specification of an utterance to be performed. The method may also include capturing a voice waveform of the user performing the utterance. The method may further include capturing video of the user performing the utterance. The method may also include analyzing the captured voice waveform and video for compliance with the specification. The method may also include determining an index of respiratory function based on the voice waveform. The method may also include comparing the index to a reference index to determine the user's health status.

[0005] In one aspect, the system may include a processor and a memory coupled to the processor. The processor may be configured to present a specification of an utterance to a user to be performed. The processor may be configured to capture an audio waveform of the user performing the utterance. The processor may be configured to capture video of the user performing the utterance. The processor may be configured to analyze the captured audio waveform and video for compliance with the specification. The processor may be configured to determine an index of respiratory function based on the audio waveform. The processor may be configured to compare the index to a reference index to determine a health status of the user.

[0006] A computer-readable storage medium storing a program of instructions executable by a machine to perform one or more of the methods described herein may also be provided.

[0007] Further features, as well as the structure and operation of various embodiments, are described in detail below with reference to the accompanying drawings, where like reference numbers indicate identical or functionally similar elements. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 illustrates an overview of multimodal spirometry for detecting possible respiratory disorders in one embodiment. [Figure 2] FIG. 1 illustrates training a machine learning model to predict vital capacity in one embodiment. [Figure 3] 1 illustrates an individual registration process in one embodiment. [Figure 4] FIG. 1 illustrates the administration of a test to detect respiratory illness to an individual, which may occur in one embodiment. [Figure 5] FIG. 1 illustrates spectral and modulation characteristics for voice activity detection (SAD) in one embodiment. [Figure 6] FIG. 1 illustrates a method for determining vital capacity in one embodiment. [Figure 7]FIG. 1 illustrates a computer-implemented method in one embodiment. [Figure 8] FIG. 1 illustrates components of an embodiment of a system for determining a user's lung capacity. [Figure 9] FIG. 1 shows a schematic diagram of an exemplary computer or processing system on which the system may be implemented in one embodiment. [Figure 10] FIG. 1 illustrates a cloud computing environment in one embodiment. [Figure 11] FIG. 1 illustrates a set of functional abstraction layers provided by a cloud computing environment in one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] For example, systems and methods for determining lung capacity for respiratory disease prediction are disclosed. In one or more embodiments, the systems and / or methods determine lung capacity based on multimedia analysis, e.g., audio and video processing. For example, the systems and / or methods can utilize protocols that correlate speech characteristics with critical lung capacity. In one aspect, a simple and low-cost application on a user device, such as a mobile device, can help users self-monitor their progress. Enabling users to reliably self-assess their lung capacity can broaden the scope of testing in terms of both population and frequency.

[0010] In one aspect, the system and method can use the production of various types of speech sounds as an indicator of the wellness or health of an individual's respiratory system. For example, vowels are sounds produced when air from the lungs passes through the mouth with minimal obstruction and without audible fricatives. As air leaves the lungs, the vocal cords vibrate to produce these sounds. Because these sounds (e.g., / a / / e / / i / / o / / u / ) are produced as air is forced out of the lungs, the ability to produce sustained, sustained vowels, for example, can be used as a measure of lung capacity. If the lungs are infected or filled with fluid such as mucus, or both, lung capacity decreases, affecting an individual's ability to produce sustained vowels. Infection of the vocal tract can also affect the production of these sounds. As another example, nasals are consonants produced when air passes through the nasal passages. If a sustained nasal sound such as / m / is produced with the mouth closed, it can be used as an indicator of how congested the nasal passages are.

[0011] Human speech can be distinguished from other vocalizations by distinct characteristics. For example, the typical speaking rate in English can be generalized to four syllables per second. Abnormalities in speech characteristics may be recognized as abnormalities in an individual's condition and can be used as indicators of an individual's respiratory health.

[0012] FIG. 1 is a diagram illustrating an overview of multimodal spirometry for detecting respiratory disorders, which may occur in one embodiment. The components illustrated in FIG. 1 may be implemented or executed on a device or computer having one or more processors, such as, for example, one or more hardware processors. For example, the device or computer may be a mobile device running an application (e.g., a mobile app) or other device. The one or more hardware processors may comprise components such as, for example, a programmable logic device, a microcontroller, a memory device, or other hardware components, or a combination thereof, configured to perform the respective tasks described in this disclosure. The associated memory device may be configured to selectively store instructions executable by the one or more hardware processors.

[0013] The processor may be a central processing unit (CPU), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or other suitable processing component or device, or one or more combinations thereof. The processor may be coupled to a memory device. The memory device may comprise a random access memory (RAM), a read-only memory (ROM), or other memory device, and may store data and / or processor instructions for implementing various functions associated with the methods and / or systems described herein. The processor may execute computer instructions stored in the memory or received from another computer device or medium.

[0014] In step 102, the processor (e.g., running an application or executing processor instructions) prompts the user to utter a series of sounds while holding their breath. For example, uttering a series of sounds may include uttering vowels such as "a-e-i-o-u" (or other sequences of vowels) multiple times until the user runs out of breath. Other examples may include uttering consonants, words, reading specific text, and / or others. This methodology may also work with other natural languages ​​and vowels, consonants, and other sounds that have similar effects to words.

[0015] In step 104, the processor can check the user's vocalization for consistency. For example, the processor can automatically and / or autonomously check for discrepancies or errors made while the user is producing the sounds that could affect the accurate measurement of lung capacity. In this process, the processor can use video analysis and / or audio analysis to determine consistency. For example, the methodology can detect rushed vocalizations, such as counting, e.g., "1234567..." versus "1 2 3 4 5 6 7...." Such rushed counting can result in less air being used for each digit, distorting the maximum count used as a measurement. In one embodiment, the methodology can examine the intensity envelope of the sounds to determine and reinforce pacing. As another example, shortened or lengthened words (e.g., "ichi ni san shi i..." or "ichi ni san shi...") may result in variations in the amount of air used in each vowel segment and variations in speech time, which may distort the measurements. In one embodiment, the methodology may measure vowel duration directly from the spectrogram to provide feedback. As yet another example, murmuring or partial whispering (e.g., "chin nin shi..." or "iiichi ni ii saan...") may use a smaller amount of exhaled air than normal. In one embodiment, the methodology may analyze the signal-to-noise ratio of vowel formants to reject speech.

[0016] Image processing of video captured while a user is making or uttering a sound can be analyzed for suitability. For example, captured images of a user making a sound can detect whether the user is using appropriate facial movements, mouth movements, or lip movements, or a combination thereof, or is in an appropriate posture when making the sound, to provide an appropriate standard basis for making measurements, e.g., to avoid distorting the measurement of the sound characteristics being calculated. For example, image processing of the video can determine continuity of identity from the video and determine that data from the same user is being tested or trained. The processor, for example, captures sounds made by the user within the suitability of making such sounds. If a mismatch is detected while the user is providing an utterance, the processor can request the user to speak again or try again.

[0017] In step 106, the processor may calculate features associated with the captured sound. For example, the processor may analyze the sound and calculate various characteristics of the sound produced by the user. The processor may also determine or obtain a spirometry measurement for the user that correlates to the calculated feature or characteristics. Considering that the user produced the sound in a healthy state (e.g., a "normal" state of the lungs), the obtained spirometry measurement (or measurement range) is associated with the user's "normal" or "healthy" lung capacity. In one embodiment, determining the user's spirometry may include running a trained machine learning model, such as a neural network model, with the calculated features as input features to the trained machine learning model. The machine learning model outputs a spirometry measurement (or measurement range) that corresponds to the input features.

[0018] The types of sounds may include vowels and consonants. Such sounds exhibit distinct spectral signatures that can be used to identify them in a spectrogram of speech. Vowels have distinct formant frequencies, and consonants have bursts of energy in distinct frequency bands. The modulation spectrum of speech exhibits distinct spectrotemporal patterns. Thus, for example, features or characteristics may include, but are not limited to, spectrotemporal characteristics such as speaking rate.

[0019] In step 108, the processor may store the calculated characteristics and spirometry measurements. The stored information may be used as a baseline or reference point. Using such baseline values, the processor may, for example, determine or assess normal or abnormal spirometry at different times or later.

[0020] The processes in steps 102, 104, 106, and 108 may be referred to as calibration processes or scaling. For example, such processes may allow for the calibration of an individual's baseline or reference point for assessing respiratory or pulmonary health, which may be specific to a particular individual.

[0021] In another aspect, the baseline calibration process (e.g., as shown in steps 102, 104, 106, and 108) can be performed on a general population group that shares common characteristics, such as demographic and / or physical characteristics. For example, sample sounds from a group of users or individuals may be captured, and corresponding spirometry measurements (or measurement ranges) can be obtained, for example, by running a trained machine learning model. The obtained measurements can be used as a baseline for that group (e.g., a "normal" spirometry value for the group's health stage).

[0022] The processes of steps 110, 112, 114, and 116 can be performed to determine a user's or individual's lung capacity at a given time (e.g., the current time). In step 110, a processor (e.g., running an application or executing processor instructions) prompts the user to utter or repeat a series of sounds. For example, the user may be prompted to take a breath and then speak or utter a series of sounds. For example, uttering a series of sounds may include uttering vowels such as "aeiou" (or other sequences of vowels) multiple times, e.g., until the user runs out of breath. Other examples may include uttering consonants, words, reading specific text, and / or the like. For example, the user may be prompted to utter sounds similar to those uttered during baseline determination, e.g., in step 102.

[0023] In step 112, the processor checks compatibility while the user is making sounds, for example as is done in step 104, using at least video analysis, for example.

[0024] In step 114, the processor calculates features or characteristics associated with the user's sound production. Based on the calculated features, the processor can determine a lung capacity, for example, the user's current lung capacity. Determining the lung capacity may include running a trained machine learning model. The trained machine learning model may be the same model used to determine a baseline value for the user or group of users.

[0025] In step 116, the processor compares the current lung capacity to, for example, the baseline volume stored in step 108, to determine whether the current lung capacity is outside of the baseline volume range. If the current lung capacity is outside of the range, the processor determines that the current lung capacity is outside of the normal range, which may indicate a respiratory disorder.

[0026] In one aspect, the machine learning model may be a neural network model or other machine learning model trained to predict spirometry measurements given a set of features or characteristics associated with a user's vocalizations. For example, training data for the machine learning model or neural network model may include labeled data containing characteristics of sounds made by the user that are associated with spirometry measurements.

[0027] For example, a neural network can be trained with input features as sound characteristics and output as spirometry data. The neural network's parameters (e.g., weights and biases) may be optimized to associate input sound features with output spirometry measurements. Simply put, an artificial neural network, or neural network, is a machine learning model that can be trained to predict or classify input data. An artificial neural network may comprise successive layers of neurons, which are interconnected so that output signals from neurons in one layer are weighted and sent to neurons in the next layer. A neuron N in a given layer may be connected to one or more neurons N in the next layer, and a different weight w may be associated with each neuron-neuron connection N-N to weight the signal sent from N to N. Neuron N generates an output signal in response to the accumulated input, and the weighted signal may propagate through successive layers of the network, from the input to the output neuron layer. An artificial neural network machine learning model may undergo a training phase in which a set of weights associated with each neuron layer is determined. The network is exposed to a set of training data in an iterative training scheme in which the weights are repeatedly updated as the network "learns" from the training data. The resulting trained model can then be applied to perform tasks based on new data, using the weights defined by the training process.

[0028] In one embodiment, the self-calibration process may include the following: Using a mobile phone or other device running an app, a user may be prompted to provide a "normal" spirometry measurement by repeatedly uttering "a," "e," "i," "o," "u" (or another utterance) until the user runs out of breath. This speech characteristic is used to create a baseline value for lung air capacity (e.g., mean time and variance across normal samples). The above steps may be repeated and averaged over a period of time (e.g., one week) to calibrate the "normal" state. The user may be asked to read a paragraph on the screen of their device. From the speech signal, a modulation spectrum may be calculated to estimate a baseline value for speaking rate. During a testing phase, the user may be asked to utter a vowel until the user runs out of breath, and a processor on the user's device may record the user making the utterance. A difference in the time averaged across the vowel from the calibrated baseline value may indicate the onset of a respiratory disorder. For example, if the difference exceeds a threshold, the processor may indicate that the spirometry is not normal. In another aspect, a user may be asked to read a paragraph of text displayed on a screen of the user's device. A processor in the user's device may calculate a modulation spectrum and estimate the user's speaking rate. The processor may compare this speaking rate with a baseline speaking rate. If the deviation or difference is greater than a threshold, the processor may signal a decrease in lung capacity.

[0029] FIG. 2 illustrates training a machine learning model for predicting vital capacity in one embodiment. A voice recording database 202 may include data representing voice recordings of multiple individuals. For example, an individual may be asked to inhale as much air as possible and begin to produce a speech sound, e.g., a vowel or consonant, or both, singly or in succession, while exhaling. The individual may be asked to repeat such sound production, e.g., multiple times, for a period of time, or for as long as the individual can comfortably produce it. For each subject or individual who provides a recording, vital capacity is also measured. The recordings and associated measured vital capacity may be stored as database 202 in one or more storage devices or systems.

[0030] In the voice processing system 204, the stored user voice is processed to extract or calculate features or characteristics of the user voice. The extracted characteristics may include characteristics related to voice / non-voice activity 206. For example, data such as spectral characteristics, modulation frequencies, and joint spectrotemporal characteristics may be extracted or calculated from the voice recording.

[0031] In the machine learning system 208, using the features or characteristics and the associated lung capacity data, a machine learning model, e.g., a neural network model, can be trained to predict lung capacity. For example, the features and associated lung capacity data are used as training data to optimize or train parameters (e.g., weights and biases) of a neural network or other machine learning model. The trained neural network (or other machine learning model) can then be used to predict lung capacity 210, for example, given a new set of features that the neural network has not seen before.

[0032] FIG. 3 illustrates an individual enrollment process in one embodiment. This process customizes or determines a reference point or level for a particular individual to determine their vital capacity. Audio data 302 is received from the individual. For example, the individual can be prompted to produce a vowel, consonant, or other letter-like sound, or a combination thereof, and the sound is captured. In an audio processing system 304, audio processing extracts or calculates features or characteristics from the captured or received sound. The extracted characteristics may include, but are not limited to, speech / non-speech activity 306. In a trained machine learning system 308, a trained machine learning model, such as a neural network model, runs using the characteristics as input features for the machine learning model to predict vital capacity 310. For example, the machine learning model can be a neural network trained according to the process shown in FIG. 2. In a data analysis system 312, predicted data (vital capacity) 310 can be designated as a baseline or reference point for the individual. The predicted vital capacity can be stored, for example, in a database 314 or a storage device.

[0033] FIG. 4 illustrates one embodiment of a system for testing an individual to detect possible respiratory illnesses. Such testing may be performed, for example, on a user's mobile device or other device running an application or app with a user interface. Audio data 402 is received from the individual. For example, the individual may be prompted to vocalize a vowel, consonant, or other letter-like sound, or a combination thereof, and the sound is captured. In an audio processing system 404, audio processing extracts or calculates features or characteristics from the captured or received sound. The extracted characteristics may include, but are not limited to, speech / non-speech activity 406. In a trained machine learning system 408, a trained machine learning model, e.g., a neural network model, runs using the characteristics as input features for the machine learning model to predict lung capacity 410. For example, the machine learning model may be a neural network trained according to the process shown in FIG. 2. In the machine learning system 414, the predicted data (vital capacity) 410 is compared to baseline values ​​or reference points received or obtained from a database 412 that has previously stored (e.g., as shown in FIG. 3) and stores baseline values ​​for this user or individual. If the difference between the predicted vital capacity 410 and the baseline vital capacity exceeds or falls outside a threshold (e.g., a predefined tolerance range), processing in the machine learning system 414 flags or signals that the user's vital capacity is outside the normal range.

[0034] Speech or non-speech activity detection (e.g., speech / non-speech activity 206 in FIG. 2 , speech / non-speech activity 306 in FIG. 3 , and / or speech / non-speech activity 406 in FIG. 4 ) can include the following: Speech is a sequence of consonants and vowels, both inharmonic and harmonic, with natural silences between them. This makes speech a complex signal with a wide range of spectrotemporal modulations. Useful temporal modulations of speech range from 0 to 20 Hz, with a peak around 4 Hz. Meanwhile, spectral modulations range from 0 to 6 cycles per octave. Pitch or voicing results in modulations in the range of 2 to 6 cycles per octave, while modulations below 2 cycles per octave reflect formant information. Effective speech / non-speech detection can be performed using various types of acoustic features that capture information based on these modulation characteristics of speech.

[0035] Figure 5 illustrates spectral and modulation features for voice activity detection (SAD) in one embodiment. Acoustic features may be generated using different signal processing techniques and can be broadly categorized by the type of modulation they capture: short-term spectral features, long-term modulation frequencies, and joint spectrotemporal features. Short-term spectral features can be extracted from power spectrum estimates of short analysis windows (e.g., 10-30 ms) of the speech signal, such as Mel-Frequency Cepstral Coefficients (MFCCs) and perceptual linear prediction (PLP) features. Long-term modulation frequency components can be estimated over longer analysis windows spanning hundreds of milliseconds from the speech subband envelope, such as subband log-Mel energy features with delta and double delta features. Joint spectrotemporal features can be extracted using 2D (two-dimensional) selective filters tuned to different rates and scales of the input spectrogram, such as multi-resolution rate / scale features.

[0036] 6 illustrates a method for determining vital capacity in one embodiment. For example, one or more hardware processors running on a user device, such as a mobile device, may perform the method. In step 602, the user is prompted to produce a vowel or series of vowels. For example, the user may be asked to produce the sound in a single exhalation. In one embodiment, the user may be prompted to recite a particular vowel or series of vowels 640. In step 604, video and audio of the user producing the sound are captured, for example, via a camera and microphone connected or coupled to the user's device.

[0037] In step 606, analysis or image processing of the captured video may be performed. Such image processing may include verifying the suitability of the speech, for example, verifying the identity of the user by analyzing the user's lip movement and / or posture while producing the sound (e.g., that the user whose lung capacity is being determined is the same user as the user who is speaking), and verifying that the user is producing the sound in an appropriate manner. In step 608, a spectral analysis of the captured audio may be performed to determine suitability related to producing the sound, for example, whether the user is speaking quickly, producing shortened or extended sounds, mumbling or whispering, or a combination thereof.

[0038] In step 610, it is determined whether both the video and audio are compatible. If the answer is NO in step 610, the procedure returns to step 602, where the user is prompted to make a sound again. If both the video and audio are compatible, in step 612, the start and end times of the utterance are recorded and transmitted along with the captured audio or speech to a capacity estimator 638. For example, the capacity estimator 638 may calculate features or characteristics of the sound to determine the user's lung capacity, e.g., as shown in FIG. 4. The process shown in FIG. 6 can be used to collect appropriate speech data for determining a user baseline (e.g., as shown in FIG. 3) and / or for training a machine learning model (e.g., as shown in FIG. 2).

[0039] The above process may also be performed while the user produces other sounds, such as consonants, or reads a given text. For example, in step 614, the user is prompted to produce a consonant or series of consonants. For example, the user may be asked to produce the sound with a single exhalation. In one embodiment, the user may be prompted to read a particular consonant or series of consonants 640. In step 616, video and audio of the user producing the sound are captured, for example, via a camera and microphone connected or coupled to the user's device.

[0040] In step 618, analysis or image processing of the captured video may be performed. Such image processing may include verifying the suitability of the speech, for example, by analyzing the user's lip movement and / or posture while producing the sound, verifying the user's identity (e.g., that the user whose lung capacity is being determined is the same user who is speaking), and verifying that the user is producing the sound in an appropriate manner. In step 620, a spectral analysis of the captured audio may be performed to determine suitability related to producing the sound, for example, whether the user is speaking quickly, producing shortened or extended sounds, mumbling or whispering, or a combination thereof.

[0041] In step 622, it is determined whether both the video and audio are compatible. If the answer is NO in step 622, the procedure returns to step 614, where the user is prompted to make a sound again. If both the video and audio are compatible, in step 624, the start and end times of the utterance are recorded and sent along with the captured audio or speech to a capacity estimator 638. For example, the capacity estimator 638 calculates features or characteristics of the sound to determine the user's lung capacity, for example, as shown in FIG. 4.

[0042] Similarly, in step 626, the user is prompted to produce one or more sounds by reading a given text aloud. For example, the user may be asked to produce the sounds with a single exhalation. In one embodiment, the user may be prompted to read a specific text 640. In step 628, video and audio of the user producing the sounds are captured, for example, via a camera and microphone connected or coupled to the user's device.

[0043] In step 630, analysis or image processing of the captured video may be performed. Such image processing may include verifying the suitability of the speech, for example, by analyzing the user's lip movement and / or posture while producing the sound, verifying the user's identity (e.g., that the user whose lung capacity is being determined is the same user who is speaking), and verifying that the user is speaking in an appropriate manner. In step 632, a spectral analysis of the captured audio may be performed to determine suitability related to producing the sound, for example, whether the user is speaking quickly, making shortened or extended sounds, mumbling or whispering, or a combination thereof.

[0044] In step 634, it is determined whether both the video and audio are compatible. If the answer is NO in step 634, the procedure returns to step 626, where the user is prompted to make a sound again. If both the video and audio are compatible, in step 636, the start and end times of the utterance are recorded and transmitted along with the captured audio or speech to a capacity estimator 638. For example, the capacity estimator 638 calculates features or characteristics of the sound to determine the user's lung capacity, for example, as shown in FIG. 4. The capacity estimator 638 may use all, one, or more, or different combinations of vowels, consonants, and text-to-speech audio 640 to determine lung capacity.

[0045] 7 illustrates a computer-implemented method in one embodiment. The method may be executed by one or more hardware processors. In step 702, a user may be presented with a specification of the utterance to be performed. The specification may include data to be uttered, such as, for example, a series of vowels, a series of consonants, a text passage to be read, a script to be read, an audio passage to be repeated, a known song to be sung, and a singing video.

[0046] At step 704, an audio waveform of the user performing the utterance is captured. For example, such audio waveform may be captured using a microphone or similar device coupled or connected to a user device, such as a mobile device. At step 706, video of the user performing the utterance is captured. For example, such video may be captured using a camera or similar device coupled or connected to a user device, such as a mobile device.

[0047] The execution of the utterance may be monitored for compliance with specifications. For example, in step 708, the captured audio waveform and video are analyzed for compliance with the specifications. For example, video analysis may determine whether the user performing the utterance is the user whose spirometry is being determined, whether the user is using appropriate or accurate lip movements to produce the appropriate acoustic characteristics for use in spirometry, whether the user is in the appropriate posture or position while speaking, and / or other factors. For example, video data may be used to confirm the continued identity of the user and / or to confirm that the user correctly enunciated the requested sequence. Audio waveforms may be analyzed to determine whether the user is speaking rapidly, whether syllables are properly enunciated (e.g., not shortened, not overly long, etc.), and / or other factors. For example, the acquired information may be compared to predefinable thresholds or standards to determine compliance.

[0048] At step 710, an index of respiratory function can be determined based on the speech waveform. For example, the index can be spirometry. In one embodiment, a neural network trained to predict lung capacity can be implemented using characteristics of the speech waveform as input features, e.g., as described above with reference to Figures 2, 3, and 4. The input features can include spectral and / or temporal features associated with speech. For example, the input feature can be the length of time it takes a user to perform a portion of an utterance in one breath.

[0049] In step 712, the indicator may be compared to a reference indicator to determine the user's health status, for example, as described above with reference to FIGS. 3 and 4. In other aspects, the indicator may be the length of time it takes the user to perform a portion of an utterance in one breath. Such an indicator may be compared to the user's baseline measurement of the time it takes the user to perform the same utterance when the user is deemed to be in good health. The reference indicator includes data related to the user obtained during a period when the user is deemed to be in good health. For example, the reference indicator may be constructed by combining multiple sessions performed by the user. In other embodiments, the reference indicator may include data related to a group of individuals within the user's demographic group or a group with similar physical characteristics to the user.

[0050] In one embodiment, one or more processes of the method, e.g., analyzing, determining, and comparing, may be performed on a remote device on a computer network, e.g., on a cloud system, using audio and / or video information collected from a user, e.g., via the user's mobile device running an app or application programming interface (API).

[0051] 8 illustrates components of a system for determining a user's lung capacity in one embodiment, for example, allowing a user to perform a self-assessment of the user's lung capacity. One or more hardware processors 802, such as a central processing unit (CPU), a graphics processing unit (GPU), and / or a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or other processors, may be coupled to a memory device 804 and may present the user with a specification of an utterance to be performed, capture an audio waveform of the user performing the utterance, and capture video of the user performing the utterance. The one or more hardware processors 802 may analyze the captured audio waveforms and video for compliance with the specification and may determine an index of respiratory function based on the audio waveform. The one or more hardware processors 802 may compare the index to a reference index to determine the user's health status. The memory device 804 may comprise a random access memory (RAM), a read-only memory (ROM), or other memory device and may store data and / or processor instructions for implementing various functions associated with the methods and / or systems described herein. The one or more processors 802 may execute computer instructions stored in memory 804 or received from other computer devices or media. The memory device 804 may, for example, store instructions and / or data for the functions of the one or more hardware processors 802 and may comprise a processing system and other programs for instructions and / or data. The one or more hardware processors 802 may receive input including user speech and / or video of a user performing speech. The one or more hardware processors 802 may generate a predictive model that predicts lung capacity. The predictive model can be used to determine a baseline or reference point for a user or group of users, for example, for comparison.The input data and / or predictive models may be stored in storage device 806 or received from a remote device via network interface 808 and temporarily loaded into memory device 804 for use. The one or more hardware processors 802 may be coupled to interface devices, such as a network interface 808 for communicating with remote systems over a network, and an input / output interface 810 for communicating with input / output devices such as a keyboard, mouse, screen, and / or the like.

[0052] 9 illustrates a schematic diagram of an exemplary computer or processing system on which a system may be implemented in one embodiment. The computer system is merely one example of a suitable processing system and is not intended to suggest any limitation as to the scope of use or functionality of embodiments of the methodologies described herein. The illustrated processing system is operational with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with the processing system illustrated in FIG. 9 may include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices.

[0053] A computer system may be described in the general context of instructions executable by a computer system, such as program modules, executed by the computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. The computer system may also be implemented in distributed cloud computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.

[0054] Components of a computer system may include, but are not limited to, one or more processors or processing units 12, a system memory 16, and a bus 14 coupling various system components that comprise the system memory 16 to the processor 12. The processor 12 may include modules 30 that perform methods described herein. The modules 30 may be programmed into integrated circuits in the processor 12 or read from the system memory 16, storage device 18, or network 24, or a combination thereof.

[0055] Bus 14 may be any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example, and without limitation, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0056] The computer system may include a variety of computer system readable media. Such media may be any available media that can be accessed by the computer system and may include both volatile and nonvolatile media, removable and non-removable media.

[0057] System memory 16 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The computer system may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example, storage device 18 may be used to read from and write to non-removable, non-volatile magnetic media (e.g., a "hard drive"). Although not shown, a magnetic disk drive may be used to read from and write to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), or an optical disk drive may be used to read from or write to a removable, non-volatile disk, such as a CD-ROM, DVD-ROM, or other optical media. In such cases, each may be connected to bus 14 by one or more data media interfaces.

[0058] The computer system may also communicate with one or more external devices 26, such as a keyboard, pointing device, screen 28, etc. The one or more devices may allow a user to communicate with the computer system and / or may be any device (e.g., a network card, a modem, etc.) that allows the computer system to communicate with one or more other computing devices. Such communication is performed through an input / output (I / O) interface 20.

[0059] It should be noted that the computer system may communicate with one or more networks 24, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet), via a network adapter 22. As shown, the network adapter 22 communicates with other components of the computer system via a bus 14. Although not shown, it should be understood that other hardware and / or software components may be used in conjunction with the computer system. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, data archive storage systems, etc.

[0060] Although this disclosure includes detailed descriptions related to cloud computing, it should be understood in advance that implementation of the teachings described herein is not limited to a cloud computing environment. Rather, embodiments of the present invention can be practiced with any other type of computing environment, now known or later developed. Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model may include at least five characteristics, at least three service models, and at least four implementation models.

[0061] The characteristics are as follows:

[0062] On-Demand Self-Service: Cloud consumers can unilaterally provision computing capacity, such as server time or network storage, automatically as needed, without the need for human interaction with the service provider.

[0063] Broad network access: Computing power is available over the network and can be accessed through standard mechanisms, facilitating use by heterogeneous thin or thick client platforms (e.g., cell phones, laptops, PDAs).

[0064] Resource Pooling: Computing resources from a provider are pooled and offered to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically allocated and reallocated based on demand. Consumers generally have no control or knowledge of the exact location of the resources they are provided with, resulting in a sense of location independence. However, consumers may be able to determine location at a higher level of abstraction (e.g., country, state, data center).

[0065] Rapid Elasticity: Computing capacity can be provisioned quickly and elastically, sometimes automatically, to instantly scale out and quickly release to instantly scale in. To the consumer, the computing power available for provisioning often appears unlimited, and can be purchased at any time and in any quantity.

[0066] Metered Services: Cloud systems leverage measurement capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, active user accounts) to automatically control and optimize resource usage. Resource usage can be monitored, controlled, and reported to provide transparency to both providers and consumers of utilized services.

[0067] The service model is as follows:

[0068] Software as a Service (SaaS): The functionality offered to the consumer is the availability of a provider's applications running on a cloud infrastructure that can be accessed from a variety of client devices through a thin client interface such as a web browser (e.g., webmail). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functionality, except for limited user-specific application configuration settings.

[0069] Platform as a Service (PaaS): The capability offered to consumers is to deploy applications they create or acquire using programming languages ​​and tools supported by the provider onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the configuration of their hosting environment.

[0070] Infrastructure as a Service (IaaS): The functionality offered to consumers is the provisioning of processors, storage, networking, and other basic computing resources on which they can deploy and run any software, including operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but has control over the operating system, storage, and deployed applications, and in some cases partial control over some network components (e.g., host firewalls).

[0071] The deployment model is as follows:

[0072] Private Cloud: This cloud infrastructure is dedicated to a specific organization and can be managed by that organization or a third party, and can exist on-premise or off-premise.

[0073] Community Cloud: This cloud infrastructure is shared by multiple organizations to support a specific community with common concerns (e.g., mission, security requirements, policies, and compliance). This cloud infrastructure can be managed by those organizations or a third party and can exist on-premises or off-premises.

[0074] Public cloud: This cloud infrastructure is available to the general public or large industry organizations and is owned by an organization that sells cloud services.

[0075] Hybrid cloud: This cloud infrastructure combines two or more cloud models (private, community, or public), each of which retains its inherent nuances but is bound by standards or specific technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).

[0076] A cloud computing environment is a service-oriented environment that emphasizes statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.

[0077] FIG. 10 illustrates an exemplary cloud computing environment 50. As illustrated, the cloud computing environment 50 includes one or more cloud computing nodes 10, with which local computing devices used by cloud consumers (e.g., PDAs or cell phones 54A, desktop computers 54B, laptop computers 54C, or automobile computer systems 54N, or combinations thereof) can communicate. The nodes 10 can communicate with each other. The nodes 10 can be physically or virtually grouped (not shown) in one or more networks, such as the private, community, public, or hybrid clouds described above, or combinations thereof. This enables the cloud computing environment 50 to provide infrastructure, platform, or software as a service, or combinations thereof, for which cloud consumers are not required to maintain resources on their local computing devices. It should be understood that the types of computing devices 54A-N illustrated in FIG. 10 are merely exemplary, and that the computing nodes 10 and the cloud computing environment 50 can communicate with any type of electronic device via any type of network or network-addressable connection (e.g., using a web browser), or both.

[0078] A set of functional abstraction layers provided by the cloud computing environment 50 (FIG. 10) is shown in FIG. 11. It should be understood in advance that the components, layers, and functions shown in FIG. 11 are merely exemplary, and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:

[0079] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, reduced instruction set computer (RISC) architecture-based server 62, server 63, blade server 64, storage device 65, and network and network components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0080] The virtualization layer 70 provides an abstraction layer from which the following virtual entities can be provided, for example: virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75.

[0081] By way of example, the management layer 80 may provide the following functionality: Resource provisioning 81 enables dynamic procurement of computing and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 82 enables cost tracking as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. By way of example, these resources may include application software licenses. Security enables identification and verification of cloud consumers and tasks, as well as protection for data and other resources. User portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 84 enables allocation and management of cloud computing resources so that requested service levels are met. Service level agreement (SLA) planning and fulfillment 85 enables advance arrangement and procurement of anticipated future cloud computing resources required in accordance with SLAs.

[0082] The workload layer 90 provides examples of functionality available in a cloud computing environment. Examples of workloads and functionality that can be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom instruction delivery 93, data analytics processing 94, transaction processing 95, and speech and lung processing 96.

[0083] The present invention may be a system, method, or computer program product, or combination thereof, integrated at any possible level of technical detail. The computer program product may include a computer-readable storage medium having stored thereon computer-readable program instructions for causing a processor to carry out aspects of the present invention.

[0084] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. The computer-readable storage medium may be, by way of example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. More specific examples of computer-readable storage media include portable computer diskettes, hard disks, RAM, ROM, EPROM (or flash memory), SRAM, CD-ROM, DVD, memory sticks, floppy disks, mechanically encoded devices having instructions recorded thereon, such as punch cards or ridge-in-groove structures, and suitable combinations thereof. Computer-readable storage devices, as used herein, should not be construed as ephemeral signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted over wires.

[0085] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computer / processing device. Alternatively, they can be downloaded to an external computer or external storage device via a network (e.g., the Internet, a LAN, a WAN, or a wireless network, or a combination thereof). The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computer / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions to a computer-readable storage medium in the respective computer / processing device for storage.

[0086] The computer-readable program instructions for carrying out the operations of the present invention can be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk and C++, and procedural programming languages ​​such as the "C" programming language and similar programming languages. The computer-readable program instructions can execute entirely on the user's computer as a stand-alone software package, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a LAN or WAN, or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry, including, for example, programmable logic circuits, field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), can execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to customize the electronic circuitry for carrying out aspects of the present invention.

[0087] Aspects of the present invention are described herein with reference to flowchart and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. Each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, can be implemented by computer-readable program instructions.

[0088] The computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, whereby the instructions, executed by the processor of such computer or other programmable data processing apparatus, create means for performing the functions / acts identified in one or more blocks of the flowcharts and / or block diagrams. The computer-readable program instructions may also be stored on a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner. The computer-readable storage medium having instructions stored thereon thereby constitutes an article of manufacture including instructions for performing aspects of the functions / acts identified in one or more blocks of the flowcharts and / or block diagrams.

[0089] Computer-readable program instructions may also be loaded into a computer, other programmable device, or other device and a series of operational steps executed on the computer, other programmable device, or other device to create a computer-implemented process, whereby the instructions executing on the computer, other programmable device, or other device perform the functions / operations identified in one or more blocks in the flowcharts and / or block diagrams.

[0090] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for performing specific logical functions. In some implementations, the functions depicted in the blocks may be performed in an order different from that depicted in the figures. For example, two blocks shown in succession may actually be accomplished as a single step, executed simultaneously or substantially simultaneously, executed in a partially or fully overlapping manner, or executed in reverse order, depending on the functionality involved. Note that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0091] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present invention. As used herein, the singular forms "a," "an," and "the" are intended to include the plural unless the context clearly dictates otherwise. As used herein, the term "or" is an inclusive operator and can mean "and / or" unless the context explicitly or clearly dictates otherwise. It will be further understood that as used herein, the terms "comprise," "comprises," "comprising," "include," "includes," "including," and / or "having" can specify the presence of stated features, integers, steps, operations, elements, or components, or combinations thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups thereof, or combinations thereof. As used herein, the phrase "in an embodiment" does not necessarily refer to the same embodiment, although it may. As used herein, the phrase "in one embodiment" does not necessarily refer to the same embodiment, although it may. As used herein, the phrase "in another embodiment" does not necessarily refer to another embodiment, although it may. Furthermore, embodiments and / or elements of embodiments may be freely combined with each other unless they are mutually exclusive.

[0092] All corresponding structure, material, acts, and equivalent or step-plus-function elements in the following claims are intended to include any structure, material, or act for performing a function in combination with other claimed elements, as specifically claimed. The description of the present invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or to limit the invention to the form disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the invention. The embodiments were chosen and described to best explain the principles and practical applications of the invention and to enable those skilled in the art to understand the invention in various embodiments with various modifications suited to the particular uses contemplated.

Claims

1. A computer comprising: presenting to the user a specification of the utterance to be performed; capturing a speech waveform of the user performing the utterance; capturing video of the user performing the utterance; analyzing the captured audio waveform and the video for compliance with the specifications; determining the continued identity of the user using the captured video; and determining an index of respiratory function based on the speech waveform; comparing the indicator to a reference indicator to determine a health status of the user; 11. A computer-implemented method comprising:

2. the metrics include the length of time it takes the user to perform the portion of the utterance in one breath; The method of claim 1.

3. At least the presenting and the capturing are performed using a mobile device. The method of claim 1.

4. the specification of the speech includes at least one of a script to be read, a spoken passage to be repeated, a known song to be sung, and a singing video; The method of claim 1.

5. the reference indicator comprises data relating to the user acquired during a period in which the user is considered healthy; The method of claim 1.

6. the reference index is constructed by combining multiple sessions performed by the user; The method of claim 1.

7. the reference indicator includes data relating to a group of individuals within the user's demographic group; The method of claim 1.

8. the analyzing, determining, and comparing are performed on a remote device via a computer network; The method of claim 1.

9. determining an index of respiratory function includes running a machine learning model trained to predict the respiratory function using input features including audio spectral and temporal features extracted from the captured audio. The method of claim 1.

10. determining that the user correctly uttered the utterance in the specification using the captured video. The method of claim 1.

11. a processor; a memory coupled to the processor; the processor presenting a specification of an utterance to be performed to a user; capturing a speech waveform of the user performing the utterance; capturing video of the user performing the utterance; analyzing the captured audio waveform and the video for compliance with the specifications; determining the continued identity of the user using the captured video; and determining an index of respiratory function based on the speech waveform; comparing the indicator to a reference indicator to determine a health status of the user; A system consisting of.

12. the metrics include the length of time it takes the user to perform the portion of the utterance in one breath; The system of claim 11.

13. the specification of the speech includes at least one of a script to be read, a spoken passage to be repeated, a known song to be sung, and a singing video; The system of claim 11.

14. the reference indicator comprises data relating to the user acquired during a period in which the user is considered healthy; The system of claim 11.

15. The reference index is constructed by combining multiple sessions performed by the user. The system of claim 11.

16. the reference indicator includes data relating to a group of individuals within the user's demographic group; The system of claim 11.

17. the processor is configured to execute a neural network trained to predict the indicator of respiratory function using input features including audio spectral and temporal features extracted from the captured audio. The system of claim 11.

18. A computer program having program instructions embodied therein, the program instructions being readable by a device, causing the device to: presenting to the user a specification of the utterance to be performed; capturing a speech waveform of the user performing the utterance; capturing video of the user performing the utterance; analyzing the captured audio waveform and the video for compliance with the specifications; determining the continued identity of the user using the captured video; and determining an index of respiratory function based on the speech waveform; comparing the indicator to a reference indicator to determine a health status of the user; A computer program that performs the following:

19. the device runs a machine learning model trained to predict the indicator of respiratory function using input features including audio spectral and temporal features extracted from the captured audio.

19. A computer program according to claim 18.

Citation Information

Patent Citations

  • System and method for measuring lung capacity and physical fitness

    JP2015536691A

  • Voice recognition device and voice recognition method

    JP2018091954A