Automated physiological and pathological assessment based on speech analysis

The method analyzes word-reading tests to derive biomarkers from voice recordings, addressing variability issues in existing voice analysis methods, enabling reliable remote monitoring and diagnosis of conditions affecting respiration, voice tone, and cognitive ability.

JP7834763B2Active Publication Date: 2026-03-24F HOFFMANN LA ROCHE & CO AG +1
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-12
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing methods for remote monitoring of physiological and pathological conditions using voice analysis, such as heart failure and asthma, suffer from inconsistency due to variability in spontaneous speech and neuropsychological effects when using predetermined standard passages, limiting their practical use.

Method used

A method for analyzing voice recordings from word-reading tests, like the Stroop test, to derive reproducible biomarkers by identifying segments and calculating metrics such as breath percentage, voiceless/voiced ratio, and voice pitch, which is language-independent and fully automated, enabling remote self-assessment and monitoring of conditions affecting respiration, voice tone, and cognitive ability.

Benefits of technology

Enables accurate and sensitive segmentation of speech segments, allowing for robust quantification of metrics like correct word rate, independent of language and recording length, thus facilitating reliable remote monitoring and diagnosis of conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007834763000006
    Figure 0007834763000006
  • Figure 0007834763000007
    Figure 0007834763000007
  • Figure 0007834763000008
    Figure 0007834763000008
Patent Text Reader

Abstract

Methods for assessing a pathological and / or physiological condition of a subject, monitoring a subject with heart failure or diagnosed with or at risk for a dyspnea and / or fatigue related condition, and diagnosing a subject with decompensated heart failure are provided. The methods include obtaining from the subject a speech recording from a word reading test that includes reading a sequence of words taken from a set of n words, and analyzing the speech recording, or a portion thereof. The analysis can include identifying a plurality of segments of the speech recording that correspond to individual words or syllables, determining values ​​of one or more metrics selected from respiration %, voiceless / voiced ratio, voice pitch, and word accuracy based at least in part on the identified segments, and comparing the values ​​of the one or more metrics to one or more respective reference values. Related systems and products are also described.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Field of the Invention The present invention relates to a computer-implemented method for the automatic assessment of a subject's physiological and / or pathological state, including in particular the analysis of audio recordings from a word reading test. A computing device for implementing such a method is also described. The method and apparatus of the present invention are applicable to the clinical assessment of pathological and physiological states that affect respiration, vocal tone, fatigue, and / or cognitive ability.

Background Art

[0002] Background of the Invention Remote monitoring of patients in various states has the potential to improve the outcome, quality, and comfort of the health management of many patients. Therefore, there is great interest in the development of devices and methods that can be used to collect biomarker data of patients that can later be evaluated by the patient's medical team. The potential benefits of remote monitoring are particularly important in the context of chronic diseases or lifelong symptomatic conditions such as heart disease or asthma. Non-invasive biomarker-based approaches are particularly desirable because they are less risky. The use of voice analysis to collect such biomarker information has been proposed, for example, in the assessment of heart failure (Maor et al., 2018), asthma, chronic obstructive pulmonary disease (COPD) (Saeed et al., 2017), and more recently COVID-19 (Laguarta et al., 2020).

[0003] However, all of these methods suffer from limitations in consistency. In fact, many of these methods rely on spontaneous speech or sounds (such as coughing) or the reading of a predetermined standard passage, such as the Rainbow Passage (Murton et al., 2017). The use of spontaneous speech or sounds is prone to significant variability, both between patients and between repeated assessments of the same patient, because the content of each audio recording can vary widely. The use of a predetermined standard passage mitigates this inherent variability due to its content, but it is still subject to interference from neuropsychological effects related to subjects becoming accustomed to the standard text as the test is repeated. This places a strong limitation on the practical use of speech analysis biomarkers in remote monitoring settings.

[0004] Therefore, there is still a need for improved methods to automatically assess pathological and physiological conditions that can be easily performed remotely while minimizing the burden on patients. [Overview of the project]

[0005] Description of the invention The inventors of the present invention have developed a novel apparatus and method for the automated assessment of a subject's physiological and / or pathological state, particularly involving the analysis of voice recordings from word-reading tests. The inventors have confirmed that using recordings from word-reading tests, such as the Stroop test, it is possible to assess a subject's pathological and / or physiological state and derive reproducible and informative biomarkers for assessing conditions affecting respiration, voice tone, fatigue, and / or cognitive ability.

[0006] The Stroop test (Stroop, 1935) is a three-part neuropsychological test (word, color, and interference) that has been used to diagnose mental and neurological disorders. For example, it forms part of a cognitive test battery used to quantify the severity of Huntington's disease (HD) according to the widely used Unified Huntington's Disease Rating Scale (UHDRS). The word and color portions of the Stroop test represent a "non-contradictory condition" where color words are printed in black ink and color patches are printed in matching ink colors. In the interference portion, color words are printed in mismatched ink colors. Patients are asked to read the words aloud or state the ink colors as quickly as possible. Clinicians interpret the responses as correct or incorrect. Scores are reported as the number of correct answers in each condition within a given 45 seconds. The non-contradictory condition is thought to measure processing speed and selective attention. The interference condition is intended to measure cognitive flexibility because it requires mental conversion between words and colors.

[0007] The method described herein is based on automatically determining one or more metrics identified as usable as biomarkers from recordings of word-reading tests inspired by the Stroop test, the metrics being selected from voice pitch, correct word rate, respiratory percentage, and voiceless / voiced ratio. The method is language-independent, fully automated, reproducible, and applicable to a variety of conditions affecting respiration, voice tone, fatigue, and / or cognitive ability. Thus, remote self-assessment and monitoring of symptoms, diagnosis, or prognosis of such conditions becomes possible in large populations.

[0008] Accordingly, according to the first aspect, a method is provided for evaluating the pathological and / or physiological state of a subject, comprising: obtaining an audio recording from the subject from a word reading test, which includes reading a sequence of words taken from a set of n words; and analyzing the audio recording or a portion of the audio recording by identifying multiple segments of the audio recording corresponding to individual words or syllables, determining values ​​for one or more metrics selected from breath percentage, voiceless / voiced ratio, voice pitch, and correct word rate, based at least partially on the identified segments, and comparing the values ​​of the one or more metrics with one or more respective reference values.

[0009] This method may have any one or more of the following characteristics.

[0010] Identifying segments of a speech recording that correspond to individual words or syllables involves obtaining a power mel spectrogram of the speech recording, calculating the maximum intensity projection of the mel spectrogram along the frequency axis, and defining the segment boundary as the point at which the maximum intensity projection of the mel spectrogram along the frequency axis intersects a threshold.

[0011] The word / syllable segmentation techniques described herein enable accurate and highly sensitive segmentation of words (and in some cases, syllables from multi-syllable words) from speech recordings, even when the pace of speech is relatively fast (i.e., with little to no pauses between words), where existing methods based on energy envelopes may not function well. Furthermore, it enables the automatic quantification of metrics derived from identified speech segments in word-reading tasks (e.g., rates such as breath percentage, voiceless / voiced ratio, and correct word rate) from data that can be easily and readily obtained remotely, for example, by the patient reading aloud and recording themselves words displayed on a computing device (e.g., a mobile computing device such as a smartphone or tablet, or a personal computer, via an application or web application, as further described herein).

[0012] A segment of a speech recording corresponding to an individual word or syllable can be defined as a segment that lies between two consecutive word / syllable boundaries. Preferably, a segment of a speech recording corresponding to an individual word / syllable can be defined as a segment between a first boundary where the maximum intensity projection of the Mel spectrogram crosses a threshold from a lower value to a higher value, and a second boundary where the maximum intensity projection of the Mel spectrogram crosses a threshold from a higher value to a lower value. Conveniently, segments of speech recordings between boundaries that do not satisfy this definition may be excluded.

[0013] Determining the value of one or more metrics may include determining the breath percentage for a recording as the percentage of time between identified segments in the audio recording, or as the ratio of the time between identified segments in the recording to the sum of the time between identified segments and the time within each identified segment.

[0014] Determining the value of one or more metrics may include determining the silent / voiced ratio for a recording as the ratio of time between identified segments in the recording to time within identified segments in the recording.

[0015] Determining the value of one or more metrics may include determining the correct word rate for an audio recording by calculating the ratio of the number of identified segments corresponding to correctly pronounced words to the duration between the start of the first identified segment and the end of the last identified segment.

[0016] Determining the value of one or more metrics may include determining the voice pitch of a recording by obtaining one or more estimates of the fundamental frequency for each of the identified segments. Determining the voice pitch value may include obtaining multiple estimates of the fundamental frequency for each of the identified segments and applying filters to these multiple estimates to obtain filtered multiple estimates. Determining the voice pitch value may include obtaining a summarized voice pitch estimate for multiple segments, such as the mean, median, or mode of the (optionally, filtered) multiple estimates for multiple segments.

[0017] Determining the value of one or more metrics may involve determining the total word rate or ground truth word rate for an audio recording by calculating the cumulative sum of the number of identified segments corresponding to spoken or correctly spoken words in an audio recording over time, and by calculating the slope of a linear regression model fitted to the cumulative sum data. Conveniently, this method yields a robust estimate of the total word rate or ground truth word rate as the number of spoken or correctly spoken words per unit time across the entire recording. The estimate thus obtained is robust to outliers (e.g., distractions that can cause isolated, momentary changes in the ground truth word rate) while being highly sensitive to true declines in the total word rate or ground truth word rate (e.g., genuine fatigue, worsening breathing, and / or cognitive decline leading to frequent segments in slow speech). Furthermore, this method is independent of the length of the recording. Therefore, it is possible to compare the total word rate or ground truth word rate obtained for audio recordings of different lengths or for different parts of the same audio recording. Furthermore, this method can be robust to external factors such as subjects taking short breaks or failing to speak for reasons unrelated to cognitive decline or respiratory deterioration (e.g., subjects initially not noticing that recording has begun). In addition, this method is conveniently robust to account for uncertainty regarding the specific timing of word initiation and / or variability in word duration.

[0018] If the method involves determining the correct word rate in an audio recording, it may include: calculating one or more Mel-frequency cepstrum coefficients (MFCCs) for each segment to obtain multiple vectors of values, each vector relating to one segment; clustering the multiple vectors of values ​​into n clusters, each cluster having n possible labels corresponding to each of n words; predicting a sequence of words in the audio recording using the labels relating to the clustered vectors of values ​​for each of the n! permutations of labels; performing a sequence alignment between the predicted sequence of words and the sequence of words used in a word-reading test; and selecting the best alignment, which is the best alignment, where the alignment matches the correctly read words in the audio recording.

[0019] Conveniently, the methods described herein for determining the correct word rate are entirely data-driven and therefore model and language-independent. In particular, since the clustering process is an unsupervised learning process, it does not require knowledge of the actual words (ground truth) that each group of segments represents. In alternative embodiments, clustering can be replaced with a supervised learning method such as a Hidden Markov Model. However, such methods would likely require retraining the model for each language.

[0020] Conveniently, the method described herein for determining the correct word rate can also address speech disorders, such as dysarthria, which can interfere with the identification of words that are read correctly but mispronounced in conventional word recognition methods. Furthermore, it enables the automatic quantification of the correct word rate in word reading tasks from data that can be easily and readily obtained remotely, for example, by having a patient read aloud words displayed on a computing device (e.g., a mobile computing device such as a smartphone or tablet) and record it themselves.

[0021] In some embodiments, predicting a sequence of words in an audio recording using labels associated with a clustered vector of values ​​involves predicting a sequence of words corresponding to each respective cluster label of the clustered vector of values, arranged according to the order of the segments from which the vector of values ​​was derived.

[0022] In some embodiments, predicting a sequence of words in an audio recording using labels associated with clustered vectors of values ​​involves predicting a sequence of words corresponding to each respective cluster label of each clustered vector of values ​​that has been assigned to a cluster with confidence that one or more predetermined criteria are met. In other words, predicting a sequence of words in an audio recording using labels associated with clustered vectors of values ​​may involve excluding predictions for clustered vectors of values ​​that are not associated with any particular cluster with confidence that one or more predetermined criteria are met. One or more predetermined criteria may be defined using thresholds for the probability that a vector of values ​​belongs to one of n clusters, the distance between a vector of values ​​and a representative vector of values ​​in one of the n clusters (e.g., the coordinates of the cluster's medoid or centroid), or a combination thereof.

[0023] In some embodiments, predicting a sequence of words in a speech recording using labels associated with a vector of clustered values ​​involves predicting a sequence of words corresponding to each cluster label of each of the clustered value vectors. In some such embodiments where polysyllabic words (in particular, polysyllabic words containing one emphasized syllable) are used, multiple segments may be identified and clustered, so multiple word predictions may be predicted for a polysyllabic word. Even in such situations, it has been found that it is still possible to determine the number of correctly pronounced words in a speech recording according to the method described herein. Indeed, as described above, it is considered that the clustering step can be robust to the presence of “noise” resulting from additional syllables so that clusters determined primarily by individual syllables can still be identified in each of the n words. Furthermore, it is considered that the sequence alignment step can be treated as an insertion into the sequence that would exist for each of the n! permutations of labels, since such additional syllables result from the presence of additional predicted words that are not expected to be present in the sequence of words used in the word reading test. Thus, the number of matches in the alignment still corresponds to the number of correctly pronounced words in the speech recording.

[0024] In some embodiments, calculating one or more MFCCs to obtain a vector of values ​​for a segment includes calculating a set of i MFCCs for each frame of the segment, and obtaining a set of j values ​​for the segment by compressing the signal formed by each of the i MFCCs across the frames in the segment to obtain a vector of ixj values ​​for the segment. For example, compressing the signal formed by each of the i MFCCs across the frames in the segment may include performing linear interpolation of the signal.

[0025] In some embodiments, calculating one or more MFCCs to obtain a vector of values for a segment involves calculating a set of i MFCCs for each frame of the segment and obtaining a set of j values for the segment for each i by interpolation, preferably linear interpolation, to obtain a vector of ixj values for the segment.

[0026] As a result, the vectors of values for each of the plurality of segments all have the same length. Such vectors of values can be conveniently used as input to any clustering technique for identifying clusters of points in a multi-dimensional space.

[0027] Calculating one or more MFCCs to obtain a vector of values for a segment can be performed as described above. As will be understood by those skilled in the art, using a fixed-length time window to obtain the MFCCs of a segment means that the total number of MFCCs per segment can vary depending on the length of the segment. In other words, a segment has some number of frames f, each frame being associated with a set of i MFCCs, and f varies depending on the length of the segment. As a result, segments corresponding to longer syllables / words are associated with a greater number of values than segments corresponding to shorter syllables / words. This can be a problem when these values are used as features to represent segments for the purpose of clustering segments in a common space. The interpolation step solves this problem. In some embodiments, calculating one or more MFCCs for a segment involves calculating a plurality of MFCCs from the 2nd to the 13th for each frame of the segment. The first MFCC is preferably not included. Without wishing to be bound by theory, the first MFCC is assumed to represent the energy in a segment that is mainly related to recording conditions and contains little information regarding the identity of a word or syllable. In contrast, the remaining 12 MFCCs cover the human audible range (by the definition of MFCC) and thus capture acoustic features related to the way words are pronounced and heard by humans.

[0028] In some embodiments, a plurality of the second through thirteenth MFCCs include at least two, at least four, at least six, at least eight, at least ten, or all twelve of the second through thirteenth MFCCs. The second through thirteenth MFCCs can advantageously include information that can be used to distinguish words from a closed set of words as points in hyperspace using a simple clustering technique. In particular, as described above, the second through thirteenth MFCCs cover the human audible range and thus are thought to capture acoustic features related to the way humans vocalize and perceive words. Thus, by using those twelve MFCCs, information that is thought to be relevant in distinguishing one word / syllable from another in a human voice recording can be advantageously captured.

[0029] When the segmentation method described herein is used, the MFCCs for each frame of the identified segments may already have been calculated as part of a step to exclude segments that represent false detections. In such embodiments, the previously calculated MFCCs can be advantageously used to obtain a vector of values for the purpose of determining the number of correctly read words in the voice recording.

[0030] In some embodiments, the parameter j is selected such that j ≤ f for all segments used in the clustering process. In other words, the parameter j may be selected so that interpolation results in compression of the signal (for each MFCC, the signal is the value of the MFCC across the frame of the segment). In some embodiments, the parameter j may be selected so that interpolation results in a 40-60% signal compression for all segments used in clustering (or at least a predetermined percentage of the segments, e.g., 90%). As will be understood by those skilled in the art, using a fixed parameter j, the level of compression applied to the segments may depend on the length of the segments. By using a 40-60% signal compression, it can be ensured that the signal within each segment is compressed to about half of its original signal density.

[0031] In a convenient embodiment, j is chosen between 10 and 15, for example, 12. While we do not wish to be bound by theory, a 25ms frame with a 10ms step size is commonly used for MFCC calculations for sound signals. Furthermore, a syllable (and a monosyllabic word) can be approximately 250ms in length on average. Thus, using j=12 can result in a compression of approximately half this number (i.e., a compression of about 40-60% on average) from an average of 25 values ​​(corresponding to 25 frames across a 250ms segment).

[0032] In some embodiments, clustering multiple vectors of values ​​into n clusters is performed using k-means. Conveniently, k-means is a simple and computationally efficient method that has proven to work well in separating words represented by vectors of MFCC values. Alternatively, other clustering methods may be used, such as partitioning around a medoid or hierarchical clustering.

[0033] Furthermore, the centroids of the acquired clusters may correspond to the corresponding word or syllable representations within the MFCC space. This can provide useful information about the process (e.g., whether segmentation and / or clustering was performed satisfactorily) and / or the speech recording (and therefore the subject). In particular, such cluster centroids can be compared between individuals and / or used as a further clinically useful measure (e.g., because they capture aspects of the subject's ability to pronounce syllables or words clearly).

[0034] In some embodiments, one or more MFCCs are normalized across segments in the record prior to clustering and / or interpolation. In particular, each MFCC can be individually centered and standardized so that each MFCC distribution has equal variance and mean 0. This can favorably improve the performance of the clustering process by preventing clustering from being "dominated" by some MFCCs that are distributed with high variance. In other words, this can ensure that all features in clustering (i.e., each MFCC used) have similar importance in clustering.

[0035] In some embodiments, performing sequence alignment involves obtaining an alignment score. In some such embodiments, the best alignment is an alignment that satisfies one or more predetermined criteria, at least one of which is applied to the alignment score. In some embodiments, the best alignment is the alignment with the highest alignment score.

[0036] In some embodiments, the sequence alignment step is performed using a local sequence alignment algorithm, preferably the Smith-Waterman algorithm.

[0037] Local sequence alignment algorithms are ideally suited for tasks that involve aligning two strings selected from a closed set, where the strings are relatively short and not necessarily of the same length (as in this case, where words may be overlooked in the reading task and / or word segmentation process). In other words, local sequence alignment algorithms such as the Smith-Waterman algorithm are particularly well suited for aligning partially overlapping sequences, which is advantageous in the context of the present invention where mismatches and gaps in alignment are expected due to the subject's inability to achieve 100% correct word counts and / or errors in the segmentation process.

[0038] In some embodiments, the Smith-Waterman algorithm is used with a gap cost between 1 and 2 (preferably 2) and a match score of 3. These parameters can result in accurate identification of words in the audio recording compared to manually annotated data. While we do not wish to be bound by theory, using a higher gap cost (e.g., 2 instead of 1) may result in a more limited search space and shorter alignment. This can conveniently capture situations where a match is expected (i.e., it is assumed that there exists a cluster label assignment such that many characters in the predicted sequence of the word can be aligned with characters in the known sequence of the word).

[0039] In some embodiments, identifying segments of the speech recording corresponding to individual words or syllables further includes normalizing the power mel spectrogram of the speech recording. Preferably, the power mel spectrogram is normalized to the frame having the highest energy in the recording. In other words, each value of the power mel spectrogram can be divided by the highest energy value in the power mel spectrogram.

[0040] As those skilled in the art will understand, a power Mel spectrogram refers to the power spectrogram of an audio signal on the Mel scale. Furthermore, obtaining a Mel spectrogram involves defining frames along an audio recording (where a frame can correspond to a signal within a fixed-width window applied along the time axis) and calculating the power spectrum on the Mel scale for each frame. This process yields a matrix of power values ​​per Mel unit for each frame (time bin). Obtaining the maximum intensity projection of such a spectrogram onto the frequency axis involves selecting the maximum intensity on the Mel spectrum for each frame.

[0041] Normalization conveniently facilitates comparisons between different audio recordings, which may relate to the same or different subjects. This can be particularly advantageous, for example, when multiple individual recordings from the same subject are combined. For instance, this may be particularly advantageous when shorter recordings are preferred (e.g., because the subject is frail), and when standard-length or other preferred-length word-reading tests are preferred. Normalizing the mel spectrogram to the frame with the highest energy in a recording conveniently results in every recording having a relative energy value of 0 dB (value after maximum intensity projection) for the loudest frame in the recording. Other frames will have relative energy values ​​below 0 dB. Furthermore, normalizing the power mel spectrogram provides a maximum intensity projection representing relative energy (dB value over time) that can be conveniently used across multiple recordings, as it allows for the comparison of relative energy (which may be predetermined or dynamically determined) between audio recordings.

[0042] Applying outlier detection methods to data derived from individual word / syllable segments conveniently allows for the removal of segments corresponding to false positives (e.g., those caused by incorrect pronunciation, breathing, and non-speech sounds). Any outlier detection method applicable to multidimensional sets of observations can be used. For example, clustering techniques can be used. In some embodiments, applying an outlier detection method to multiple vectors of values ​​involves excluding all segments where the vector of values ​​exceeds a predetermined distance from the rest of the vector of values.

[0043] Identifying segments of a speech recording corresponding to individual words or syllables may further include performing onset detection for at least one of the segments by computing a spectral flux function across the segment's mel spectrogram, and forming two new segments by defining further boundaries whenever an onset is detected within the segment.

[0044] In some embodiments, identifying segments of a speech recording corresponding to individual words / syllables further includes excluding segments representing false positives by removing segments shorter than a predetermined threshold and / or segments whose average relative energy falls below a predetermined threshold. For example, segments shorter than 100 ms may be conveniently excluded. Similarly, segments with an average relative energy of less than -40 dB may be conveniently excluded. Such techniques can easily and efficiently exclude segments corresponding to words or syllables. Preferably, segments are filtered to exclude short segments and / or low-energy segments prior to the calculation of the MFCC of the segments and the application of the outlier detection method as described above. In fact, this conveniently avoids the unnecessary step of calculating the MFCC for false segments and prevents such false segments from introducing further noise into the outlier detection method.

[0045] In some embodiments of any part of the method, the audio recording includes a reference tone. For example, the recording may be acquired using a computing device configured to emit a reference tone immediately after the start of a recording by a user performing a reading test. This may be useful to provide the user with instructions on when to begin the reading task. In embodiments in which the audio recording includes a reference tone, one or more parameters of the method may be selected such that the reference tone is identified as a segment corresponding to a single word or syllable, and / or segments containing the reference tone are excluded in a process for removing false positives. For example, a set of MFCCs used in the false positive removal process and / or a predetermined distance used in this process may be selected such that segments corresponding to the reference tone are removed in each audio recording (or at least a selected percentage of the audio recordings).

[0046] Identifying segments of a speech recording corresponding to individual words or syllables may further include calculating one or more Mel-frequency cepstrum coefficients (MFCCs) for each segment to obtain multiple vectors of values ​​in which each vector relates to one segment, and then excluding segments representing false positives by applying an outlier detection method to these multiple vectors of values. Identifying segments of a speech recording corresponding to individual words or syllables may further include excluding segments representing false positives by removing segments shorter than a predetermined threshold and / or segments whose average relative energy falls below a predetermined threshold.

[0047] The n words may be one-syllable or two-syllable. Each of the n words may contain one or more vowels within itself. Each of the n words may contain a single stressed syllable. The n words may be color words, and optionally, the words may be displayed in a single color in the word reading test, or the words may be displayed in a color independently selected from a set of m colors in the word reading test.

[0048] In the context of this invention, a subject is a human subject. The terms “subject,” “patient,” and “individual” are used interchangeably throughout this disclosure.

[0049] Obtaining audio recordings from a subject from a word reading test includes obtaining audio recordings from a first word reading test and audio recordings from a second word reading test, the word reading test includes reading a sequence of words taken from a set of n words which are color words, the words are displayed in a single color in the first word reading test and in a color independently selected from a set of m colors in the second word reading test, and optionally the sequence of words in the second word reading test is the same as the sequence of words in the first word reading test.

[0050] A word sequence can contain a predetermined number of words, the number of which is chosen to ensure that the recording contains enough information to estimate one or more metrics and / or to allow comparison of one or more metrics with previously obtained baseline values. A word sequence may contain at least 20, at least 30, or about 40 words. For example, the inventors of the present invention have found that a word reading test containing a 40-word sequence provides sufficient information to estimate all metrics of interest, while being manageable even for subjects with severe dyspnea and / or fatigue, such as patients with decompensated heart failure.

[0051] The predetermined number of words may depend on the expected physiological and / or pathological condition of the subject. For example, the predetermined number of words may be selected so that subjects with a particular disease, disorder, or condition are expected to be able to read the sequence of words within a predetermined time. The predicted number of words per predetermined period may be determined using a comparative training cohort. Preferably, the comparative training cohort consists of individuals with similar conditions, diseases, or disorders and / or similar levels of fatigue and / or dyspnea to the intended user. The predetermined time length is conveniently less than 120 seconds. If the test is too long, it may be affected by external parameters such as boredom or physical exhaustion and / or may be less convenient for the user, potentially leading to reduced engagement. The predetermined time length may be selected from 30 seconds, 35 seconds, 40 seconds, 45 seconds, 50 seconds, 55 seconds, or 60 seconds. The predetermined time length and / or number of words may be selected based on the existence of standard and / or comparative tests.

[0052] Preferably, the recording is of a length necessary for the subject to read aloud the sequence of displayed words. Thus, the computing device can record an audio recording until the subject indicates that recording should be stopped and / or until the subject reads aloud the entire sequence of displayed words. For example, the computing device can record an audio recording until the subject provides input through the user interface indicating completion of the test. As another example, the computing device can record an audio recording for a predetermined length of time and crop the recording to include a number of segments corresponding to the expected number of words in the sequence of words. Alternatively, the computing device may record an audio recording until it detects that the subject has not spoken for a predetermined period of time. In other words, the method may include having a computing device associated with a subject record an audio recording from the time the computing device receives a start signal until the computing device receives a stop signal. The start and / or stop signals may be received from the subject through the user interface. Alternatively, the start and / or stop signals may be generated automatically. For example, the start signal may be generated when the computing device begins displaying words. A stop signal may be generated when a computing device determines that no audio signal has been detected for a set minimum period, such as 2, 5, 10, or 20 seconds. While we do not wish to be bound by theory, the use of an audio recording expected to contain a known number of words (corresponding to the number of words in the set of words) may be particularly advantageous in any embodiment of the present invention. Indeed, such embodiments can conveniently simplify the alignment process because a known sequence of words is assumed to have a known length with respect to any recording.

[0053] The recording may include multiple recordings. Each recording may be from a word-reading test that involves reading a sequence of at least 20, at least 25, or at least 30 words. For example, a word-reading test that involves reading a sequence of, for example, 40 words may be split into two tests, each involving reading a sequence of 20 words. This may allow recording from a word-reading test that involves reading a sequence of a predetermined length if, due to the subject's pathological or physiological condition, the subject is unable to read the predetermined length sequence in a single test. In embodiments using multiple separate audio recordings, the step of identifying segments corresponding to individual words / syllables is conveniently performed at least partially separately for each separate audio recording. For example, steps including normalization, dynamic thresholding, scaling, etc., are conveniently performed separately for each recording. In embodiments using multiple separate audio recordings, the alignment step may be performed separately for each recording. In contrast, the clustering step may conveniently be performed on combined data from multiple recordings.

[0054] The steps of displaying a sequence of words for a word reading test and recording the word recording may be performed by a computing device located away from the computing device performing the analysis steps. For example, the display and recording steps may be performed by the user's personal computing device (which may be a PC or a mobile device such as a mobile phone or tablet), while the analysis of the voice recording may be performed by a remote computer, such as a server. This allows for the remote acquisition of clinically relevant data, for example, from a patient's home, while leveraging the high computing power of the remote computer for analysis.

[0055] In some embodiments, the computing device associated with the subject is a mobile computing device such as a mobile phone or tablet. In some embodiments, displaying a sequence of words and recording an audio recording on the computing device associated with the subject is done via a software application that runs locally on the computing device associated with the subject (sometimes referred to as a “mobile app” or “native app” in the context of a mobile device), a web application that runs in a web browser, or a hybrid application that embeds a mobile website within a native app.

[0056] In some embodiments, acquiring an audio recording includes recording the audio recording and performing the steps of analyzing the audio recording, with the acquisition and analysis performed by the same computing device (i.e., locally). This conveniently eliminates the need to connect to a remote device for analysis and the need to transfer sensitive information. The results of the analysis (e.g., correct word rate, pitch, etc.) as well as the audio recording or a compressed version thereof may still be communicated to a remote computing device for storage and / or meta-analysis in such embodiments.

[0057] This method can be used to assess the condition of a subject who has been diagnosed with, or is at risk of having, a condition affecting respiration, voice tone, fatigue, and / or cognitive ability. This method can be used to diagnose a subject as having a condition affecting respiration, voice tone, fatigue, and / or cognitive ability. In the context of the present invention, an individual can be considered to have a condition affecting respiration, voice tone, fatigue, and / or cognitive ability when the performance of a task, such as a word-reading test, by the individual is influenced by psychological, physiological, neurological, or respiratory factors. Examples of conditions, diseases, or disorders that may affect a subject's respiration, voice tone, fatigue, or cognitive ability include: (i) Cardiovascular diseases such as heart failure, coronary heart disease, myocardial infarction (heart attack), atrial fibrillation, arrhythmia (heart rhythm disorder), and heart valve disease; (ii) respiratory diseases, disorders, or conditions such as obstructive pulmonary disease (e.g., asthma, chronic bronchitis, bronchiectasis, and chronic obstructive pulmonary disease (COPD)), chronic respiratory disease (CRD), respiratory tract infections, and lung tumors, respiratory infections (e.g., COVID-19, pneumonia, etc.), obesity, dyspnea (e.g., dyspnea associated with heart failure), panic attacks (anxiety disorders), pulmonary embolism, physical limitations or damage to the lungs (e.g., rib fractures, lung collapse, pulmonary fibrosis, etc.), pulmonary hypertension, or any other disease, disorder, or condition affecting lung / cardiopulmonary function (e.g., measurable by spirometry); (iii) Neurovascular diseases or disorders such as stroke, neurodegenerative diseases, myopathy, and diabetic neuropathy; (iv) Psychiatric disorders or conditions such as depression, drowsiness, attention deficit disorder, and chronic fatigue syndrome; (v) Conditions that affect an individual's fatigue state or cognitive ability through systemic mechanisms, such as pain, abnormal glucose levels (e.g., resulting from diabetes mellitus), or renal dysfunction (e.g., in the context of chronic renal failure or renal replacement therapy).

[0058] Therefore, the methods described herein can be used for the diagnosis, monitoring, or treatment of any of the conditions, diseases, or disorders described above.

[0059] In the context of the present invention, a word reading test (also referred to herein as a “word reading task”) refers to a test that requires an individual to read aloud a set of words (also referred to herein as a “sequence of words”) that are not connected to form a sentence, and the words are taken from a predetermined set (for example, the words may be taken randomly or pseudo-randomly from the set). For example, all words in the set of words may be nouns, such as a set of words about colors in a selected language.

[0060] As those skilled in the art will understand, the method for analyzing voice recordings from subjects is a computer-implemented method. Indeed, the analysis of voice recordings described herein, including, for example, syllable detection, classification, and alignment, requires the analysis of large amounts of data through complex mathematical operations that go beyond the realm of mental activity.

[0061] A second embodiment provides a method for monitoring a subject with heart failure or diagnosing a subject with worsening heart failure or decompensated heart failure, the method comprising: obtaining an audio recording from the subject from a word reading test, which includes reading a sequence of words taken from a set of n words; and analyzing the audio recording or a portion of the audio recording by identifying multiple segments of the audio recording corresponding to individual words or syllables, determining values ​​for one or more metrics selected from respiratory %, voiceless / voiced ratio, voice pitch, and correct word rate, based at least partially on the identified segments, and comparing the values ​​of the one or more metrics to one or more respective reference values. The method further includes any of the features of the first embodiment.

[0062] According to a third aspect, a method is provided for treating a subject with worsening heart failure or decompensated heart failure, comprising using the method of the preceding aspects to diagnose the subject with worsening heart failure or decompensated heart failure, and treating the subject with respect to heart failure. The method may further include monitoring disease progression, monitoring the subject's treatment and / or recovery, using any of the methods of the preceding aspects. The method may include monitoring the subject at a first time point and subsequent time points and increasing or otherwise modifying the treatment if a comparison of the values ​​of one or more metrics related to the first time point and subsequent time points indicates that the subject's heart failure condition has not improved. The method may also include monitoring the subject at a first time point and subsequent time points and maintaining or decreasing the treatment if a comparison of the values ​​of one or more metrics related to the first time point and subsequent time points indicates that the subject's heart failure condition has improved.

[0063] A fourth aspect provides a method for monitoring a subject diagnosed with or at risk of a dyspnea and / or fatigue-related condition, comprising: obtaining an audio recording from the subject from a word-reading test, which includes reading a sequence of words taken from a set of n words; and analyzing the audio recording or a portion of the audio recording by identifying multiple segments of the audio recording corresponding to individual words or syllables, determining values ​​for one or more metrics selected from breath percentage, voiceless / voiced ratio, voice pitch, and correct word rate, based at least partially on the identified segments, and comparing the values ​​of the one or more metrics to one or more respective reference values. The method may have any of the features described in relation to the first aspect.

[0064] A fifth aspect provides a method for evaluating the level of dyspnea and / or fatigue in a subject, comprising: obtaining an audio recording from the subject from a word-reading test, which includes reading a sequence of words taken from a set of n words; and analyzing the audio recording or a portion of the audio recording by identifying multiple segments of the audio recording corresponding to individual words or syllables, determining values ​​for one or more metrics, preferably including the correct word rate, selected from breath percentage, voiceless / voiced ratio, voice pitch, and correct word rate, based at least partially on the identified segments, and comparing the values ​​of the one or more metrics with one or more respective reference values. The method may have any of the features described in relation to the first aspect.

[0065] According to the sixth aspect, a method is provided for treating a subject diagnosed with or at risk of having a dyspnea and / or fatigue-related condition, comprising: assessing the subject's level of dyspnea and / or fatigue using the method of the preceding aspect; and, depending on the results of the assessment, treating the subject for the condition or adjusting the subject's treatment for the condition. The method may include performing assessments at a first time point and subsequent time points, and increasing or otherwise modifying treatment if a comparison of values ​​of one or more metrics related to the first time point and subsequent time points indicates that the subject's level of fatigue and / or dyspnea has increased or not improved. The method may also include performing assessments at a first time point and subsequent time points, and maintaining or decreasing treatment if a comparison of values ​​of one or more metrics related to the first time point and subsequent time points indicates that the subject's level of fatigue and / or dyspnea has improved or not increased. The method may have any of the features described in relation to the first aspect.

[0066] A seventh aspect provides a method for diagnosing a subject with a respiratory infection such as COVID-19, or for treating a patient diagnosed with a respiratory infection, the method comprising: obtaining an audio recording from the subject from a word-reading test, which includes reading a sequence of words taken from a set of n words; and analyzing the audio recording or a portion of the audio recording by identifying multiple segments of the audio recording corresponding to individual words or syllables, determining values ​​for one or more metrics selected from breath percentage, voiceless / voiced ratio, voice pitch, and correct word rate, based at least partially on the identified segments, and comparing the values ​​of one or more metrics, including voice pitch, to one or more respective reference values. The method may further include any of the features of the first aspect.

[0067] The method may include treating a subject for a respiratory infection if comparison indicates that the subject has a respiratory infection. The method may further include monitoring the subject's treatment and / or recovery using any of the methods described above. The method may include monitoring the subject at a first time point and subsequent time points and increasing or otherwise modifying treatment if a comparison of the values ​​of one or more metrics related to the first time point and subsequent time points indicates that the subject's respiratory infection has not improved. The method may also include monitoring the subject at a first time point and subsequent time points and maintaining or decreasing treatment if a comparison of the values ​​of one or more metrics related to the first time point and subsequent time points indicates that the subject's respiratory infection has improved.

[0068] According to the eighth aspect, a system is provided comprising at least one processor and at least one non-temporary computer-readable medium containing instructions that, when executed by the at least one processor, cause the at least one processor to perform an operation including a step of any embodiment of the method of any of the above aspects.

[0069] A non-transient computer-readable medium storing instructions that cause at least one processor to perform an operation including a step of any embodiment of the method described above when executed by at least one processor.

[0070] A computer program product that includes instructions causing at least one processor to perform an operation including a step of any embodiment of the method described above when executed by at least one processor. [Brief explanation of the drawing]

[0071] [Figure 1] An exemplary computing system in which embodiments of the present invention can be used is shown. [Figure 2] This flowchart shows a method for evaluating a subject's physiological and / or pathological state by determining the correct word rate from a word reading test. [Figure 3] This flowchart shows a method for assessing a subject's physiological and / or pathological state by determining voice pitch, breath percentage, and / or voiceless / voiced ratio from a word reading test. [Figure 4] This outlines methods for diagnosing, prognosing, or monitoring subjects. [Figure 5A-5B] A two-step method for identifying word boundaries according to an exemplary embodiment is shown. (A) A coarse word boundary was identified on a relative energy scale. A Mel-frequency spectrogram of the input audio input was constructed, and the maximum intensity projection of the Mel-frequency spectrogram along the frequency axis generated the relative energy. (B) One coarsely segmented word (highlighted in gray) was split into two estimated words based on onset intensity. [Figure 6]An exemplary embodiment of an outlier removal technique is shown. All segmented words are parameterized using the first three MFCCs (Mel-Frequency Cepstrum Coefficients), and inliers (estimated words, n=75) shown in gray and outliers (non-speech sounds, n=3) shown in black are plotted in a 3D scatter plot. [Figures 7A-7B] This paper illustrates a clustering method for identifying words according to an exemplary embodiment. Estimated words from a single recording (three different words were presented in a word-reading test) were grouped into three distinct clusters by applying K-means clustering. The visual appearance of the words within the three characteristic clusters is shown in the upper graph (one word per row), and the corresponding cluster centers are shown in the lower graph. In particular, (A) represents three word clusters from a test spoken in English (words = 75), and (B) represents three word clusters from another test spoken in German (words = 64). [Figure 8] This document illustrates a word sequence alignment method in an exemplary embodiment. In particular, it shows the application of the Smith-Waterman algorithm to a 10-word sequence. Aligning the displayed sequence RRBGGRGBRR with the predicted sequence BRBGBGBRRB reveals partially overlapping sequences and yields five correct words: matches (|), gaps (-), and mismatches (:). [Figure 9A-9B]This shows the classification accuracy of a model-less word recognition algorithm in an exemplary embodiment. The classification accuracy of each word is displayed as a normalized confusion matrix (row sum = 1). Rows represent the true labels from manual annotations, and columns represent the predicted labels from the automated algorithm. Correct predictions are on the diagonal with a black background, and incorrect predictions are on the gray background. (A) English words: / r / for / red / (n=582), / g / for / green / (n=581), and / b / for / blue / (n=553). (B) German words: / r / for / rot / (n=460), / g / for / gruen / (n=459), and / b / for / blau / (n=429). [Figure 10] This graph shows a scatter plot comparison between clinical Stroopword scores obtained using UHDRS for a series of patients with Huntington's disease and automated rating scales according to exemplary embodiments. Linear relationships between variables were determined by regression. The resulting regression lines (black lines) and 95% confidence intervals (gray shaded areas) are plotted. The significance levels of Pearson correlation coefficients (r) and p-values ​​are shown in the graph. [Figures 11A-11B] The distribution of the number of correctly pronounced words (A) and the number of individual word / syllable segments identified in sets of English, French, Italian, and Spanish recordings (B) is shown. The data demonstrates that the number of correctly pronounced words identified according to the method described herein is robust to variations in word length, even when multiple syllables within individual words are identified as separate entities (Figure 13B). [Figures 12A-12B]The results of matched Stroop word reading (A, non-consistent condition) and Stroop color word reading (B, interference condition) tests from healthy individuals analyzed as described herein are shown. Each subfigure shows the set of words displayed in each test (upper panel), the normalized signal amplitude of each recording (middle panel) (with segment identification and word prediction (shown as the color of each segment) superimposed), and the Mel spectrogram and accompanying scale of the signal shown in the middle panel (lower panel). The data demonstrate that segment identification and correct word counting processes function equally well under both the non-consistent and interference conditions. [Figure 13] The image shows a screenshot of a web-based word reading application according to an exemplary embodiment. Participants were asked to record themselves performing five different reading tasks: (i) reading a fixed, predetermined passage of text (patient consent statement) – also referred to herein as the “reading task”; (ii) reading a set of increasing consecutive numbers – also referred to herein as the “counting task”; (iii) reading a set of decreasing consecutive numbers – also referred to herein as the “reverse counting task”; (iv) Stroop word reading test (non-contradictory portion) – reading a fixed number of randomly selected colored words displayed in black; (v) Stroop colored word reading test (interfering portion) – reading a fixed number of randomly selected colored words displayed in randomly selected colors. [Figure 14A-14D]This section shows the results of an analysis of audio recordings from the Stroop reading test performed by healthy individuals after rest (light gray series) or moderate exercise (climbing four steps - dark gray series), as analyzed as described herein. Each subfigure shows the result for one of the biomarker metrics described herein. Each pair of points with the same “TEST DAY” (x-axis) shows the rest and post-exercise results for the same individual on the same day (results for the same test are shown across subfigures on the same “TEST DAY,” n=15 days). (A) Pitch - Estimated average pitch (Hz) across all audio segments of the Stroop color word reading test (interference condition) recording, Cohen's d=2.75. (B) Correct word rate (number of correct words per second in the Stroop color word reading test recording), Cohen's d=-1.57. (C) Silent / Voiced Ratio (unitless - total time between voiced segments versus total time from voiced segments in the Stroop color word reading test recording), Cohen's d=1.44. (D) Breath % (% - total time between voiced segments versus total time between and within voiced segments in the Stroop color word reading test recording), Cohen's d=1.43. (A')~(D') show the same metrics as (A)~(D), but are obtained using combined results from the Stroop color word reading test recordings for which data is shown in (A)~(D) and from the Stroop word reading test recordings from the same test session. (A') Pitch - combined test, Cohen's d=3.47. (B') Correct Word Rate - combined test, Cohen's d=-2.26. (C') Silent / Voiced Ratio - combined test, Cohen's d=1.25. (D') Breath % - combined test, Cohen's d=1.26. [Figure 15A-15J]This shows the results of an analysis of voice recordings from the Stroop reading test (A-D, interference conditions; A'-D', combinations of interference and non-contradictory conditions), reading tasks (E-G), and digit counting tasks (H-J, reverse digit counting; H'-J', combinations of forward and reverse counting) in three groups of heart failure patients: patients with decompensated heart failure at admission (labeled "HF: Admitted," n=25), the same decompensated heart failure patients at discharge (labeled "HF: Discharged," n=25), and stable outpatients (labeled "OP: Stable," n=19). (A) Box plot of respiration % (%, calculated as 100 * (voiceless / (voiceless + voiced))) overlaid on patient data. (B) Box plot of voiceless / voiced ratio (unitless, calculated as voiceless / voiced) overlaid on patient data. Voiceless / voiced ratio in the word reading test (word color reading test, interference condition) differed significantly between each of the two groups of decompensated HF patients and stable patients (HF: inpatient vs. OP: stable: Cohen d=1.75, permutation test p=0.0000; HF: discharged vs. OP: stable: Cohen d=1.77, permutation test p=0.0000). (B) Box plot of voiceless / voiced ratio (unitless, calculated as voiceless / voiced) overlaid on patient data. Voiceless / voiced ratio in the word reading test (word color reading test, interference condition) differed significantly between each of the two groups of decompensated HF patients and stable patients (HF: inpatient vs. OP: stable: Cohen d=1.31, permutation test p=0.0000; HF: discharged vs. OP: stable: Cohen d=1.52, permutation test p=0.0000). (C) Box plot of correct word rate (number of correct words per second) overlaid on patient data. The correct word rate in the word reading test (word color reading test, interference condition) differed significantly between each of the two groups of decompensated HF patients and stable patients (HF: inpatient vs. OP: stable: Cohen d=-1.14, permutation test p=0.0001; HF: discharged vs. OP: stable: Cohen d=-0.87, permutation test p=0.0035). (D) Box plot of speech rate (number of words per second) overlaid on patient data.Speech rates in the word reading test (word color reading test, interference condition) differed significantly between each of the two groups of decompensated HF patients and stable patients (HF: inpatient vs. OP: stable: Cohen d=-0.89, permutation test p=0.0019; HF: discharged vs. OP: stable: Cohen d=-0.98, permutation test p=0.0011). (A') Box plot of respiratory % overlaid on patient data. Respiratory % in the word reading test (word color reading test, combination of interference and non-contradictory conditions) differed significantly between each of the two groups of decompensated HF patients and stable patients (HF: inpatient vs. OP: stable: Cohen d=1.71, permutation test p=0.0000; HF: discharged vs. OP: stable: Cohen d=1.85, permutation test p=0.0000). (B') Box plot of voiceless / voiced ratio overlaid on patient data. The voiceless / voiced ratio in the word reading test (word color reading test, interference and non-contradictory conditions) differed significantly between each of the two groups of decompensated HF patients and stable patients (HF:inpatient vs. OP:stable:Cohen d=1.41, permutation test p=0.0000; HF:discharged vs. OP:stable:Cohen d=1.71, permutation test p=0.0000). (C') Box plot of correct word rate overlaid on patient data. The correct word rate in the word reading test (combination of word color reading test, interference and non-contradictory conditions) differed significantly between each of the two groups of decompensated HF patients and stable patients (d=-1.09 for HF:inpatient vs. OP:stable:Cohen, permutation test p=0.0002; d=-0.81 for HF:discharged vs. OP:stable:Cohen, permutation test p=0.0053). (D') Box plot of speech rate (words per second) overlaid on patient data. The speech rate in the word reading test (combination of word color reading test, interference and non-contradictory conditions) differed significantly between each of the two groups of decompensated HF patients and stable patients (d=-0.92 for HF:inpatient vs. OP:stable:Cohen, permutation test p=0.0019; d=-0.95 for HF:discharged vs. OP:stable:Cohen, permutation test p=0.0013). (E) Box plot of respiratory %(%) overlaid on patient data.(F) Box plot of voiceless / voiced ratio (unitless) overlaid on patient data. The voiceless / voiced ratio in the reading task differed significantly between each of the two groups of decompensated HF patients and stable patients (HF: inpatient vs. OP: stable: Cohen d=1.54, permutation test p=0.0000; HF: discharged vs. OP: stable: Cohen d=1.28, permutation test p=0.0000). (F) Box plot of voiceless / voiced ratio (unitless) overlaid on patient data. The voiceless / voiced ratio in the reading task differed significantly between each of the two groups of decompensated HF patients and stable patients (HF: inpatient vs. OP: stable: Cohen d=1.35, permutation test p=0.0000; HF: discharged vs. OP: stable: Cohen d=0.89, permutation test p=0.0002). (G) Box plot of speech rate (words per second) overlaid on patient data. Speech rates in the reading task differed significantly between each of the two groups of decompensated HF patients and stable patients (HF:inpatient vs. OP:stable:Cohen d=-1.60, permutation test p=0.0000; HF:discharged vs. OP:stable:Cohen d=-0.64, permutation test p=0.0190). (H) Box plot of respiration %(%) overlaid on patient data. Respiration % in the reverse counting task did not differ significantly between the decompensated HF patient group and the stable patient group (HF:inpatient vs. OP:stable:Cohen d=-0.24, permutation test p=0.2251; HF:discharged vs. OP:stable:Cohen d=-0.21, permutation test p=0.2537). (I) Box plot of voiceless / voiced ratio (unitless) overlaid on patient data. The silent / voiced ratio in the reverse counting task did not differ significantly between the two groups of decompensated HF patients and stable patients (HF:inpatient vs. OP:stable:Cohen d=-0.19, permutation test p=0.2718; HF:discharged vs. OP:stable:Cohen d=-0.26, permutation test p=0.2126). (J) Box plot of speech rate (words per second) overlaid on patient data. The speech rate in the reverse counting task did not differ significantly between the two groups of decompensated HF patients and stable patients (HF:inpatient vs. OP:stable:Cohen d=0.19, permutation test p=0.2754; HF:discharged vs. OP:stable:Cohen d=0.22, permutation test p=0.2349).(H') Box plot of respiration %(%) overlaid on patient data. Respiration % in the combination counting task did not differ significantly between at least one of the decompensated HF patient groups and stable patients. (I') Box plot of voiceless / voiced ratio (unitless) overlaid on patient data. Voiceless / voiced ratio in the combination counting task did not differ significantly between at least one of the two groups of decompensated HF patients and stable patients. (J') Box plot of speech rate (words per second) overlaid on patient data. Speech rate in the combination counting task did not differ significantly between the two groups of decompensated HF patients and stable patients. *p value (permutation test) < 0.05, **p value (permutation test) < 0.01, ***p value (permutation test) < 0.001, ****p value (permutation test) < 0.0001 ns = not significant (> 0.05). All permutation tests were performed using 10,000 permutations. [Figure 16] The results of the analysis of voice recordings from the Stroop reading test in three groups of heart failure patients—namely, patients with decompensated heart failure at admission (black data series, n=25) and the same decompensated heart failure patients at discharge (dark gray data series, n=25) (data series on the left of the plot, 2 points per patient (admission and discharge))—and stable outpatients (light gray data series, n=19 - data series on the right of the plot) are shown in terms of mean pitch (points) and standard deviation (error bars). The error bars show the standard deviation between the normal condition and the interference condition. [Figure 17] This shows the results of an analysis of the average pitch of voice recordings from the Stroop reading test in selected decompensated heart failure patients from admission (marked "admission") to discharge (the last data point for each patient). A. Female patients (n=7). B. Male patients (n=17). [Figure 18]The Bland-Altman plots assess the level of agreement in pitch measurements between the Stroop word reading test and the Stroop color reading test in 48 heart failure patients (A, 162 sets of records analyzed) and between the digit count test and the inverse digit count test in 48 heart failure patients (B, 161 sets of records analyzed). Each data point represents the difference in mean pitch (Hz) estimated using the respective test. The dashed line shows the mean difference (center line) and a standard deviation (SD) interval of ±1.96. Reproducibility was quantified using a consensus report (CR=2*SD) and was 27.76 for the digit count test and 17.64 for the word reading test. [Figure 19] This shows the results of an analysis of voice recordings (estimated voice pitch) from the same subjects during COVID-19 isolation (A, B) and on the day they returned to work (C), using the Stroop reading test (interference condition). (A-C) shows the pitch contour (white dots) overlaid on the Mel spectrogram. (D) For subjects diagnosed with COVID-19, data for the days they self-reported mild fatigue symptoms (estimated pitch = 247 Hz shown in vertical line A) and mild dyspnea symptoms (estimated pitch = 223 Hz shown in vertical line B) during isolation, as well as asymptomatic data on the day they returned to work (estimated pitch = 201 Hz shown in line C), are shown on a histogram showing data for 10 healthy female volunteers (n=1026 voice samples) and the estimated normal distribution probability density function (mean=183, sd=11; estimated by fitting these 1026 samples using the fit function from scipy.stats.norm).

[0072] Where drawings described herein illustrate embodiments of the invention, they should not be construed as limiting the scope of the invention. Where necessary, similar reference numerals are used in different figures to relate to the same structural features of the illustrated embodiments. [Modes for carrying out the invention]

[0073] Detailed explanation Specific embodiments of the present invention will be described below with reference to the drawings.

[0074] Figure 1 shows an exemplary computing system in which embodiments of the present invention can be used.

[0075] A user (not shown) has a first computing device, which is typically a mobile computing device such as a mobile phone 1 or a tablet. Alternatively, computing device 1 may be fixed, such as a PC. Computing device 1 has at least one processor 101 and at least one memory 102 that collaboratively provide at least one execution environment. Typically, the mobile device has firmware, and the application runs in at least one normal execution environment (REE) with an operating system such as iOS, Android, or Windows. Furthermore, computing device 1 may include means 103 for communicating with other elements of the computing infrastructure, for example, via the public internet 3. These may include a radio telecommunications device for communicating with a radio telecommunications network and a local radio communication device for communicating with the public internet 3 using, for example, Wi-Fi technology.

[0076] The computing device 1 typically includes a user interface 104, which includes a display. The display 104 may be a touchscreen. Other types of user interfaces may be provided, such as a speaker, a keyboard, or one or more buttons (not shown). Furthermore, the computing device 1 may be equipped with sound capture means, such as a microphone 105.

[0077] Furthermore, a second computing device 2 is also shown in Figure 1. The second computing device 2 can, for example, form part of an analytics provider computing system. The second computing device 2 typically comprises one or more processors 201 (e.g., servers), multiple switches (not shown), and one or more databases 202, and further details of the second computing device 2 used are not necessarily required to understand the functional aspects and possible implementation methods of the embodiments of the present invention and will not be described further here. The first computing device 1 can be connected to the analytics provider computing device 2 by a network connection, such as via the public internet 3.

[0078] Figure 2 is a flowchart illustrating a method for assessing a subject's physiological and / or pathological condition by determining the correct word rate from a word reading test. The method includes step 210 of obtaining an audio recording from the subject from a word reading test. The audio recording is from a word reading test that involves reading a sequence of words taken from a (closed) set of n words.

[0079] In some embodiments, a word is a color word. In some such embodiments, a word is displayed in a single color in a word reading test. In such a setting, the total number of words correctly read over a given period may match the Stroop word count from the first part (under the “non-consistency condition”) of a three-part Stroop test. In some embodiments, a word is a color word displayed in a color that does not necessarily correspond to the meaning of the individual word. For example, a word may be randomly or pseudo-randomly selected from a set of color words, and each word may be displayed in a color randomly or pseudo-randomly selected from a set of colors. In some embodiments, a word is a color word displayed in a color that does not (or does not necessarily correspond to, i.e., is selected independently of the meaning of the individual word) correspond to the meaning of the individual word. For example, a word may be randomly or pseudo-randomly selected from a set of color words, and each word may be displayed in a color randomly or pseudo-randomly selected from a set of colors excluding the color that matches the color word to be displayed. The colors included in the set of colors for display may be the same as or different from the colors included in the set of color words. In such embodiments, the total number of words correctly read aloud over a given period may coincide with the Stroop word count from the third part ("inconsistency condition") of the three-part Stroop test. In some embodiments, the audio recording includes a first recording from a word reading test, which involves reading aloud a sequence of words taken from a (closed) set of n words, where the words are color words displayed in a single color; and a second recording from a word reading test, which involves reading aloud a sequence of words taken from a (closed) set of n words, where the words are color words displayed in a color that does not necessarily correspond to the meaning of the individual words (e.g., selected independently of the meaning of the individual words). The sequences of words used in the first and second recordings may be identical. Thus, the words for the first and second word reading tests only need to be taken once from the set of n words.This conveniently increases the amount of information available to identify segments and clusters (see below), resulting in two records that can be used to measure one or more biomarkers, and such biomarkers can later be compared between the two records (for example, to assess the stability of the measurement and / or to investigate the effects that are more likely to influence one or more of the measurements for the first and second word-reading tests).

[0080] In some embodiments, n is 2 to 10, preferably 2 to 5, for example, 3. The number of different words n in the word sequence is preferably at least 2, otherwise the subject will not need to read any further words after reading the first word. The number of different words n for generating a set of words is preferably 10 or less, otherwise the number of times each word is expected to appear in the audio recording may be so small as to negatively impact the accuracy of the clustering process (see below). Preferably, the number of different words n is selected such that the number of times each word is expected to appear in the set of words read aloud by the subject is at least 10. As those skilled in the art will understand, this may depend at least on the length of the set of words and the expected length of the recording that the subject is expected to be able to undertake in light of their condition (e.g., level of fatigue and / or shortness of breath). Appropriate choices regarding the number of different words n and the length of the set of words can be obtained, for example, by using equivalent training cohorts.

[0081] The n words may be color words, such as words for each of the colors “red”, “green”, and “blue” (i.e., [’RED’, ’GREEN’, ’BLUE’] in English, [’ROT’, ’GRUEN’, ’BLAU’] in German, [’ROJO’, ’VERDE’, ’AZUL’] in Spanish, [’ROUGE’, ’VERT’, ’BLEU’] in French, [’RφD’, ’GRφN’, ’BLÅ’] in Danish, [’CZERWONY’, ’ZIELONY’, ’NIEBIESKI’] in Polish, [’КРАСНЫЙ’, ’ЗЕЛЕНЫЙ’, ’СИНИЙ’] in Russian, [’赤’, ’緑’, ’青’] in Japanese, [’ROSSO’, ’VERDE’, ’BLU’] in Italian, [’ROOD’, ’GROEN’, ’BLAUW’] in Dutch, etc.). Color words are commonly used in the word reading part of the Stroop reading test. Words for each of the colors “red”, “green”, and “blue” are common choices for this test, and thus can enable comparison or integration of the test results with existing embodiments of the Stroop test in a clinical setting.

[0082] In some embodiments, the n words are selected such that each contains a single vowel. In some embodiments, the n words are selected such that each word contains one or more vowels within the word. In some embodiments, the word contains a single stressed syllable.

[0083] In any preferred embodiment, a word is either a one-syllable word or a two-syllable word. It may be even more advantageous if all words have the same number of syllables. For example, it may be advantageous if all words are either one-syllable or two-syllable. Embodiments using only one-syllable words may be particularly advantageous because in such embodiments each segment corresponds to a single word. Conveniently, such embodiments result in a count of the number of segments corresponding to the number of words read aloud, and / or the timing of the segments which can be directly used to obtain the speech rate (or any other feature relating to the rhythm of speech). Furthermore, n words that are one-syllable can improve the accuracy of clustering because a single vector of values ​​is expected for each word, resulting in n clusters which are expected to be relatively uniform. Moreover, the use of one-syllable words can improve the accuracy of speech rate determination because it eliminates potential problems that may be associated with identifying syllables belonging to the same word.

[0084] Embodiments that use only two-syllable words can conveniently result in a count of the number of segments that can be related to the number of words read aloud (and therefore the speech rate / correct word rate) and / or compared between audio recordings from word-reading tests having the same characteristics.

[0085] In some embodiments using two-syllable words, the method may further include excluding a segment corresponding to a specified one of the two syllables in a word before counting the number of segments identified in the speech recording and / or determining the number of words correctly pronounced in the speech recording. A segment corresponding to one of the two syllables in a word can be identified based on the relative timing of two consecutive segments. For example, segments that are closely adjacent to each other, such as segments whose total time is less than a certain time (e.g., 400 milliseconds) and / or whose interval is less than a certain time (e.g., 10 milliseconds), can be assumed to belong to the same word. Furthermore, a particular segment to be excluded can be identified as the first or second segment of two segments assumed to belong to the same word. Alternatively, a particular segment to be excluded may be identified based on the characteristics of the sound signals in the two segments. For example, the segment with the lowest energy can be excluded. Another alternative is to identify a particular segment to be excluded based on the relative length of the two segments. For example, the segment with the shortest length can be excluded. Alternatively, the method may include merging a segment corresponding to a specified one of two syllables within a word with a closely following or preceding segment, such as segments that are within a specified time interval (e.g., 10 milliseconds) of each other. While we do not wish to be bound by any particular theory, merging segments corresponding to syllables in the same word can be extremely difficult when analyzing fast speech. Therefore, merging segments that are within a specified time interval of each other is considered particularly suitable for speech that is at a speed similar to or slower than free speech. In embodiments where the speech is expected to be relatively fast, it may be preferable to use segments that are presumed to directly correspond to a single syllable rather than merging or excluding segments.

[0086] In embodiments using two-syllable words (or, more generally, multi-syllable words), the two-syllable word preferably has one emphasized syllable. While we do not wish to be bound by theory, it is thought that clustering (see below) may be more robust to the presence of “noise” arising from segments corresponding to syllables rather than words when one of the syllables is emphasized. Indeed, in such cases, signals from unemphasized syllables can be considered noise in the clustering process, which still produces clusters that are uniform in terms of the identity of the emphasized syllable assigned to each cluster.

[0087] In some embodiments, a sequence of words includes at least 20, at least 30, at least 40, at least 50, or about 60 words. In some embodiments, a set of words is randomly selected from a set of n words. In some embodiments, the method includes randomly selecting a set of words from a set of n words and displaying the set of words on a computing device associated with the subject. In some embodiments, a set of words is displayed in groups of m words on a line, where m may be, for example, 4. Displaying four words per line has been found to be convenient in typical smartphone screen display situations as described herein. As will be understood by those skilled in the art, the number of words displayed as a group (m) can be adjusted according to the size of the screen / window on which the words are displayed and / or according to the user's preference (e.g., preferred font size). Such adjustment may be automatic, for example, through automatic detection of the screen or window size. Preferably, groups of m words are displayed simultaneously. For example, all words in a line of, for example, four words are preferably displayed simultaneously. This reduces the risk that the test results are influenced by external parameters (i.e., parameters that do not represent the user's ability to perform the word-reading test), such as delays in the display of consecutive words. In some embodiments, portions of n words can be displayed simultaneously, and these portions can be updated as the user progresses through the test, for example, by individual downward scrolling. In some embodiments, all n words are displayed simultaneously. Such embodiments can conveniently reduce the influence of external parameters, such as delays in the display of consecutive words, delays in the display of new words, or delays in the user's downward or upward scrolling to restart from the beginning of a set of words.

[0088] In some embodiments of any part of the process, acquiring an audio recording involves evaluating the quality of the audio recording by determining the noise level and / or signal-to-noise ratio of the recording. The signals (or noise) in the recording can be estimated based on (e.g., by averaging) relative energy values ​​assumed to correspond to the signals (or noise). The assumed relative energy values ​​corresponding to the signals may be, for example, the upper x (where x may be, for example, 10%) relative energy values ​​observed in the recording. Similarly, the assumed relative energy values ​​corresponding to background noise may be, for example, the lower x (where x may be, for example, 10%) relative energy values ​​observed in the recording. Conveniently, when relative energy is used, the values ​​of the signals and / or noise in decibels can be expressed as 10 * It can be calculated as log10(relE), where relE is a relative energy value such as the average relative energy value of the top 10% or bottom 10% of relative energy values ​​observed in the recording. As will be further explained below, the relative energy value may also be obtained by normalizing the observed power (also called energy) value relative to the highest value observed in the recording. This results in the highest observed energy having a relative energy of 0 dB. In such embodiments, the signal-to-noise ratio can be determined as the ratio of the estimated signal (e.g., the average relE of the top x% of relE observed in the recording) as described above to the noise (e.g., the average relE of the top x% of relE observed in the recording) as described above. This is then calculated as log of this ratio. 10The result can be obtained by calculating and multiplying the result by 10 to provide a value in dB units. In some such embodiments, the method may include analyzing an audio recording if the noise level is below a predetermined threshold and / or the signal level is above a predetermined threshold and / or the signal-to-noise ratio is above a predetermined threshold. Suitable thresholds for the noise level may be selected as -70 dB, -60 dB, -50 dB, or -40 dB (preferably about -50 dB). Suitable thresholds for the signal-to-noise ratio may be selected as 25 dB, 30 dB, 35 dB, or 40 dB (preferably above 30 dB). In some embodiments, acquiring an audio recording may include applying one or more preprocessing steps to a previously acquired audio recording audio file. In the context of the present invention, “preprocessing steps” refers to any steps applied to the audio recording data prior to the analysis according to the present invention (i.e., identification of individual word segments). In some embodiments, acquiring an audio recording may include applying one or more preprocessing steps to reduce the size of a previously acquired audio recording audio file. For example, downsampling can be used to reduce the size of the audio file used. The inventors of this invention have found that audio recordings can be downsampled to 16 Hz without impairing the performance of the method. This may be particularly advantageous when analysis is performed on a remote computing device and the recording is acquired on the user's computing device, as it facilitates the transmission of audio recordings from the user's computing device to a remote computing device.

[0089] In step 220, multiple segments of the speech recording corresponding to individual words or syllables are identified. Step 220 may be performed as described below in relation to Figure 3 (step 320).

[0090] In steps 230-270, the correct word rate (number of correctly read words per unit of time) in the audio recording is determined.

[0091] In particular, in step 230, one or more Mel-frequency cepstrum coefficients (MFCCs) are calculated for each of the segments identified in step 220. As a result, multiple vectors of values ​​are obtained, each vector relating to one segment. In the embodiment shown in Figure 2, an optional step 232 is shown in which the MFCCs are normalized across segments in the recording, and an optional step 234 is shown in which each of the multiple vectors is compressed to a common size. In particular, a set of i MFCCs (e.g., 12 MFCCs: MFCC 2 to 13) is calculated for each frame of the segment, and a set of j values ​​(e.g., 12 values) is obtained for the segment by compressing the signal formed by each of the i MFCCs across frames in the segment, resulting in a vector of ixj values ​​(e.g., 144 values) for the segment.

[0092] In step 240, the multiple vectors of values ​​are clustered into n clusters (for example, using k-means), where n is the expected number of different words in the word reading test. No specific label (i.e., word identity) is associated with each cluster. Instead, it is assumed that segments corresponding to the same word (for one-syllable words) or the same syllable of the same word (for two-syllable words) are taken up by the MFCC to form the same cluster. For two-syllable words, one of the syllables in the word may be dominant in the clustering, and segments corresponding to the same dominant syllable are assumed to be taken up by the MFCC to form the same cluster. Non-dominant syllables can effectively act as noise in the clustering. According to these assumptions, each cluster should group values ​​corresponding to segments primarily containing one of the n words, and one of the n! possible permutations of the n labels for these clusters corresponds to the (unknown) true label.

[0093] In step 250, the sequence of words in the audio recording is predicted for each of the n! possible permutations of the n labels. For example, for each possible assignment of the n labels, clusters are predicted for the identified segments, and the corresponding labels are predicted as words incorporated into the identified segments. Some identified segments may not be associated with a cluster, for example, because the MFCC of the segment is not predicted to belong to a particular cluster with sufficient confidence. In such cases, no words can be predicted for this segment. This may be the case, for example, for a segment corresponding to a syllable / word misdetection, or for a segment corresponding to an unemphasized syllable of a multisyllable word.

[0094] In step 260, sequence alignment is performed (for example, using the Smith-Waterman algorithm) between each of the predicted word sequences and the word sequences used in the word reading test. The word sequences used in the word reading test may be retrieved from memory or received (for example, along with the audio recording) by the processor performing each step of the method.

[0095] In step 270, the label that yields the best alignment (e.g., the label that yields the highest alignment score) is selected and assumed to be the true label of the cluster. The alignment matches are assumed to correspond to words that were read correctly in the audio recording and can be used to calculate the ground truth word rate. The ground truth word rate can be obtained, for example, by dividing the total number of words read correctly (matches) by the total recording time. Alternatively, the ground truth word rate can be obtained by calculating multiple local means within each time window and then considering the resulting multiple ground truth word rate estimates, or by finding a metric (e.g., mean, median, mode) for summarizing the multiple ground truth word estimates. Preferably, the ground truth word rate can be estimated as the gradient of a linear model fitted to the cumulative number of ground truth words read as a function of time. Such a count may be incremented by one at the time corresponding to the start of any segment identified as corresponding to a correctly read word. In yet another embodiment, determining the correct word rate for an audio recording involves dividing the recording into several equal time bins, calculating the total number of correctly pronounced words in each time bin, and calculating a summarized measure of the correct word rate across the time bins. For example, the mean, trimmed mean, or median of the correct word rate across the time bins can be used as a summarized measure of the correct word rate. Using the median or trimmed mean can conveniently reduce the impact of outliers, such as bins that contain no words at all.

[0096] If multiple audio recordings are obtained, they may be analyzed separately or at least partially together. In some embodiments, multiple audio recordings are obtained for the same subject, and at least steps 220 and 230 are performed individually for each audio recording. In some embodiments, multiple audio recordings are obtained for the same subject, and at least step 240 is performed together using values ​​from multiple of the recordings. In some embodiments, steps 250–270 are performed individually for each recording using the results of clustering step 240, which is performed using values ​​from one or more (e.g., all) of the multiple recordings.

[0097] Figure 3 is a flowchart illustrating a method for assessing a subject's physiological and / or pathological condition by determining voice pitch, breath percentage, and / or voiceless / voiced ratio from a word-reading test. The method includes step 310 of obtaining an audio recording from the subject from a word-reading test. The audio recording may be an audio recording from a word-reading test that involves reading a sequence of words taken from a (closed) set of n words. In particular, the words preferably do not have any particular logical connection.

[0098] In step 320, multiple segments of the audio recording corresponding to individual words or syllables are identified. In such cases, it is particularly advantageous that the words used in the reading test are single-syllable, as each segment can be assumed to correspond to a single word, and therefore the timing of the segments can be directly associated with the speech rate. If two-syllable words (or other multi-syllable words) are used, it may be advantageous that all words have the same number of syllables, as this simplifies the calculation and / or interpretation of the speech rate.

[0099] In step 330, the breath percentage and / or voiceless / voiced ratio and / or voice pitch associated with the voice recording are determined using at least partially the segments identified in the voice recording.

[0100] The breathing percentage reflects the proportion of time in the recording that includes voiced segments. This can be calculated as the ratio between the time between segments identified in step 320 and the total time in the recording, or the sum of the time within a segment identified in step 320 and the time between segments identified in step 320. The voiceless / voiced ratio represents the time in the recording during which the subject is breathing, or is assumed to be breathing, relative to the time in the recording during which the subject is vocalizing. The voiceless / voiced ratio can be determined as the ratio of (i) the time between segments identified in step 320 and (ii) the time within a segment identified in step 320.

[0101] The audio pitch for an audio recording or a segment thereof refers to an estimate of the fundamental frequency of the sound signal in the recording. Therefore, the audio pitch may also be referred to herein as F0 or f0, where "f" refers to frequency and the index "0" indicates that the estimated frequency is assumed to be the fundamental frequency. The fundamental frequency of a signal is the reciprocal of the fundamental period of the signal, which is the minimum repetition interval of the signal. Various computational methods are available for estimating the pitch (or its fundamental frequency) of a signal, and all such methods may be used herein. Numerous computational pitch estimation methods estimate the pitch of a signal by dividing the signal into time windows, and then, for each window, (i) estimating the spectrum of the signal (e.g., using a short-time Fourier transform), (ii) calculating a score for each pitch candidate within a given range (e.g., by computing an integral transform on the spectrum), and (ii) selecting the candidate with the highest score as the estimated pitch. Such methods may yield multiple pitch estimates (one for each time window). Therefore, the pitch estimate of a signal may be provided as a summarized estimate spanning a window (e.g., the mean, mode, or median pitch spanning the window) and / or range. More recently, deep learning-based methods have been proposed, some of which determine the pitch estimate of a signal (i.e., provide as output a predicted pitch for the signal, rather than for each of several windows in the signal). Determining the speech pitch may include obtaining a speech pitch estimate or a speech pitch estimate range for each segment identified in step 320. The speech pitch of a segment may be a summarized estimate of the speech pitch spanning the segment, such as the mean, median, or mode of several speech pitch estimates for the segment. The speech pitch range of a segment may be a speech pitch range that is expected to encompass a certain proportion of several speech pitch estimates for the segment. For example, the speech pitch range of a segment may be the interval between the lowest and highest pitch estimates from several speech pitch estimates for the segment.Alternatively, the segment's pitch range may be the interval between the x-th and y-th percentiles of the segment's multiple pitch estimates. Another alternative is that the segment's pitch range may be the interval corresponding to the confidence interval around the mean pitch of the segment's multiple pitch estimates. Such a confidence interval can be obtained by applying a range centered on the mean, which is expressed in units of the estimated standard deviation centered on the mean (e.g., mean ± n SD (wherein SD is the standard deviation and n may be any predetermined value)). Determining the pitch may include obtaining a summarized pitch estimate or summarized pitch range across segments identified in step 320 from which pitch estimates or pitch range estimates have been obtained. A summarized pitch estimate across multiple segments may be obtained as the mean, median, or mode of the multiple pitch estimates for each segment. The estimated range of the summarized speech pitch across segments can be obtained as described above, using the estimated speech pitch for each segment (which may include one, for example, summarized speech pitch estimate per segment, or multiple speech pitch estimates).

[0102] The speech pitch of a segment (or multiple speech pitches) can be estimated using any method known in the art. In particular, the speech pitch of a segment can be estimated using the SWIPE or SWIPE' method described by Camacho and Harris (2008). Preferably, the estimated speech pitch of a segment is obtained by applying SWIPE' to the segment. This method has been shown to provide a good balance between computational accuracy and speed. Compared to SWIPE, SWIPE' reduces harmonic errors by using only the first and principal harmonics of the signal. Alternatively, pitch estimation can be performed using a deep learning method such as the CREPE method described by Kim et al. (2018). This method has been shown to provide a robust pitch estimate, although it is computationally more demanding than methods such as SWIPE or SWIPE'. Other methods can also be used, such as PYIN (as described by Mauch and Dixon (2014)) or the method described by Ardaillon and Roebel (2019). Pitch estimation is typically applied using signals from a time window (also called a "frame," as mentioned above). Therefore, pitch estimation for a segment can generate multiple estimates, each corresponding to one frame. Appropriately, multiple pitch estimates (e.g., corresponding to multiple frames within a segment) may be further processed to reduce estimation errors, for example, by applying a median filter. The inventors of this invention have found that a median filter applied using a 50ms window is particularly suitable. The average of such filtered estimates for a segment can be used as the pitch estimate for the segment.

[0103] Next, a method used to identify multiple segments of a speech recording corresponding to individual words or syllables is described. Other methods exist in the art and such other methods may also be used in other embodiments. In the embodiment shown in Figure 3, in step 322, a power mel spectrogram of the speech recording is obtained. This is typically achieved by defining frames along the speech recording (frames may correspond to signals in a fixed-width sliding window applied along the time axis) and calculating the power spectrum of each frame on a mel scale (typically by obtaining a spectrogram for each frame and then mapping the spectrogram to a mel scale using overlapping triangular filters along a range of frequencies assumed to correspond to the human hearing range). This process yields a matrix of power values ​​per mel unit for each time bin (a time bin corresponds to one of the positions in the sliding window). Thus, in some embodiments of any aspect, obtaining a power mel spectrogram of a speech recording involves applying 138 triangular filters over a sliding window (preferably having a size of 15 ms and a step size of 10 ms) and a range of 25.5 Hz to 8 kHz. While we do not wish to be bound by theory, using a relatively narrow time window (e.g., 10-15ms, as opposed to 25ms or more) may be useful in the context of identifying segments corresponding to individual words or syllables, particularly for the purpose of identifying segment boundaries corresponding to the beginning of a word or syllable. This is because using a relatively narrow time window may improve the sensitivity of detection, while using a wider time window may smooth out small signals that could be rich in information.

[0104] As those skilled in the art will understand, overlapping triangular filters (typically 138) applied to a frequency spectrogram (Hz scale) are commonly used to obtain a Mel-scale spectrogram. Furthermore, the fact that it covers the range from 25.5 Hz to 8 kHz has been found to be advantageous for adequately capturing the range of human hearing.

[0105] Optionally, the power mel spectrogram may be normalized, for example, by dividing the value of each frame by the highest energy value observed in the recording (323). In step 324, the maximum intensity projection of the mel spectrogram along the frequency axis is obtained. Segment boundaries are identified as the points in time when the maximum intensity projection of the mel spectrogram along the frequency axis intersects a threshold (326). In particular, two sets of consecutive boundaries, where the maximum intensity projection of the mel spectrogram intersects the threshold from lower to higher values ​​at the first boundary and from higher to lower values ​​at the second boundary, may be considered to define a segment corresponding to a single word or syllable. The threshold used in step 326 may optionally be determined dynamically in step 325 (the term "dynamically determined" means that the threshold for a particular speech recording is determined according to the characteristics of that particular speech recording, rather than being predetermined independently of that particular recording).

[0106] Therefore, in some embodiments, the threshold is determined dynamically for each recording. Preferably, the threshold is determined as a function of the maximum intensity projection value for the recording. For example, the threshold may be determined as a weighted average of the relative energy values ​​assumed to correspond to the signal and the relative energy values ​​assumed to correspond to the background noise. The relative energy values ​​assumed to correspond to the signal may be, for example, the top x (where x may be, for example, 10%) relative energy values ​​observed in the recording. Similarly, the relative energy values ​​assumed to correspond to the background noise may be, for example, the bottom x (where x may be, for example, 10%) relative energy values ​​observed in the recording. It may be particularly convenient to use the average of the top 10% relative energy values ​​and the average of the bottom 10% relative energy values ​​across frames. Alternatively, a predetermined value of relative energy assumed to correspond to the signal (i.e., the audio signal) may be used. For example, a value of about -10 dB has been commonly observed by the inventors of the present invention and can be usefully selected. Similarly, a predetermined value of relative energy assumed to correspond to the background noise may be used. For example, a value of approximately -60 dB has been commonly observed by the inventors of this invention and can be usefully selected.

[0107] If the threshold is determined as a weighted average of the relative energy values ​​assumed to correspond to the signal and the relative energy values ​​assumed to correspond to the background noise, the weight of the latter may be selected between 0.5 and 0.9, and the weight of the former may be selected between 0.5 and 0.1. In some embodiments, the weight for the contribution of background noise may be greater than the weight for the contribution of the signal. This may be particularly advantageous when the audio recording has been preprocessed by performing one or more noise-canceling steps. Indeed, in such cases, the lower end of the signal (low relative energy) may contain more information than would be expected for a signal that has not been preprocessed with respect to noise cancellation. Many modern computing devices, including mobile devices, can produce audio recordings that are preprocessed to some extent in this manner. Therefore, it may be useful to emphasize the lower end of the relative energy values ​​to some extent. Weights of about 0.2 and about 0.8 for the contributions of the signal and background noise, respectively, may be advantageous. Furthermore, an advantageous threshold may be determined by trial and error and / or formal training with training data. While we do not wish to be bound by theory, the use of dynamically determined thresholds may be particularly advantageous when the audio recording includes a reference tone and / or when the signal-to-noise ratio is good (e.g., above a predetermined threshold such as 30 dB). Conversely, the use of predetermined thresholds may be particularly advantageous when the audio recording does not include a reference tone and / or when the signal-to-noise ratio is poor.

[0108] In other embodiments, the threshold is predetermined. In some embodiments, the predetermined threshold is selected between -60dB and -40dB, such as -60dB, -55dB, -50dB, -45dB, or -40dB. Preferably, the predetermined threshold is about -50dB. The inventors of the present invention have found that this threshold provides a good balance between sensitivity and specificity in the identification of word / syllable boundaries in high-quality audio recordings, particularly audio recordings preprocessed using one or more noise cancellation steps.

[0109] Optionally, the segments may be “refined” by analyzing the separate segments identified in step 326 and determining whether further (internal) boundaries can be found. Thus, identifying segments of a speech recording corresponding to individual words or syllables may further include performing onset detection for each segment and forming two new segments by defining further boundaries whenever an onset is detected within a segment.

[0110] This can be done by performing onset detection on at least one of the segments by calculating a spectral flux function for the segment's mel spectrogram (327), and then forming two new segments by defining further (internal) boundaries whenever an onset is detected within a segment (328). Onset detection using spectral flux functions is commonly used in the analysis of music recordings for beat detection. As those skilled in the art will understand, onset detection using spectral flux functions is a method of examining the derivative of an energy signal. In other words, the spectral flux function measures how quickly the power spectrum of a signal is changing. It can therefore be particularly useful for identifying “valleys” (sudden changes in the energy signal) in a signal that may correspond to the beginning of a new word or syllable within a segment. This can conveniently “refine” the segmentation as needed. This technique can be particularly useful as a “refinement step” when word / syllable boundaries have already been identified using less sensitive techniques that result in “coarse” segments. This is because, at least in part, this technique can be applied independently to segments using parameters appropriate for those segments (e.g., thresholds for onset detection).

[0111] Performing onset detection (327) may include calculating a spectral flux function or onset intensity function (327a), normalizing the segment's onset intensity function to a value between 0 and 1 (327b), smoothing the (normalized) onset intensity function (327c), and applying a threshold to the spectral flux function or a function derived therefrom (327d), where an onset is detected if the function increases beyond the threshold. Thus, performing onset detection may include applying a threshold to the spectral flux function or a function derived therefrom, where an onset is detected if the function increases beyond the threshold. In some embodiments, performing onset detection includes normalizing the segment's onset intensity function to a value between 0 and 1, and separating the segment into subsegments if the normalized onset intensity exceeds a threshold. Thresholds between 0.1 and 0.4, such as between 0.2 and 0.3, may result in particularly low false positive rates when applied to the normalized onset intensity function. An appropriate threshold can be defined as the threshold that minimizes the false positive detection rate when the method is applied to training data.

[0112] In some embodiments, the onset detection process includes calculating the onset intensity over time from a powermel spectrogram using the superflux method described by Boeck S and Widmer G (2013) (based on the spectral flux function, but including a spectral trajectory tracking step to a common spectral flux calculation method). In some embodiments, the onset detection process includes calculating the onset intensity function over time from a powermel spectrogram using the superflux method, such as the one implemented in the LibROSA library (https: / / librosa.github.io / librosa / , see function librosa.onset.onset_strength; McFee et al. (2015)). Preferably, the onset detection process further includes normalizing the onset intensity function of a segment to a value between 0 and 1. This can be achieved, for example, by dividing each value of the onset intensity function by the maximum onset intensity in the segment. Normalizing the onset intensity function can result in a reduction in the number of false positive detections.

[0113] In some embodiments, performing onset detection further includes smoothing the (optionally, normalized) onset intensity function of the segment. For example, smoothing can be achieved by calculating a moving average with a fixed window size. For example, a window size of 10–15 ms may be useful, such as 11 ms. Smoothing can further reduce the rate of false positives detected.

[0114] A discretionary false detection removal step 329 is shown in Figure 3. The process described herein for identifying correctly pronounced words is conveniently tolerant of the presence of falsely detected segments, at least to some extent. This is because the alignment step can include, at least in part, a gap for false detections that does not significantly affect the overall accuracy of the method. Therefore, in some embodiments, the false detection removal step may be omitted. In the embodiment shown in Figure 3, the false detection removal step includes calculating one or more Mel-frequency cepstrum coefficients (MFCCs) for a segment (preferably the first three MFCCs, as they are expected to capture features that distinguish noise from true speech) to obtain a plurality of vectors of values, each vector relating to one segment (329a), and excluding all segments whose vectors of values ​​are beyond a predetermined distance from the remaining vectors of values ​​(329b). This technique assumes that the majority of segments are correct detections (i.e., corresponding to true speech), and that segments that do not contain true speech have MFCC features that differ from correct detections. Other outlier detection methods may be applied to exclude some of the vectors of values ​​that are suspected to be associated with false positives.

[0115] In some embodiments, identifying segments of a speech recording corresponding to individual words / syllables further includes excluding segments representing false positives by removing segments shorter than a predetermined threshold and / or segments whose average relative energy falls below a predetermined threshold. For example, segments shorter than 100 ms may be conveniently excluded. Similarly, segments with an average relative energy of less than -40 dB may be conveniently excluded. Such techniques can easily and efficiently exclude segments that do not correspond to a word or syllable. Preferably, segments are filtered to exclude short segments and / or low-energy segments prior to the calculation of the MFCC of the segments and the application of the outlier detection method as described above. In fact, this conveniently avoids the unnecessary step of calculating the MFCC for false segments and prevents such false segments from introducing further noise into the outlier detection method.

[0116] Calculating one or more Mel-frequency cepstrum coefficients (MFCCs) for a segment typically involves defining frames along a segment of an audio recording (where a frame can correspond to a signal within a fixed-width window applied along the time axis). The window is typically a sliding window, i.e., a window of a predetermined length (e.g., 10–25 ms, 25 ms, etc.) that moves along the time axis in defined step lengths (e.g., 3–10 ms, 10 ms, etc.), resulting in partially overlapping frames. Calculating one or more MFCCs typically involves, for each frame, calculating the Fourier transform (FT) of the signal within the frame, mapping the power of the thus obtained spectrum to a Mel scale (e.g., using a triangular overlapping filter), finding the logarithm of the power at each Mel frequency, and performing the discrete cosine transform of the thus obtained signal (i.e., obtaining the spectrum of the spectrum). The resulting spectral amplitude represents the MFCC of the frame. As mentioned above, a set of 138 Mel values ​​is generally obtained for a power Mel spectrum (i.e., the frequency range is generally mapped to 138 Mel scale values ​​using 138 overlapping triangular filters). However, through the process of calculating the MFCC, this information is compressed into a smaller set of values ​​(MFCC), typically to 13 values. Often, the information contained in the majority of the 138 Mel values ​​is correlated so that this compression of the signal does not result in a detrimental loss of information in the signal.

[0117] In particular, the calculation of one or more Mel-frequency cepstrum coefficients (MFCCs) for a segment can be performed as described by Rusz et al. (2015). The calculation of one or more Mel-frequency cepstrum coefficients (MFCCs) for a segment can be performed as implemented in the LibROSA library (https: / / librosa.github.io / librosa / ; McFee et al. (2015); see librosa.feature.mfcc). Alternatively, the calculation of one or more MFCCs for a segment can be performed as implemented in the library "python_speech_features" (James Lyons et al., 2020).

[0118] In some embodiments, the calculation of one or more Mel-frequency cepstrum coefficients (MFCCs) of a segment includes calculating at least the first three MFCCs for each frame of the segment (optionally, all 13 MFCCs) and obtaining a vector of at least three values ​​for the segment (one for each MFCC used) by calculating a summarized scale for each MFCC across the frames in the segment. The number and / or identity of the at least three MFCCs used in the outlier detection method can be determined using training data and / or internal control data. For example, at least three MFCCs may be selected as the minimum set of MFCCs sufficient to remove a certain percentage of error detections in the training data (e.g., at least 90%, or at least 95%). As another example, at least three MFCCs may be selected as the minimum set of MFCCs sufficient to remove segments corresponding to internal controls (e.g., a reference tone, as further described below). Preferably, only the first three MFCCs are used in the outlier detection method. This conveniently captures information that allows us to separate true words / syllables from false detections (e.g., breaths, non-speech sounds) without introducing information that could result in different words forming separate distributions of points that could potentially confuse the outlier detection process.

[0119] In some embodiments, applying an outlier detection method to multiple vectors of values ​​involves excluding all segments where the vector of values ​​exceeds a predetermined distance from the rest of the vector of values. The distance between a particular vector of values ​​and the rest of the vector of values ​​can be quantified using the Mahalanobis distance. The Mahalanobis distance is a convenient measure of distance between a point and a distribution. It has the advantages of being unitless, scale-invariant, and taking into account the correlation of the data. Alternatively, the distance between a particular vector of values ​​and the rest of the vector of values ​​can be quantified using the distance between the particular vector of values ​​and a representative value of the rest of the vector of values ​​(e.g., mean or medoid) (e.g., Euclidean distance, Manhattan distance). The values ​​may optionally be scaled to have a unit variance along each coordinate, for example, before applying outlier detection. The predetermined distance may be selected depending on the observed variability in the multiple vectors of values. For example, the predetermined distance may be a multiple of a measure of data variability, such as the standard deviation, or the value of a selected quantile. In such embodiments, the predetermined distance may be selected depending on the expected proportion of false detections. A threshold of 1 to 3 times the standard deviation centered on the mean of multiple vectors of values ​​can be selected, enabling accurate removal of outliers. In particular, a threshold of 2 times the standard deviation proved advantageous when the expected false positive rate is approximately 5%.

[0120] A nearly identical method for false positive removal is described by Rusz et al. (2015). However, the method described in that document is significantly more complex than the method of this disclosure. In particular, it relies on an iterative process in which, in each iteration, inliers and outliers are identified using quantile-based thresholds on the distribution of mutual distances, and then outliers are excluded using quantile-based thresholds on the distribution of distances between inliers and outliers, as defined earlier. A simpler method as described herein may be advantageous in the context of the present invention. While we do not wish to be bound by theory, the false positive removal method described herein is considered particularly advantageous in this context because of its low false positive rate. This may be partly due to the extremely high accuracy of the segment detection method described herein. While we do not wish to be bound by theory, the method for syllable segmentation used in Rusz et al. (2015) (which relies on parameterizing the signal into 12 MFCCs within a sliding window of length 10 ms and step of 3 ms, searching for a low-frequency spectral envelope that can be described using the first three MFCCs, then calculating the mean of each of the three MFCCs within each envelope, and separating these points into syllables and interludes using k-means method) may not be as precise as the method described herein. This is because, at least in part, it is designed to identify the contrast between interludes and words, since all words are identical, and in part, the method in Rusz et al. (2015) relies heavily on an outlier detection process of repetitions to improve the overall accuracy of the process of identifying true positive segments. In fact, the method in Rusz et al. (2015) was specifically developed to handle syllable detection using speech recordings in which patients are asked to repeat the same syllable at a comfortable pace. Therefore, the data consists only of segments (intervals and syllables) of two expected categories of homogeneous content. In such cases, the first three MFCCs can be used in combination with a complex iterative error detection process for segment identification to achieve good accuracy.However, this may be less accurate in the context of analyzing audio recordings from word reading tests, as at least two or more types of syllables are expected.

[0121] The segments identified in step 320 can be used to determine the words that were read correctly and therefore the correct word rate in a word reading test, as described in relation to Figure 2 (steps 230-270).

[0122] The inventors of the present invention have identified that respiratory %, voiceless / voiced, voice pitch, and ground truth word rate, as determined in relation to Figures 2 and 3, can be used as biomarkers indicating a subject's physiological or pathological state. In particular, the biomarkers measured as described herein, especially the biomarkers of respiratory %, voiceless / voiced, and ground truth word rate, have been found to be extremely sensitive indicators of a subject's level of dyspnea and / or fatigue. Furthermore, the method for obtaining voice pitch estimates as described herein has been found to yield highly reliable estimates that can be used as biomarkers or indicators of any physiological or pathological state associated with voice pitch variation. Thus, the methods described herein can be used for the diagnosis, monitoring, or treatment of any condition, disease, or disorder associated with dyspnea, fatigue, and / or voice pitch variation.

[0123] Figure 4 schematically illustrates how monitoring, diagnosis, or prognosis of a subject's disease, disorder, or condition is provided. A disease, disorder, or condition is a disease, disorder, or condition that affects breathing, voice tone, fatigue, and / or cognitive ability.

[0124] The method includes step 410 of obtaining an audio recording from a subject from a word reading test. In the illustrated embodiment, obtaining an audio recording includes causing a computing device associated with the subject (e.g., computing device 1) to display a set of words (e.g., on display 104) (310a) and causing computing device 1 to record an audio recording (e.g., via microphone 105) (310b). Optionally, obtaining an audio recording may further include causing the computing device to emit a reference tone (310c). Alternatively, or in addition, step 310 of obtaining an audio recording from a subject from a word reading test may include receiving the audio recording from a computing device associated with the subject (e.g., computing device 1).

[0125] The method further includes step 420, which identifies multiple segments of the speech recording corresponding to individual words or syllables. This may be performed as described in relation to Figure 3. The method further includes step 430, which optionally includes determining the speech rate for the speech recording by counting the number of segments identified in the speech recording, at least in part. The method further includes step 470, which determines the correct word rate in the speech recording, as described in relation to Figure 2 (steps 230-270). The correct word rate derived from the speech recording may indicate the subject's level of cognitive impairment, fatigue, and / or shortness of breath. The method further includes (430a), which optionally includes determining the respiratory percentage in the speech recording, as described in relation to Figure 3 (steps 320 and 330). The respiratory percentage derived from the speech recording may indicate the subject's level of cognitive impairment, fatigue, and / or shortness of breath. The method further includes (430b), which optionally includes determining the voiceless / voiced ratio in the speech recording, as described in relation to Figure 3 (steps 320 and 330). The respiratory percentage derived from the voice recording may indicate the subject's level of cognitive impairment, fatigue, and / or shortness of breath. The method optionally includes determining the voice pitch in the voice recording (430c), as described in relation to Figure 3 (steps 320 and 330). The voice pitch derived from the voice recording may indicate the subject's physiological and / or pathological condition, such as a subject with dyspnea, heart failure decompensation, or infection (particularly a pulmonary infection). The method may further include step 480, which compares the metrics obtained in steps 430 and 470 with one or more previously obtained values ​​for the same subject, or one or more reference values. One or more reference values ​​may include one or more values ​​of one or more previously obtained metrics for the same subject. Thus, any method described herein may include a step of repeating the method for the same subject at one or more linkage points (e.g., repeating steps 410-480).One or more reference values ​​may include one or more values ​​of one or more metrics previously obtained from one or more reference populations (e.g., one or more training cohorts).

[0126] A disease, disorder, or condition can be monitored in a subject diagnosed with a disease, disorder, or condition, or a subject can be diagnosed with the likelihood of having a condition, including symptoms such as dyspnea and / or fatigue, by comparing with previously obtained values ​​for the same subject. Alternatively, a disease, disorder, or condition can be diagnosed by comparing with previously obtained values ​​for the same subject. A subject can be diagnosed with a disease, disorder, or condition, or the progression, recovery, or treatment of a disease, disorder, or condition, by comparing with one or more reference values. For example, reference values ​​may correspond to disease populations and / or healthy populations. Monitoring of a disease, disorder, or condition in a subject can be used to automatically evaluate the progress of a treatment, for example, to determine whether the treatment is effective.

[0127] Step 420, which identifies multiple segments of the speech recording corresponding to individual words or syllables; Step 430, which determines the breath percentage, voiceless / voiced ratio, or pitch of the speech recording; and Step 470, which determines the correct word rate in the speech recording, may be performed by the user computing device 1 or the analysis provider computer 2.

[0128] Accordingly, the present disclosure relates to a method for monitoring a subject diagnosed with or at risk of having a condition affecting respiration, voice tone, fatigue, and / or cognitive ability, in some embodiments, the method comprising: obtaining an audio recording from the subject from a word reading test; identifying a plurality of individual word / syllable segments; determining the value of one or more biomarkers selected from respiration %, voiceless / voiced ratio, voice pitch, and correct word rate, at least in part on the identified segments; and comparing the value of one or more biomarkers to one or more respective reference values. In some embodiments of any aspect, the method further comprises treating the subject for a disease, disorder, or condition.

[0129] A subject may be undergoing or receiving a particular set of treatments. Accordingly, reference to monitoring a subject may include, for example, monitoring a subject's treatment by measuring one or more biomarkers disclosed herein at a first time point and a subsequent time point, and comparing the biomarkers measured at the first time point and the subsequent time point to determine whether one or more of the subject's symptoms have improved between the first time point and the subsequent time point. Such a method may further include modifying or recommending modification of the subject's set of treatments if the comparison indicates that one or more of the subject's symptoms have not improved, or have not improved sufficiently.

[0130] Furthermore, a method is disclosed for diagnosing a subject having a condition affecting respiration, voice tone, fatigue, and / or cognitive ability, comprising: obtaining an audio recording from the subject from a word-reading test; identifying a plurality of individual word / syllable segments; determining the value of one or more biomarkers selected from respiration%, voiceless / voiced ratio, voice pitch, and correct word rate, at least partially based on the identified segments; and comparing the value of one or more biomarkers to one or more respective reference values. In some embodiments, one or more biomarkers are selected from respiration%, voiceless / voiced ratio, and correct word rate, and one or more reference values ​​are predetermined values ​​associated with patients having the condition and / or patients not having the condition (e.g., healthy individuals). The predetermined values ​​associated with patients having the condition and / or patients not having the condition may have been previously obtained using one or more training cohorts. In some embodiments, one or more biomarkers include voice pitch, and one or more reference values ​​are values ​​previously obtained from the same subject.

[0131] The condition may be one related to dyspnea and / or fatigue. Accordingly, the Disclosure also provides a method for monitoring a subject diagnosed with or at risk of having a condition related to dyspnea and / or fatigue, comprising: obtaining an audio recording from the subject from a word-reading test; identifying a plurality of individual word / syllable segments; determining the value of one or more biomarkers selected from respiration%, voiceless / voiced ratio, voice pitch, and correct word rate, based at least partially on the identified segments; and comparing the value of one or more biomarkers to one or more respective reference values. Similarly, the Disclosure also provides a method for assessing the level of dyspnea and / or fatigue in a subject, comprising: obtaining an audio recording from the subject from a word-reading test; identifying a plurality of individual word / syllable segments; determining the value of one or more biomarkers selected from respiration%, voiceless / voiced ratio, voice pitch, and correct word rate, based at least partially on the identified segments; and comparing the value of one or more biomarkers to one or more respective reference values.

[0132] The condition may be a cardiovascular disease such as heart failure, coronary heart disease, myocardial infarction (heart attack), atrial fibrillation, arrhythmia (cardiac disorder), and valvular heart disease. In certain embodiments, the condition is heart failure. Accordingly, the disclosure also provides a method for identifying a subject with heart failure as having decompensated heart failure, comprising: obtaining an audio recording from the subject from a word-reading test; identifying a plurality of individual word / syllable segments; determining the value of one or more biomarkers selected from respiration%, voiceless / voiced ratio, voice pitch, and correct word rate, at least in part on the identified segments; and comparing the value of one or more biomarkers to one or more respective reference values. In some embodiments, one or more biomarkers are selected from respiration%, voiceless / voiced ratio, and correct word rate, and one or more reference values ​​are predetermined values ​​associated with patients with decompensated heart failure and / or patients with stable heart failure. Predefined values ​​associated with patients with decompensated heart failure and / or stable heart failure may have been previously obtained using one or more training cohorts. In some embodiments, one or more biomarkers include voice pitch, and one or more reference values ​​are values ​​previously obtained from the same subjects.

[0133] In some embodiments, the Disclosure also provides a method for monitoring subjects with decompensated heart failure, comprising: obtaining an audio recording from the subject from a word-reading test; identifying a plurality of individual word / syllable segments; determining the value of one or more biomarkers selected from respiration%, voiceless / voiced ratio, voice pitch, and correct word rate, at least in part on the identified segments; and comparing the value of one or more biomarkers to one or more respective reference values. In some embodiments, the one or more biomarkers are selected from respiration%, voiceless / voiced ratio, and correct word rate, and the one or more reference values ​​are predetermined values ​​associated with patients with decompensated heart failure and / or patients with stable heart failure and / or patients with decompensated heart failure in recovery. The predetermined values ​​associated with patients with decompensated heart failure and / or patients with stable heart failure and / or patients with decompensated heart failure in recovery may have been previously obtained using one or more training cohorts. In some embodiments, one or more biomarkers include voice pitch, and one or more reference values ​​are values ​​previously obtained from the same subject. For example, one or more reference values ​​may include one or more values ​​obtained when the subject was diagnosed with decompensated heart failure.

[0134] In some embodiments, one or more biomarkers include respiratory% and respiratory% above a predetermined baseline or range indicates that the subject is likely to have a condition associated with dyspnea and / or fatigue, where the baseline or range relates to a subject or group of subjects that is less likely to have the condition. In some embodiments, one or more biomarkers include respiratory% and respiratory% below a predetermined baseline or range indicates that the subject is likely to be recovering from a condition associated with dyspnea and / or fatigue, where the subject has been diagnosed with the condition, and the baseline or range has been previously obtained from the same subject, for example, when the subject was diagnosed with the condition. In some embodiments, a subject has been diagnosed with a condition associated with dyspnea and / or fatigue and is receiving treatment for that condition, and one or more biomarkers, including respiratory%, indicate that respiratory% below a predetermined baseline or range of values ​​is likely to indicate that the subject is responding well to the treatment. The predetermined baseline or range of values ​​may have been previously obtained from the same subject, for example, at the time the subject was diagnosed with the condition, or from a group of subjects known to have the condition. In some embodiments, a subject has been diagnosed with a condition associated with dyspnea and / or fatigue and is receiving treatment for that condition, and one or more biomarkers, including respiratory%, indicate that respiratory% within or above a predetermined baseline or range of values ​​is likely to indicate that the subject is not responding well to the treatment.The predetermined reference values ​​or ranges may have been previously obtained from the same subject, for example, at the time the subject was diagnosed with the condition, or from a group of subjects known to have the condition. The condition may be decompensated heart failure.

[0135] In some embodiments, one or more biomarkers include a voiceless / voiced ratio, where a voiceless / voiced ratio above a predetermined baseline or range indicates that the subject is likely to have a condition associated with dyspnea and / or fatigue, and the predetermined baseline or range relates to a subject or group of subjects that is less likely to have the condition. In some embodiments, one or more biomarkers include a voiceless / voiced ratio, where a voiceless / voiced ratio below a predetermined baseline or range indicates that the subject is likely to be recovering from a condition associated with dyspnea and / or fatigue, and the predetermined baseline or range relates to a subject or group of subjects that has the condition. In some embodiments, one or more biomarkers include a voiceless / voiced ratio, where a voiceless / voiced ratio below a predetermined baseline or range indicates that the subject is likely to be recovering from a condition associated with dyspnea and / or fatigue, where the subject has been diagnosed with the condition, and the predetermined baseline or range has been previously obtained from the same subject, for example, when the subject was diagnosed with the condition. In some embodiments, a subject has been diagnosed with a condition associated with dyspnea and / or fatigue and is receiving treatment for that condition, and one or more biomarkers include a voiceless / voiced ratio, where a voiceless / voiced ratio below a predetermined baseline or range of values ​​indicates that the subject is likely responding well to the treatment. The predetermined baseline or range of values ​​may have been previously obtained from the same subject, for example, at the time the subject was diagnosed with the condition, or from a group of subjects known to have the condition. In some embodiments, a subject has been diagnosed with a condition associated with dyspnea and / or fatigue and is receiving treatment for that condition, and one or more biomarkers include a voiceless / voiced ratio, where a voiceless / voiced ratio within or above a predetermined baseline or range of values ​​indicates that the subject is likely not responding well to the treatment.The predetermined reference values ​​or ranges may have been previously obtained from the same subject, for example, at the time the subject was diagnosed with the condition, or from a group of subjects known to have the condition. The condition may be decompensated heart failure.

[0136] In some embodiments, one or more biomarkers include a ground truth word rate, where a ground truth word rate below a predetermined baseline or range indicates that the subject is likely to have a condition associated with dyspnea and / or fatigue, and the predetermined baseline or range relates to subjects or groups of subjects who are not likely to have the condition. In some embodiments, one or more biomarkers include a ground truth word rate, where a ground truth word rate above a predetermined baseline or range indicates that the subject is likely to be recovering from a condition associated with dyspnea and / or fatigue, and the predetermined baseline or range relates to subjects or groups of subjects who have the condition. In some embodiments, one or more biomarkers include a ground truth word rate, where a ground truth word rate above a predetermined baseline or range indicates that the subject is likely to be recovering from a condition associated with dyspnea and / or fatigue, where the subject has been diagnosed with the condition, and the predetermined baseline or range has been previously obtained from the same subject, for example, when the subject was diagnosed with the condition. In some embodiments, a subject has been diagnosed with a condition associated with dyspnea and / or fatigue and is receiving treatment for that condition, and one or more biomarkers, including a correct word rate, indicate that a correct word rate within or above a predetermined baseline or range is likely to indicate that the subject is responding well to the treatment. The predetermined baseline or range may have been previously obtained from the same subject, for example, at the time the subject was diagnosed with the condition, or from a group of subjects known to have the condition. In some embodiments, a subject has been diagnosed with a condition associated with dyspnea and / or fatigue and is receiving treatment for that condition, and one or more biomarkers, including a correct word rate, indicate that a correct word rate within or below a predetermined baseline or range is likely to indicate that the subject is not responding well to the treatment.The predetermined reference values ​​or ranges may have been previously obtained from the same subject, for example, at the time the subject was diagnosed with the condition, or from a group of subjects known to have the condition. The condition may be decompensated heart failure.

[0137] In some embodiments, one or more biomarkers include voice pitch, where a voice pitch significantly different from a predetermined baseline or range indicates that the subject is likely to have a condition associated with dyspnea and / or fatigue, and the predetermined baseline or range relates to a subject or group of subjects that are not likely to have the condition. In some embodiments, one or more biomarkers include voice, where a voice significantly different from a predetermined baseline or range indicates that the subject is likely to be recovering from a condition associated with dyspnea and / or fatigue, and the predetermined baseline or range relates to a subject or group of subjects that have the condition. In some embodiments, one or more biomarkers include voice pitch, where a voice pitch significantly different from a predetermined baseline or range indicates that the subject is likely to be recovering from a condition associated with dyspnea and / or fatigue, where the subject has been diagnosed with the condition, and the predetermined baseline or range has been previously obtained from the same subject, for example, when the subject was diagnosed with the condition. In some embodiments, a subject has been diagnosed with a condition associated with dyspnea and / or fatigue and is receiving treatment for that condition, and one or more biomarkers, including voice pitch, indicate that a voice pitch significantly different from a predetermined baseline or range of values ​​is likely to indicate that the subject is responding well to the treatment. The predetermined baseline or range of values ​​may have been previously obtained from the same subject, for example, at the time the subject was diagnosed with the condition, or from a group of subjects known to have the condition. In some embodiments, a subject has been diagnosed with a condition associated with dyspnea and / or fatigue and is receiving treatment for that condition, and one or more biomarkers, including voice pitch, indicate that a voice pitch significantly different from a predetermined baseline or range of values ​​is likely to indicate that the subject is not responding well to the treatment.The predetermined reference values ​​or ranges may have been previously obtained from the same subject, for example, at the time the subject was diagnosed with the condition, or from a group of subjects known to have the condition. Preferably, the predetermined reference values ​​or ranges are previously obtained from the same subject.

[0138] Conditions may include obstructive pulmonary diseases (e.g., asthma, chronic bronchitis, bronchiectasis, and chronic obstructive pulmonary disease (COPD)), respiratory diseases such as chronic respiratory diseases (CRD), respiratory tract infections, and lung tumors, respiratory infections (e.g., COVID-19, pneumonia, etc.), obesity, dyspnea (e.g., dyspnea associated with heart failure, panic attacks (anxiety disorders), pulmonary embolism, physical limitations or damage to the lungs (e.g., rib fractures, lung collapse, pulmonary fibrosis, etc.), pulmonary hypertension, or any other disease, disorder, or condition affecting lung / cardiopulmonary function (e.g., measurable by spirometry).

[0139] Accordingly, a method for evaluating the lung or cardiopulmonary function of a subject is also disclosed herein, comprising: obtaining an audio recording from the subject from a word-reading test; identifying a plurality of individual word / syllable segments; determining the value of one or more biomarkers selected from respiratory %, voiceless / voiced ratio, voice pitch, and correct word rate, based at least partially on the identified segments; and comparing the value of one or more biomarkers to one or more respective reference values. Furthermore, a method for diagnosing a subject having a respiratory disease is also disclosed herein, comprising: obtaining an audio recording from the subject from a word-reading test; identifying a plurality of individual word / syllable segments; determining the value of one or more biomarkers selected from respiratory %, voiceless / voiced ratio, voice pitch, and correct word rate, based at least partially on the identified segments; and comparing the value of one or more biomarkers to one or more respective reference values. In some embodiments, one or more biomarkers are selected from respiratory %, voiceless / voiced ratio, and correct word rate, and one or more reference values ​​are predetermined values ​​associated with patients with respiratory disease and / or patients without respiratory disease (e.g., healthy individuals). The predetermined values ​​may have been previously obtained using one or more training cohorts. In some embodiments, one or more biomarkers include voice pitch, and one or more reference values ​​are previously obtained values ​​from the same subjects. Alternatively, or in addition to this, one or more biomarkers may include voice pitch, and one or more reference values ​​may include values ​​associated with patients with respiratory disease and / or patients without respiratory disease (e.g., healthy individuals). The respiratory disease is preferably a disease related to dyspnea. In some embodiments, the disease is COVID-19.

[0140] Any condition affecting a subject's respiratory capacity (including, for example, mental disorders such as anxiety disorders), any condition affecting a subject's fatigue (including, for example, mental disorders such as depression and chronic fatigue syndrome), and / or any condition affecting cognitive ability (including, for example, mental disorders such as attention deficit disorder) can be conveniently diagnosed or monitored using the method of the present invention. In particular, a condition may be a neurovascular disease or disorder such as stroke, neurodegenerative disease, myopathy, diabetic neuropathy, mental disorders or disorders such as depression, drowsiness, attention deficit disorder, chronic fatigue syndrome, or a condition affecting an individual's fatigue state or cognitive ability through systemic mechanisms such as pain, abnormal blood glucose levels (e.g., due to diabetes), or renal dysfunction (e.g., in situations such as chronic renal failure or renal replacement therapy). [Examples]

[0141] Example 1: Development of an automated smartphone-based Stroop word reading test for remote monitoring of disease symptoms. In this embodiment, the inventors of the present invention developed an automated smartphone-based Stroop word reading test (SWR) to test the feasibility of remote monitoring of disease symptoms in Huntington's disease. In the smartphone-based SWR test, color words were displayed in black on the screen according to a randomly generated sequence (4 words per line, for a total of 60 words). Speech data was recorded with a built-in microphone and uploaded to the cloud via WiFi. The inventors of the present invention developed a language-independent method for segmenting and classifying individual words from the speech signal. Finally, by comparing the displayed word sequence with a predicted word sequence, they were able to reliably estimate the number of correct words using the Smith-Waterman algorithm, which is commonly used for genome sequence alignment.

[0142] method Participants and relative clinical evaluation: As part of the HD OLE (Open-Blinded Extension) study (NCT03342053), 46 patients were recruited from three locations, including Canada, Germany, and the United Kingdom. All patients underwent extensive neurological and neuropsychological examinations at baseline visits. Disease severity was quantified using the Unified Huntington's Disease Rating Scale (UHDRS). In particular, the Stroop Word Reading Test (SCWT1-Word Raw Score) is part of the UHDRS cognitive assessment, and dysarthria (UHDRS Dysarthria Score) is part of the UHDRS motor assessment. Local languages ​​were used at each location (i.e., English in Canada and the UK (n=27), and German in Germany (n=19)).

[0143] Smartphone App and Self-Managed Speech Recording: A smartphone-based Stroop word reading test was developed as a custom Android application (Galaxy S7; Samsung, Seoul, South Korea). Upon baseline visit, patients received a smartphone and completed the test during a teaching session. Subsequently, speech tests were conducted remotely at home once a week. Speech signals were acquired at 44.1 kHz with 16-bit resolution and downsampled to 16 kHz for analysis. Data was securely transferred to a remote location via WiFi for processing and analysis. The data presented in this embodiment was from the initial self-managed home test (n=46) only. A total of 60 color words (4 words per line) were displayed in black according to a randomly generated sequence and explicitly stored as metadata. Patients read the words after a short reference tone (1.1 kHz, 50 ms) for a given 45 seconds. Patients were instructed to restart reading the words from the beginning if they completed reading all 60 words within the 45-second time limit. All the records analyzed here had low ambient noise levels (-56.7±7.4dB, n=46) and good signal-to-noise ratios (44.5±7.8dB, n=46).

[0144] A language-independent method for analyzing Stroop word reading tests: In light of its potential use in multilingual and diverse patient population contexts, the algorithm was designed without any pre-trained models. Words were segmented directly from the speech signal without any contextual cues. In the classification stage, word labels were selected to maximize the partial overlap between the displayed sequence and the predicted sequence. The fully automated method for Stroop word reading tests can be divided into four parts. In summary, the inventors of this invention first introduced a two-step method to obtain highly sensitive segmentation of individual words. Next, the inventors of this invention deployed an outlier removal step to eliminate false positives mainly caused by inaccurate pronunciation, breathing, and non-speech sounds. Then, they converted these into individual estimated words represented by 144 (12 × 12) Mel-frequency cepstrum coefficient (MFCC) features and performed 3-class K-mean clustering. Finally, the inventors of this invention employed the Smith-Waterman algorithm, a local sequence alignment method, to estimate the number of correct words. Each of these steps will be explained in more detail below.

[0145] Word Boundary Identification: In this particular example, each color word used consisted of a single syllable, namely the English / red / , / green / , / blue / and the German / rot / , / gruen / , / blau / . Thus, word segmentation becomes a general syllable detection problem. According to phonology, the syllable nucleus, also called the peak, is the central part of the syllable (most commonly the vowel), while consonants form the boundaries between them (Kenneth, 2003). Several automatic syllable detection methods have been described for connected speech (see, e.g., Xie and Niyogi, 2006; Wang and Narayanan, 2007; Rusz et al., 2016). For example, syllable nuclei are identified primarily based on either broadband energy envelopes (Xie and Niyogi, 2006) or subband energy envelopes (Wang and Narayanan, 2007). However, in the case of high-speed speech, it is difficult to distinguish transitions between different syllables solely by the energy envelope. Considering the fast tempo and syllable repetition in word-reading tasks, more sensitive syllable nucleus identification is still needed.

[0146] The two-stage method was inspired by how syllable boundaries are performed by manual labeling, i.e., by visual inspection of the intensity and spectral flux of the spectrogram. In summary, the power Mel spectrogram was first calculated using 138 triangular filters spanning the range of 25.5 Hz to 8 kHz with a sliding window size of 15 ms and a step size of 10 ms, and normalized for the strongest frame energy over a 45-second period. Next, the maximum energy of the speech frame was derived to represent the intensity corresponding to the maximum intensity projection of the Mel spectrogram along the frequency axis. In this way, the frame with the loudest sound has a relative energy value of 0 dB, and the other frames have values ​​smaller than that. For example, as shown in Figure 5A, all syllable nuclei have a relative energy above -50 dB. Coarse word boundaries were identified by using a threshold for the relative energy scale.

[0147] Next, the spectral flux of the Mel spectrogram was calculated to identify the precise boundaries of each word. This corresponds to vertical edge detection in the Mel spectrogram. Onset intensities were calculated using the superflux method developed by Boeck and Widmer (2013) and normalized to values ​​between 0 and 1. If the onset intensity exceeds a threshold, i.e., 0.2, the segment is divided into subsegments. One coarsely segmented word (highlighted in gray) was divided into two estimated words based on the onset intensities shown in Figure 5B.

[0148] All calculations were performed in Python using the Librosa library (https: / / librosa.github.io / librosa / , McFee et al., 2015) or the python_speech_features library (https: / / github.com / jameslyons / python_speech_features, James Lyons et al., 2020). For calculating onset strength, the function librosa.onset.onset_strength was used with the parameters lag=2 (time lag for calculating the difference) and max_size=3 (size of the local max filter). In the example shown in Figures 5A and 5B, 68 coarse segments were identified in the first step, and an additional 10 were identified in the refinement step.

[0149] An outlier removal step was performed to eliminate false positives primarily caused by inaccurate pronunciation, breathing, and non-speech sounds. Observations shorter than 100 ms and average relative energy values ​​shorter than -40 dB were removed first. Mel-frequency cepstrum coefficients (MFCCs) are commonly used as features in speech recognition systems (Davis and Mermelstein, 198; Huang et al., 2001). Here, a matrix of 13 MFCCs was calculated for each estimated word with a sliding window size of 25 ms and a step size of 10 ms. Audible noise is expected to differ from the true word by the first three MFCCs (Rusz et al., 2015). Therefore, words were parameterized using the mean of the first three MFCCs. Outlier detection was performed on these based on the Mahalanobis distance. Outliers were identified using a cutoff value of twice the standard deviation. Figure 6 illustrates this step, where inliers (estimated words) are shown in gray and outliers (non-speech sounds) are shown in black in the 3D scatter plot.

[0150] K-means clustering: K-means is an unsupervised clustering algorithm that divides observations into k clusters (Lloyd, 1982). The inventors of this invention assumed that words pronounced by a subject in a given recording have similar spectral representations within word clusters and different patterns between word clusters. In this way, a word can be divided into n clusters, where n is equal to the number of individual color words (here, n=3). However, the durations of the words may vary from one another (average duration of 0.23 to 0.35 ms). The steps to generate feature representations of equal size for each word are as follows: Starting with a matrix of 13 previously calculated MFCCs, the first MFCC (related to power) was removed from the matrix. The remaining 12 MFCC matrices, with various frame numbers, were treated as an image and resized to a fixed-size image (12x12 pixels, reduced to 40% to 60% of its width) by linear interpolation along the time axis. As a result, each word, regardless of its duration, was converted into a total of 144 MFCC values ​​(12 × 12 = 144). By applying K-means clustering, the estimated words from a single record were classified into three distinct clusters. Figure 7 shows the visual appearance of the words in the three discriminative clusters shown in the upper graph (one word per row) and the corresponding cluster centers shown in the lower graph, in particular, Figure 7A representing three word clusters extracted from one test in English (words = 75), and Figure 7B representing three word clusters extracted from one test in German (words = 64).

[0151] Word Sequence Alignment: Speech recognition refers to understanding the content of speech. In principle, speech recognition can be performed using deep learning models (e.g., Mozilla's free speech recognition project DeepSpeech) and hidden Markov models (e.g., Carnegie Mellon University's Sphinx toolkit). However, such pre-trained models are built on healthy populations, are language-dependent, and may not be very accurate when applied to patients with speech disorders. In this study, the inventors of the present invention introduced an end-to-end model-less solution for inferring speech content. Such a word recognition task was transformed into a genome sequence alignment problem. A closed set of color words is like the letters of a DNA code. Reading errors, as well as system errors introduced in the segmentation and clustering steps, are similar to mutations, deletions, or insertions that occur in the DNA sequence of a gene. Instead of performing isolated word recognition, the objective was to maximize the overlapping sequences between the displayed sequence and the predicted sequence so that the entire speech content is utilized as a whole.

[0152] The Smith-Waterman algorithm is suitable for partially overlapping sequences because it performs local sequence alignment (i.e., some characters may not be considered) (Smith and Waterman, 1981). The algorithm allows comparing segments of all possible lengths and optimizes a similarity metric based on a scoring metric such as gap cost = 2 and match score = 3. In this study, the number of segmented words defines the search space within the given sequence. In a 3-class situation, there are 6 (3!=6) possible permutations of word labels. For each permutation, it is possible to generate a predicted sequence, align it with the given sequence, and trace back to the segment with the highest similarity score. The inventors of this invention assumed that the subject would read the words exactly as they were presented in most cases. Therefore, segment length becomes the measure to maximize in the problem. In other words, the optimal selection of labels for a given cluster is found in a way that maximizes overlapping sequences. This allows each word to be classified according to its respective cluster label. Furthermore, the exact matches found in partially overlapping sequences provide a good estimate of the correct word read by the subject. Figure 8 shows an example of the alignment between the displayed sequence RRBGGRGBRRG and the predicted sequence BRBGBGBRRB, and returns five correct words out of the ten words read aloud.

[0153] Manual-level ground truth: Manual annotation of all segmented words (1938 words from 27 English recordings, 1452 words from 19 German recordings) was performed blindly via audio playback. Manual labeling was performed after the algorithm was designed and was not used for parameter tuning. The start / end times of each word were obtained using the proposed two-stage method. Words were labeled with / r / for / red / and / rot / , / g / for / green / and / gruen / , and / b / for / blue / and / blau / , respectively, in their respective texts. Words that were difficult to annotate for any reason (e.g., inaccurate syllable divisions, breathing, other words, etc.) were labeled as "garbage class" with / n / .

[0154] Outcome Determination Methods: Based on the results of word segmentation and classification, two complementary test-level outcome determination methods were designed: the number of correct words to quantify processing speed as part of a cognitive indicator, and the speech rate to quantify speech motor ability. Specifically, the speech rate was defined as the number of words per second and calculated as the slope of the regression line with respect to the cumulative sum of segmented words over time.

[0155] Statistical analysis: The Shapiro-Wilk test was used to test for normal distribution. Pearson correlation was applied to examine significant relationships. To evaluate the Pearson correlation coefficient, the following criteria were used: good (values ​​between 0.25 and 0.5), moderate to good (values ​​between 0.5 and 0.75), and excellent (values ​​greater than or equal to 0.75). For comparisons between groups, independent sample ANOVA and unpaired t-tests were performed. Effect size was measured using Cohen's d, where d=0.2 represents a small effect, d=0.5 represents a moderate effect, and d=0.8 represents a large effect.

[0156] result Evaluation of Word Classification Performance: To evaluate the classification accuracy of the proposed model-less word recognition algorithm, labels obtained by manual annotation and automated algorithms were compared. Overall classification accuracy was high, with average scores of 0.83 for English and 0.85 for German. The normalized confusion matrix in Figure 9 shows the performance of the model-less word classifier at the word level. The high classification accuracy suggests that the proposed word classifier can learn all components of a speech classifier, including pronunciation, acoustics, and linguistic content, directly from a 45-second speech recording. It leverages an unsupervised classifier and a dynamic local sequence alignment strategy to tag each word. This means that there is no need to carry a language model during deployment, making it highly practical for application to multilingual and diverse disease populations.

[0157] Clinical validation of two complementary outcome measures: The number of correct words, determined by a fully automated method, was compared to the standard clinical UHDRS-Stroop word score. In general, with respect to the number of correct words, smartphones and clinical measures were highly correlated, as shown in Figure 10 (Pearson correlation coefficient r=0.81, p<0.001).

[0158] Performance evaluation in additional languages: The results obtained in this study were further extended to a study including HD patients speaking 10 different languages. Specifically, the method described in this example was applied to this multilingual cohort using the following words: "English": ['RED', 'GREEN', 'BLUE'], "German": ['ROT', 'GRUEN', 'BLAU'], "Spanish": ['ROJO','VERDE', 'AZUL'], "French": ['ROUGE','VERT', 'BLEU'], "Danish": ['RφD', 'GRφN', 'BLÅ'], "Polish": ['CZERWONY', 'ZIELONY', 'NIEBIESKI'], "Russian": ['КРАСНЫЙ', 'ЗЕЛЕНЫЙ', 'СИНИЙ'], "Japanese": ['赤', '緑', '青'], "Italian": ['ROSSO','VERDE', 'BLU'], "Dutch": ['ROOD', 'GROEN', 'BLAUW']. Notably, for some of these languages, all of the words used were 1 syllable (e.g., English, German), while for other languages, some of the words were 2 syllables (e.g., Italian, Spanish). Figure 11A shows the distribution of the number of correctly pronounced words determined from a set of recordings in English, French, Italian, and Spanish, and Figure 11B shows the distribution of the number of segments (immediately before clustering, i.e., after refinement and outlier removal) identified in each of these languages. The data show that the number of correctly pronounced words identified according to the above method is robust to variations in word length (Figure 11A), even if multiple syllables within an individual word are identified as separate entities (Figure 11B).

[0159] Conclusion This embodiment describes and demonstrates the clinical applicability of an automated (smartphone-based) Stroop word reading test that can be remotely self-administered from the patient's home. The fully automated method allows for offline analysis of speech data. The method is language-independent and uses an unsupervised classifier and dynamic local sequence alignment strategy to tag each word in relation to its linguistic content. Without reliance on pre-trained models, words were classified with high overall accuracy: 0.83 for English-speaking patients and 0.85 for German-speaking patients. This method was shown to enable the assessment of cognitive and speech-motor function in HD patients. Two complementary outcome assessments, namely an assessment for assessing cognitive ability and an assessment for assessing speech-motor impairment, were clinically validated in 46 patients from the HD OLE study. In summary, the method described herein successfully establishes a basis for self-assessment of disease symptoms using a smartphone-based speech test in a large population. This can ultimately provide significant benefits to patients in improving their quality of life in most clinical trials to find effective treatments.

[0160] Example 2: Automated Stroop word reading test - Interference conditions In this embodiment, the inventors of the present invention tested whether the interference portion of the Stroop word reading test could be automatically performed using the method outlined in Example 1. Both the Stroop word reading test and the Stroop color word reading test, as described in relation to Example 1, were performed on a healthy volunteer cohort. Furthermore, the inventors of the present invention tested the performance of the method by analyzing recordings of the Stroop word reading test and the Stroop color word reading test (where the words are displayed in black in the former and in different colors in the latter) using the same sequence of words (see Figures 12A and 12B). The results of applying the method of Example 1 to two audio recordings obtained from individuals performing these paired tests are shown in Figures 12A and 12B. In these figures, segments are highlighted as colored sections of signals in the central panel of each figure, and word predictions are shown in the central panel of each figure by the color of the segments. The data show that segment identification and correct word counting work equally well under both consistent and interference conditions. In fact, despite the presence of incorrect words read aloud by individuals in the interference test, there was no discrepancy in cluster assignment between the word reading test and the interference test. Furthermore, as can be seen in Figure 12B, the predicted number of correctly read words obtained using the described automated evaluation method correlated highly with the ground truth data obtained through manual annotation of the audio recordings.

[0161] Example 3: Automated web-based Stroop word reading test for remote monitoring of respiratory symptoms and monitoring of disease symptoms in heart failure patients In this embodiment, the inventors of the present invention performed the above-described automated Stroop word reading test (SWR) in the context of remote monitoring of disease symptoms in patients with dyspnea and heart failure.

[0162] The same mechanism as in Example 1 was used, except that this solution was deployed through a web-based application. The web-based test mechanism is shown in Figure 13. Participants were asked to record themselves using their computing devices while performing the following tasks: (i) a reading task (reading a patient consent statement, see the top panel in Figure 13), (ii) a digit counting task (reading a number between 1 and 10), (iii) a reverse digit counting test (reading a number between 10 and 1), and (iv) two word reading tests: a Stroop word reading test (non-contradictory condition, i.e., a color word is randomly selected from a set of three color words as described in Example 1 and displayed in black) and a Stroop color word reading test (interference condition, i.e., a color word is randomly selected from a set of three color words and displayed in a randomly selected color).

[0163] In contrast to Example 1, the word-reading test recordings were not of a fixed length. Instead, each recording was the length required for the individual to read all the displayed words (40 words in this case). This is advantageous in that many patients with cardiac abnormalities or respiratory difficulties may not have the stamina to perform long tests. Furthermore, the words displayed in the Stroop word-reading test and the Stroop color word-reading test were identical, with the color only being changed in the Stroop color word-reading test. This conveniently allowed for comparison of recordings from the two tests because their audio content should be similar, enabling the acquisition of additional data for better accuracy in the clustering process. In fact, to ensure that the clustering process was performed with enough words to have good accuracy, the two recordings (i.e., 40 words each from the Stroop word-reading test and the Stroop color word-reading test, totaling 80 words) were used in combination for each patient in the clustering process. The segment identification process was performed separately for the two recordings, as was the alignment process. Furthermore, the segment identification process described in Example 1 was also applied to the reading task and digit count / reverse digit count recording. Next, the results of the alignment process were used along with the segment information to calculate the correct word rate (calculated as the number of correct words per second) for each of the Stroop word reading test and the Stroop color word reading test. The correct word rate was estimated as the number of correct words read divided by the test duration. The cumulative number of words read was incremented by 1 at the point corresponding to the start of all segments identified as corresponding to correctly read words. As described in Example 1, the speech rate (i.e., all words, not just correct words) was also calculated using the gradient of a linear model fitted to the cumulative number of words read.

[0164] Next, using segment information, the respiratory percentage (respiratory %, 100%) is calculated individually for each test.* We evaluated the following: (time between segments) / (time within a segment + time within a segment), the voiceless / voiced ratio (calculated as (time between segments / time within a segment)), and the average speech pitch (calculated as the average of the individual speech pitches estimated for each segment). For each segment, the speech pitch was estimated using SWIPE' implemented in the Speech Signal Processing Toolkit (http: / / sp-tk.sourceforge.net / ) via the r9y9 Python wrapper (https: / / github.com / r9y9 / pysptk). We also tested an alternative method (CREPE) implemented in the Python package available at https: / / github.com / marl / crepe. The results shown here use SWIPE'. To reduce pitch estimation error, a median filter with a size of 5 (corresponding to a 50ms time window) was applied to the pitch estimates from the speech segments. Finally, a single mean was obtained for a given recording.

[0165] This method was first tested in healthy subjects who were tested several days before and after moderate exercise (climbing four stairs). This situation simulates the effects of dyspnea and therefore tests whether the metrics described above can function as biomarkers of dyspnea. The results of this analysis are shown in Table 1 and Figure 14 below for multi-day (row) Stroop color word test records (interference conditions - panels A-D, and the mean of the results of the interference and coherent conditions - panels A'-D'), where panels A and A' show pitch estimates, panels B and D' show correct word rates, panels C and C' show voiceless / voiced ratios, and panels D and D' show respiratory percentages. Cohen's d was calculated for each metric between pre-exercise and post-exercise results to quantify the effect size associated with dyspnea for each metric. For the pitch metric, the effect size (Cohen's d) was 3.47 for the combined test data and Cohen's d = 2.75 for the interference condition only. Regarding the correct word rate, Cohen's d was -2.26 for the combined test data and Cohen's d=-1.57 for the interference condition. For voiceless / voiced, Cohen's d was 1.25 for the combined test data and Cohen's d=1.44 for the interference condition. For respiratory %, Cohen's d was 1.26 for the combined test data and Cohen's d=1.43 for the interference condition. Thus, each of these metrics shows a significant difference between the resting state and the dyspnea state (whether using data from the color word test recordings in the interference condition alone or combining data from the color word test recordings in the interference and coherent conditions), and is therefore usable for monitoring dyspnea. [Table 1]

[0166] The data in Table 1 show that each of the tested metrics exhibits a significant difference between the resting and shortness-of-breath states, and this is consistent across the word test (color words, coherent state) and the color word test (color words, interference condition) (naturally, it is likely to be higher in the coherent state, and apart from the correct word rate, which may provide further indication of cognitive ability when comparing the coherent and interference states). Therefore, these metrics can be used to monitor dyspnea (either the word test or the color word test alone, or a combination of both).

[0167] Therefore, the inventors of the present invention set out to determine whether these biomarkers could also be used to monitor patients with heart failure. Metrics were obtained as described in two cohorts of heart failure patients: a cohort of patients hospitalized due to decompensation (n=25) and a cohort of outpatients with stable heart failure (n=19). The former was evaluated both at admission (HF: admission) and at discharge (HF: discharge). The results of this analysis are shown in Tables 2 and 3, as well as in Figures 15, 16, and 17. The data in panels A-D and A'-D' of Figure 15 show that each metric derived from the Stroop word reading test (A-D: interference condition only, A'-D': mean of interference and coherent conditions) differed significantly between patients with decompensated heart failure and stable outpatients. Furthermore, the metrics of respiratory %, voiceless / voiced, and correct word rate were particularly sensitive metrics for distinguishing these patient groups. The characteristics of the data in Figures 15A' to 15D' and Figures 15A to 15D are shown below.

[0168] Stroop score: Number of correct words per second (combination color word reading test, Figure 15C'): HF: Hospitalization (mean ± standard deviation): 1.5 ± 0.4, n = 25 HF: Discharge (mean ± standard deviation): 1.6 ± 0.4, n=25 OP: Stable (mean ± standard deviation): 1.9 ± 0.2, n = 19 HF: Hospitalization vs. OP: Stable: Cohen's d: -1.09, Permutation test p-value = 0.0002 HF: Discharge vs. OP: Stable: Cohen's d: -0.81, permutation test p-value = 0.0053 HF: Hospitalization vs. HF: Discharge: Cohen's d: -0.21, permutation test p-value = 0.2276 Stroop score: Number of correct words per second (color word reading test, interference conditions, Figure 15C): HF: Hospitalization (mean ± standard deviation): 1.5 ± 0.4, n = 25 HF: Discharge (mean ± standard deviation): 1.6 ± 0.4, n=25 OP: Stable (mean ± standard deviation): 1.9 ± 0.2, n = 19 HF: Hospitalization vs. OP: Stable: Cohen's d: -1.14, permutation test p-value = 0.0001 HF: Discharge vs. OP: Stable: Cohen's d: -0.87, permutation test p-value = 0.0035 HF: Hospitalization vs. HF: Discharge: Cohen's d: -0.28, permutation test p-value = 0.1600

[0169] This data demonstrates that the correct word rate from word-reading test recordings can be used to distinguish patients with decompensated heart failure from those with stable heart failure. Furthermore, this metric can also be used to monitor patients' recovery from the decompensated state.

[0170] RST (Speech Rate): Number of words per second (combination color word reading test, Figure 15D'): HF: Hospitalization (mean ± standard deviation): 1.8 ± 0.3, n = 25 HF: Discharge (mean ± standard deviation): 1.8 ± 0.3, n=25 OP: Stable (mean ± standard deviation): 2.0 ± 0.2, n = 19 HF: Hospitalization vs. OP: Stable: Cohen's d: -0.92, Permutation test p-value = 0.0019 HF: Discharge vs. OP: Stable: Cohen's d: -0.95, permutation test p-value = 0.0013 HF: Hospitalization vs. HF: Discharge: Cohen's d: -0.07, permutation test p-value = 0.4033 RST (Speech Rate): Number of words per second (Color word reading test, interference conditions, Figure 15D): HF: Hospitalization (mean ± standard deviation): 1.8 ± 0.3, n = 25 HF: Discharge (mean ± standard deviation): 1.7 ± 0.4, n=25 OP: Stable (mean ± standard deviation): 2.0 ± 0.2, n = 19 HF: Hospitalization vs. OP: Stable: Cohen's d: -0.89, Permutation test p-value = 0.0019 HF: Discharge vs. OP: Stable: Cohen's d: -0.98, permutation test p-value = 0.0011 HF: Hospitalization vs. HF: Discharge: Cohen's d: 0.11, permutation test p-value = 0.3374

[0171] This data demonstrates that speech rate (speech timing rate, RST) from word reading test recordings can be used to distinguish patients with decompensated heart failure from those with stable heart failure. However, this metric cannot be used to monitor patient recovery from a decompensated state to a state of recovery where the patient is discharged, and it is not as sensitive as ground truth word rate. Speech rate was determined by calculating the cumulative sum of the number of identified segments in speech recordings over time and then calculating the slope of a linear regression model fitted to the cumulative sum data.

[0172] Therefore, this data suggests that by combining not only shortness of breath but also fatigue-related effects (using a metric that is more sensitive to cognitive ability while also capturing shortness of breath-related effects), a more sensitive biomarker for heart failure status can be obtained.

[0173] Breathing percentage in word reading tests (combination color word reading test, Figure 15A'): HF: Hospitalization (mean ± standard deviation): 41.9 ± 8.2, n = 25 HF: Discharge (mean ± standard deviation): 42.0 ± 7.5, n=25 OP: Stable (mean ± standard deviation): 29.6 ± 5.1, n = 19 HF: Hospitalization vs. OP: Stable: Cohen's d: 1.71, permutation test p-value = 0.0000 HF: Discharge vs. OP: Stable: Cohen's d: 1.85, permutation test p-value = 0.0000 HF: Hospitalization vs. HF: Discharge: Cohen's d: -0.02, permutation test p-value = 0.4767 Breathing percentage in word reading tests (color word reading tests, interference conditions, Figure 15A): HF: Hospitalization vs. OP: Stable: Cohen's d: 1.75, permutation test p-value = 0.0000 HF: Discharge vs. OP: Stable: Cohen's d: 1.77, permutation test p-value = 0.0000 HF: Hospitalization vs. HF: Discharge: Cohen's d: -0.00, permutation test p-value = 0.4973

[0174] Voiceless / voiced ratio in word reading tests (combination color word reading test, Figure 15B'): HF: Hospitalization (mean ± standard deviation): 0.8 ± 0.3, n = 25 HF: Discharge (mean ± standard deviation): 0.8 ± 0.2, n=25 OP: Stable (mean ± standard deviation): 0.4 ± 0.1, n = 19 HF: Hospitalization vs. OP: Stable: Cohen's d: 1.41, permutation test p-value = 0.0000 HF: Discharge vs. OP: Stable: Cohen's d: 1.70, permutation test p-value = 0.0000 HF: Hospitalization vs. HF: Discharge: Cohen's d: 0.02, permutation test p-value = 0.4760 Voiceless / voiced ratio in word reading tests (color word reading test, interference conditions, Figure 15B): HF: Hospitalization vs. OP: Stable: Cohen's d: 1.31, permutation test p-value = 0.0000 HF: Discharge vs. OP: Stable: Cohen's d: 1.52, permutation test p-value = 0.0000 HF: Hospitalization vs. HF: Discharge: Cohen's d: 0.03, permutation test p-value = 0.4659

[0175] The data above demonstrates that respiratory percentage and voiceless / voiced ratio from word-reading test recordings can be used to distinguish patients with decompensated heart failure from those with stable heart failure. Both of these metrics are highly sensitive to differences between patients with decompensated and stable heart failure, but do not significantly change between admission and discharge. Note that these two metrics are related in a quadratic relationship.

[0176] Therefore, the above metrics can be used together to identify whether a patient has decompensated heart failure or stable heart failure (using either the correct word rate, respiratory % and / or voiced / voiced ratio), to identify patients with decompensated heart failure requiring hospitalization (using the correct word rate), to identify patients with heart failure who have recovered sufficiently to be discharged but are still unstable (and therefore may require further / more extensive monitoring) (using the correct word rate in an optional combination with respiratory % and / or voiceless / voiced ratio), and to monitor recovery during and after hospitalization (using the correct word rate during hospitalization and either the correct word rate, respiratory % and / or voiced / voiced ratio after discharge).

[0177] Furthermore, biomarkers from the word reading test were compared with corresponding metrics obtained from the digit count and reading tests. These results are shown in Figures 15E to 15J and Figure 18. The characteristics of the data in Figures 15E to 15J are described below.

[0178] Breathing percentage during reading aloud tasks (Figure 15E): HF: Hospitalization vs. OP: Stable: Cohen's d: 1.54, permutation test p-value = 0.0000 HF: Discharge vs. OP: Stable: Cohen's d: 1.28, permutation test p-value = 0.0000 HF: Hospitalization vs. HF: Discharge: Cohen's d: 0.09, permutation test p-value = 0.3810 Silent / voiced ratio in reading tasks (Figure 15F): HF: Hospitalization vs. OP: Stable: Cohen's d: 1.35, permutation test p-value = 0.0000 HF: Discharge vs. OP: Stable: Cohen's d: 0.89, permutation test p-value = 0.0002 HF: Hospitalization vs. HF: Discharge: Cohen's d: -0.03, permutation test p-value = 0.4734 Speech rate (words per second) in reading aloud tasks (Figure 15G): HF: Hospitalization vs. OP: Stable: Cohen's d: -1.60, permutation test p-value = 0.0000 HF: Discharge vs. OP: Stable: Cohen's d: -0.64, permutation test p-value = 0.0190 HF: Hospitalization vs. HF: Discharge: Cohen's d: -0.40, permutation test p-value = 0.0848 Respiration % in the reverse counting task (Figure 15H): HF: Hospitalization vs. OP: Stable: Cohen's d: -0.24, Permutation test p-value = 0.2151 HF: Discharge vs. OP: Stable: Cohen's d: -0.21, permutation test p-value = 0.2537 HF: Hospitalization vs. HF: Discharge: Cohen's d: -0.05, permutation test p-value = 0.4321 Silent / voiced ratio in the reverse counting task (Figure 15I): HF: Hospitalization vs. OP: Stable: Cohen's d: -0.19, Permutation test p-value = 0.2718 HF: Discharge vs. OP: Stable: Cohen's d: -0.26, permutation test p-value = 0.2126 HF: Hospitalization vs. HF: Discharge: Cohen d: 0.04, permutation test p-value = 0.4472 Speech rate in the reverse counting task (Figure 15J): HF: Hospitalization vs. OP: Stable: Cohen's d: 0.19, permutation test p-value = 0.2754 HF: Discharge vs. OP: Stable: Cohen's d: 0.22, permutation test p-value = 0.2349 HF: Hospitalization vs. HF: Discharge: Cohen d: 0.01, permutation test p-value = 0.4797

[0179] The data above indicates that the respiratory percentage, voiceless / voiced ratio, and speech rate in the reading test can each be used to distinguish patients with decompensated heart failure from patients with stable heart failure. However, none of these metrics can be used to distinguish patients with decompensated heart failure at admission from those at discharge. Furthermore, due to the nature of the task, this test cannot be used to obtain metrics equivalent to the correct word rate. Thus, the set of biomarkers derived from the reading test is not as sensitive as the set of biomarkers derived from the word reading test.

[0180] The data further demonstrate that respiratory percentage, voiceless / voiced ratio, and speech rate in the digit counting test cannot be used to distinguish patients with decompensated heart failure from those with stable heart failure. Thus, the set of biomarkers derived from the digit counting test is not as sensitive as the set of biomarkers derived from the word reading test. [Table 2] TIFF0007834763000003.tif176170 [Table 3] TIFF0007834763000005.tif83170

[0181] The data in Figure 16 shows voice pitch estimates from a word-reading test (average of estimates from the color word-reading test in interference and coherent conditions, with error bars representing the standard deviation between the normal and interference conditions) for patients with decompensated heart failure (shown as two points on the left: admission (black) and discharge (dark gray)) and outpatients with stable heart failure (light gray point on the right). The data in Figure 17 shows voice pitch estimates (average of estimates from the color word-reading test in interference and coherent conditions) for patients with decompensated heart failure at various days from admission (registration). The data shows that for most patients with decompensated heart failure, recovery in the hospital is associated with a change in pitch estimates from the word-reading test. However, individual trends may vary among heart failure patients, with some showing an increase in pitch during hospitalization and others showing a decrease. Note that most patients showed a decrease in pitch during recovery. Therefore, voice pitch derived from the word-reading test can be used to monitor recovery during hospitalization for heart failure.

[0182] The data in Figure 18 shows Bland-Altman plots evaluating the agreement between pitch measurements from the digit count test and inverse digit count test for 48 heart failure patients (B, analyzing a total of 161 pairs of records), and the agreement between pitch measurements from the Stroop word reading test (color words, coherent condition) and the Stroop color reading test (color words, interference condition) for 48 heart failure patients (A, analyzing 162 pairs of records). Each data point represents the difference in mean pitch (Hz) estimated using the respective test. The dashed line shows the mean difference (center line) and a standard deviation (SD) interval of ±1.96. Reproducibility is measured in the consensus report (CR=2). *Quantified using SD, the values ​​are 27.76 for the digit count test and 17.64 for the word reading test. A smaller CR value indicates a higher level of reproducibility. Thus, this data indicates that pitch estimates obtained from audio recordings of the word reading test are more reliable (less variable) than pitch estimates obtained from audio recordings of other reading tests, such as the digit count test. The inventors of this invention believe that this is at least in part because the word reading test is less susceptible to influences related to the subject becoming familiar with the sequence of words and / or the cognitive content of the text being read aloud. Furthermore, the words used in this example (color words) conveniently contain a single vowel within the context of the word, and the pitch related to how the same subject pronounces the vowel in the word is less susceptible to external factors than, for example, vowel repetition tests commonly used to assess pitch. In other words, the use of a limited set of words that contains sounds suitable for pitch estimation, but where these sounds reside within a normalized word context and are not accompanied by a biased context of a set of sentences with cognitive content or logical connections (all of which can affect speech pitch and therefore act as confusing factors when pitch is used as a biomarker), conveniently yields more reliable speech biomarkers.

[0183] Similar conclusions apply (to varying degrees) to the metrics of respiration %, speech rate, and voiceless / voiced ratio, and these metrics are more consistent when derived from digit counting versus inverse digit counting tasks (respiration % CR=19.39, N=161; speech rate CR=1.00, N=161; voiceless / voiced CR=0.60, N=161) than when derived from word reading tests versus color word reading tests (i.e., coherent vs. interfering color word reading; respiration % CR=13.06, N=162; speech rate CR=0.50, N=162; voiceless / voiced CR=0.56, N=162).

[0184] Finally, we evaluated the potential of this method for diagnosing or monitoring the status of COVID-19. The biomarker was obtained as described above in a cohort of 10 healthy volunteers and in patients diagnosed with COVID-19. In patients diagnosed with COVID-19, the biomarker was measured over multiple days, including days when the patient had not yet shown any symptoms, and over multiple days during periods when the patient reported only mild fatigue or dyspnea. The results of this analysis are shown in Figure 19. These data indicate that the voice pitch estimates of patients with very mild or no symptoms differed (significantly higher) from those of the healthy volunteer cohort, and that the voice pitch estimates of patients with mild symptoms also differed from those of asymptomatic, recovered patients.

[0185] Thus, the data in Figure 19 suggests that voice pitch biomarkers can be used to identify COVID-19 patients, even if they are asymptomatic, and to monitor disease progression (e.g., recovery).

[0186] References 1.Maor et al.(2018).Vocal Biomarker Is Associated With Hospitalization and Mortality Among Heart Failure Patients.Journal of the American Heart Association.2020;9:e013359. 2.Laguarta et al.(2020).COVID-19 Artificial Intelligence Diagnosis using only Cough Recordings.Open Journal of Engineering in Medicine and Biology.DOI:10.1109 / OJEMB.202.3026928. 3. Mauch and Dixon (2014) 4.Murton et al.(2017).Acoustic speech analysis of patients with decompensated heart failure:A pilot study.J.Acoust.Soc.Am.142(4). 5.Saeed et al.(2018),Study of voice disorders in patients with bronchial asthmas and chronic obstructive pulmonary disease.Egyptian Journal of Bronchology,Vol.12,No.1,pp 20-26. 6.Camacho and Harris(2008).A sawtooth waveform inspired pitch estimator for speech and music.The Journal of the Acoustical Society of America,124(3),pp.1638-1652. 7.Ardaillon and Roebel(2019).Fully-Convolutional Network for Pitch Estimation of Speech Signals.Insterspeech 2019,Sep 2019,Graz,Austria.ff10.21437 / Interspeech.2019-2815ff.ffhal-02439798 8.Kim et al.(2018).CREPE:A Convolutional Representation for Pitch Estimation.2018 IEEE International Conference on Acoustics,Speech and Signal Processing(ICASSP),Calgary,AB,2018,pp.161-165,doi:10.1109 / ICASSP.2018.8461329 9.Kenneth,D.J.,Temporal constraints and characterising syllable structuring.Phonetic Interpretation:Papers in Laboratory Phonology VI.,2003:p.253-268. 10.Xie,Z.M.and P.Niyogi,Robust Acoustic-Based Syllable Detection.Interspeech 2006 and 9th International Conference on Spoken Language Processing,Vols 1-5,2006:p.1571-1574. 11.Wang,D.and S.S.Narayanan,Robust speech rate estimation for spontaneous speech.Ieee Transactions on Audio Speech and Language Processing,2007.15(8):p.2190-2201. 12.Rusz,J.,et al.,Quantitative assessment of motor speech abnormalities in idiopathic rapid eye movement sleep behaviour disorder.Sleep Med,2016.19:p.141-7. 13.Boeck,S.and G.Widmer,Maximum filter vibrato suppression for onset detection.16th International Conference on Digital Audio Effects,Maynooth,Ireland,2013. 14.Davis,S.B.and P.Mermelstein,Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences.Ieee Transactions on Acoustics Speech and Signal Processing,1980.28(4):p.357-366. 15.Huang,X.,A.Acero,and H.Hon,Spoken Language Processing:A guide to theory,algorithm,and system development.Prentice Hall,2001. 16.Rusz,J.,et al.,Automatic Evaluation of Speech Rhythm Instability and Acceleration in Dysarthrias Associated with Basal Ganglia Dysfunction.Front Bioeng Biotechnol,2015.3:p.104. 17.Lloyd,S.P.,Least-Squares Quantization in Pcm.Ieee Transactions on Information Theory,1982.28(2):p.129-137. 18.Smith,T.F.and M.S.Waterman,Identification of common molecular subsequences.J Mol Biol,1981.147(1):p.195-7. 19.Hlavnicka, J., et al., Automated analysis of connected speech reveals early biomarkers of Parkinson's disease in patients with rapid eye movement sleep behavior disorder.Sci Rep, 2017.7(1):p.12. 20.Stroop, JR, Studies of interference in serial verbal reactions.Journal of Experimental Psychology, 1935.General(18):p.19. 21.McFee,B.et al.,librosa:Audio and Music Signal Analysis in Python.PROC.OF THE 14th PYTHON IN SCIENCE CONF.(SCIPY 2015). 22.James Lyons et al.(2020,January 14).jameslyons / python_speech_features:release v0.6.1(Version 0.6.1).Zenodo.http: / / doi.org / 10.5281 / zenodo.3607820

[0187] All documents referenced herein are incorporated herein by reference in their entirety.

[0188] The term "computer system" includes hardware, software, and data storage devices for the embodiment of the system or the execution of the method according to the embodiments described above. For example, a computer system may comprise a central processing unit (CPU), input means, output means, and data storage units, which may be embodied as one or more connected computing devices. Preferably, a computer system comprises a computing device having a display or a display that provides a visual output display (for example, in the design of a business process). The data storage unit may comprise RAM, a disk drive, or other computer-readable media. A computer system may include a plurality of computing devices connected by a network and capable of communicating with each other through that network.

[0189] The methods of the embodiments described above may be provided as a computer program, or as a computer program product or computer-readable medium carrying a computer program configured to perform the above-described method when executed on a computer.

[0190] The term “computer-readable media” includes, but is not limited to, any non-temporary media that can be directly read and accessed by a computer or computer system. Examples of media include, but are not limited to, magnetic storage media such as floppy disks, hard disks, and magnetic tapes; optical storage media such as optical disks or CD-ROMs; electrical storage media such as RAM, ROM, and flash memory; and hybrids and combinations of the above, such as magnetic / optical storage media.

[0191] Unless otherwise indicated by the context, the above-described descriptions and definitions of features apply equally to all aspects and embodiments described herein, and are not limited to any particular aspect or embodiment of the present invention.

[0192] Where used herein, “and / or” should be interpreted as a specific disclosure of each of the two features or components specified therein, which may or may not be present. For example, “A and / or B” should be interpreted as a specific disclosure of (i) A, (ii) B, and (iii) A and B, as if each were described separately herein.

[0193] When used herein and in the appended claims, the singular forms “a” and “an” include plural references unless the context makes otherwise clear. Ranges may be expressed herein as “approximately” from a certain value and / or “approximately” another specific value. Where such ranges are expressed, alternative embodiments include those from and / or those other specific values. Similarly, the use of the antecedent “approximately” will be understood to mean that a particular value forms an alternative embodiment when the value is expressed as an approximation. The term “approximately” with respect to numbers is arbitrary and means, for example, + / - 10%.

[0194] Throughout this Specification, including the following claims, unless the context requires otherwise, the terms “comprise” and “include,” as well as variations such as “comprises,” “comprising,” and “including,” shall be understood to mean that they include the thing or process or group of things or processes described herein, but not to exclude any other thing or process or group of things or processes.

[0195] Other aspects and embodiments of the present invention provide the above aspects and embodiments in which the term "comprising" is replaced with the term "consisting of" or "consisting essentially of" unless it is clear from the context otherwise.

[0196] Features disclosed in the above description, or in the following claims, or in the accompanying drawings, and expressed in concrete form, or with respect to means for performing the disclosed functions or methods or processes for obtaining the disclosed results, can be used, individually or in any combination of such features, as necessary to realize the present invention in a variety of forms.

[0197] While the present invention has been described in conjunction with the exemplary embodiments described above, numerous equivalent modifications and variations will be apparent to those skilled in the art in view of this disclosure. Therefore, the exemplary embodiments of the present invention described above are illustrative and not limiting. Various modifications can be made to the described embodiments without departing from the spirit and scope of the invention.

[0198] To avoid misunderstanding, any theoretical explanations provided herein are provided for the purpose of improving the reader's understanding. The inventors of the present invention do not wish to be bound by any of these theoretical explanations.

[0199] Any headings used in this specification are for structural purposes only and should not be construed as limiting the subjects described.

Claims

1. A computer implementation method for evaluating the pathological and / or physiological state of a subject, Obtain an audio recording from a word reading test, which includes the reading aloud of a sequence of words randomly or pseudo-randomly selected from a set of n words, and which are not connected to form a sentence. Analyzing the aforementioned audio recording or a portion thereof, Identify multiple segments of the audio recording corresponding to individual words or syllables, Based at least partially on the identified segment, a value is determined for one or more metrics selected from the respiratory percentage, voiceless / voiced ratio, speech pitch, and word accuracy. The value of one or more of the aforementioned metrics is compared with one or more respective reference values. By doing so, the audio recording or a portion of the audio recording is analyzed. Computer implementation methods, including those mentioned above.

2. Identifying the segment of the speech recording corresponding to individual words or syllables is: To obtain a power meter spectrogram of the aforementioned audio recording, Calculating the maximum intensity projection of the power mel spectrogram along the frequency axis, The segment boundary is defined as the point in time when the maximum intensity projection of the power mel spectrogram along the frequency axis intersects the threshold. The computer implementation method according to claim 1, including the method described in claim 1.

3. Identifying the segment of the speech recording corresponding to an individual word or syllable is: (i) Normalizing the PowerMel spectrogram of the audio recording, and / or (ii) Performing onset detection for at least one of the segments by calculating the spectral flux function for the power mel spectrogram of the segment, and each time an onset is detected within the segment, further boundaries are defined to form two new segments, and / or (iii) Calculate one or more Mel-frequency cepstrum coefficients for the segment to obtain multiple vectors of values ​​in which each vector relates to one segment, and exclude segments that represent false detections by applying an outlier detection method to the multiple vectors of values, and / or (iv) Excluding segments representing false positives by removing segments shorter than a predetermined threshold and / or segments whose average relative energy falls below a predetermined threshold. The computer implementation method according to claim 2, further comprising:

4. The computer implementation method according to claim 3, wherein identifying the segments of the speech recording corresponding to individual words or syllables includes normalizing the power mel spectrogram of the speech recording to the frame having the highest energy in the recording.

5. The computer implementation method according to any one of claims 1 to 4, wherein determining the value of one or more metrics includes determining the breathing percentage with respect to the recording as a percentage of time between the identified segments in the audio recording, or as a ratio of the time between the identified segments in the recording to the sum of the time between the identified segments and the time within the identified segments in the recording.

6. The computer implementation method according to any one of claims 1 to 5, wherein determining the value of one or more metrics includes determining the silent / voiced ratio for the recording as the ratio of the time between the identified segments in the recording to the time within the identified segments in the recording.

7. The computer implementation method according to any one of claims 1 to 6, wherein determining the value of one or more metrics includes determining the audio pitch relating to the recording by obtaining one or more estimates of fundamental frequencies for each of the identified segments.

8. The computer implementation method according to claim 7, wherein determining the value of the voice pitch includes obtaining a plurality of estimates of the fundamental frequency for each of the identified segments, applying a filter to the plurality of estimates to obtain a plurality of filtered estimates, and / or determining the value of the voice pitch includes obtaining summarized voice pitch estimates for a plurality of the plurality of segments.

9. The computer implementation method according to claim 8, wherein the summarized speech pitch estimates for the plurality of segments are the mean, median, or mode of the plurality of estimates for the plurality of segments, or the filtered plurality of estimates for the plurality of segments.

10. A computer implementation method according to any one of claims 1 to 9, wherein determining the value of one or more metrics includes determining the word accuracy with respect to the audio recording by calculating a ratio of the number of identified segments corresponding to correctly read words divided by the time between the start of the first identified segment and the end of the last identified segment, or by calculating a cumulative sum over time of the number of identified segments corresponding to correctly read words in the audio recording and calculating the slope of a linear regression model fitted to the data of the cumulative sum.

11. Determining the value of one or more of the aforementioned metrics includes determining the word accuracy rate for the aforementioned record, Determining the correct answer rate for the aforementioned word means For each of the identified segments, one or more Mel-frequency cepstrum coefficients (MFCCs) are calculated to obtain multiple vectors of values ​​in which each vector relates to one segment. The process involves clustering the multiple vectors of the aforementioned values ​​into n clusters, each cluster having n possible labels corresponding to each of the n words, For each of the n! permutations of the labels, predict the sequence of words in the audio recording using the labels relating to the clustered vector of the values, and perform sequence alignment between the predicted sequence of words and the sequence of words used in the word reading test. The best alignment is the selection of a label that results in the best alignment, where the match in the alignment corresponds to the correctly pronounced word in the audio recording. A computer implementation method according to any one of claims 1 to 10, including the method described in any one of claims 1 to 10.

12. The computer implementation method according to claim 11, wherein calculating one or more MFCCs to obtain a vector of values ​​for a segment comprises calculating a set of i MFCCs for each frame of the segment with respect to each i, obtaining a set of j values ​​for the segment by interpolation or linear interpolation, and obtaining a vector of ixj values ​​for the segment, and / or clustering the multiple vectors of values ​​into n clusters, performed using k-means, and / or the sequence alignment step is performed using a local sequence alignment algorithm or a Smith-Waterman algorithm, and / or performing sequence alignment comprises obtaining an alignment score, the best alignment is the alignment with the highest alignment score.

13. The aforementioned n words are, (i) It is one syllable or two syllables and / or (ii) Each contains one or more vowels within the word and / or (iii) Each contains a single stressed syllable and / or (iv) A color word, the computer implementation method according to any one of claims 1 to 12.

14. The computer implementation method according to claim 13, wherein the word is a color word displayed in a single color in the word reading test, or the word is a color word displayed in a color independently selected from a set of m colors in the word reading test.

15. (i) Obtaining an audio recording from the subject from a word reading test comprises obtaining an audio recording from a first word reading test and an audio recording from a second word reading test, wherein the word reading test comprises reading aloud a sequence of words taken from a set of n words which are color words, the words being displayed in a single color in the first word reading test and in a color independently selected from a set of m colors in the second word reading test, and / or (ii) The computer implementation method according to any one of claims 1 to 14, wherein obtaining an audio recording from the subject from a word reading test includes receiving the word recording from a computing device associated with the subject.

16. The computer implementation method according to claim 15, wherein obtaining an audio recording from a subject of a word reading test comprises receiving the word recording from a computing device associated with the subject, and obtaining the audio recording further comprises causing the computing device associated with the subject to display the sequence of words and / or record an audio recording and / or emit a fixed-length tone before recording the audio recording.

17. The computer implementation method according to any one of claims 1 to 16, wherein the sequence of words includes a predetermined number of words, or the sequence of words includes a predetermined number of words, which is at least 20, at least 30, or about 40.

18. A computer implementation method for monitoring a subject with heart failure, or for diagnosing a subject with worsening heart failure or decompensated heart failure, A computer implementation method comprising evaluating the pathological and / or physiological state of a subject using a computer implementation method according to any one of claims 1 to 17, wherein each of the one or more reference values ​​is a value of the same metric previously obtained for the same subject, or each of the one or more reference values ​​includes a value of the same metric associated with a patient with decompensated heart failure, a patient with stable heart failure, and / or a patient with decompensated heart failure in recovery.

19. A computer implementation method for monitoring a subject who is diagnosed with or at risk of being in a condition related to dyspnea and / or fatigue, or for evaluating the level of dyspnea and / or fatigue in a subject, A computer implementation method comprising evaluating the pathological and / or physiological state of a subject using a computer implementation method according to any one of claims 1 to 17, wherein each of the one or more reference values ​​is a value of the same metric previously obtained for the same subject, or each of the one or more reference values ​​includes the same metric value associated with a patient having the condition and / or the same metric value associated with a patient not having the condition.

20. It is a system, At least one processor, At least one non-temporary computer-readable medium containing instructions and Includes, A system wherein, when the instruction is executed by the at least one processor, the at least one processor causes the at least one processor to perform an operation including the operation described in any one of claims 1 to 19.

Citation Information

Patent Citations

  • Fatigue level recognition method and device, computer equipment, and storage medium

    CN109119095A

  • Evaluation of pulmonary disease by voice analysis

    JP2018534026A

  • Regular verbal screening for heart disease

    JP2020507437A

  • Systems and methods for determining physiological conditions

    JP2021500209A

  • Diagnostic techniques based on speech-sample alignment

    US20200294531A1