Method for confidence estimation of speech-processing applications

A multi-criteria speech-quality assessment system improves the reliability of speech-to-respiration systems by quantifying confidence levels, addressing accuracy issues in casual speech environments.

WO2025209894A1PCT designated stage Publication Date: 2025-10-09KONINKLIJKE PHILIPS NV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/058215
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-02
Filing Date
2025-03-26
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing speech-to-respiration systems face accuracy issues due to variations in speech quality, particularly in casual conversations, as they are typically trained on controlled environments, leading to unreliable predictions.

Method used

Implement a comprehensive speech-quality assessment system that includes multiple criteria such as SNR, human perception, ASR quality, and OOD assessments to determine confidence levels in speech processing results, providing feedback for improving speech quality and reliability.

Benefits of technology

Enhances the accuracy and reliability of speech-to-respiration systems by quantifying confidence levels, allowing users to adjust speech quality for better results and ensuring reliable respiratory parameter estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025058215_09102025_PF_FP_ABST
    Figure EP2025058215_09102025_PF_FP_ABST
Patent Text Reader

Abstract

In a speech-processing system, the quality of the input speech affects the accuracy and / or reliability of the results. In a speech-to-respiration system, wherein a subject's speech is used to determine / estimate measures of the subject's respiration process, for example, the quality of speech directly impacts the accuracy of the estimated respiratory signal. A combination of different technologies is used to assess the quality of the subject's speech, including for example, signal analysis, audio analysis, automated speech-recognition (ASR), and others. Based on these speechquality assessments, a level of confidence in the speech-processing result is determined and provided to the subject. Neural networks are used in the determination of the quality of the input speech.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR CONFIDENCE ESTIMATION OF SPEECH-PROCESSING

[0002] APPLICATIONS

[0003] FIELD OF THE INVENTION

[0004] This invention relates to the field of medical instruments, and in particular to a method and system for analyzing speech to provide a signal representing pulmonary activity and a quality metric associated with this respiration signal.

[0005] BACKGROUND OF THE INVENTION

[0006] In recent years, techniques for assessing a subject's pulmonary state or condition, or other physiological parameters, based on an analysis of the subject's speech ("speech-to-respiration") have advanced significantly. Such techniques have become of increasing importance in the field of telehealth, particularly in view of Covid-19 and other contagious illnesses, wherein the subject has a "remote visit" with a medical practitioner. While speaking with the subject, the practitioner may enable a speech-to-respiration application and receive a real-time assessment of the subject's respiration process.

[0007] Speech production involves a systematic outflow of air during exhalation characterized by linguistic content and prosodic factors of the utterance. Accordingly, there is a strong correlation between speech and respiration, and techniques have been developed to model this relationship and sense respiratory dynamics directly from the speech. These techniques generally apply "deep learning" to sense a breathing signal and breathing parameters from the speech. Estimating the breathing pattern from the speech provides information about the subject's respiratory parameters, thus enabling an understanding of the subject's respiratory health based on the subject's speech.

[0008] As noted above, the analysis of a subject's respiration based on the subject's speech during a telehealth session enables a medical practioner to identify respiratory ailments, such as a respiratory infection, Chronic Obstructive Pulmonay Disease (COPD), asthma, and others. In like manner, monitoring the subject's speech over multiple telehealth sessions enables the practioner to assess whether the subject's respiration may be improving or degrading over time.

[0009] In addition to telehealth applications, speech-to-respiration techniques may be used in other applications. For example, as disclosed in U.S. patent application 17 / 071 ,312 (USPA 2021 / 0146082) by Aki Sakari Harma, Francesco Vicario, and Venkata Srikanth Nallanthighal, and incorporated by reference herein, a speech-to- respiration process may be used during the delivery of gas to a patient by predicting the inhalation cycle of the patient based on the patient's speech patterns.

[0010] However, as with any process that attempts to predict unknown values of a first variable based on a model of the correlation between values of the first variable and observed values of a second variable, the accuracy of the predicted values of the first variable will be based on the quality / accuracy of the observed values of the second variable.

[0011] Additionally, the models used in speech-to-respiration systems are generally trained in a controlled environment, such that the subject will tend to speak clearly and definitely. However, in casual conversation, the same person's speech may not be reflective of the controlled speech that was used to train the speech-to-respiration model.

[0012] SUMMARY OF THE INVENTION

[0013] It would be advantageous to provide a speech processing system, such as a speech-to-respiration system, that includes a comprehensive speech-quality assessment. Such an assessment would increase the likelihood of rejecting poor speech processing results, and / or provide notice that the speech-processing process should be repeated. The speech-quality assessment may also provide feedback with regard to speech changes that would improve the speech-quality and resultant improvement in the speech processing results. Additionally, providing a comprehensive speech-quality assessment will serve to increase the user's confidence in the resultant speech processing results.

[0014] To better address one or more of these concerns, in embodiments of this invention, a plurality of speech-quality assessments are performed, using quality criteria from a variety of fields of speech processing and analysis. For example, in speech-to-respiration systems, in addition to criteria commonly used to assess speech-to-respiration quality, such as pause durations, count of phonemes, rate of delivered words, etc., a quality measure may be based on the electrical characteristics of the audio signal, such as a Signal-to-Noise ratio (SNR). Another quality measure may be based on an assessment of factors related to the human perception of speech or sounds. In like manner, from the field of automated speech recognition (ASR), a quality measure may be based on a word-error rate. Another quality measure may be based on a comparison between the spectral characteristics of the current speech and the spectral characteristics of the speech used to train the speech-to-respiration model to identify "Out-of-Distribution" (OOD) words or phrases.

[0015] Embodiments of this invention include a system comprising: an audio receiver that is configured to provide a speech signal; a speech processing circuit that is configured to receive the speech signal and to provide a speech-dependent result; a speech-quality circuit that is configured to receive the speech signal and to provide a measure of confidence in the speech-dependent result; and a user interface circuit that provides the speech-dependent result and the measure of confidence to a user; wherein the speech-quality circuit is configured to determine the measure of confidence based on a plurality of speech-quality assessments from among the following; an SNR (Signal-to-Noise Ratio) quality assessment; a human perception quality assessment; an ASR (Automated Speech Recognition) quality assessment; a speech-to-respiration quality assessment; and an OOD (Out-of-Distribution) assessment.

[0016] Preferably, at least three of the above speech-quality assessments are performed to provide the measure of confidence to the user.

[0017] Further features of the invention include assessing each of a pluality of segments of the speech signal; providing a binary result for each of the plurality of speech-quality assessments for each segment; determining a Speech Quality Index (SOI) for each segment based on the binary results; and determining the measure of confidence based on the SOI of each segment. In embodiments of the invention, the measure of confidence may be quantized into one of a plurality of quantization levels, and the user interface circuit is configured to display the measure of confidence using a distinct display characteristic for each quantization level.

[0018] Embodiments of this invention include a method comprising: receiving a speech signal; processing the speech signal to provide speech-dependent result; processing the speech signal to provide a measure of confidence in the respiration signal; and providing the speech-dependent result and the measure of confidence to a user; wherein the measure of confidence is based on a plurality of speechquality assessments from among the following; an SNR (Signal-to-Noise Ratio) quality assessment; a human perception quality assessment; an ASR (Automated Speech Recognition) quality assessment; a speech-to-respiration quality assessment; and an OOD (Out-of-Distribution) assessment.

[0019] Preferably, at least three of the above speech-quality assessments are performed to provide the measure of confidence to the user.

[0020] Embodiments of the invention include a computer program that, when executed by a processor, causes the processor to perform the above method.

[0021] BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The invention is explained in further detail, and by way of example, in embodiments that include a speech-to-respiration system with reference to the accompanying drawings wherein:

[0023] FIGs. 1 A-1 B illustrate an example embodiment of this invention, and an example display output.

[0024] FIGs. 2A-2C illustrate example configurations of a speech-to-respiration circuit, and an example comparison of a predicted respiration signal and the actual respiration signal. FIG. 3 illustrates an example flow diagram for assessing a measure of confidence in a respiration result based on a plurality of quality assessments of the input speech.

[0025] Throughout the drawings, the same reference numerals indicate similar or corresponding features or functions. The drawings are included for illustrative purposes and are not intended to limit the scope of the invention.

[0026] DETAILED DESCRIPTION

[0027] In the following description, for purposes of explanation rather than limitation, specific details are set forth such as the particular architecture, interfaces, techniques, etc., in order to provide a thorough understanding of the concepts of the invention. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments, which depart from these specific details. In like manner, the text of this description is directed to the example embodiments as illustrated in the Figures, and is not intended to limit the claimed invention beyond the limits expressly included in the claims. For purposes of simplicity and clarity, detailed descriptions of well-known devices, circuits, and methods are omitted so as not to obscure the description of the present invention with unnecessary detail.

[0028] As used herein, the term "circuit" refers to a configuration of electrical elements that perform a given function or set of functions. The circuit may include discrete elements or integrated elements, and may include, for example, a processor configured to execute a program to perform the function(s).

[0029] For illustrative purposes and ease of understanding, the invention is described using an example embodiment of a speech-processing system comprising a speech-to-respiration system, wherein the speech-dependent result is a respiration signal and / or associated parameters. One of skill in the art will recognize that the application of quality assessments from multiple fields of speech processing may be applied to applications in each of these particular fields, as well as other speech processing applications.

[0030] The term "quality" has a variety of definitions; the Merriam-Webster (M-W) dictionary defines eight distinctly different definitions of "quality". US Patent Application 2016 / 0379638, "INPUT SPEECH QUALITY MATCHING", filed 26 June 2015 for Basye et al., uses the term "quality" in the context of "a distinguishing attribute or characteristic" (M-W definition 4a), including "the manner in which the utterance was spoken" ('638 Abstract), such as quiet, loud, energetic, etc. As used herein, the term "quality" is used in the context of "degree of excellence: grade" (M- W definition 2a), as in "poor quality", "high quality", "sufficient quality", etc.

[0031] FIG. 1A illustrates an example embodiment of this invention.

[0032] An audio receiver 120 receives an audio signal comprising speech. This speech may come from any of a variety of sources 110, such as via a telephone 110i over a telephony network, a local microphone 1 W2, a smartphone 1 3 via an internet connection, etc. The receiver 120 may include circuitry that converts radiofrequency signals into audio-frequency signals. As used herein, the terms audio signal and speech signal are used synonymously, and comprise an electronic signal that corresponds to the input speech.

[0033] The speech signal is provided to a speech-to-respiration circuit 130, and a speech-quality circuit 140. As detailed further below, the speech-to-respiration circuit 130 is configured to estimate respiration parameters, such as breaths per minute, breath volume, etc. based on the speech signal. This speech-to-respiration circuit may include a deep-learning neural network that is trained and tested to provide accurate estimates of these respiration parameters.

[0034] Although the following descriptions address the training and use of a single deep-learning neural network model, one of skill in the art will recognize that the speech-to-respiration circuit 130 may include a plurality of models, each trained using a different class of speakers, such as models based on gender, age, or other parameters found to influence the speech-to-respiration process. These parameters would be input to the speech-to-respiration circuit 130 at the start of the speech-to- respiration process to enable the speech-to-respiration circuit to select and apply the appropriate model.

[0035] As also detailed below, the speech-quality circuit 140 is configured to conduct a variety of assessments of the speech signal corresponding to the input speech. These assessments may include conventional signal processing parameters, such as signal-to-noise ratio (SNR), and conventional audio processing parameters, such as audibility. These assessments may also include assessments in fields outside of conventional signal quality assessments, such as speech recognition, as well as other assessments specifically related to the speech-to-respiration process. The output of the speech-to-respiration circuit 130 and the speech-quality circuit 140 provides respiration estimates with an associated confidence in these estimates 150, based on the assessed quality of the input speech. This output 150 may be displayed on a user-interface 160, or subsequently processed via a respiration analysis circuit 170. The respiration analysis circuit 170 may be configured to provide diagnostic information based on the respiration parameters. Preferably, any diagnostic information will be provided along with a confidence factor based on the quality of the input speech, to caution the user when the input speech is assessed to be potentially unreliable for estimating the respiration parameters that were used to determine the diagnostic information.

[0036] FIG. 1 B illustrates an example display of the estimated respiration waveform 166 based on the input speech signal 162. Also displayed is the determined quality measure 164 based on the plurality of quality assessments. As detailed below, this quality measure 164 is determined as a "Speech Quality Index" (SQI) that is determined for discrete segments of the input speech signal, and displayed as a continuous waveform. A composite parameter, such as the estimated breaths per minute (BPM) 167 may also be determined and displayed continuously.

[0037] Of particular note, in preferred embodiments, the overall quality of the speech input is quantized into one of a number of confidence levels, and a visual indication 168 of the overall confidence level is also displayed. This visual indication 168 may comprise a choice of three colors: green, yellow, and red; wherein green is displayed when the overall SQI is high, red is displayed when the overall SQI is low, and yellow is displayed when the overall SQI is between a low threshold and a high threshold. In general, the user may be advised to disregard the results when the indication 168 is red. Guidance (not illustrated) may be provided on the display based on the assessments for improving the speech quality.

[0038] As noted above, if a respiration analysis 170 is performed, the results of this analysis may also be displayed, along with a confidence estimate based on the speech quality assessments.

[0039] FIGs. 2A-2C illustrate example configurations of a speech-to-respiration circuit, and an example comparison of a predicted respiration signal and the actual respiration signal. FIG. 2A illustrates an example configuration for training a deep neural network 270 that may be used to model the correlation between input speech 225 and a respiration signal 235 while the speech is being input. Typically, dozens of subjects, or more, are used to train the neural network 270.

[0040] Respiration measurement belts 215 are strapped to a subject 210, and a transducer 230 converts the expansions and contractions into a respiration signal 235. The transducer 230 may also perform signal processing to provide a suitable respiration signal 235. In some embodiments, the transducer 230 may include the belts 215 that provide an electrical signal directly in response to being stretched.

[0041] One of skill in the art will recognize that any of a variety of devices may be used to monitor the subject's respiration 235 while the subject is instructed to speak into a microphone 220 to provide a speech signal 225. In some embodiments, the subject 210 is requested to read a standard passage, such as "The Rainbow Passage" from the book "Voice and Articulation" by G. Fairbanks. The subject may also be engaged in conversation with the person conducting this training. Other sets of speech and respiration signals may be used to train the network 270, such as recorded speech-respiration datasets that are provided to developers of speech processing applications. In some embodiments, the network 270 may be "pretrained" using such recorded datasets, then "fine-tuned" for the particular subject 210 or select groups of subjects.

[0042] A trainer circuit 240 receives the speech signal 225 and corresponding respiration signal 235. Although the "raw" speech signal 225 may be used directly as an input to network 270, in some embodiments, an input processor 250 pre- processes the speech signal 225 to provide a set of inputs ("features") to the network 270. Such pre-processing may include spectral analysis, Mel-frequency scaling, normalization, etc., such as disclosed in "Deep learning architectures for estimating breathing signal and respiratory parameters from speech recordings", by Venkata Srikanth Nallanthighal, Zohreh Mostaani, Aki Harma, Helmer Strik, and Mathew Magimai-Doss; Neural Networks 141 (2021 ) pgs. 211-224, which is incorporated by reference herein. In some embodiments both the raw speech signal and the pre-processed signals are input to the network 270. As used further herein, the term "speech signal" refers to raw speech signal, the pre-processed speech signal, or a combination of both, depending upon the particular embodiment of the input processor 250 and the deep neural network 270. The deep neural network 270 is designed to provide an estimate of a respiration signal corresponding to the input speech 225. The network 270 may comprise a deep recurrent neural network or other sequential regression algorithms, such as disclosed in the aforementioned U.S. Patent Application 17 / 071 ,312 (USPA 2021 / 0146082). Although illustrated as a single network 270, the network 270 may comprise a plurality of neural networks and an overall architecture that "fuses" the outputs of these neural networks to provide a potentially more accurate respiration signal.

[0043] In embodiments, the input 242 to the network 270 may be discrete segments of the input speech 225, such as segments of a few seconds each. In some embodiments, to minimize transition anomalies, the segments may overlap, such that the beginning of a segment overlaps with the ending of a prior segment, and the ending of the segment overlaps with the beginning of a next segment, as described in EP23155738, "Speech Processing of Audio Signal" by Aki Harma, filed 9 February 2023, which is incorporated by reference herein.

[0044] Each of these segments may be pre-processed as detailed above. In a training mode, the network 270 receives the speech input 242 and provides an estimate of the respiration signal 272 to the trainer 240. The trainer 240 includes a feedback circuit 260 that compares the predicted respiration signal 272 with the known respiration signal 235, and provides feedback 244 to the network 270 based on the correspondence, or lack of correspondence, between the predicted 272 and known 235 respiration signals.

[0045] Typically, the training of the network 270 continues for each subject until the feedback circuit 260 indicates that a sufficient level of correspondence between the predicted 272 and known 235 respiration signals has been achieved.

[0046] Subsequent tests may be performed using variations of speech signals 225 in a "test mode", wherein the feedback circuit 260 provides a measure of the correspondence between the predicted 272 and known 235 respiration signals, but the trainer circuit 240 does not provide the feedback signal 244 to the network 270. That is, the network 270 remains unchanged throughout the test period.

[0047] FIG. 2B illustrates an example configuration of the speech-to-respiration system in normal operation, after the training and testing of the network 270. In this configuration, the network 270 remains unchanged from its trained and tested state. The subject 210 speaks and produces a (new) speech signal 225'. This signal 225' is processed via the input processor 250, which performs the same processing as performed during the training process. Similarly, the same segmentation of the speech signal 225 applied during training is applied to the current speech signal 225'. The processed input speech 252 is provided to the network 270, which applies the input speech 252 to the trained model of the network 270. In response to the input speech, the network 270 produces an estimated respiration signal 235' corresponding to the input speech signal 225'.

[0048] FIG. 2C illustrates an example comparison of an estimated respiration signal 235' and an actual ("ground truth") respiration signal 235 that was measured while the speech signal 225' was input to the speech-to-respiration system. As can be seen, the correlation between the estimated respiration and the actual respiration is significant.

[0049] FIG. 3 illustrates an example flow diagram for assessing a measure of confidence in a respiration result based on a plurality of quality assessments of the input speech, such as provided by the speech quality tests circuit 140 of FIG. 1 .

[0050] A speech detector circuit 315 receives an input audio signal 310 and determines whether speech is present, for example by detecting frequencies in the human vocal range (e.g. 50-500Hz), inhalation pauses, and other speech characteristics. This speech detector 315 may be common to both the speech-to- respiration circuit 130 and the speech quality tests circuit 140 of FIG. 1 . If speech is detected, the speech signal is communicated to each of a plurality of speech quality assessment circuits 320-345. One of skill in the art will recognize that fewer or more assessment circuits may be used, each assessment circuit assessing the quality of the speech signal based on different evaluation measures.

[0051] In embodiments, each assessment circuit may comprise an artificial intelligence (Al) model, such as a neural network, that is trained to provide one or more quality metrics based on input speech. A single Al model may be used to provide the quality metrics for multiple assessment circuits 320-345, and may also include the threshold circuit 350 as well as the speech-quality index (SQI) circuit 360 and the overall confidence circuit 370.

[0052] In embodiments where there are multiple speakers, a source-separation circuit may be used to p re-process the 'mixed' speech input to provide an input audio signal 310 specific to each speaker. Additional assessments may be applied to also assess the quality of the source-separation results.

[0053] As noted above, the input speech may be segmented for the speech-to- respiration process. This same segregation may be applied to the input speech 310 in the speech detector 315. Optionally, the unsegmented speech 310 may be provided to some or all of the assessment circuits, and each of these assessment circuits may apply a segmentation that is most appropriate for the given assessment.

[0054] A signal-quality assessment circuit 320 determines one or more electrical characteristics of the speech signal. These electrical characteristics are generally common to all electrical signals, and not specific to speech signals per se, and may include such parameters as noise, jitter, distortion, and others characteristics common in the field of signal processing. In preferred embodiments, the assessment will include at least a measure of the Signal to Noise Ratio (SNR).

[0055] An audio quality assessment circuit 325 determines characteristics specific to audio characteristics of the speech signal. These characteristics may be based on the audibility of the speech based on human perception factors. Using tools such as openSMILE ("open-source Speech & Music Interpretation by Large-space Extraction"), the speech signal can be analyzed to determine audio characteristics based on the human perception of speech and sound, such as loudness, psychoacoustic sharpness, spectral harmonicity, cepstrum coefficients, and so on. Additionally, "quality models" have been developed based on human assessments of speech quality to automatically assess the quality of a codec output after lossy audio compression, and may be used to provide a MUSHRA (Multiple Stimuli with Hidden Reference and Anchor)-like Mean Opinion Score for the input speech.

[0056] The quality of the speech signal is also assessed with regard to how well the words in the speech can be recognized by a speech-recognition system, at 330. Any of a variety of automated speech-recognition (ASR) applications may be used, such as Carnegie Mellon University's "Sphinx" toolkit, Mozilla's "Common Voice", Google's "Gboard", and others. The subject is prompted to read scripts in the language being modeled. In some embodiments, specific protocols are followed for recording speech data for a holistic collection of relevant linguistic content required to estimate respiratory parameters accurately (e.g. the aforementioned rainbow passage in English). The known texts from these protocols are used as ground truth to obtain the ASR metrics, such as word-error-rate and semantic distance.

[0057] In embodiments, the speech-recognition quality assessment may include a neural network that is trained by having humans assign a quality level to various speech inputs, based on perceptual assessments of intelligibility and other characteristics.

[0058] At 335, the speech is analyzed with regard to parameters specific to the speech-to-respiration process, such as pause durations, count of phonemes, rate of delivered words, and so on. As noted above, tools such as openSMILE may be used to obtain other measures that may be related to characteristics that affect the speech-to-respiration process, such as zero-crossing rate, jitter, shimmer, spectral features, and others.

[0059] Additionally, as noted above, the input speech may be processed as overlapping segments as disclosed in EP23155738. It has been found that if the input speech is of high quality, the overlapping segments will be highly correlated. In embodiments of this invention, the quality of the input speech may include the "distance" (Euclidean or cosine) between each of the overlapping segments.

[0060] At 340, the speech is analyzed with regard to how well the input speech corresponds to the speech that was used to train the neural network 270. Tools such as openSMILE may be used to characterize the training speech to determine such measures such as Mel / Bark Frequency Cepstral Coefficients, Formont frequencies and bandwidths, and other spectral features. The current input speech is similarly characterized and statistical tests, such as the Kolmogorov-Smirnov test, may be used to determine whether the input speech differs substantially from the training speech ("Out-of-Distribution").

[0061] Other quality assessment circuits may be provided, at 345. For example, in embodiments that use a Bayesian neural network model comprising a stochastic deep learning architecture that is trained using Bayesian methods, the model itself may provide a confidence level in its output, and this confidence level may be included in the plurality of quality assessments.

[0062] The results of each of the assessment circuits are used to determine a speech-quality index (SQI) for each of the input speech segments, and for determining an overall speech-quality measure. As noted above, the SQI and overall speech quality may be determined using an Al model that is trained by assessing the correspondence between the estimated respiration and the known respiration during the test phase of the speech-respiration model for each set of test speech samples, wherein the input to the Al model are the results of the quality assessments of each set of test speech samples.

[0063] Alternatively, in some embodiments, a more deterministic approach may be taken, such as illustrated in FIG. 3. In FIG. 3, the individual assessments are compared to threshold values, at 350, and each comparison results in a binary determination (pass / fail, 0 / 1). These binary determinations (0 / 1 ) for each segment of the input speech are averaged to provide the speech-quality index (SQI) for the segment, which is a value between 0 and 1 , as illustrated in the graph 164 of FIG. 3.

[0064] At 370, the overall confidence in the accuracy of the speech-to-respiration process is determined by averaging the SQIs of all of the input speech segments, and quantizing the result into one of a set of confidence levels, such a "high", "medium", and "low" confidence, and displayed at 168 of FIG. 3. As noted above, this overall confidence result may be used to caution the user with regard to reliance on the respiration result, to provide guidance for improving the input speech quality to improve the speech-to-respiration process, and so on.

[0065] While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive; the invention is not limited to the disclosed embodiments.

[0066] For example, one of skill in the art will recognize that any of a variety of techniques known in the art may be used to determine an aggregate value from a set of individual values. For example, rather than defining threshold values for each assessment value, the assessment values may be normalized to a fixed range of values, such as 0 to 100, and these values may be averaged to determine the SQI of each segment. In some embodiments, this averaging may be weighted if it is found that some assessment results are more reliable indicators of the likely accuracy of the speech-to-respiration results. In some embodiments, a mix of binary and nonbinary values may be used, wherein, in the example range of 0 to 100, a binary "pass" would have a value of 100, and a "fail" would have a value of 0. Such a mix may be warranted if some measures, such as the word-error rate from the speech- recognition assessment 330, lend themselves to a non-threshold determination, while other measures lend themselves to threshold tests.

[0067] As noted above, although the invention is disclosed in the context of a speech-to-respiration system, the principles and features of this invention may be applied to assess the speech quality in various other speech-processing systems. For example, in speech processing applications outside the field of speech-to- respiration system, such as speaker-recognition, the quality assessment for these applications may include quality assessment parameters specific to the speech-to- respiration process, such as pause durations, count of phonemes, rate of delivered words, and so on. In this manner, combined with quality assessments from other technologies, the confidence of a correct identification of the speaker can be provided to a user of the speaker-recognition system.

[0068] Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. Also in the claims, the expression "at least one of A, B, and C" means "at least one of A, B, and / or C". A single processor or other unit may fulfill the functions of several items recited in the claims; in like manner, multiple processors may be used in lieu of a single processor. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. A computer program may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. Any reference signs in the claims should not be construed as limiting the scope.

Claims

1 . A system comprising: an audio receiver that is configured to provide a speech signal; a speech processing circuit that is configured to receive the speech signal and to provide a speech-dependent output; a speech-quality circuit that is configured to receive the speech signal and to provide a measure of confidence in the speech-dependent output; and a user interface circuit that provides the speech-dependent output and the measure of confidence to a user; wherein the speech-quality circuit is configured to determine the measure of confidence based on a plurality of speech-quality assessments; wherein the plurality of speech-quality assessments comprises at least three of the following: an SNR (Signal-to-Noise Ratio) quality assessment; a human perception quality assessment; an ASR (Automated Speech Recognition) quality assessment; a speech-to-respiration quality assessment; and an OOD (Out-of-Distribution) assessment.

2. The system of claim 1 , wherein the speech processing circuit comprises a speech- to-respiration system.

3. The system of claim 1 , wherein the ASR quality assessment is based on word- error-rate and semantic distance.

4. The system of claim 1 , wherein the speech-quality assessment is based on at least one of: pause durations; count of phonemes; and rate of words delivered.

5. The system of claim 1 , wherein the speech processing circuit is trained using a set of training data, and the OOD assessment is based on a correspondence between the speech signal and the training data.

6. The system of claim 1 , wherein each of the plurality of speech-quality assessments comprises a binary result, and the speech-quality circuit is configured to: determine a Speech Quality Index (SQI) for each segment of a plurality of segments of the speech signal based on the binary results; and determine the measure of confidence based on the SQI of each segment.

7. The system of claim 6, wherein the measure of confidence is quantized into one of a plurality of quantization levels, and wherein the user interface circuit is configured to display the measure of confidence using a distinct display characteristic for each quantization level.

8. The system of claim 1 , wherein one or more of the plurality of speech-quality assessments comprise a neural network.

9. The system of claim 1 , wherein the speech processing circuit comprises a deep neural network.

10. The system of claim 10, wherein the deep neural network comprises a Bayesian neural network model, and the measure of confidence is further based on a confidence level provided by the Bayesian neural network model.11 . A speech-quality circuit that is configured to receive a speech signal and to provide a measure of confidence in one or more qualities of the speech signal, comprising: a plurality of speech-quality assessment circuits, and an aggregate circuit that is configured to combine outputs of the plurality of speech-quality assessments to determine the measure of confidence; and a user interface that is configured to provide the measure of confidence to a user;wherein the plurality of speech-quality assessments comprises at least three of the following: an SNR (Signal-to-Noise Ratio) quality assessment; a human perception quality assessment; an ASR (Automated Speech Recognition) quality assessment; a speech-to-respiration quality assessment; and an OOD (Out-of-Distribution) assessment.

12. The speech-quality circuit of claim 11 , wherein the ASR quality assessment is based on word-error-rate and semantic distance.

13. The speech-quality circuit of claim 11 , wherein the speech-quality assessment is based on at least one of: pause durations; count of phonemes; and rate of words delivered.

14. The speech-quality circuit of claim 11 , wherein the speech processing circuit is trained using a set of training data, and the OOD assessment is based on a correspondence between the speech signal and the training data.

15. A non-transitory computer-readable medium comprising a program that, when executed by a processing system, causes the processing system to execute the method of claim 11 .

Citation Information

Patent Citations

  • Breathing signal-dependent speech processing of an audio signal

    EP4414984A1

  • Speech-based breathing prediction

    US20210146082A1

  • Detection and Use of Acoustic Signal Quality Indicators

    US20090299741A1

  • Systems and methods for identifying patient talking during measurement of a physiological parameter

    US20140276165A1

  • Input speech quality matching

    US20160379638A1