Audio signal processing
The apparatus generates respiratory waveform signals from audio segments to adapt speech processing, addressing the limitations of existing technologies by enhancing flexibility and efficiency in handling varying speaker properties and conditions.
Patent Information
- Application Number
- JP2025542351
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-09
- Filing Date
- 2024-02-01
- Publication Date
- 2026-02-25
AI Technical Summary
Existing speech processing technologies are not optimally adaptable to varying speaker properties and environmental conditions, leading to suboptimal performance, complexity, and resource inefficiency, particularly in scenarios involving physical activity or breathing difficulties.
An apparatus that generates respiratory waveform signals from audio segments using overlapping fragments, applying weighted combinations and adapting speech processing based on these signals to improve flexibility and efficiency.
Enhances speech processing by reducing complexity and resource usage while improving accuracy and adaptability to varying speaker conditions, especially during physical activities.
Smart Images

Figure 2026506482000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to performing speech processing of audio signals, particularly but not exclusively to automatic speech recognition or speech enhancement of audio signals capturing a speaker performing an activity such as exercising. [Background technology]
[0002] Speech processing of audio signals is widely applied in a variety of practical applications and is becoming increasingly important and meaningful for many everyday activities and devices.
[0003] For example, speech enhancement is widely applied to improve the intelligibility and reproducibility of speech captured by audio signals, e.g., sounds captured in real-world environments. Speech coding is also frequently performed from captured audio. For example, speech coding of captured audio signals is an integral part of smartphones. Another speech processing application that has become increasingly frequent in recent years is the application of speech recognition, e.g., to provide a user interface for devices.
[0004] For example, a personal assistant or home speaker voice interface allows one to control media, navigate functions, track, obtain various information services, etc., simply by using the voice interface and while performing other activities, such as exercising.
[0005] Voice interfaces based on Automatic Speech Recognition (ASR) have also become very important in controlling various devices, such as wearable devices, wearable devices, IoT consumer devices, health devices, and fitness devices. Voice is the most natural interface for interacting with devices during activities involving large user movements and fluctuations. As an example, US20090098981A1 discloses a fitness device with a voice interface used for communication between a user and a virtual fitness coach. US 7 529 670B1 discloses a voice processing system in which processing depends on the determined lung state of the speaker. The article by Nallanthighal et al., "Deep learning architectures for estimating breathing signal and respiratory parameters from speech recordings", Neural Networks, vol. 141, 1 September 2021, pages 211-224, XP093054576, ISSN: 0893-6080, DOI: 10.1016 / j.neunet.2021.03.29, presents an approach to deriving breathing signal samples from speech using neural networks.
[0006] However, while significant effort has been expended in developing and optimizing speech processing algorithms for various applications, resulting in many highly advantageous and efficient approaches, these may not be optimal in all situations. For example, many speech processing operations are developed for specific nominal conditions, such as nominal speaker properties or nominal acoustic environments. Many speech processing applications may not provide optimal performance in scenarios where actual conditions differ from the expected nominal conditions. Additionally, some speech processing algorithms may be unnecessarily complex or resource-intensive. Adapting speech processing to compensate for various properties tends to result in less flexible, more complex, and more resource-intensive implementations, often resulting in suboptimal performance.
[0007] Therefore, improved approaches to speech processing would be advantageous, particularly approaches that allow for greater flexibility, greater adaptability, improved performance, improved quality such as speech enhancement, encoding, and / or recognition, reduced complexity and / or resource usage, improved remote control of audio processing, improved adaptation to variations in speaker properties and / or activity, reduced computational burden, improved user experience, easier implementation, and / or an improved spatial audio experience. Summary of the Invention [Problem to be solved by the invention]
[0008] Accordingly, the Invention seeks to preferably mitigate, reduce or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]
[0009] According to one aspect of the present invention, there is provided an apparatus for speech processing of an audio signal, the apparatus comprising: an input configured to receive an audio signal comprising a speech audio component of a speaker; a segmenter configured to generate segments of the audio signal, each segment having a segment duration and consecutive segments having an inter-segment duration that is the time between the consecutive segments, the segment duration exceeding the inter-segment duration such that consecutive segments overlap; a first generator (107) configured to generate, for each segment of the audio signal, fragments of a respiratory waveform signal from each segment of the audio signal, at least some fragments having fragment durations that exceed the inter-segment duration; a second generator (109) configured to combine the fragments of the respiratory waveform signal to generate the respiratory waveform signal, the second generator (109) configured to combine fragments from the at least some fragments by applying a weighted combination to samples of different fragments at the same time point, the weights of the weighted combination being determined from a fragment window function; and a speech processor configured to apply speech processing to the audio signal, the speech processing being dependent on the respiratory waveform signal.
[0010] This approach provides improved speech processing in many embodiments and scenarios. For many signals and scenarios, this approach can provide speech processing that more closely reflects variations in the properties and characteristics of a speaker's voice. For example, it can adapt to a speaker's level of fatigue, breathing difficulties, etc. This approach, particularly the use of segment / fragment-based processing, can provide accelerated adaptation and, in particular, improved adaptive speech processing. This approach provides efficient implementation and, in many embodiments, allows for reduced complexity and / or resource usage.
[0011] This approach can generally provide a more accurate and improved respiratory waveform signal, and in many scenarios can reduce or mitigate the edge effects and errors present in many segment-based processes.
[0012] The audio processing may include audio enhancement processing, audio recognition processing, and / or audio encoding processing. The segment duration of a segment may be the duration of the time interval of the audio signal contained in the segment. The inter-segment duration may be the period / time offset / difference / segment time interval between subsequent segments.
[0013] At least some of the fragments have a fragment duration that exceeds the inter-segment duration.
[0014] This results in particularly advantageous behavior and performance in many scenarios and applications.
[0015] In many embodiments, at least some of the fragments have a duration that is no less than 10, 20, 30, 50% or 100% of the inter-segment duration. The fragment duration of a fragment can be the duration of the time interval of the respiratory waveform signal represented by the fragment.
[0016] The second generator is configured to combine the fragments by applying a weighted combination to samples of different fragments at the same time point, the weights of the weighted combination being determined from a fragment window function.
[0017] This provides improved operation and performance in many embodiments, which in many embodiments allows for accelerated, improved, and / or more efficient operations. The fragment window function can be such that the weights for each sample time point have a constant combined value (specifically, the sum of the weights for each sample / fragment is constant at each time point).
[0018] In accordance with an optional feature of the invention, the segment duration exceeds the inter-segment duration by 50% or more.
[0019] This provides improved operation and performance in many embodiments. The segment duration can be twice (100% longer) than the inter-segment duration in many embodiments.
[0020] In accordance with an optional feature of the invention, a segment of the audio signal includes at least a first portion of the audio signal that is also included in a previous segment of the audio signal and a second portion of the audio signal that is also included in a subsequent segment of the audio signal.
[0021] This provides improved operation and performance in many embodiments.
[0022] In many embodiments, one or more segments include samples from both the previous and next segments.
[0023] In accordance with an optional feature of the invention, the sample rate of the respiration waveform signal is at least 10 times, or in some embodiments 20, 50 or 100 times lower than the sample rate of the audio signal.
[0024] This provides improved operation and performance in many embodiments, including, in particular, much reduced complexity and reduced resource usage, for example, enabling real-time processing using devices with limited computing resources in many embodiments.
[0025] According to an optional feature of the invention, the first generator may be configured to generate fragments having a lower sample rate than the segments of the audio signal.
[0026] This provides improved operation and performance in many embodiments.
[0027] In accordance with an optional feature of the invention, the segmenter is configured to generate segments having a lower sample rate than the audio signal.
[0028] This provides improved operation and performance in many embodiments.
[0029] According to an optional feature of the invention, the first generator includes a trained artificial neural network having an input node for receiving a sample of a segment of the audio signal and an output node configured to provide a sample of a fragment of the respiration waveform signal for the segment of the audio signal.
[0030] This approach can, in many embodiments and scenarios, provide a particularly advantageous configuration that allows for the utilization, facilitation, and / or improvement of artificial neural networks in generating respiration waveform signals for use in adapting and optimizing voice processing, typically including voice enhancement and / or recognition.
[0031] This approach provides an efficient implementation and, in many embodiments, allows for reduced complexity and / or resource usage.
[0032] An artificial neural network is a trained artificial neural network.
[0033] The artificial neural network can be an artificial neural network trained with training data including training speech audio signals and training respiratory waveform signals generated from measurements of respiratory waveforms, the training using a cost function that compares the training respiratory waveform signals to the respiratory waveform signals generated by the artificial neural network for the training speech audio signals. The artificial neural network is a trained artificial neural network that has been trained with training data including training speech audio signals representing various relevant speakers in various different states and performing different activities.
[0034] The artificial neural network can be an artificial neural network trained with training data having training input data including a training speech audio signal, and trained with a cost function including a contribution indicative of the difference between the measured training respiration waveform signal and the respiration waveform signal generated by the artificial neural network in response to the training speech audio signal.
[0035] According to an optional feature of the invention, the first generator includes a model for human respiration, the first generator being configured to generate the fragments from the model and to adapt parameters of the model in response to the segments.
[0036] This provides improved operation and performance in many embodiments. The model can take as input a given segment of an audio signal and produce as output a fragment of a respiration waveform signal for the given segment.
[0037] According to an optional feature of the invention, the segmenter is configured to apply a frequency transform to the audio signal to generate each segment as a matrix of time-frequency samples including at least samples of two different time intervals and two different frequency ranges.
[0038] This provides improved operation and performance in many embodiments. This approach facilitates operation and provides a more efficient representation of the characteristics of audio signals that are suitable for determining respiration waveform signals.
[0039] According to an optional feature of the invention, the device has an adapter configured to adapt at least one of the inter-segment duration and the segment duration in response to at least one of a characteristic of the audio signal and a characteristic of the respiration waveform signal.
[0040] This provides improved operation and performance in many embodiments.
[0041] In some embodiments, the speech processing includes speech recognition processing that generates a plurality of candidate recognition terms from the audio signal, and the speech processor is configured to select from among the candidate recognition terms in response to the respiratory waveform signal.
[0042] This provides advantageous speech recognition in many scenarios and allows it to be achieved with reduced complexity and resource demands. For example, it allows adaptation and consideration of respiratory waveform signals implemented as post-processing applied to the results of existing speech recognition algorithms. This approach can provide improved backward compatibility.
[0043] In some embodiments, the audio processor is configured to determine a respiration rate estimate from the respiration waveform signal and to adapt audio processing in dependence on the respiration rate.
[0044] This provides improved operation and performance in many embodiments. The breathing rate parameter is a particularly suitable parameter for adapting voice, as there is typically a high correlation between some voice characteristics and breathing rate.
[0045] In some embodiments, the audio processor is configured to determine timing characteristics of the inspiration interval from the respiratory waveform signal and to adapt audio processing in dependence on the timing characteristics of the inspiration interval.
[0046] This provides improved operation and performance in many embodiments. The timing characteristics are typically duration, repetition time / frequency, and / or inhale / breath time during the interval.
[0047] In some embodiments, the audio processor is configured to determine an inhale / breath in a time segment of the audio signal from the timing characteristics, and to attenuate the audio signal during the inhale / breath in the time segment.
[0048] This provides improved operation and performance in many embodiments.The attenuation can be a partial attenuation or a complete attenuation (mute) of the audio signal.
[0049] In some embodiments, the voice processor is configured to determine inhalations / breaths in a time segment of the audio signal from the timing characteristics and to exclude the inhalations / breaths in the time segment from voice recognition processing applied to the audio signal.
[0050] This provides improved operation and performance in many embodiments.
[0051] In some embodiments, the audio processor is configured to determine a breath sound level during inspiration / inhalation from the respiratory waveform signal and to adapt audio processing depending on the breath sound level during inspiration / inhalation.
[0052] This provides improved operation and performance in many embodiments.
[0053] In some embodiments, the voice processor includes a human breathing model for determining breathing characteristics of a speaker, and the voice processor is further configured to adapt voice processing in response to the breathing characteristics and to adapt the human breathing model in dependence on speaker data.
[0054] This can provide improved operation and / or performance in many embodiments.
[0055] A method according to one aspect of the present invention includes the steps of receiving an audio signal including a voice audio component of a speaker; generating segments of the audio signal, each segment having a segment duration and an inter-segment time interval, the inter-segment duration being the time between successive segments and the segment duration exceeding the inter-segment duration so that successive segments overlap; generating fragments of a respiratory waveform signal for each segment of the audio signal from each segment of the audio signal; and generating a respiratory waveform signal by combining the fragments of the respiratory waveform signal, the combining including combining fragments of at least some of the fragments by applying a weighted combination to samples of different fragments at the same time, the weights of the weighted combination being determined from a fragment window function; and applying voice processing to the audio signal, the voice processing being dependent on the respiratory waveform signal.
[0056] The audio processing may include audio enhancement processing.
[0057] This approach provides improved speech enhancement and allows for the generation of improved speech signals in many scenarios.
[0058] The voice processing may include voice recognition processing.
[0059] This approach improves speech recognition, enabling more accurate detection of words, terms, and sentences in many scenarios, and in particular improving speech recognition of speakers in a variety of situations, conditions, and activities, such as improving the voice detection of people exercising.
[0060] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief explanation of the drawings]
[0061] Embodiments of the present invention will now be described, by way of example only, with reference to the drawings in which: [Figure 1] 1 illustrates some elements of an example audio device according to some embodiments of the present invention. [Figure 2] 1 illustrates an example of segmentation of an audio signal. [Figure 3] FIG. 1 shows an example of the structure of an artificial neural network. [Figure 4] FIG. 1 shows an example of a node in an artificial neural network. [Figure 5] Diagram showing some elements of an example training setup for an artificial neural network. [Figure 6] FIG. 10 is a diagram showing an example of a prediction error of a respiratory waveform signal. [Figure 7] 1A-1C illustrate an example process for estimating a respiration waveform signal from an audio signal according to some embodiments of the present invention. [Figure 8] FIG. 1 illustrates some elements of a possible configuration of a processor for implementing elements of an apparatus according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0062] 1 illustrates some elements of a speech processing device according to some embodiments of the present invention. The speech processing device is suitable for providing improved speech processing to a speaker in many different environments and scenarios, such as when the speaker is exercising.
[0063] The audio processing device comprises a receiver / input unit 101 configured to receive an audio signal including a speech audio component. The audio signal may in particular be a microphone signal representing audio captured by a microphone. The audio signal may in particular be a voice signal, and in many embodiments a voice signal captured by a microphone configured to capture the voice of a single person, such as a headset microphone, a smartphone, or a body-worn microphone. Thus, the audio signal may include not only the speech audio component (hereinafter also referred to as the speech signal) but also other sounds, such as background sounds from the environment.
[0064] The audio processing device further comprises an audio processing 103, which is configured to apply audio processing to the received audio signal and thus to the audio audio component of the audio signal. The audio processing can in particular be a speech enhancement process, a speech coding process or a speech recognition process.
[0065] In this approach, the audio processing circuitry 103 is configured to perform audio processing in response to a respiration waveform signal. The audio processing device is configured to generate a respiration waveform signal from the audio signal. The respiration waveform signal indicates the speaker's lung air volume as a function of time, and thus reflects the speaker's breathing and the flow of air in and out of the speaker's lungs.
[0066] The audio processing device is configured to process the audio signal to extract information indicative of the speaker's breathing captured by the audio signal. The audio device can be configured to receive samples of the audio input signal and generate a time series of parameters representing a time series of measured / estimated lung air volumes.
[0067] Thus, the respiration waveform signal can be a time-varying signal that reflects variations in a speaker's breathing, and in particular, can reflect changes in lung volume due to the speaker's breathing. In some embodiments, the respiration waveform signal can directly represent current lung volume. In other embodiments, the respiration waveform signal can represent changes in current lung volume, such as when the respiration waveform signal represents breathing in / breathing out airflow.
[0068] Traditionally, a respiration waveform signal can be captured using a separate sensor device, such as a breathing belt worn by the subject while they are speaking. However, in the current approach, a determiner can determine a respiration waveform signal even when such sensor data is unavailable by deriving the respiration waveform signal from speech data. As a low-complexity example, the determiner can use a voice activity detection (SAD) algorithm to segment the speech data into speech and non-speech segments. Based on the knowledge that inspiration occurs during pauses in speech, the determiner 105 can use the duration and frequency of the pauses as a proxy for breathing activity and determine the respiration waveform signal accordingly. In many embodiments, more accurate determinations can be used, including using artificial neural networks as described below. This can, for example, compensate for or take into account pauses in speech that are more frequent than true inspiration / inhalation events, which can lead to errors in respiration rate estimation.
[0069] The audio processing device 103 is configured to adapt its audio processing in response to the respiration waveform signal. Thus, rather than providing predetermined audio processing based solely on nominal or expected characteristics, the audio processing device of Figure 1 is configured to adapt its audio processing to reflect the speaker's current respiration.
[0070] This approach reflects the inventors' recognition that a speaker's breathing affects their speaking style, and that not only is it possible to determine the time-varying characteristics due to breathing from captured audio, but that these time-varying characteristics can further be used to adapt speech processing of the same audio signal to improve performance, and in particular to better adapt performance to different scenarios and user behavior. The specific voice processing and adaptations performed can depend on the particular implementation. However, in many embodiments, this approach can be used to adapt operation to reflect different breathing patterns a person may exhibit, for example, due to performing different activities. For example, when exercising, breathing may become labored, causing speech to become distorted or slow. The voice processing device of FIG. 1 can automatically adapt to be optimized for slower voices with longer silence intervals.
[0071] This approach can provide significantly improved speech processing in many scenarios, and can provide benefits, for example, particularly for speakers engaged in strenuous activities such as exercise, or speakers who have breathing or speech difficulties.
[0072] This approach is particularly advantageous for automatic speech recognition, e.g., for providing voice interfaces. For example, a personal assistant's voice interface allows for control of media, navigation functions, sports tracking, and various other information services while exercising. However, voice recognition is typically optimized for normal, stationary speech, and in some cases personalized. During and immediately after intense exercise, oxygen needs increase significantly, resulting in changes in breathing and affecting speech. This improves word error rate (WUR) and enhances the usability and user experience provided by the voice interface. Voice processing devices can detect the user's breathing patterns and adapt voice recognition to take the changed voice characteristics into account.
[0073] Speech recognition algorithms are optimized for groups of users, and the voice prompts that initiate interactions are often also optimized for individual users.
[0074] More specifically, when a user is exposed to physical stress, such as during sports or physically demanding work, and greater respiratory effort is required, this tends to affect the user's speech. During strenuous exercise, such as interval training, speech is severely affected. For example, a user may typically need to take a breath every one to three words or may frequently swallow the end of words. This state of breathlessness is known as dyspnea, and speech under such conditions is hereinafter referred to as dyspneic speech. Although dyspneic speech patterns begin to appear already under mild physical stress, human listeners generally unconsciously adapt to this and experience no difficulty understanding such speech.
[0075] However, automatic speech recognition algorithms trained using resting speech may have difficulty accurately understanding dyspneic speech, and in fact may have difficulty even accurately detecting predefined voice prompts or trigger terms. During voice commands and conversations using an automatic speech recognition-based interface, increased breathing rates, rapid audible inhalations / inhalations, and final swallows can significantly degrade performance. This can make the voice interface difficult and frustrating to use.
[0076] The speech processing approach of Figure 1 not only detects such conditions, but allows the system to automatically adapt speech processing to suit the current conditions.
[0077] As a low-complexity example, the audio processor 103 can be configured to select from among several different audio processing algorithms / processes, each of which can be optimized for a particular breathing scenario. For example, one process can be optimized / trained for a resting breathing pattern; another process can be optimized / trained for a more strained but controlled breathing pattern; and yet another process can be optimized / trained for a maximally labored breathing pattern. Based on the respiration waveform signal, the signal processor 103 can select the audio processing algorithm / process that best matches the current breathing pattern indicated by the respiration waveform signal and apply this process to the received audio signal.
[0078] Thus, in some embodiments, the audio processor 103 can be configured to switch between different audio processing algorithms / functions depending on the respiration waveform signal. The audio processor 103 can analyze the respiration waveform signal to determine a parameter value, such as respiration rate. Each processing algorithm / function can be associated with a range of respiration rates, and the audio processor 103 can select an algorithm that includes the rate determined from the respiration waveform signal and apply it to the audio signal.
[0079] In some embodiments, the selectable individual algorithms / functions may include, for example, approaches that use substantially the same signal processing but with parameters optimized for given conditions indicated by the respiration waveform signal. For example, the same audio processing may be used but the operating parameters may be different for each respiration rate. The operating parameters may be determined by training or by individual manual optimization, for example, based on testing using captured audio of people whose respiration characteristics have been measured by other means.
[0080] In some embodiments, the individual algorithms / functions that can be selected may involve fundamentally different approaches, e.g., approaches that use completely different approaches and principles, e.g., different speech recognition approaches may be used for a person at rest, i.e., someone who speaks rapidly and continuously, and for a person who is extremely labored in breathing, where most individual words may be interrupted by loud breathing noises.
[0081] As an example, in many embodiments, the audio processor 103 can be configured to determine timing characteristics of the inspiration / inhalation interval from the respiratory waveform signal.
[0082] Breathing involves a series of alternating inhalations (inhalation of air into the lungs) and exhalations (exhalation of air from the lungs), and the audio processor 103 can be configured to evaluate the audio / breath waveform signal to determine timing characteristics of the intervals during which the speaker inhales / inhales. Specifically, the audio processor 103 can be configured to evaluate the audio / breath waveform signal to determine the frequency of the inhalation / inhalation intervals, the duration of the inhalation / inhalation intervals, and / or when these intervals begin and / or end. Such parameters can be determined, for example, by a peak-picking algorithm that identifies local minima and maxima in the respiratory waveform signal corresponding to the onset of the breathing in and breathing out phases. The respiratory rate can be determined from the time between two consecutive inhalations or exhalations.
[0083] In some embodiments, the voice processor 103 may, for example, select different algorithms or adapt voice processing parameters in response to such timing indicators. For example, a longer duration of inspiration or inhalation (e.g., compared to exhalation) may indicate that the person is taking a deep breath, which may reflect a different speech pattern than during normal breathing.
[0084] However, in some embodiments, the audio processor 103 can be configured to modify audio processing specifically for inhalation / inhalation time intervals. In particular, the audio processor 103 can be configured in many embodiments to inhibit or stop audio processing for inhalation / inhalation time intervals. For example, audio processing in the form of voice recognition can simply ignore all inhalation / inhalation time intervals and only be applied to non-inhalation / inhalation time intervals.
[0085] The voice processor 103 can be configured to determine inhalation / inhalation time segments of the audio signal. These time segments would therefore reflect time intervals of the audio signal during which the speaker is estimated to be inhaling. The voice processor 103, in some embodiments, can be configured to attenuate the audio signal during the inhalation / inhalation time segments. For example, a voice enhancement algorithm configured to reduce noise by emphasizing the voice component of the audio signal can attenuate the audio signal during the inhalation / inhalation time intervals, and in many embodiments, can attenuate the signal completely.
[0086] Such an approach tends to provide significantly improved performance, allowing speech processing to be adapted and applied specifically to portions of the audio signal where speech is present, while enabling a different approach, particularly attenuation, to portions of the audio signal where speech is less likely to be present. This approach can, for example, allow for significant reduction of breathing sounds by removing many breathing sounds during the inhalation / inhalation phase of the breathing waveform signal. This can significantly improve the perceived clarity of speech and provide a correspondingly better perceived speech. In some embodiments, this can be facilitated by allowing for the removal of a portion of the speech data occurring during inhalation / inhalation (as estimated from the breathing waveform) prior to automatic speech recognition, reducing the risk of false word detection. In speech coding embodiments, for example, coding efficiency can be increased because no speech coding occurs during the inhalation / inhalation time interval. For example, an indicator of a silent or non-speech interval can simply be inserted into the coded data stream to indicate that no speech data is provided for this interval.
[0087] The audio processing device of Figure 1 uses a particular segmented approach to generate a respiration waveform signal based on an input audio signal. The input unit 101 is coupled to a segmenter 105 configured to generate segments of the audio signal. The segmenter 105 can be specifically configured to generate the segments periodically, where the segments are generated to have a duration that exceeds the duration between segments, thereby generating at least partially overlapping segments.
[0088] For example, as illustrated in FIG. 2, an audio signal 201 can be segmented into overlapping segments 203, each corresponding to an audio signal within a time interval f1-f5. Each segment has a predetermined duration, i.e., each segment has a segment duration. The segment duration is the duration of the time interval of the audio signal represented by the segment. Segments are repeatedly generated to represent the audio signal over a time interval (usually much longer) than the segment duration. The inter-segment duration (the duration of the time interval between successive segments) is shorter than the segment duration so that the segments overlap. Thus, at least some portion / time interval of the audio signal (especially some samples) is included in two segments (or possibly more than two segments in some situations). The inter-segment duration can be determined, for example, as the time between the start times and / or end times of successive segments, or indeed, the time between any other corresponding relative times in successive segments. Typically, the inter-segment duration can be determined as the time / duration between the (time) midpoints of consecutive segments (this can be taken as the nearest sample instant, or, for example, as the time between sample instants if an even number of samples is contained in the segment).
[0089] Typically, because the segments are generated periodically, the inter-segment duration is constant for at least a plurality of segments (e.g., 5, 10, or more segments). Likewise, the segment duration is also typically constant for at least a plurality of segments (e.g., 5, 10, or more segments). However, it will be appreciated that in some embodiments the inter-segment duration and / or the segment duration may be variable, and in particular the audio processing device may be configured in some embodiments to adapt one or both of these values.
[0090] In many embodiments, there may be a certain relationship / ratio between the segment duration and the inter-segment duration. In many embodiments, this ratio may be an integer ratio. For example, in the approach illustrated in FIG. 2, the segment duration is twice the inter-segment duration. This results in a 50% overlap between two consecutive segments, with each part / sample of the audio signal being contained in two segments.
[0091] The overlapping segments of the audio signal are provided to a first generator 107 configured to generate a fragment of a respiration waveform signal for each segment of the audio signal. The first generator 107 is configured to process each segment individually and generate a fragment of a respiration waveform signal from the individual segments. The first generator 107 can accordingly perform segment-based processing, where each segment of the audio signal is processed individually to generate one fragment of a respiration waveform signal. Thus, for each segment of the audio signal, a fragment of a respiration waveform signal can be generated from that segment of the audio signal.
[0092] In some embodiments, each segment of the audio signal can be processed to generate a fragment of the respiratory waveform signal. In some embodiments, only the audio signal within one segment can be considered / processed when generating a fragment of the respiratory waveform signal. In many embodiments, each fragment of the respiratory waveform signal can include 2, 5, 10, 20, 50, or more samples.
[0093] Each generated fragment can be generated to have a duration that matches the inter-segment duration. For example, a fragment can be generated to have the same duration as the inter-segment duration and cover a time interval centered on the midpoint of the segment time interval. This allows consecutive fragments to be adjacent, but not overlap, so that a continuous respiratory waveform signal is generated. However, in the approach of the device of FIG. 2, fragments can have durations that exceed the segment duration, and the fragments can overlap. The fragments can be timed to be centered on each of the segment time intervals, i.e., so that their midpoints coincide. In many embodiments, each fragment can be generated to cover a time interval that matches the segment time interval of the underlying segment from which the fragment is generated. Thus, in many embodiments, the duration of a fragment is the same as the duration of the segment; indeed, in many embodiments, the time interval over which a fragment is generated is the same as the time interval of the segment used to generate the fragment. For completeness, it is also possible for the fragment duration to be shorter than the inter-segment duration, although this is not the case with the device of FIG. 2. This may lead to a respiratory waveform signal with gaps, but this may be acceptable in some circumstances. In particular, the speech processing circuit 103 can be configured to determine parameters of the respiratory waveform signal (eg, word rate) based on, for example, the fragment alone, and then apply these to the entire segment.
[0094] The terms fragment and segment may be used interchangeably.
[0095] As an example, the first generator 107 may include a model of human respiration that generates a respiration waveform signal based on several model parameters. Some of the model parameters may be set to reflect a particular speaker, for example, based on manually provided data or predetermined data, for example, from an audio signal (or a previous session). Such model parameters may reflect, for example, whether the speaker is female or male, a child or adult, or elderly. The model parameters may also reflect, for example, normal speaking pitch, average word rate, etc.
[0096] The model can be configured to generate respiratory waveform signal fragments having a predetermined size, i.e., specifically, fragments having a given duration each time the model is evaluated. The model can be evaluated for each segment of the audio signal based on the model parameters. Furthermore, before evaluating the model, one or more parameters can be specifically determined for the current segment based on an evaluation of the segment itself. For example, the current pitch rate, word length, speech activity level, etc. can be determined and used as inputs for the model that generates the respiratory waveform signal.
[0097] In one implementation, the model is based on using a voice activity detection (SAD) algorithm to find breaks in each voice input segment. For each time sample of audio data, SAD provides an indication of whether the sample data is voice audio data or non-voice audio data. Based on the SAD signal, sequences of voice breaks, e.g., longer than 200 ms, can be selected. Next, assuming that the breaks represent inhalation events during breathing, a respiratory waveform fragment is shaped such that each edge of the break is selected as a local minimum and each edge of the break is a local maximum. All intermediate points can then be formed by linear (or polynomial) interpolation to generate a "sawtooth" waveform pattern corresponding to the respiratory waveform fragment.
[0098] The first generator 107 is coupled to a second generator 107, which is configured to combine fragments of the respiratory waveform signal to generate the respiratory waveform signal. In an embodiment, each fragment has a duration corresponding to the segment time interval for which the fragment was generated, and the combination is by simply concatenating the fragments into a signal of extended duration. As will be explained in more detail below, in the case of overlapping fragments, the combination can be, for example, a weighted addition of signal levels / samples of the respiratory waveform signal generated for the same time point.
[0099] This respiration waveform signal is supplied to the audio processing circuit 103 together with the audio signal, and the audio processing circuit 103 performs audio processing based on the generated respiration waveform signal.
[0100] In many embodiments, the determination of the fragments of the respiratory waveform signal from the segments can be performed by an appropriately trained artificial neural network.
[0101] In many embodiments, the first generator 107 includes a trained artificial neural network having an input node for receiving a sample of a segment of an audio signal and an output node configured to provide a sample of a fragment of a respiration waveform signal for the segment of the audio signal.
[0102] In many embodiments, the determiner 105 can include a trained artificial network configured to determine samples of the respiration waveform signal based on samples of the audio signal. The artificial neural network can include input nodes that receive the samples of the audio signal and output nodes that generate the samples of the respiration waveform signal. The artificial neural network can be trained using, for example, speech data and sensor data representing measurements of lung air volume during speech.
[0103] In this example, the artificial neural network has input nodes that receive samples of one segment of the audio signal, the number of input nodes matching the number of samples in the signal. The artificial neural network can then process the segment to generate output samples that are samples of a fragment of the respiratory waveform signal. The number of output nodes typically matches the number of samples in the fragment, with each output node providing a sample of a fragment of the respiratory waveform signal. Thus, the artificial neural network is run segment by segment to generate the fragments of the respiratory waveform signal.
[0104] The artificial neural network used in the described functions can be a network of nodes organized as layers, with each node holding a node value. Figure 3 shows an example of a section of an artificial neural network.
[0105] The node value of a given node can be calculated to include contributions from some, or often all, nodes in previous layers of the artificial neural network. Specifically, the node value of a node can be calculated as a weighted sum of the node values of all node outputs in the previous layer. Typically, a bias is added, and the result is applied to an activation function. The activation function typically provides nonlinearity to account for the key role of each neuron. Such nonlinearity and activation function have a significant impact on the learning and adaptation process of the neural network. Thus, the node value is generated as a function of the node values in the previous layer.
[0106] The artificial neural network may specifically include an input layer 301 that includes a plurality of nodes that receive input data values to the artificial neural network. Thus, the node values of the nodes in the input layer may typically be direct input data values to the artificial neural network and therefore may not be calculated from other node values.
[0107] An artificial neural network may further include zero, one, or more hidden layers 303 or processing layers. For each such layer, node values are typically generated as a function of the node values of the nodes in the previous layer; in particular, an activation function (such as a sigmoid, ReLU, or Tanh function) may be applied, followed by a weighted combination and an added bias. Specifically, as shown in Figure 3, each node, sometimes called a neuron, can receive input values (from nodes in the previous layer) and then calculate the node value as a function of these values. Often, this involves first generating a value as a linear combination of the input values, each weighted by a weight:
number
[0108] An activation function can then be applied to the resulting combination. For example, the node value l can be determined as l = f(k).
[0109] This function can be, for example, a modified linear identity function, as described in Xavier Glorot, Antoine Bordes, Yoshua Bengio Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, PMLR 15:315-323, 2011. f(k) = ReLU(k) = max(0,k)
[0110] Other commonly used functions include the sigmoid function or the tanh function. In many embodiments, a node output or value can be calculated using multiple functions. For example, both ReLU and Sigmoid functions can be combined using an activation function such that f(k)=ReLU(k)+σ(k).
[0111] Such operations can be performed by each node of the artificial neural network (typically except for the input node).
[0112] The artificial neural network further comprises an output layer 305 that provides output from the artificial neural network, i.e., the output data of the artificial neural network are the node values of the output layer. As for the hidden / processing layers, the output node values are generated by functions of the node values of the previous layer. However, in contrast to the hidden / processing layers, where the node values are typically not accessible or further used, the node values of the output layer are accessible and provide the results of the operation of the artificial neural network.
[0113] Several different network structures and toolboxes for artificial neural networks have been developed, and in many embodiments, artificial neural networks can be based on the adaptation and customization of such networks. An example of a network architecture suitable for the above applications is the long short-term memory (LSTM) by Sepp Hochsreiter, as described in Hochreiter, Sepp, and Jurgen Schmidhuber. "Long short-term memory." Neural computation 9.8 (1997): 1735-1780.
[0114] LSTM is an architecture used for classification and regression of time-domain signals using iterative causal or bidirectional evaluation, and has been successfully applied to audio signals. For example: t = σg (W f * x t + U f * h t-1 + V f oc t-1 + b f ) where * is matrix multiplication, o is the Hadamard product, x is the input vector, and h t-1 is the output vector of the previous time step, W, V, u are the network weights, b is the bias vector, σ g is typically a sigmoid function or other compressive nonlinear function, c t-1 is the previous state vector, and the output f t is the activation vector corresponding to the forget gate of the LSTM network, and later the current state vector c t Affects.
[0115] In theory, classical (or "vanilla") artificial neural networks can track arbitrary long-term dependencies in input sequences. The problem with vanilla artificial neural networks is computational (or practical) in nature. When training vanilla artificial neural networks using backpropagation, the long-term gradients that are backpropagated can "vanish" (i.e., tend to zero) or "explode" (i.e., tend to infinity) due to the calculations involved in the process, which use finite-precision numbers. Artificial neural networks that use LSTM units partially solve the vanishing gradient problem because the LSTM units allow the gradients to flow unchanged. However, LSTM networks can still suffer from the exploding gradient problem.
[0116] In some cases, an artificial neural network can be further configured to include additional contributions that allow the artificial neural network to be dynamically adapted or customized to specific desired characteristics or features of the generated output. For example, a set of values can be provided to adapt the artificial neural network. These values can be included by providing contributions to several nodes of the artificial neural network. These nodes can specifically be input nodes, but typically can be nodes in hidden or processing layers. Such adaptation values can be weighted and summed, for example, as contributions to a weighted sum / correlation value for a given node.
[0117] The above description relates to a neural network approach that may be suitable for many embodiments and implementations. However, it will be understood that many other types and structures of neural networks can be used. Indeed, many different approaches for generating neural networks have been developed, including neural networks that use complex structures and processes different from those described above. This approach is not limited to any particular neural network approach, and any suitable approach can be used without detracting from the invention.
[0118] Artificial neural networks are adapted for specific purposes through a training process used to adapt / tune / modify the weights and other parameters (e.g., biases) of the artificial neural network. It will be appreciated that many different training process algorithms for training artificial neural networks are known. Typically, training is based on a large training set, in which a large number of examples of input data are provided to the network. Furthermore, the output of the artificial neural network is typically compared (directly or indirectly) to expected or ideal results. A cost function can be generated to reflect the desired outcome of the training process. In a typical scenario known as supervised learning, the cost function often represents the distance between the prediction for particular input data and the ground truth. Based on the cost function, the weights can be modified, and by repeating the process with the modified weights, the artificial neural network can be adapted to a state where the cost function is minimized.
[0119] More specifically, during the training step, a neural network may have two distinct information flows: from input to output (forward pass) and from output to input (backward pass). In the forward pass, data is processed by the neural network as described above, and in the backward pass, weights are updated to minimize the cost function. Typically, such backward propagation follows the gradient direction of the cost function landscape. In other words, by comparing the predicted output for a batch of data input with the ground truth, the direction in which the cost function is minimized can be estimated, and backward propagation can be performed by updating the weights accordingly. Other known approaches for training artificial neural networks include, for example, the Levenberg-Marquardt algorithm, the conjugate gradient method, and the Newton method.
[0120] In this case, training can include, inter alia, a training set including a potentially large number of pairs of captured audio signals and measured respiratory waveform signals. The measured respiratory waveform signals can be generated from sensor signals of a sensor configured to measure lung air volume-dependent characteristics. Thus, training of the artificial neural network can be performed using a training set comprising linked audio / speech data / signals and sensor data / signals representing lung air volume measurements during speech.
[0121] In some embodiments, the training data is an audio signal in a time segment corresponding to the processing interval of the artificial neural network being trained; for example, the number of samples in the training audio signal may correspond to the number of samples corresponding to the input nodes of the artificial neural network being trained. Thus, each training example may correspond to one operation of the artificial neural network being trained. However, typically, to speed up the training process, a batch of training samples is considered for each step. Furthermore, many upgrades to gradient descent are possible to speed up convergence or avoid local minima in the cost function landscape.
[0122] FIG. 5 shows an example of how training data is generated by dedicated testing. A speech audio signal 501 can be captured by a microphone 503 during a time interval in which a person is speaking. The resulting captured audio signal is provided as input data to an artificial neural network 505 (appropriate processing such as amplification, digitization, filtering, etc. can be performed before generating the test audio signal provided to the artificial neural network). Additionally, a respiratory waveform signal 507 is determined as a sensor signal, for example, from a suitable lung volume sensor 509 placed on the user. For example, the sensor 509 can be a respiratory inductive plethysmography (RIP) sensor or a nasal cannula flow sensor.
[0123] A number of such measurements can be performed to generate a number of pairs of training audio signals and respiration waveform signals. These signals are then applied to train an artificial neural network, where a cost function is determined as the difference between the respiration waveform signal generated by the artificial neural network for the training audio signals and the measured training respiration waveform signal. The artificial neural network is then adapted based on the cost function, as is known in the art. For example, a cost value can be determined for each training audio signal and / or training downmix audio signal combination set (e.g., an average cost value for the training set is determined). Typically, the cost function includes at least one component reflecting how close the generated signal is to a reference signal, i.e., the so-called reconstruction error. In some embodiments, the cost function includes at least one component reflecting how close the generated signal is to the reference signal from a perceptual perspective. In the example of FIG. 5, the reference signal can specifically be the measured respiration waveform signal.
[0124] The described approach can provide particularly advantageous performance in many scenarios and applications. This allows for significantly improved speech processing based on a more accurate respiration waveform signal, typically estimated from an audio signal. Segment-based processing is highly advantageous because it enables efficient and practical processing and implementation. However, segmented operations are typically disadvantageous and associated with deficiencies and issues. In particular, segments tend to introduce edge effects, and the inventors have recognized that determining a respiration waveform signal from an audio signal through segmented processing can result in many edge effects. The inventors have further recognized that improved performance can be achieved in many scenarios by extending the duration of segments and generating fragments based on speech information and data / segments that extend beyond the interval between segments, i.e., based on overlapping segments.
[0125] Figure 6 shows an example of the distribution of reconstruction error within a 128-sample respiratory waveform signal reconstruction window accumulated over several thousand frames. The error is clearly smallest in the center of the frame and increases toward the beginning and end of the segment. In the approach described for the device in Figure 1, reconstruction can focus on the central region with lower regions and / or consider multiple segments / fragments for reconstruction toward the edges.
[0126] In many embodiments, the segment duration can exceed the inter-segment duration by a significant amount, and in many embodiments, can exceed it by 25%, 50%, 75%, or 100% or more. In particular, for example in the described example of Figure 2, the segment duration is twice (100% longer) than the inter-segment duration, resulting in each sample or time point in the audio signal being contained in two segments, allowing for a particularly advantageous operation that is well suited to efficient implementation.
[0127] In many embodiments, a segment can include a portion of the audio signal that is also included in a previous segment and another portion of the audio signal that is included in a subsequent segment of the audio signal. For example, in Figure 2, segment f2 includes a first portion that is also included in the previous segment f1 and a second portion that is also included in the subsequent segment f3. Indeed, in this example, the first half of segment f2 is included in segment f1, and the second half of segment f2 is included in segment f3.
[0128] Such overlapping segments, and including at least some portions of the audio signal in multiple segments, provides particularly advantageous performance in many embodiments, as it can significantly reduce segment edge effects in each fragment while allowing the fragments to continue one another without gaps, i.e., generating a continuous respiratory waveform signal while particularly mitigating edge effects on the respiratory waveform signal.
[0129] In many embodiments, the fragments can be generated to be at least partially overlapping. The combining performed by the second generator 107 can in some cases be a simple combining such as concatenating consecutive fragments into the respiratory waveform signal, or in some cases involves selective combining of overlapping portions (e.g., selecting the respiratory waveform signal sample closest to the sample midpoint).
[0130] However, in many embodiments, the combination can be a weighted combination, where a sample at a given time point is generated as a weighted combination of samples at that time point from different fragments. In particular, if two fragments overlap, a sample of the respiratory waveform signal at a given time point can be generated by weighting and adding together the two samples of the two fragments for that time point.
[0131] The weighting factor for each sample can be determined as a function of the time difference between the sample's time point and the midpoint of the fragment, in particular as a monotonically decreasing function of that difference.
[0132] In many embodiments, the weighting values can be determined to reflect a fragment window function. For example, in some embodiments, each sample of a fragment can be multiplied by a weighting factor that is a function of the sample time within the fragment. The fragment window function can be, for example, a raised sine window function, a Hanning window function, or the like.
[0133] In most embodiments, the weights can be determined so that the total / combined / accumulated weighting factor at a given time point (taking into account all fragments contributing to that time point) is constant. For example, in many embodiments, the weighting factors sum to 1 for all time points. Such an approach is typically achieved by selecting an appropriate window for a given fragment size and overlap. For example, in the particular example given above, a raised sine window function applied to the fragments essentially applies a constant combined weighting factor to the fragments.
[0134] After scaling by the weighting factors of the window function, the results can simply be added by the second generator 107 .
[0135] In many embodiments, the audio processor is configured to generate a respiration waveform signal having a much lower sample rate than the audio signal. Indeed, in many embodiments, the audio processor may generate a respiration waveform signal with a sample rate that is 10, 50, or 100 times or more lower than the sample rate of the audio signal. For example, the audio signal often has a sample rate of 16 kHz, 32 kHz, or 64 kHz (or higher), while the respiration waveform signal has a sample rate of 50 Hz, 100 Hz, 200 Hz, or 500 Hz or lower.
[0136] In many embodiments, the first generator 107 and / or the segmenter 105 can be configured to generate segments having a lower sample rate than the audio signal. Thus, in many embodiments, the audio processing device can be configured to decimate the audio signal by a factor of 10 or more, often 100 or more, prior to generating the respiration waveform signal.
[0137] Such decimation significantly simplifies the computation and often allows for a significant reduction in complexity and required resource usage. In particular, a significant reduction in the number of samples in each segment in a given time interval facilitates processing to extract model parameters that can be used, for example, to fit a model for generating a respiratory waveform signal. In implementations based on the use of artificial neural networks, a significant reduction in samples in a segment can significantly reduce complexity and resource usage. In particular, the number of input nodes can be reduced by a factor corresponding to the decimation factor for a given segment duration. This further reduces the number of nodes and calculations required to perform the operation of the artificial neural network, significantly reducing computational resources.
[0138] This approach reflects the inventors' recognition that the dynamics of respiratory function are much slower and can be represented at a much lower sample rate than the dynamics of audio and speech, and the counterintuitive insight that this not only allows for decimation of the generated respiratory waveform signal, but also allows all computations to be performed in the decimated domain while retaining the information necessary to generate the respiratory waveform signal.
[0139] In some embodiments, the decimation can be performed as part of the process of generating the respiration waveform signal. In some embodiments, the first generator 107 is configured to generate fragments with a lower sample rate than the segments of the audio signal. Thus, the segments with a higher sample rate can be provided to the first generator 107, which can then directly generate the respiration waveform signal at a much lower sample rate.
[0140] For example, samples of a segment at a higher sample rate can be fed into an artificial neural network, i.e., the artificial neural network can have multiple input nodes corresponding to the number of samples in the segment. Based on this, a fragment of the respiratory waveform signal containing samples at a much lower sample rate can be generated. The artificial neural network can then have multiple output nodes corresponding to the samples (at the lower sample rate) of the fragment. The artificial neural network can therefore have many fewer output nodes than input nodes. This approach allows for computationally efficient operation, potentially enabling improved determination of the respiratory waveform signal by taking into account higher frequency information.
[0141] In many embodiments, the segment samples provided to the first generator 107 may be in the form of time-domain samples. However, in some embodiments, the segment samples may be provided in the frequency domain, and in particular, the segmenter 107 may apply a frequency transform to generate a frequency representation of the signal. Accordingly, in many embodiments, each segment may be represented by a set of frequency samples, such as the output of an FFT operation or a (e.g., QMF) filter bank.
[0142] In many embodiments, each segment can advantageously be represented by a set of samples including both frequency samples and time samples, and in particular, each segment can be represented by a matrix of time-frequency samples including at least samples from two different time intervals and two different frequency ranges.
[0143] For example, an FFT or QMF filter bank can be applied to the input signal in a given block size that is smaller than the entire segment. For example, a given segment can be divided into a number of N blocks, with each block containing M samples. A frequency transform can be applied to each block, e.g., to generate M frequency bins for each block. The frequency transform can be applied repeatedly to generate N samples for each bin (i.e., N time-domain samples for each frequency bin). The resulting samples can be used to represent the segment by an NxM sample matrix, with N time-domain samples for each of the M frequency bins.
[0144] Thus, in many embodiments, each segment can represent a matrix of time-frequency samples that includes at least samples from two different time intervals and two different frequency ranges. For each frequency bin, there are typically multiple sample values at different times that correspond to each frequency transform block.
[0145] In many embodiments, such a matrix representation of a segment can be fed to, for example, an artificial neural network, i.e., the artificial neural network can include input nodes for receiving blocks of samples that include both frequency and time separated samples.
[0146] In many embodiments, the matrix representation may have a reduced sample rate, as described above, and in particular, in many embodiments, the frequency transformation may be performed after decimation has been applied to the time-domain representation of the audio signal, as described above.
[0147] Representing segments by combined time-frequency sets of samples provides, in many embodiments, improved determination of the respiratory waveform signal and allows for accelerated and improved implementation, for example, by enabling frequency transforms with significantly reduced complexity and resource usage, and has been shown to improve operation and performance, particularly when used in conjunction with artificial neural networks.
[0148] Below, a specific example of deriving a respiration waveform signal from an audio signal is described. This approach includes most of the features and approaches previously described, but it will be understood that this does not mean that these features must be applied together, but rather that they may be applied individually / separately in different embodiments, and indeed that in many embodiments only some of these features will be implemented / included.
[0149] In this example, the speech processing device may predict fragments b(k) = k = 0, ..., B of the respiratory waveform signal for a segment of an audio signal x(t), t = 0, ..., N-1, where t and k are time indexes, typically at different sampling rates, the length of the segment is N samples, and the generated fragments are B samples long.
[0150] This operation can be performed within a frame / segment of N samples (i.e., with a segment duration of N samples) or in steps of S samples (i.e., with an inter-segment duration of S samples), where typically S = N / 2. We use the subscripts α and β to denote the signal and frame sizes for audio and VRB signals, respectively.
[0151] And the pth audio vector is x p = (x(τ), τ = pS α , ..., pS α + N α - 1). The estimation model is given by N βa respiratory waveform signal segment b of samples p For the reconstruction, k = [pS β , pS β + N β A window function w that is non-zero in [- 1] and zero elsewhere p Define (k). The window function has the following properties:
number
number
[0152] However, in other embodiments, other window functions may be used that, for example, meet the above properties. In some embodiments, the shape of the window function may be optimized based on the shape of the frame reconstruction error function shown in FIG. A respiration waveform signal can then be generated as follows:
number
[0153] In practical implementation, p has limited support, such as p = 0, ..., P - 1, so the initial and final S β The reconstruction of these samples is not perfect, but tapers off due to the tails of the window function. In one embodiment, the first and last reconstruction windows of the processed signal can be modified to avoid tapering.
[0154] This approach is illustrated in Figure 7, where figure (a) discloses an example of a sample frequency and time matrix, figure (b) shows the generated fragments of the respiratory waveform signal, figure (c) shows the fragment window function, and figure (d) shows the resulting combined respiratory waveform signal.
[0155] In many embodiments, the audio processing device may include an adapter 113 configured to adapt timing characteristics of the segmentation, in particular adapting the inter-segment duration and / or the segment duration. Thus, in many embodiments, the adapter 113 may dynamically change the inter-segment duration and / or the segment duration to reflect current conditions.
[0156] The adapter 213 can be configured to adapt the timing parameters in response to, among other things, characteristics of the audio signal and / or characteristics of the respiration waveform signal. Thus, in some embodiments, the adapter 213 can be configured to evaluate the audio signal to extract characteristics such as, for example, pitch or word rate, and modify the inter-segment duration and / or segment duration accordingly, for example, by shortening both the inter-segment duration and segment duration for increasing word rate. In some embodiments, the adapter 213 can be configured to evaluate the respiration waveform signal to extract characteristics such as, for example, respiration rate, and modify the inter-segment duration and / or segment duration accordingly, for example, by shortening both the inter-segment duration and segment duration for increasing respiration rate.
[0157] In some embodiments, the adapter 213 can be configured to adapt the inter-segment duration depending on the characteristics of the audio signal. For example, if the level of background noise is high, the model uses a longer inter-segment time for reconstructing the respiratory waveform. In another embodiment, a voice classification method can be used to detect whether the speaker is speaking quietly, normally, or loudly, and the inter-segment duration can be changed accordingly, such as selecting a shorter duration for louder (e.g., shouting) voices.
[0158] In some embodiments, adapter 213 can be configured to adapt the inter-segment duration depending on the characteristics of the respiratory waveform signal. For example, if the respiratory rate calculated from the respiratory waveform is low, the inter-segment duration can be increased and the segment durations can be increased to accommodate more respiratory events in each segment.
[0159] In some embodiments, the adapter 213 can be configured to adapt the segment duration depending on the characteristics of the audio signal, for example, if the level of background noise is high, the model may use a long segment duration for reconstructing the respiratory waveform.
[0160] In some embodiments, adapter 213 can be configured to adapt the segment duration depending on the characteristics of the respiratory waveform signal, for example, if the average respiratory rate calculated from the respiratory waveform signal is low, the segment length is increased to accommodate more respiratory events.
[0161] A typical (initial) segment duration N is 4 seconds, and the inter-segment can be set to have a duration of 2 seconds (initially).
[0162] As previously explained, this approach can improve performance in a variety of voice processing applications.
[0163] In particular, in the case of automatic speech recognition, this approach can provide significantly improved performance by adapting the process to the speaker's current characteristics.
[0164] For example, as described above, the voice processor 103 can be configured to determine inhalation / inhalation time segments of the audio signal that correspond to times when the speaker is breathing in and therefore not speaking. The voice processor 103 can then be configured to exclude these inhalation / inhalation time segments from the voice recognition processing applied to the audio signal. Thus, times when the respiration waveform signal may indicate very little likelihood of speech can be excluded from voice recognition, thereby reducing the risk of falsely detecting words when no speech is present. The ability to estimate when words are likely to be spoken can also improve the accuracy of detection at other times.
[0165] As another example, the respiration waveform signal can be used to adapt speech recognition, for example, by estimating word rate from the level of effort and distress in the breathing pattern estimated from the respiration waveform signal.
[0166] In some embodiments, the speech processor 103 can be configured to perform speech recognition operations without (initially) considering the respiration waveform signal or any parameters derived therefrom. However, rather than simply providing an estimated term or sentence, the speech recognition can generate multiple different candidates for estimation of what may have been spoken. The speech processor 103 can then be configured to select among the different candidates based on the respiration waveform signal. For example, the activity waveform for the estimated sentence can be compared to the respiration waveform signal and selected based on how closely they match. For example, a sentence that does not match an inhale / inhale interval can be selected over a sentence that does match an inhale / inhale interval (e.g., including speech during the estimated inhale / inhale interval).
[0167] As another example, the selection between candidates can be performed based on the level of respiratory distress. For example, when a user is relaxed, the speech word rate is typically much higher than when the user is in great distress and experiencing respiratory distress (e.g., due to strenuous exercise). Thus, the speech processor 103 can be configured to select among the candidates based on these word rates. For example, if the respiratory waveform signal indicates that the speaker is relaxed, the candidate with the highest word rate can be selected, and if the respiratory waveform signal indicates that the speaker is experiencing respiratory distress, the candidate with the lowest word rate can be selected.
[0168] Of course, in many scenarios, the respiration waveform signal is only a factor in selecting candidates, and other parameters may be considered. For example, speech recognition may estimate a confidence or belief value for each candidate / guessing in addition to candidate words, and selection may further take such confidence levels into account.
[0169] A particular advantage of such an approach is that improved detection can be achieved by additionally taking into account the speaker's breathing characteristics, while still allowing the use of conventional speech recognition modules. Thus, post-processing of results from existing speech recognition algorithms can be used to improve performance, while ensuring backward compatibility and allowing the use of existing recognition algorithms.
[0170] In many embodiments, the audio processing can be, for example, audio encoding, where the captured audio is encoded for efficient transmission or distribution over an appropriate communication channel. As mentioned above, more efficient encoding can typically be achieved by simply not encoding the signal in the inspiration / inhalation time segment or by reducing the number of bits allocated to the signal in the entropy encoding during the inspiration / inhalation period.
[0171] As another example, information about breathing waveforms can be used to control voice manipulation, e.g., to mask a speaker's fatigue. For example, a system can detect a state of respiratory distress by detecting an increase in the speaker's breathing rate. A speech synthesizer, such as a Wavenet model, can then be used to resynthesize a new speech signal that resembles the same speaker's voice but with a different breathing waveform pattern corresponding to a speaker with a slower breathing rate.
[0172] In some embodiments, the audio processing can be audio enhancement. Such audio enhancement can be used, for example, to provide a clearer audio signal with reduced noise from other audio sources and in the audio signal. For example, as previously described, the audio signal can be attenuated during inspiration / inhalation time segments. As another example, different frequencies can be attenuated depending on the respiratory waveform signal, thereby differentially attenuating frequencies that tend to be more dominant to some respiratory noises. As another example, the respiratory waveform signal can indicate transitions or phases in the respiratory cycle that tend to be associated with particular sounds, and the audio processing can be configured to generate corresponding signals and apply them in antiphase to provide an audio canceling effect for such sounds.
[0173] In many embodiments, the speech processor 103 can be configured to evaluate the respiration waveform signal to determine one or more parameters of the speaker's respiration, and can further be configured to adapt speech processing based on the determined parameter values.
[0174] For example, in many embodiments, the audio processor 103 is configured to determine a respiration rate estimate from the respiration waveform signal and adapt audio processing in dependence on this respiration rate.
[0175] Because breathing is an inherently repetitive / periodic process, the respiratory waveform signal is also a periodic signal (or at least has a strong periodic component), and the respiratory rate can be determined by determining the repetition rate / duration of the periodic component of the respiratory waveform signal. For example, the respiratory waveform signal can be correlated with itself, and the period between correlation peaks can be determined and used as a measure of the respiratory rate. It will be appreciated that many different techniques and algorithms are known for detecting periodic components in a signal, and any approach can be used without detracting from the present invention.
[0176] Breathing rate may be indicative of several parameters that may affect a person's speech. For example, breathing rate may be a strong indicator of how relaxed, distressed, or tired a speaker is, which, as mentioned above, may affect speech in various ways, e.g., word rate, amount of breath sounds, etc. As mentioned above, voice processing can be adapted to reflect these parameters.
[0177] The respiratory rate can also provide information related to the timing of inspiration / inhalation time intervals / segments. In particular, the frequency of inspiration / inhalation time intervals and the duration between them can be directly given by the respiratory rate.
[0178] As another example, respiration rate can be used to adapt speech processing by selecting a particular speech pre-processing method based on the speaker's respiration rate, e.g., using different parameters of a voice activity detection or speech enhancement algorithm when a speaker's respiration rate is low than when the respiration rate is elevated, i.e., high.
[0179] In some embodiments, the speech processor 103 is configured to determine a speech word rate from the respiration waveform signal and adapt speech processing in dependence on this speech word rate.
[0180] The speech word rate can be determined, for example, by the output of an automatic speech recognition method.
[0181] As mentioned above, a speaker's word rate can vary greatly depending on how relaxed or stressed the speaker is. The respiration waveform signal can be used to determine this level, and therefore the degree of dyspnea the speaker is currently experiencing. A word rate for a given level of dyspnea can be determined (e.g., as a predetermined value for a given degree of dyspnea). Speech processing can then be performed based on the determined word rate. For example, as mentioned above, speech processor 103 can perform speech recognition to provide multiple possible candidates, and the candidate with the word rate closest to the word rate determined from the respiration waveform signal can be selected. In some embodiments, speech recognition thus provides several candidates for a speech utterance, and based on the expected word rate determined from the respiration waveform signal, speech processor 103 can select the utterance corresponding to the lowest word rate and with long speech pauses if breathing is very difficult.
[0182] In fact, many speech recognition systems have problems when word rates are low and breathless speech pauses occur, but in this approach, the parameters of the speech recognition can be adjusted to address this, or the output of the speech recognition system can be processed accordingly.
[0183] As another example, the word rate can be used to adapt speech processing by having specific settings in speech pre-processing or speech coding algorithms that are set based on the word rate, which can trigger, for example, in a speech coder, the use of different quantization vector codebooks in the encoder and decoder.
[0184] In some embodiments, the audio processor 103 is configured to determine the breath sound level during inspiration / inhalation from the respiratory waveform signal and adapt audio processing depending on the speech breath sound level during inspiration / inhalation.
[0185] The respiratory sound level can be determined, for example, by calculating the root mean square amplitude level of the inspiratory sound signal (audio signal detected in the inhalation / inhalation time interval), and if necessary, weighting can be applied to account for the relative loudness of the sound as perceived by the human ear.
[0186] For example, in an audio recording, the sound of breathing can distract a listener from what is being spoken or sung, in which case the sound level of the detected breathing can be a measure of how much the breathing needs to be attenuated to become inaudible.
[0187] In some embodiments, the audio processor 103 can be configured to take into account other data than just the audio signal (and the respiratory waveform signal derived therefrom).
[0188] This approach can provide particularly advantageous and synergistic benefits for embodiments in which audio processing is adapted based on parameters of the respiration waveform signal that may not directly indicate a person's physical movement but may instead be due to other problems, for example, being able to distinguish between dyspnea caused by strenuous exercise and dyspnea caused by a medical respiratory problem and therefore occurring in a relaxed state.
[0189] As a specific example, the speech processor 103 can make a prediction of respiration rate or word rate based on the detected heart rate and use this estimate to set parameters for speech preprocessing or encoding algorithms, as described above. This is based on the consideration that heart rate and respiration rate are often correlated. This is advantageous for obtaining the best performance of the system from the speaker's first utterances, before accessing a sufficiently long speech recording necessary for a model for estimating respiration waveforms to be able to provide an initial estimate. This is also advantageous when the speaker only uses very short utterances, for example, when using voice commands to control a music playback application on a portable audio player.
[0190] In some embodiments, the voice processor 103 can include a human respiration model that can be used to determine the speaker's breathing characteristics. Such a respiration model can be adapted based on received heart rate and / or activity. In some embodiments, the voice processor 103 can derive a computational model that generates an estimate of respiration characteristics from activity characteristics corresponding to current physical activity and vitals measured using the wearable.
[0191] For example, a human respiration model may model specific biophysical processes in the body's organs, or it may be a functional, high-level metabolic model. For example, the metabolic equivalent (MET) relates the body's oxygen uptake rate for a particular activity as a multiple of resting VO2. Based on the known MET for an activity, effective lung capacity, VO2, and the oxygen concentration in air, an indirect estimate of the breathing rate required to meet the subject's oxygen needs can be obtained. This can be adapted based on heart rate by a more fine-grained model that includes specific MET values for different intensities of a given physical activity, from a light jog to full-speed uphill running, based on measurements of heart rate and activity intensity.
[0192] This approach is applicable to a variety of languages. In many embodiments, the same approach can be used for different languages, e.g., the same trained neural network can be used. Thus, in some embodiments, the processing can be language-independent. This reflects the fact that many linguistic properties are the same across different languages, and therefore the relationship between many speech parameters and breathing may be less sensitive to language changes. For example, across all languages, speech occurs on the exhale, not the inhale.
[0193] Indeed, experiments conducted by the applicant have shown that the effect of breathing on speech has similar temporal statistics in most languages, making the approaches described in many embodiments and scenarios suitable for speech in many (if not all) languages, and therefore in many cases no language adaptations or limitations are employed or envisioned.
[0194] However, in some embodiments, the processing may be specific to a particular language. For example, training a neural network using only a particular language may optimize / train it for that particular language. Thus, in such cases, it may perform better for speech in that language.
[0195] In some embodiments, the approach can be configured to first detect the language and then apply settings specific to that language, such as loading the neural network with coefficients and parameters for that particular language. In other embodiments, the neural network can be trained using, for example, different languages, and the neural network can automatically adapt to different languages.
[0196] The audio device may in particular be implemented as one or more suitably programmed processors. For example, the artificial neural network may be implemented in one or more such suitably programmed processors. Different functional blocks, in particular the artificial neural network, may be implemented in separate processors and / or may, for example, be implemented in the same processor. Examples of suitable processors are given below:
[0197] 8 is a block diagram illustrating an exemplary processor 800 according to an embodiment of the disclosure. Processor 800 may be used to implement one or more processors that implement the apparatus or elements thereof (including, inter alia, one or more artificial neural networks) as previously described. Processor 800 may be any suitable processor type, including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA) (wherein the FPGA is programmed to form a processor), a graphics processing unit (GPU), an application specific integrated circuit (ASIC) (wherein the ASIC is designed to form a processor), or a combination thereof.
[0198] The processor 800 may include one or more cores 802. The cores 802 may include one or more arithmetic logic units (ALUs) 804. In some embodiments, the cores 802 may include a floating point logic unit (FPLU) 806 and / or a digital signal processing unit (DSPU) 808 in addition to or instead of the ALUs 804.
[0199] The processor 800 may include one or more registers 812 communicatively coupled to the cores 802. The registers 812 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 812 may be realized using static memory. The registers may provide data, instructions, and addresses to the cores 802.
[0200] In some embodiments, processor 800 may include one or more levels of cache memory 810 communicatively coupled to cores 802. Cache memory 810 may provide computer-readable instructions to cores 802 for execution. Cache memory 810 may provide data for processing by cores 802. In some embodiments, computer-readable instructions may be provided to cache memory 810 by local memory, e.g., local memory attached to external bus 816. Cache memory 810 may be implemented using any suitable cache memory type, such as, for example, static random access memory, dynamic random access memory, and / or any other suitable memory technology.
[0201] The processor 800 may include a controller 814, which may control input to the processor 800 from other processors and / or components included in the system and / or control output from the processor 800 to other processors and / or components included in the system. The controller 814 may control data paths within the ALU 804, the FPLU 806, and / or the DSPU 808. The controller 814 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of the controller 814 may be realized as standalone gates, FPGAs, ASICs, or any other suitable technology.
[0202] Registers 812 and cache 810 may communicate with controller 814 and core 802 via internal connections 820A, 820B, 820C, and 820D. The internal connections may be implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection technology.
[0203] Inputs and outputs of processor 800 are provided via bus 816, which may include one or more conductive lines. Bus 816 is communicatively coupled to one or more components of processor 800, such as controller 814, cache 810, and / or registers 812. Bus 816 may be coupled to one or more components of the system.
[0204] The bus 816 may be coupled to one or more external memories. The external memory may include read-only memory 832. The ROM 832 may be masked ROM, electrically programmable read-only memory (EPROM), or any other suitable technology. The external memory may include random access memory 833. The RAM 833 may be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory may include electrically erasable programmable read-only memory (EEPROM) 835. The external memory may include flash memory 834. The external memory may include a magnetic storage device such as a disk 836. In some embodiments, external memory may be included in the system.
[0205] The foregoing description (and claims / inventions) are directed to generating a respiratory waveform signal and adapting speech processing dependent on the respiratory waveform signal. However, it will be understood that generating a respiratory waveform signal is not limited to this application. The generated respiratory waveform signal can be provided to a determiner configured to determine a respiratory parameter / characteristic from the respiratory waveform signal. The respiratory parameter / characteristic can be indicative of a characteristic of the speaker's breathing. The respiratory parameter can be, for example, a respiratory rate, a respiratory depth, etc. As mentioned above, such respiratory parameters can be used for speech processing. However, they can also be used for many other purposes and, indeed, can be provided to a user via a suitable user interface. For example, the user may be a medical professional who can utilize the provided information for diagnostic purposes.
[0206] Thus, although the described approach includes an audio processor configured to apply audio processing to an audio signal, the audio processing relies on the respiration waveform signal, and therefore the respiration waveform signal can be used for other purposes.
[0207] The present invention may be implemented in any suitable form including hardware, software, firmware or any combination of these. The present invention may optionally be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the present invention may be physically, functionally, and logically implemented in any suitable way. Indeed, functionality may be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the present invention may be implemented in a single unit, or may be physically and functionally distributed between different units, circuits, and processors.
[0208] Although the present invention has been described in connection with several embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, while certain features may appear to be described in connection with particular embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0209] Furthermore, although individually listed, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Furthermore, although individual features may be included in different claims, these may be advantageously combined in some cases, and their inclusion in different claims does not imply that a combination of features is not feasible and / or advantageous. Furthermore, the inclusion of a feature in one category of claims does not imply limitation to this category, but rather indicates that the feature is equally applicable to other claim categories, where appropriate. Furthermore, the order of features in the claims does not imply a particular order in which the features must operate, and in particular the order of individual steps in method claims does not imply that the steps must be performed in this order. Rather, steps may be performed in any suitable order. Furthermore, a reference to the singular does not exclude a plurality; thus, references to "a," "an," "first," "second," etc. do not exclude a plurality. Reference signs in the claims are provided merely as a clarifying example and should not be construed as limiting the scope of the claims in any way.
[0210] Generally, examples of apparatus and methods for audio processing are illustrated by the following embodiments. Embodiment 1. An apparatus for speech processing of an audio signal, said apparatus comprising: an input (101) configured to receive an audio signal, said audio signal comprising a speech audio component of a speaker; a segmenter (105) configured to generate segments of the audio signal, each segment having a segment duration, consecutive segments having an inter-segment duration that is the time between consecutive segments, the segment duration exceeding the inter-segment duration such that consecutive segments overlap; a first generator (107) configured to generate a fragment of a respiration waveform signal for each segment of the audio signal; a second generator (109) configured to generate a respiratory waveform signal by combining fragments of the respiratory waveform signal; an audio processor (103) configured to apply audio processing to the audio signal, the audio processing being dependent on the respiration waveform signal. Embodiment 2. The apparatus of embodiment 1, wherein at least some of the fragments have a fragment duration that exceeds the inter-segment duration. Embodiment 3. The apparatus of embodiment 2, wherein the fragment duration is equal to the segment duration. Embodiment 4. An apparatus according to embodiment 2 or 3, wherein the second generator (109) is configured to combine fragments by applying a weighted combination to samples of different fragments at the same time point, the weights of the weighted combination being determined from a fragment window function. Embodiment 5. The apparatus of any preceding embodiment, wherein the segment duration exceeds the inter-segment duration by 50% or more. Embodiment 6. The apparatus of any preceding embodiment, wherein the segment of the audio signal includes at least a first portion of the audio signal that is also included in a previous segment of the audio signal, and a second portion of the audio signal that is also included in a subsequent segment of the audio signal. Embodiment 7. The apparatus of any preceding embodiment, wherein the sample rate of the respiration waveform signal is at least 10 times lower than the sample rate of the audio signal. Embodiment 8. The apparatus of any preceding embodiment, wherein the first generator (107) is configured to generate fragments having a lower sample rate than the segments of the audio signal. Embodiment 9. The apparatus of any preceding embodiment, wherein the segmenter (105) is configured to generate segments having a lower sample rate than the audio signal. Embodiment 10. The apparatus of any preceding embodiment, wherein the first generator (107) includes a trained artificial neural network having an input node for receiving a sample of a segment of an audio signal and an output node configured to provide a sample of a fragment of a respiration waveform signal for the segment of the audio signal. Embodiment 11. The apparatus of any preceding embodiment, wherein the first generator (107) includes a model for human respiration, and wherein the first generator is configured to generate fragments from the model and adapt parameters of the model according to the segments. Embodiment 12. The apparatus of any preceding embodiment, wherein the segmenter (105) is configured to apply a frequency transform to the audio signal to generate each segment as a matrix of time-frequency samples that includes samples in at least two different time intervals and two different frequency ranges. Embodiment 13. The device of any preceding embodiment, further comprising an adapter (113) configured to adapt at least one of an inter-segment duration and a segment duration in response to at least one of a characteristic of the audio signal and a characteristic of the respiration waveform signal. Embodiment 14. A method for processing an audio signal, comprising: receiving an audio signal, the audio signal including a speaker's voice audio component; generating segments of the audio signal, each segment having a segment duration, successive segments having an inter-segment duration that is the time between successive segments, the segment duration exceeding the inter-segment duration such that successive segments overlap; generating a fragment of a respiration waveform signal for each segment of the audio signal; 1. A method comprising the steps of combining fragments of a respiration waveform signal to generate a respiration waveform signal, and applying audio processing to an audio signal, wherein the audio processing is dependent on the respiration waveform signal.
Claims
1. 1. An apparatus for sound processing of an audio signal, the apparatus comprising: an input configured to receive the audio signal, the audio signal including a speech audio component of a speaker; a segmenter configured to generate segments of the audio signal, each segment having a segment duration, consecutive segments having an inter-segment duration that is the time between consecutive segments, the segment duration exceeding the inter-segment duration such that consecutive segments overlap; a first generator configured to generate, for each segment of the audio signal, fragments of a respiration waveform signal from the segment, wherein at least some fragments have fragment durations that exceed the inter-segment duration; and a second generator configured to combine fragments of the respiratory waveform signal to generate the respiratory waveform signal, the second generator configured to combine fragments of at least some of the fragments by applying a weighted combination to samples of different fragments at the same time point, the weights of the weighted combination being determined from a fragment window function; an audio processor configured to apply audio processing to the audio signal, the audio processing being dependent on the respiration waveform signal; and A device having:
2. The apparatus of claim 1 , wherein the segment duration exceeds the inter-segment duration by 50% or more.
3. 3. The apparatus of claim 1, wherein a segment of the audio signal includes at least a first portion of the audio signal that is also included in a previous segment of the audio signal and a second portion of the audio signal that is also included in a subsequent segment of the audio signal.
4. 4. The apparatus of claim 1, wherein the sample rate of the respiratory waveform signal is at least 10 times lower than the sample rate of the audio signal.
5. 5. The apparatus of claim 1, wherein the first generator is configured to generate the fragment to have a lower sample rate than the segment of the audio signal.
6. 6. The apparatus of claim 1, wherein the segmenter is configured to generate the segments to have a lower sample rate than the audio signal.
7. 7. The apparatus of claim 1, wherein the first generator comprises a trained artificial neural network having an input node for receiving a sample of a segment of the audio signal and an output node configured to provide a sample of a fragment of the respiration waveform signal for the segment of the audio signal.
8. 8. The apparatus of claim 1, wherein the first generator comprises a model for human respiration, and the first generator is configured to generate the fragments from the model and to adapt parameters of the model depending on the segment.
9. 9. The apparatus of claim 1, wherein the segmenter is configured to apply a frequency transform to the audio signal and generate each segment as a matrix of time-frequency samples comprising at least samples of two different time intervals and two different frequency ranges.
10. 10. The apparatus of claim 1, further comprising an adapter configured to adapt at least one of the inter-segment duration and the segment duration in response to at least one of characteristics of the audio signal and characteristics of the respiratory waveform signal.
11. 1. A method for processing an audio signal, the method comprising: receiving the audio signal, the audio signal including a speaker's voice audio component; generating segments of the audio signal, each segment having a segment duration, consecutive segments having an inter-segment duration that is the time between consecutive segments, the segment durations exceeding the inter-segment durations such that consecutive segments overlap; generating a fragment of a respiration waveform signal for each segment of the audio signal from each segment of the audio signal; generating the respiratory waveform signal by combining the fragments of the respiratory waveform signal, the combining being configured to combine fragments of at least some of the fragments by applying a weighted combination to samples of different fragments at the same time point, the weights of the weighted combination being determined from a fragment window function; applying audio processing to the audio signal, wherein the audio processing is dependent on the respiration waveform signal.
12. A computer program product which, when run on a computer, causes the computer to carry out the method of claim 11.