Speech processing of audio signals
By performing segmented processing on the audio signal and weighted combination of the respiratory waveform signal, and using artificial neural networks to adapt speech processing, the problem of insufficient performance of speech processing in actual conditions in existing technologies is solved, and more efficient and flexible speech enhancement and recognition are achieved.
Patent Information
- Application Number
- CN202480011729.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-09
- Filing Date
- 2024-02-01
- Publication Date
- 2025-09-19
AI Technical Summary
Existing speech processing algorithms cannot provide optimal performance when actual conditions differ from expected nominal conditions, and have problems such as insufficient flexibility, high complexity, large resource requirements, and heavy computational load. They perform particularly poorly when speaker attributes and activities change.
The audio signal is segmented to generate respiratory waveform signal segments, and speech processing is performed using weighted combination and artificial neural network. The speech processing is adapted to reflect the speaker's current breathing state, including segment processing with segment duration exceeding inter-segment duration and weighted combination based on respiratory waveform signal.
It improves the adaptability and flexibility of speech processing, reduces complexity and resource usage, improves the quality of speech enhancement, recognition and encoding, reduces edge effects and errors, and adapts to speech changes in different activities and speaker states.
Smart Images

Figure CN120677525A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to performing speech processing of audio signals, and in particular, but not exclusively, to automatic speech recognition or speech enhancement of audio signals used to capture a speaker (e.g., performing an activity such as an exercise activity). Background Art
[0002] Speech processing of audio signals is widely used in various practical applications and is becoming increasingly important and significant for many daily activities and devices.
[0003] For example, speech enhancement can be widely used to provide improved clarity and reproducibility of speech captured by audio signals (such as sounds captured in real-world environments). Speech encoding is also often performed from captured audio. For example, speech encoding of captured audio signals is an integral part of smartphones. Another speech processing application that has become increasingly frequent in recent years is speech recognition, for example, to provide a user interface to a device.
[0004] For example, a voice interface for a personal assistant or home speaker enables controlling media, navigation functions, tracking, accessing various information services, etc. while performing other activities (e.g., during exercise) simply by using the voice interface.
[0005] Voice interfaces based on automatic speech recognition (ASR) have also become very important in controlling various devices such as wearable, proximity and IoT consumer, health and fitness devices. Voice is the most natural interface to interface with devices during activities that involve a lot of user movement and change. As an example, US20090098981 A1 discloses a fitness device with a voice interface for communicating with a virtual fitness trainer to the user. US 7529670 B1 discloses a speech processing system in which the processing depends on the determined lung state of the speaker. The article by Nallanthighal et al., "Deep learning architecture for estimating respiratory signals and respiratory parameters from speech recordings", Neural Networks, Vol. 141, September 1, 2021, pp. 211-224, XP093054576, ISSN: 0893-6080, DOI: 10.1016 / j.neunet.2021.03.29 discloses a method for determining samples of respiratory signals from audio using a neural network.
[0006] However, while considerable effort has been invested in developing and optimizing speech processing algorithms for a variety of applications, and many very advantageous and efficient methods have been developed, they are often not optimal in all situations. For example, many speech processing operations are developed for specific nominal conditions, such as nominal speaker attributes, nominal acoustic environments, etc. Many speech processing applications may not provide optimal performance when actual conditions differ from the expected nominal conditions. Many speech processing algorithms may also be more complex or more resource-intensive than desired. Adjusting speech processing to try to compensate for different attributes often results in less flexible, more complex, and / or more resource-intensive implementations, which generally do not provide optimal performance.
[0007] Therefore, an improved method for speech processing would be advantageous. In particular, a method that allows for increased flexibility, improved adaptability, improved performance, improved quality of, for example, speech enhancement, encoding, and / or recognition, reduced complexity and / or resource usage, improved remote control of audio processing, improved adaptation to changes in speaker attributes and / or activity, reduced computational load, improved user experience, simplified implementation, and / or improved spatial audio experience would be advantageous. Summary of the Invention
[0008] Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages, singly or in any combination.
[0009] According to one aspect of the present invention, there is provided an apparatus for speech processing of an audio signal, the apparatus comprising: an input arranged to receive the audio signal, the audio signal comprising a speech audio component for a speaker; a segmenter arranged to generate segments of the audio signal, each segment having a segment duration and consecutive segments, the consecutive segments having an inter-segment duration, the inter-segment duration being the time between the consecutive segments, the segment duration exceeding the inter-segment duration such that the consecutive segments overlap; a first generator (107) arranged to generate, for each segment of the audio signal, a segment from the first generator; segments of the respiratory waveform signal of each segment of the audio signal, at least some of the segments having a segment duration exceeding the inter-segment duration; a second generator (109), the second generator (109) being arranged to generate the respiratory waveform signal by combining the segments of the respiratory waveform signal, the second generator (109) being arranged to combine segments of the at least some segments by applying a weighted combination to samples of different segments for the same moment, wherein the weights of the weighted combination are determined according to a segment window function; and a speech processor, the speech processor being arranged to apply speech processing to the audio signal, the speech processing being dependent on the respiratory waveform signal.
[0010] The method can provide improved speech processing in many embodiments and scenarios. For many signals and scenarios, the method can provide speech processing that more closely reflects changes in the properties and characteristics of the speaker's speech. For example, it can adapt to the speaker's level of fatigue or, for example, breathing difficulties. The method, particularly using segmentation / fragment-based processing, can provide facilitated adaptation and can particularly provide improved adaptive speech processing. The method can provide an efficient implementation and, in many embodiments, can allow for reduced complexity and / or resource usage.
[0011] This approach may generally provide a more accurate / improved respiratory waveform signal and, in particular, may in many scenarios reduce or mitigate edge effects and errors present in many segmentation-based processes.
[0012] Speech processing may include speech enhancement processing, speech recognition processing, and / or speech coding processing. The segment duration of a segment may be the duration of a time interval of the audio signal included in the segment. The inter-segment duration may be the duration / time offset / difference between subsequent segments / segment time intervals.
[0013] At least some of the segments have a segment duration that exceeds the inter-segment duration.
[0014] This can provide particularly advantageous operation and performance in many scenarios and applications.
[0015] In many embodiments, the duration of at least some segments exceeds the inter-segment duration by no less than 10%, 20%, 30%, 50%, or 100%.The segment duration of a segment may be the duration of a time interval of the respiratory waveform signal represented by the segment.
[0016] The second generator is arranged to combine the segments by applying a weighted combination to samples of different segments for the same time instant, wherein the weights of the weighted combination are determined according to the segment window function.
[0017] This can provide improved operation and performance in many embodiments. In many embodiments, this allows for facilitated, improved, and / or more efficient operation. The segment window function can be such that the weights at each sample instant have a constant combined value (and in particular, the sum of the weights of different samples / segments can be constant for different instants).
[0018] According to an optional feature of the invention, the segment duration exceeds the inter-segment duration by not less than 50%.
[0019] This may provide improved operation and performance in many embodiments.The segment duration may in many embodiments be twice as long as the inter-segment duration.
[0020] According to an optional feature of the invention, a segment of the audio signal comprises at least a first portion of the audio signal also included in a previous segment of the audio signal and a second portion of the audio signal also included in a subsequent segment of the audio signal.
[0021] This may provide improved operation and performance in many embodiments.
[0022] In many embodiments, one or more segments include samples of both previous and subsequent segments.
[0023] According to an optional feature of the invention, the sampling rate of the respiratory waveform signal is no less than 10 times lower than the sampling rate of the audio signal, or in some embodiments no less than 20 times, 50 times or 100 times.
[0024] This can provide improved operation and performance in many embodiments, and in particular can include reduced complexity and reduced resource usage. For example, in many embodiments, it allows real-time processing using devices with limited computing resources.
[0025] According to an optional feature of the invention, the first generator may be arranged to generate the segments to have a lower sampling rate than the segments of the audio signal.
[0026] This may provide improved operation and performance in many embodiments.
[0027] According to an optional feature of the invention, the segmenter is arranged to generate the segments to have a lower sampling rate than the audio signal.
[0028] This may provide improved operation and performance in many embodiments.
[0029] According to an optional feature of the invention, the first generator comprises a trained artificial neural network having input nodes for receiving samples of the segment of the audio signal and output nodes arranged to provide samples of the segments of the respiratory waveform signal for the segment of the audio signal.
[0030] The method may provide a particularly advantageous arrangement which, in many embodiments and scenarios, allows for the possibility of exploiting the boost and / or improvements of artificial neural networks when generating respiratory waveform signals for adapting and optimizing speech processing (including speech enhancement and / or recognition in general).
[0031] This approach may provide an efficient implementation and may allow for reduced complexity and / or resource usage in many embodiments.
[0032] An ANN is a trained artificial neural network.
[0033] The artificial neural network may be a trained artificial neural network trained using training data, the training data including training speech audio signals and training breathing waveform signals generated from measurements of breathing waveforms; the training employing a cost function that compares the training breathing waveform signals with breathing waveform signals generated by the artificial neural network for the training speech audio signals. The artificial neural network may be a trained artificial neural network trained using training data including training speech audio signals representing a range of relevant speakers in different state ranges and performing different activities.
[0034] The artificial neural network may be a trained artificial neural network trained by training data having training input data including a training speech audio signal, and using a cost function including a contribution indicative of a difference between a measured training breathing waveform signal and a breathing waveform signal generated by the artificial neural network in response to the training speech audio signal.
[0035] According to an optional feature of the invention, the first generator comprises a model for human breathing, and wherein the first generator is arranged to generate segments from the model and to adapt parameters of the model in response to the segmentation.
[0036] This may provide improved operation and performance in many embodiments.The model may receive as input a given segment of an audio signal and generate as output a segment of a breathing waveform signal for the given segment.
[0037] According to an optional feature of the invention, the segmenter is arranged to apply a frequency transform to the audio signal and generate each segment as a matrix of time-frequency samples comprising at least samples for two different time intervals and for two different frequency ranges.
[0038] This may provide improved operation and performance in many embodiments.The method may allow for facilitated operation and provide a more efficient representation of the properties of the audio signal, such that these properties are suitable for determining a respiration waveform signal.
[0039] According to an optional feature of the invention, the apparatus further comprises an adapter arranged to adapt at least one of the inter-segment duration and the segment duration in response to at least one of a property of the audio signal and a property of the respiratory waveform signal.
[0040] This may provide improved operation and performance in many embodiments.
[0041] In some embodiments, the speech processing comprises a speech recognition processing which generates a plurality of recognition term candidates from the audio signal; and the speech processor is arranged to select between the recognition term candidates in response to the breathing waveform signal.
[0042] This can provide advantageous speech recognition in many scenarios and can allow this to be implemented with low complexity and resource requirements. For example, it can allow the adaptation and consideration of respiratory waveform signals to be implemented as post-processing applied to the results of existing speech recognition algorithms. This approach can provide improved backward compatibility.
[0043] In some embodiments, the speech processor is arranged to determine a breathing rate estimate from the breathing waveform signal and to adapt the speech processing in dependence on the breathing rate.
[0044] This may provide improved operation and performance in many embodiments.The breathing rate parameter may be a particularly suitable parameter for adapting speech, as there may often be a high correlation between some speech properties and breathing rate.
[0045] In some embodiments, the speech processor is arranged to determine timing properties of an inspiratory interval from the respiratory waveform signal and to adapt the speech processing in dependence on the timing properties of the inspiratory interval.
[0046] This may provide improved operation and performance in many embodiments.The timing properties may typically be the duration, repetition time / frequency and / or time of inspiration / respiration at intervals.
[0047] In some embodiments, the speech processor is arranged to determine inspiration / inhalation time segments of the audio signal based on the timing properties and to attenuate the audio signal during the inspiration / inhalation time segments.
[0048] This may provide improved operation and performance in many embodiments.The attenuation may be a partial attenuation or a complete attenuation (muting) of the audio signal.
[0049] In some embodiments, the speech processor is arranged to determine inspiration / inhalation time segments of the audio signal based on the timing properties and to exclude the inspiration / inhalation time segments from the speech recognition processing applied to the audio signal.
[0050] This may provide improved operation and performance in many embodiments.
[0051] In some embodiments, the speech processor is arranged to determine a level of breathing sounds during inspiration / inhalation from the respiratory waveform signal and to adapt the speech processing in dependence on the level of breathing sounds during inspiration / inhalation.
[0052] This may provide improved operation and performance in many embodiments.
[0053] In some embodiments, the speech processor comprises a human breathing model for determining breathing properties of the speaker, and the speech processor is further arranged to adapt the speech processing in response to the breathing properties; and to adapt the human breathing model in accordance with the speaker data.
[0054] This may provide improved operation and / or performance in many embodiments.
[0055] According to one aspect of the present invention, a method is provided, comprising: receiving an audio signal, the audio signal comprising a speech audio component for a speaker; generating segments of the audio signal, each segment having a segment duration and a time interval between consecutive segments, the consecutive segments having an inter-segment duration, the inter-segment duration being the time between the consecutive segments, the segment duration exceeding the inter-segment duration such that the consecutive segments overlap; generating a segment of a breathing waveform signal for each segment of the audio signal according to each segment of the audio signal; generating the breathing waveform signal by combining the segments of the breathing waveform signal, the combining comprising combining segments of at least some of the segments by applying a weighted combination to samples of different segments for the same moment, wherein the weights of the weighted combination are determined according to a segment window function; and applying speech processing to the audio signal, the speech processing depending on the breathing waveform signal.
[0056] Speech processing may include speech enhancement processing.
[0057] The approach may provide improved speech enhancement and may allow the generation of improved speech signals in many scenarios.
[0058] Speech processing may include speech recognition processing.
[0059] The method can provide improved speech recognition. It can allow more accurate detection of words, terms, and sentences in many scenarios, and can particularly allow improved speech recognition for speakers in a variety of situations, states, and activities. For example, it can allow improved speech detection for people exercising.
[0060] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0062] Figure 1 shows some elements of an example of an audio device according to some embodiments of the present invention;
[0063] Figure 2An example of segmentation of an audio signal is shown;
[0064] Figure 3 An example of the structure of an artificial neural network is shown;
[0065] Figure 4 An example of a node of an artificial neural network is shown;
[0066] Figure 5 Some elements of an example of a training setup for an artificial neural network are shown;
[0067] Figure 6 An example of prediction error for a respiratory waveform signal is shown;
[0068] Figure 7 shows an example of a process for estimating a breathing waveform signal from an audio signal according to some embodiments of the present invention; and
[0069] Figure 8 Some elements of a possible arrangement of a processor for implementing elements of an apparatus according to some embodiments of the present invention are shown. DETAILED DESCRIPTION
[0070] Figure 1 Some elements of a speech processing apparatus according to some embodiments of the present invention are shown.The speech processing apparatus may be adapted to provide improved speech processing for speakers in many different environments and scenarios, such as speakers who are exercising.
[0071] The speech processing device comprises a receiver / input 101, which is arranged to receive an audio signal comprising a speech audio component. The audio signal may specifically be a microphone signal representing audio captured by a microphone. The audio signal may specifically be a speech signal, and in many embodiments may be a speech signal captured by a microphone arranged to capture the speech of a single person, such as a microphone of a headset, a smartphone, or a body-worn microphone. Thus, the audio signal will comprise a speech audio component (which will also be referred to as a speech signal) and possibly other sounds, such as background sounds from the environment.
[0072] The speech processing apparatus further comprises a speech processing 103, which is arranged to apply speech processing to the received audio signal (particularly the speech audio component of the audio signal). The speech processing may be speech enhancement processing, speech coding or speech recognition process.
[0073] In this method, speech processing circuitry 103 is configured to perform speech processing in response to a respiration waveform signal. The speech processing circuitry is configured to generate a respiration waveform signal from an audio signal. The respiration waveform signal indicates the volume of air in the speaker's lungs as a function of time. Therefore, it reflects the speaker's breathing and the flow of air into and out of the speaker's lungs.
[0074] The speech processing means is arranged to process the audio signal so as to extract information indicative of the speaker's breathing captured by the audio signal.The audio means may be arranged to receive samples of the audio input signal and generate a time series or parameters representing a time series of lung air volume measurements / estimates.
[0075] Therefore, the breathing waveform signal can be a time-varying signal that reflects changes in the speaker's breathing, and specifically can reflect changes in lung volume caused by the speaker's breathing. In some embodiments, the breathing waveform signal can directly represent the current lung volume. In other embodiments, the breathing waveform signal can, for example, represent changes in the current lung volume, such as when the breathing waveform signal represents inhaled / exhaled air flow (inhalation / exhalation).
[0076] Conventionally, a separate sensor device (such as a respiratory belt worn when the subject is speaking) can be used to capture the respiratory waveform signal. However, in the current method, even without such sensor data, the determiner can determine the respiratory waveform signal by deriving the respiratory waveform signal from the voice data. As a low-complexity example, the determiner can use a voice activity detection (SAD) algorithm to segment the voice data into voice segments and non-voice segments. Based on the knowledge that inhalation occurs during voice pauses, the determiner 105 can use the duration and frequency of the pauses as a proxy for respiratory behavior, and therefore determine the respiratory waveform signal accordingly. In many embodiments, a more accurate determination can be used, including the use of an artificial neural network, as will be described later. In the event that an error in the estimation of, for example, respiratory rate may result, this can, for example, compensate for or take into account the presence of more pauses in the voice than the true inhalation / breathing.
[0077] The speech processor 103 is arranged to adapt the speech processing in response to the respiratory waveform signal. Thus, rather than performing predetermined speech processing based solely on nominal or expected characteristics, Figure 1 The speech processing means is arranged to adapt the speech processing to reflect the current breathing of the speaker.
[0078] This method reflects the inventors' recognition that a speaker's breathing affects speech, and that not only can the time-varying characteristics produced by breathing be determined from the captured audio, but that these time-varying characteristics can also be used to adapt speech processing of the same audio signal to provide improved performance, in particular performance that can better adapt to different scenarios and user behaviors.
[0079] The specific speech processing and adaptation performed may depend on the individual embodiment. However, in many embodiments, the method can be used to adapt the operation to reflect the different breathing patterns that a person may exhibit due to performing different activities. For example, when a person exercises, breathing may become strained, resulting in very interrupted and slowed speech. Figure 1 The speech processing device can automatically adapt to optimize for slow speech with long silence intervals.
[0080] The approach may provide significantly improved speech processing in many scenarios, and may, for example, provide particularly advantageous operation for speakers who engage in strenuous activities (such as exercise) or who have breathing or speaking difficulties.
[0081] The method may be particularly advantageous for automatic speech recognition, for example, to provide a voice interface. For example, a voice interface for a personal assistant enables control of media, navigation features, exercise tracking, and various other information services during exercise. However, typically speech recognition is optimized for normal speech at rest and may be personalized. During and after strenuous exercise, the demand for oxygen increases significantly, resulting in changes in breathing that affect speech. This in turn increases the word error rate (WUR) and, therefore, the usability and user experience provided by the voice interface. The speech processing device may detect the user's inhalation / breathing pattern and adapt the speech recognition to take into account the changed speech characteristics.
[0082] Speech recognition algorithms are optimized for user groups, and the voice prompts that initiate interactions are often optimized for individual users.
[0083] In more detail, when a user is under physical stress, such as during exercise or physically demanding work, more breathing effort is required, which tends to affect the user's speech. During intense workouts, such as during interval training, speech can be severely affected. For example, typically, a user may need to take a breath after every 1-3 words, or may frequently swallow the ends of words. This condition of rapid breathing is known as dyspnea, and speech in this condition will hereinafter be referred to as dysphagia speech. Dyspnea speech patterns begin to appear under even minor physical stress, although human listeners usually subconsciously adapt to this condition and have no problem understanding such speech.
[0084] However, automatic speech recognition algorithms that have been trained using speech from resting people have serious difficulties accurately understanding breathless speech. In fact, even detecting the intended voice prompt or trigger word may be difficult. During voice commands and conversations with automatic speech recognition-based interfaces, increased breathing rate, rapid and audible inhalation / exhalation, and swallowing of word ends can significantly degrade performance. This can make voice interfaces difficult and frustrating to use.
[0085] exist Figure 1 In the method of the speech processing device, such conditions can not only be detected, but the system can also automatically adapt the speech processing to the current conditions.
[0086] As a low complexity example, the speech processor 103 can be arranged to select between a plurality of different speech processing algorithms / processes. These speech processing algorithms / processes may have been optimized for a given breathing scenario. For example, one process may have been optimized / trained for a resting breathing pattern. Another process may have been optimized / trained for a more tense but still controllable breathing pattern. Yet another process may have been optimized / trained for a most painful breathing pattern. Based on the breathing waveform signal, the signal processor 103 can select the speech processing algorithm / process that most closely matches the current breathing pattern indicated by the breathing waveform signal, and then apply that process to the received audio signal.
[0087] Therefore, in some embodiments, the speech processor 103 can be arranged to switch between different speech processing algorithms / functions based on the respiratory waveform signal. The speech processor 103 can analyze the respiratory waveform signal to determine a parameter value, such as respiratory rate. Each processing algorithm / function can be associated with a range of respiratory rates, and the speech processor 103 can select an algorithm that includes the rate determined based on the respiratory waveform signal and apply it to the audio signal.
[0088] In some embodiments, the individual algorithms / functions that can be selected between can include, for example, methods that use substantially the same signal processing but with parameters optimized for a given condition indicated by the respiratory waveform signal. For example, the same speech processing can be used, but the operating parameters are different for different respiratory rates. The operating parameters may have been determined, for example, through training or based on separate manual optimization using audio captured from a person (e.g., whose respiratory characteristics have been measured by other means).
[0089] In some embodiments, the individual algorithms / functions that can be selected between may include methods that are fundamentally different and, for example, use completely different approaches and principles. For example, different speech recognition methods may be used for a person who is breathing at rest and therefore speaking in rapid succession, and for a person who is breathing extremely hard (who can primarily speak single words interrupted by a lot of breathing noise).
[0090] As an example, in many embodiments the speech processor 103 may be arranged to determine timing properties for inspiration / respiration from the respiratory waveform signal.
[0091] Breathing comprises a series of alternating inhalations (drawing air into the lungs) and exhalations (exhaling air from the lungs), and the speech processor 103 can be arranged to evaluate the respiratory waveform signal to determine the timing properties of the intervals in which the speaker inhales / breaths. Specifically, the speech processor 103 can be arranged to evaluate the audio signal / respiratory waveform signal to determine the frequency of the inhalation / exhalation intervals, the duration of the inhalation / exhalation intervals and / or the time when these intervals start and / or stop. Such parameters can be determined, for example, by a peak picking algorithm that locates the local minima and maxima in the respiratory waveform signal that correspond to the beginning of the inhalation (inhalation) and exhalation (exhalation) phases. Respiratory rate can be determined based on the time between two consecutive inhalations or exhalations.
[0092] In some embodiments, the speech processor 103 can, for example, select between different algorithms or adapt parameters of speech processing based on such timing indications. For example, if the duration of inspiration / inhalation increases (e.g., relative to exhalation / exhalation), it can indicate that the person is taking a deep breath, which can reflect a different speech pattern than during normal breathing.
[0093] However, in some embodiments, the speech processor 103 may be specifically configured to modify speech processing for inhalation / inhalation intervals. In particular, in many embodiments, the speech processor 103 may be configured to suppress or stop speech processing during inhalation / inhalation intervals. For example, speech processing in the form of speech recognition may simply ignore all inhalation / inhalation intervals and only apply to non-inhalation / inhalation intervals.
[0094] The speech processor 103 can be configured to determine inhalation / exhalation time segments of the audio signal. These time segments can therefore reflect time intervals in the audio signal during which the speaker is estimated to be inhaling. In some embodiments, the speech processor 103 can be configured to attenuate the audio signal during the inhalation / exhalation time segments. For example, a speech enhancement algorithm configured to reduce noise by emphasizing the speech component of the audio signal can attenuate the audio signal during the inhalation / exhalation time intervals, including, in many embodiments, completely attenuating the signal.
[0095] Such a method can tend to provide significantly improved performance and can allow speech processing to adapt to and be specifically applied to the part of the audio signal where speech is present, while allowing different methods, specifically different attenuations, for the part of the audio signal where speech is unlikely to be present. The method can, for example, significantly reduce breathing sounds by eliminating heavy breathing sounds during the inhalation / inhalation phase of the respiratory wave signal. This can significantly improve perceived speech clarity and can accordingly provide perceived speech. In certain embodiments, it can be allowed to remove the part of the voice data occurring during the inhalation / inhalation period (estimated according to the respiratory wave) before automatic speech recognition, thereby promoting this and reducing the risk of erroneous detection of words. In a speech coding embodiment, for example, coding efficiency can be increased because speech coding can not be performed during the inhalation / inhalation time interval. For example, an indication of a silent or non-speech time interval can simply be inserted into the coded data stream to indicate that speech data is not provided within the interval.
[0096] Figure 1 The speech processing apparatus of the present invention uses a specific segmentation method to generate a respiratory waveform signal based on an input audio signal. Input 101 is coupled to a segmenter 105, which is arranged to generate segments of the audio signal. Segmenter 105 can be specifically arranged to generate segments periodically. However, the segments are generated to have a duration that exceeds the duration between segments, and thus at least partially overlap segments.
[0097] For example, Figure 2 As shown, audio signal 201 can be divided into overlapping segments 203, each of which corresponds to an audio signal in time intervals f1-f5. Each segment has a given duration, i.e., each segment has a segment duration. The segment duration is the duration of the time interval of the audio signal represented by the segment. Segments are repeatedly generated to represent the audio signal over time intervals that are greater than (usually much greater than) the segment duration. The inter-segment duration (which is the duration of the time interval between consecutive segments) is less than the segment duration, causing the segments to overlap. Therefore, at least some portions / time intervals of the audio signal (specifically, some samples) are included in two (or in some cases, more than two) segments. The inter-segment duration can be determined, for example, as the time between the start and / or end times of consecutive segments (or indeed any other matching relative time in consecutive segments). Typically, the inter-segment duration can be determined as the time / duration between the (temporal) midpoints of the consecutive segments (which can be considered as the nearest sample instant, or, for example, if the segment contains an even number of samples, the time between sample instants).
[0098] Typically, the segments are generated periodically, and thus the inter-segment duration is constant for at least a plurality of segments (e.g., not less than five, ten, or more segments). Similarly, typically, the segment duration is also constant for at least a plurality of segments (e.g., not less than five, ten, or more segments). However, it will be appreciated that in some embodiments, the inter-segment duration and / or the segment duration may be variable, and in particular, in some embodiments, the speech processing apparatus may be arranged to adapt one or both of these values.
[0099] In many embodiments, there may be a fixed relationship / ratio between the segment duration and the inter-segment duration. In many embodiments, the ratio may be an integer ratio. For example, in Figure 2 In the approach shown, the segment duration is twice the inter-segment duration. This results in a scenario where there is a 50% overlap between two consecutive segments and each part / sample of the audio signal is included in both segments.
[0100] The overlapping segments of the audio signal are fed to a first generator 107, which is arranged to generate a segment of the breathing waveform signal for each segment of the audio signal. The first generator 107 is arranged to process each segment separately and generate a segment of the breathing waveform signal from each segment. The first generator 107 can accordingly perform segment-based processing, wherein each segment of the audio signal is processed separately to generate a segment of the breathing waveform signal. Thus, for each segment of the audio signal, a segment of the breathing waveform signal can be generated from the segment of the audio signal.
[0101] In some embodiments, each segment of the audio signal may be processed to generate a segment of the breathing waveform signal. In some embodiments, when generating a segment of the breathing waveform signal, only the audio signal in one segment may be considered / processed. In many embodiments, each segment of the breathing waveform signal may include no fewer than 2, 5, 10, 20, or even 50 samples.
[0102] Each segment may be generated to have a duration that matches the duration of the inter-segment duration. For example, a segment may be generated to cover a time interval having the same duration as the inter-segment duration and centered at the midpoint of the segmented time interval. This may result in subsequent segments not overlapping but contiguous, such that a continuous respiratory waveform signal is generated. However, in Figure 2In the method of the apparatus of claim 1, the segments have a duration that exceeds the segment duration, and the segments may overlap. The segments may still be timed to be centered around each segment time interval, i.e., the midpoints may coincide. In many embodiments, each segment may be generated to cover a time interval that matches the segment time interval of the segment from which it was generated. Thus, in many embodiments, the segment duration may be the same as the segment duration, and indeed in many embodiments, the time interval within which a segment is generated is the same as the time interval of the segment from which it was generated. For the sake of completeness, it should also be noted that, although not Figure 2 In the case of a device with a segment duration, the segment duration may even be shorter than the inter-segment duration. This may result in a breathing waveform signal with gaps, but this may be acceptable in some cases. In particular, the speech processing circuit 103 may be arranged to determine parameters of the breathing waveform signal (e.g., word rate) based only on the segments, and then apply these parameters to the entire segment.
[0103] The terms fragment and segment are used interchangeably.
[0104] As an example, the first generator 107 may include a model for human breathing that generates a breathing waveform signal based on a plurality of model parameters. Some model parameters may be set to reflect a specific speaker, for example, based on manually provided data or data previously determined, for example, based on an audio signal (or based on a previous conversation). Such model parameters may, for example, reflect whether the speaker is female or male, a child, an adult, or an elderly person, etc. Model parameters may also reflect, for example, normal speaking pitch, average word rate, etc.
[0105] The model can be arranged to generate segments of a breathing waveform signal of a predetermined size, i.e., specifically, it can generate segments of a given duration each time the model is evaluated. The model can be evaluated for each segment of the audio signal based on the model parameters. Furthermore, before evaluating the model, one or more parameters can be determined specifically for the current segment based on the evaluation segment itself. For example, the current pitch rate, word length, voice activity level, etc. can be determined and used as input to the model that generates the breathing waveform signal.
[0106] In one embodiment, the model can be based on finding pauses in each speech input segment using a speech activity detection (SAD) algorithm. For each time sample of the audio data, the SAD gives an indication of whether the sample data is speech audio data or non-speech audio data. Based on the SAD signal, we can select sequences of speech pauses that are, for example, longer than 200ms. Next, assuming that the pauses represent inspiratory events in breathing, respiratory waveform segments are formed such that each end of the pause is selected as a local minimum point and each end of the pause is a local maximum point. All intermediate points can then be formed by linear (or polynomial) interpolation to produce a "sawtooth" wave pattern corresponding to the respiratory waveform segment.
[0107] First generator 107 is coupled to second generator 107, which is configured to generate a respiratory waveform signal by combining segments of the respiratory waveform signal. In embodiments where each segment has a duration corresponding to the segmented time interval within which it was generated, combining can be performed simply by concatenating the segments into a signal of extended duration. As will be described in greater detail later, for overlapping segments, combining can, for example, comprise a weighted summation of the signal levels / samples of the respiratory waveform signal generated at the same instant.
[0108] The respiration waveform signal is fed together with the audio signal to the voice processing circuit 103 , and the voice processing circuit 103 proceeds to perform voice processing based on the generated respiration waveform signal.
[0109] In many embodiments, the determination of segments of the respiratory waveform signal based on segmentation may be performed by a suitably trained artificial neural network.
[0110] In many embodiments, the first generator 107 comprises a trained artificial neural network having input nodes for receiving segmented samples of the audio signal and output nodes arranged to provide segmented samples of the segmented respiratory waveform signal of the audio signal.
[0111] In many embodiments, the determiner 105 may include a trained artificial neural network that is arranged to determine samples of the respiratory waveform signal based on samples of the audio signal. The artificial neural network may include an input node that receives samples of the audio signal and an output node that generates samples of the respiratory waveform signal. For example, the artificial neural network may have been trained using speech data and sensor data representing lung air volume measurements during speech.
[0112] In this example, the artificial neural network can have input nodes that receive samples of a segment of the audio signal, and the number of input nodes can match the number of samples in the signal. The artificial neural network can then process the segment to generate output samples, which are samples of the segment of the respiratory waveform signal. The number of output nodes can generally be the same as the number of samples in the segment, with each output node providing a sample of the segment of the respiratory waveform signal. Thus, the artificial neural network can be executed segment by segment to generate segments of the respiratory waveform signal.
[0113] An artificial neural network as used in the described functions may be a network of nodes arranged in layers and each node holding a node value. Figure 3 An example of a portion of an artificial neural network is shown.
[0114] The node value of a given node can be calculated to include contributions from some or generally all nodes in the previous layer of the artificial neural network. Specifically, the node value of a node can be calculated as a weighted sum of the node values output by all nodes in the previous layer. Typically, a bias can be added, and the result can be subjected to an activation function. The activation function provides the necessary components of each neuron by generally providing nonlinearity. This nonlinearity and activation function have a significant impact on the learning and adaptation process of the neural network. Therefore, the node value is generated based on the node values of the previous layer.
[0115] The artificial neural network may specifically include an input layer 301, which includes a plurality of nodes that receive input data values of the artificial neural network. Therefore, the node values of the nodes of the input layer can generally be directly the input data values of the artificial neural network, and therefore, calculations may not be performed based on other node values.
[0116] The artificial neural network may also include zero, one or more hidden layers 303 or processing layers. For each of such layers, a node value is typically generated based on the node values of the nodes of the previous layer, specifically based on a weighted combination and an added bias, followed by an activation function (such as a sigmoid, ReLU or Tanh function). Specifically, Figure 3 As shown, each node (which can also be called a neuron) can receive input values (from nodes in the previous layer) and thus calculate the node value based on these values. Typically, this involves first generating a value that is a linear combination of the input values, where each of these input values is weighted by a weight:
[0117]
[0118] Where w refers to weights, x refers to nodes of the previous layer, and n refers to the index of different nodes of the previous layer.
[0119] An activation function can then be applied to the resulting combination. For example, the node value 1 can be determined as:
[0120] l=f(k)
[0121] The function may be, for example, the rectified linear unit function described in Proceedings of the 14th International Conference on Artificial Intelligence and Statistics (PMLR 15:315-323, 2011) by Xavier Glorot, Antoine Bordes, and Yoshua Bengio:
[0122] f(k)=ReLU(k)=max(0,k)
[0123] Other commonly used functions include sigmoid function or tanh function. In many embodiments, multiple functions can be used to calculate node output or value. For example, an activation function such as the following can be used to combine both ReLU function and Sigmoid function:
[0124] f(k)=ReLU(k)+σ(k)
[0125] Such an operation can be performed by every node of the artificial neural network (except the usual input nodes).
[0126] The artificial neural network also includes an output layer 305, which provides the output from the artificial neural network. That is, the output data of the artificial neural network are the node values of the output layer. For hidden / processing layers, the output node values are generated by a function of the node values of the previous layer. However, in contrast to hidden / processing layers, where the node values are generally not accessible or further used, the node values of the output layer are accessible and provide the results of the operation of the artificial neural network.
[0127] A number of different network structures and toolboxes have been developed for artificial neural networks, and in many embodiments, artificial neural networks can be based on adapting and customizing such networks. One example of a network architecture that may be suitable for the above applications is the long short-term memory (LSTM) [by Sepp Hochsreiter, in Hochreiter, Sepp, and Jürgen Schmidhuber, "Long Short-Term Memory." Neural Computation 9.8 (1997): 1735-1780.
[0128] LSTM is an architecture for classification and regression of time-domain signals using recurrent causal or bidirectional evaluation and has been successfully applied to audio signals.
[0129]
[0130] Where * represents matrix multiplication, o represents Hadamard product, x is the input vector, h t-1 represents the output vector of the previous time step, W, V, U are the network weights, and b is the bias vector, σ ɡ Usually a sigmoid function or other compressed nonlinear function, and c t-1 is the previous state vector, and the output f t is the activation vector corresponding to the forget gate of the LSTM network, which will subsequently affect the current state vector c t .
[0131] In theory, classical (or "vanilla") artificial neural networks can track arbitrary long-term dependencies in input sequences. The problem with vanilla artificial neural networks is computational (or practical) in nature: when training a vanilla artificial neural network using backpropagation, the backpropagated long-term gradients can "vanish" (i.e., they can tend to zero) or "explode" (i.e., they can tend to infinity) because the process involves computations with finite-precision numbers. Artificial neural networks using LSTM cells partially solve the vanishing gradient problem because LSTM cells allow the gradients to also remain constant. However, LSTM networks can still suffer from the exploding gradient problem.
[0132] In some cases, the artificial neural network can be further configured to include additional contributions that allow the artificial neural network to be dynamically adapted or customized for specific desired properties or characteristics of the generated output. For example, a set of values can be provided to adapt the artificial neural network. These values can be included by providing contributions to certain nodes of the artificial neural network. These nodes can specifically be input nodes, but can generally be nodes of hidden layers or processing layers. Such adaptation values can, for example, be weighted and added as contributions to the weighted sum / correlation value of a given node.
[0133] The above description relates to a neural network approach that may be applicable to many embodiments and implementations. However, it should be understood that many other types and structures of neural networks may be used. Indeed, many different methods for generating neural networks have been and are being developed, including neural networks that use complex structures and processes different from those described above. The method is not limited to any specific neural network approach, and any suitable approach may be used without departing from the present invention.
[0134] Artificial neural networks are adapted to specific purposes through a training process that is used to adapt / tune / modify the weights and other parameters (e.g., bias) of the artificial neural network. It should be understood that many different training processes and algorithms are known for training artificial neural networks. Typically, training is based on a large training set, in which examples of a large amount of input data are provided to the network. In addition, the output of the artificial neural network is typically compared (directly or indirectly) with an expected or ideal result. A cost function can be generated to reflect the desired outcome of the training process. In a typical scenario known as supervised learning, a cost function typically represents the distance between the prediction of specific input data and the ground truth. Based on the cost function, the weights can be changed, and by re-iterating the process for the modified weights, the artificial neural network can be adapted to a state where the cost function is minimized.
[0135] In more detail, during the training step, a neural network can have two different information flows, from input to output (forward pass) and from output to input (backward pass). In the forward pass, as described above, the data is processed by the neural network, while in the backward pass, the weights are updated to minimize the cost function. Typically, this backward propagation follows the gradient direction of the cost function landscape. In other words, by comparing the predicted output with the ground truth of a batch of data inputs, the direction in which the cost function is minimized and propagated backward can be estimated by updating the weights accordingly. Other methods for training artificial neural networks include, for example, the Levenberg-Marquardt algorithm, the conjugate gradient method, and Newton's method.
[0136] In the present case, training may specifically include a training set that may include a large number of pairs of captured audio signals and measured respiratory waveform signals. The measured respiratory waveform signals may be generated from sensor signals of a sensor that is arranged to measure a property that depends on lung air volume. Thus, training of the artificial neural network may be performed using a training set that includes linked audio / speech data / signals and sensor data / signals representing lung air volume measurements during speech.
[0137] In some embodiments, the training data can be audio signals in time segments corresponding to the processing time intervals of the trained artificial neural network. For example, the number of samples in the training audio signal can correspond to the number of samples corresponding to the input nodes of the trained artificial neural network. Each training example can therefore correspond to one operation of the trained artificial neural network. However, typically, a batch of training examples is considered for each step to speed up the training process. In addition, many upgrades to gradient descent are possible to accelerate convergence or avoid local minima in the cost function landscape.
[0138] Figure 5An example of how training data can be generated through a dedicated test is shown. During time intervals when a person is speaking, a microphone 503 can capture a speech audio signal 501. The resulting captured audio signal can be fed as input data to an artificial neural network 505 (suitable processing, such as amplification, digitization, and filtering, can be performed before generating a test audio signal fed to the artificial neural network). In addition, a respiratory waveform signal 507 is determined as a sensor signal from a suitable lung volume sensor 509, for example, located on the user. For example, sensor 509 can be a respiratory inductance plethysmography (RIP) sensor or a nasal cannula flow rate sensor.
[0139] A large number of such measurements can be made to generate a large number of pairs of training audio signals and breathing waveform signals. These signals can then be applied and used to train an artificial neural network, where the cost function is determined as the difference between the breathing waveform signal generated by the artificial neural network for the training audio signal and the measured training breathing waveform signal. The artificial neural network can then be adapted based on such cost functions known in the art. For example, a cost value can be determined for each training audio signal and / or for a combined set of training downmix audio signals (e.g., determining an average cost value for the training set). Typically, the cost function will include at least one component that reflects how close the generated signal is to a reference signal, the so-called reconstruction error. In some embodiments, the cost function will include at least one component that reflects how close the generated signal is to a reference signal of the perceptual perspective. In Figure 5 In the example of , the reference signal may specifically be a measured respiratory waveform signal.
[0140] The described method can provide particularly advantageous performance in many scenarios and applications. It can generally allow for improved speech processing based on a more accurate respiratory waveform signal estimated from the audio signal. Segmentation-based processing can be highly advantageous because it allows for efficient and practical processing and implementation methods. However, segmentation operations are also generally disadvantageous and associated with defects and problems. In particular, they tend to provide segmentation edge effects, and the inventors have particularly recognized that determining the respiratory waveform signal from the audio signal by segmentation processing can result in many edge effects. The inventors have also recognized that in many scenarios, improved performance can be achieved by extending the segmentation duration and by generating segments based on speech information and data / segments that extend beyond the interval between segments (i.e., based on overlapping segments).
[0141] Figure 6 An example of the distribution of reconstruction errors within a 128-sample respiratory waveform signal reconstruction window accumulated over thousands of frames is shown. The error is clearly smallest in the middle of the frame and increases towards the beginning and end of the segment. Figure 1In the method of the apparatus, the reconstruction may be focused on a central region having a lower area and / or multiple segments / fragments may be considered for reconstruction towards the edges.
[0142] In many embodiments, the segment duration may exceed the inter-segment duration by a substantial amount, and in many embodiments, by no less than 25%, 50%, 75%, or 100%. In particular, in e.g. Figure 2 In the described example, the segment duration is twice the inter-segment duration (100% longer), resulting in each sample or moment of the audio signal being included in two segments and allowing for particularly advantageous operation that is highly suitable for efficient implementation.
[0143] In many embodiments, a segment may include a portion of the audio signal that is also included in a previous segment, and may also include another portion of the audio signal that is included in a subsequent segment of the audio signal. Figure 2 , segment f2 includes a first portion also included in the preceding segment f1 and a second portion also included in the subsequent segment f3. In fact, in this example, the first half of segment f2 is also included in segment f1, and the second half of segment f2 is also included in segment f3.
[0144] In many embodiments, such overlapping segmentation and including at least some portions of the audio signal in multiple segments can provide particularly advantageous performance. This can significantly reduce segment edge effects for each segment while allowing the segments to be continuous with each other without any gaps, i.e., this can particularly allow for the generation of a continuous breathing waveform signal while mitigating edge effects of the breathing waveform signal.
[0145] In many embodiments, the segments can be generated to at least partially overlap. In some cases, the combination performed by the second generator 107 can be a simple combination, such as concatenating sequential segments into a respiratory waveform signal, possibly including selective combination with respect to overlapping portions (e.g., selecting the respiratory waveform signal sample closest to the midpoint of the sample).
[0146] However, in many embodiments, the combination can be a weighted combination, where the sample at a given moment is generated as a weighted combination of samples at that moment of different segments. In particular, for two overlapping segments, the sample of the respiratory waveform signal at that moment can be generated by weighting the two samples of the two segments at a given moment and adding them together.
[0147] The weight factor of each sample may be determined according to the time difference between the instant of the sample and the midpoint of the segment, and specifically determined as a monotonically decreasing function of the difference.
[0148] In many embodiments, the weight values can be determined to reflect a segment window function. For example, in some embodiments, each sample of a segment can be multiplied by a weight factor that is a function of the sampling time within the segment. The segment window function can be, for example, a squared sine window function, a Hanning window function, or the like.
[0149] In most embodiments, the weights can be determined such that the total / combined / accumulated weight factor for a given moment in time (considering all segments contributing to that moment in time) is constant. For example, in many embodiments, the weight factors may sum to 1 for all moments in time. This approach can generally be achieved by selecting an appropriate window for a given segment size and overlap. For example, for the specific example provided previously, a squared sine window function applied to a segment will inherently result in a constant combined weight factor being applied to the segment.
[0150] After scaling by the window function weighting factor, the second generator 107 may simply add the results together.
[0151] In many embodiments, the speech processing device is arranged to generate the breathing waveform signal to have a much lower sampling rate than the audio signal. In fact, in many embodiments, the speech processing device can generate the breathing waveform signal to have a sampling rate that is no less than 10, 50, or even 100 times lower than the sampling rate of the audio signal. For example, the sampling rate of the audio signal can typically be (no less than) 16kHz, 32kHz, or 64kHz, while the breathing waveform signal has a sampling rate of no more than 50Hz, 100Hz, 200Hz, or 500Hz.
[0152] In many embodiments, the first generator 107 and / or the segmenter 105 may be arranged to generate segments having a lower sampling rate than the audio signal. Thus, in many embodiments, the speech processing apparatus may be arranged to decimate the audio signal by a factor of no less than 10 and typically greater than 100 before generating the respiratory waveform signal.
[0153] Such decimation can result in substantially facilitated operation and can generally allow for a significant reduction in complexity and required resource usage. In particular, the significantly reduced number of samples for each segment of a given time interval can result in facilitated processing, for example to extract model parameters that can be used to adapt a model used to generate a respiratory waveform signal. In embodiments based on the use of artificial neural networks, the significant reduction in samples in a segment can allow for a significant reduction in complexity and resource usage. In particular, the number of input nodes for a given segment duration can be reduced by a multiple corresponding to the decimation factor. This can further reduce the number of nodes and computations required to perform artificial neural network operations, resulting in a significant reduction in computational resources.
[0154] This approach may reflect the inventors' recognition that the dynamics of respiratory function are much slower than those of audio and speech, and can be represented at a much lower sampling rate. This further reflects the counterintuitive insight that not only does this allow the generated respiratory waveform signal to be decimated, but the entire operation can be performed in the decimated domain while still retaining the information required to generate the respiratory waveform signal.
[0155] In some embodiments, decimation may be performed as part of the process of generating the respiratory waveform signal. In some embodiments, the first generator 107 may be arranged to generate segments having a lower sampling rate than the segments of the audio signal. Thus, segments having a higher sampling rate may be provided to the first generator 107, which may then directly generate the respiratory waveform signal at a much lower sampling rate.
[0156] For example, samples of a segment at a higher sampling rate can be fed to an artificial neural network, i.e., the artificial neural network can have a number of input nodes corresponding to the number of samples in the segment. It can then generate a segment of the respiratory waveform signal based on this, which includes samples at a much lower sampling rate. Therefore, the artificial neural network can have a number of output nodes corresponding to the samples of the segment (at a lower sampling rate). Therefore, the artificial neural network can have far fewer output nodes than input nodes. Such an approach can allow for computationally efficient operation, but in some cases provides an improved determination of the respiratory waveform signal because more high-frequency information can be taken into account.
[0157] In many embodiments, the segment samples fed to the first generator 107 may be in the form of time domain samples. However, in some embodiments, the segment samples may be provided in the frequency domain, and in particular, the segmenter 107 may apply a frequency transform to generate a frequency representation of the signal. In many embodiments, each segment may be represented by a set of frequency samples, such as the output of an FFT operation or a (e.g., QMF) filter bank.
[0158] In many embodiments, each segment can advantageously be represented by a set of samples comprising both frequency samples and time samples. Specifically, each segment can be represented by a time-frequency sample matrix comprising at least samples for two different time intervals and for two different frequency ranges.
[0159] For example, an FFT or QMF filter bank can be applied to the input signal with a given block size that is smaller than the full segment. For example, a given segment can be divided into multiple (N) blocks, where each block includes M samples. A frequency transform can be applied to each block, resulting in, for example, M frequency bins generated for each block. Repeated application of the frequency transform results in N samples being generated for each bin, i.e., N time domain samples are generated for each frequency bin. The resulting samples can be used to represent the segment by an NxM sample matrix, where there are N time domain samples for each of the M frequency bins.
[0160] Thus, in many embodiments, each segment may represent a matrix of time-frequency samples comprising at least samples for two different time intervals and for two different frequency ranges. For each frequency bin, there may typically be multiple sample values at different instants corresponding to each frequency transform block.
[0161] In many embodiments, such a matrix representation of the segmentation may be fed to, for example, an artificial neural network, ie the artificial neural network may comprise an input node for receiving samples of blocks comprising both frequency separated samples and time separated samples.
[0162] In many embodiments, the matrix representation may be at a reduced sampling rate, as previously described. Specifically, in many embodiments, the frequency transform may be performed after applying decimation to the time domain representation of the audio signal, as previously described.
[0163] In many embodiments, the representation of the segments by a combined set of time-frequency samples can provide improved determination of the respiratory waveform signal and can facilitate and improve implementation, for example, by allowing frequency transforms with significantly reduced complexity and resource usage. In particular, it has been found to provide improved operation and performance when used with artificial neural networks.
[0164] A specific example of determining a respiratory waveform signal from an audio signal is described below. This method includes most of the previously described features and methods, but it should be understood that this does not imply that these features must be applied together. Instead, they can be applied individually / separately in different embodiments, and in fact, in many embodiments, only a subset of these features is implemented / included.
[0165] In this example, the speech processing device can predict segments of the respiratory waveform signal b(k) (k=0,..,B) for segments of the audio signal x(t) (t=0,..,N-1), where t and k are time indices typically at different sampling rates, the length of the segment is N samples, and the length of the generated segment is B samples.
[0166] The operation can be performed in frames / segments of N samples (i.e., with a segment duration of N samples) with a step size of S samples (i.e., with an inter-segment duration of S samples). Typically, S = N / 2. We use the subscripts α and β to denote the signal and frame sizes in audio and VRB signals, respectively.
[0167] Then by x p =(x(τ),τ=pS α ,…,pS α +N α -1) Given the pth audio vector. The estimation model can generate N β The respiratory waveform signal of samples is segmented into b p For reconstruction, we define the window function w p (k), the window function w p (k) When k = [pS β ,pS β +N β -1], otherwise zero. The window function has the following properties:
[0168]
[0169] A typical example is a square sine window.
[0170]
[0171] However, in other embodiments, other window functions that satisfy the above properties may be used. Figure 7 The shape of the window function is optimized based on the shape of the frame reconstruction error function shown.
[0172] The respiratory waveform signal can then be generated as:
[0173]
[0174] In a practical implementation, p has some finite support, say p=0, .., P-1, and thus the first S of the respiratory waveform signal β samples and the final S β The reconstruction of the samples is not complete, but is tapered off by the tail of the window function. In one embodiment, the reconstruction window at the beginning and end of the processed signal can be modified so that the tapering can be avoided.
[0175] This method is Figure 7 , where plot (a) discloses an example of a frequency and time matrix of samples, plot (b) shows segmentation of the generated respiratory waveform signal, plot (c) shows the segment window function, and plot (d) shows the resulting combined respiratory waveform signal.
[0176] In many embodiments, the speech processing apparatus may include an adapter 113 that is arranged to adapt the timing properties of the segments, and in particular the inter-segment duration and / or the segment duration. Thus, in many embodiments, the adapter 113 may dynamically change the inter-segment duration and / or the segment duration to reflect current conditions.
[0177] The adapter 213 may specifically be arranged to adapt the timing parameters in response to at least one of a property of the audio signal and a property of the respiratory waveform signal. Thus, in some embodiments, the adapter 213 may be arranged directly to evaluate the audio signal to extract a property such as pitch or word rate, and in response thereto, continue to change the inter-segment duration and / or segment duration, for example, by making both the inter-segment duration and the segment duration shorter for an increased word rate. In some embodiments, the adapter 213 may be arranged to evaluate the respiratory waveform signal to extract a property such as, for example, respiratory rate, and in response thereto, continue to change the inter-segment duration and / or segment duration, for example, by making both the inter-segment duration and the segment duration shorter for an increased respiratory rate.
[0178] In some embodiments, the adapter 213 may be arranged to adapt the inter-segment duration in response to properties of the audio signal. For example, if the level of background noise is high, the model may use a longer inter-segment duration in the reconstruction of the respiratory waveform. In another embodiment, a speech classification method may be used to detect whether the speaker is speaking softly, speaking normally, or shouting, and the inter-segment duration may be varied accordingly, such as selecting a shorter duration for louder (e.g., shouting) speech.
[0179] In some embodiments, the adapter 213 may be arranged to adapt the inter-segment duration in response to properties of the respiratory waveform signal. For example, if the respiratory rate calculated from the respiratory waveform is low, the inter-segment duration may be increased to allow for increased segment durations in order to accommodate more respiratory events in each segment.
[0180] In some embodiments, the adapter 213 may be arranged to adapt the segment duration in response to properties of the audio signal.For example, if the level of background noise is high, the model may use longer segment durations in the reconstruction of the respiratory waveform.
[0181] In some embodiments, the adapter 213 may be arranged to adapt the segment duration in response to properties of the respiratory waveform signal. For example, when the average respiratory rate calculated from the respiratory waveform signal is low, the segment length is increased to accommodate more respiratory events.
[0182] A typical (initial) segment duration N is 4 seconds, and segments may (initially) be set to have a duration of 2 seconds.
[0183] As previously mentioned, this approach can provide improved performance for a range of speech processing applications.
[0184] In particular, for automatic speech recognition, the method can provide significantly improved performance through a process that adapts to the current properties of the speaker.
[0185] For example, as previously described, the speech processor 103 can be configured to determine inhalation / exhalation time segments of the audio signal that correspond to times when the speaker is inhaling and, therefore, not speaking. The speech processor 103 can then be configured to exclude these inhalation / exhalation time segments from the speech recognition processing applied to the audio signal. Thus, times when the respiratory waveform signal may indicate that speech is highly unlikely can be excluded from speech recognition, thereby reducing the risk of falsely detecting words when speech is absent. This can also improve accuracy detection at other times because an estimate can be made of when a word may be spoken.
[0186] As another example, the breathing waveform signal may be used to adapt speech recognition by, for example, estimating word rate based on effort and pain in the breathing pattern estimated from the breathing waveform signal.
[0187] In some embodiments, the speech processor 103 can be arranged to (first) perform a speech recognition operation without considering the respiratory waveform signal or any parameters derived therefrom. However, rather than providing only estimated terms or sentences, the speech recognition can generate a plurality of different candidates to estimate what may have been said. The speech processor 103 can then be arranged to select between different candidates based on the respiratory waveform signal. For example, the activity waveform of the estimated sentence can be compared with the respiratory waveform signal, and the selection can be based on the closeness of these matches. For example, a sentence that is aligned with the inhalation / inhalation interval (and which, for example, includes speech during the estimated inhalation / inhalation interval) can be selected, while a sentence that is not aligned with the inhalation / inhalation interval is not selected.
[0188] As another example, the selection between candidates can be based on the level of breathing distress. For example, when a user is relaxed, the word rate of speech is typically much higher than when the user is in high distress and experiencing breathing difficulties (e.g., due to strenuous physical exercise). Therefore, the speech processor 103 can be arranged to select between candidates based on the word rate of such content. For example, if the breathing waveform signal indicates that the speaker is relaxed, the candidate with the highest word rate can be selected, and if the breathing waveform signal indicates that the speaker is experiencing breathing difficulties, the candidate with the lowest word rate can continue to be selected.
[0189] Of course, in many scenarios, the respiratory waveform signal may be only one factor in selecting between candidates, and other parameters may also be considered. For example, in addition to the candidate items, the speech recognition may also estimate a confidence or reliability value for each candidate / estimate, and the selection may also take such confidence level into account.
[0190] A particular advantage of this approach is that improved detection can be achieved by further taking into account the speaker's breathing properties, but conventional speech recognition modules can be used. Thus, post-processing of results from existing speech recognition algorithms can be used to provide improved performance while allowing backward compatibility and allowing the use of existing recognition algorithms.
[0191] In many embodiments, the speech processing may be speech encoding, wherein the captured speech is, for example, encoded for efficient transmission or distribution over a suitable communication channel. As previously mentioned, more efficient encoding can often be achieved simply by not encoding any signal during the inspiration / inhalation time segments or by allocating fewer bits to the signal in the entropy encoding during the inspiration / inhalation period.
[0192] As another example, information about breathing waves can be used to control speech manipulation, such as to hide a speaker's fatigue. For example, the system can detect a state of dyspnea by detecting an increased breathing rate in the speaker. A speech synthesizer (such as a Wavenet model) can then be used to resynthesize a new speech signal that is similar to the same speaker's speech, but with a different breathing wave pattern corresponding to a speaker with a lower breathing rate.
[0193] In some embodiments, the speech processing can be speech enhancement. Such speech enhancement can, for example, be used to provide a clearer speech signal, wherein, for example, other audio sources and noise in the audio signal are reduced. For example, as previously described, the audio signal can be attenuated during the inhalation / exhalation time segment. As another example, different frequencies can be attenuated based on the respiratory waveform signal, thereby specifically attenuating frequencies that tend to be more dominant for some respiratory noises. As another example, the respiratory waveform signal can indicate a transition or phase in the respiratory cycle that tends to be associated with a particular sound, and the speech processing can be arranged to generate corresponding signals and invert them to provide an audio cancellation effect for such sounds.
[0194] In many embodiments, the speech processor 103 may be arranged to evaluate the respiration waveform signal to determine one or more parameters of the speaker's respiration. It may further be arranged to adapt the speech processing based on the determined parameter values.
[0195] For example, in many embodiments the speech processor 103 is arranged to determine a breathing rate estimate from the breathing waveform signal and to adapt the speech processing according to the breathing rate.
[0196] Since respiration is an inherently repetitive / periodic process, the respiration waveform signal will also be a periodic signal (or at least have a strong periodic component), and the respiration rate can be determined by determining the repetition rate / duration of the periodic component of the respiration waveform signal. For example, the respiration waveform signal can be correlated with itself, and the duration between correlation peaks can be determined and used as a measure of the respiration rate. It should be understood that many different techniques and algorithms are known for detecting periodic components in a signal, and any method can be used without departing from the present invention.
[0197] Respiration rate can indicate a number of parameters that may affect a person's speech. For example, breathing rate can be a strong indicator of the speaker's level of relaxation / distress, and as previously discussed, this can affect speech in different ways, including, for example, word rate, the amount of breathing noise, etc. As previously discussed, speech processing can be adapted to reflect these parameters.
[0198] The respiratory rate can also provide information related to the timing of the inspiration / inhalation time intervals. In particular, the respiratory rate can directly give the frequency and time between inspiration / inhalation time intervals.
[0199] As another example, breathing rate can be used to adapt speech processing by selecting a particular speech preprocessing method based on the speaker's breathing rate. For example, different parameters for a voice activity detection or speech enhancement algorithm may be used when the speaker has a low breathing rate compared to when the breathing rate is elevated or high.
[0200] In some embodiments, the speech processor 103 is arranged to determine the speech word rate from the respiratory waveform signal and to adapt the speech processing according to the speech word rate.
[0201] For example, the speech word rate may be determined by the output of an automatic speech recognition method.
[0202] As described, the speaker's word rate can vary based on the degree of relaxation or distress of the speaker. The respiratory waveform signal can be used to determine this degree, and therefore the degree of dyspnea currently experienced by the speaker. A word rate for a given degree of dyspnea can be determined (e.g., as a predetermined value for a given degree of respiratory distress). Speech processing can then be performed based on the determined word rate. For example, as previously described, the speech processor 103 can perform speech recognition that provides multiple possible candidates, and can select a candidate with a word rate that most closely matches the word rate determined based on the respiratory waveform signal. In some embodiments, speech recognition can therefore provide several candidates for speech utterances, and based on the expected word rate determined based on the respiratory waveform signal, the speech processor 103 can select utterances corresponding to the lowest word rate and long speech pause duration in the case of high dyspnea.
[0203] In fact, many speech recognition systems have difficulty with low word rates and breathless speech pauses. However, in this method, the parameters of the speech recognition can be adjusted to accommodate this, or the output of the speech recognition system can be processed to accommodate this condition.
[0204] As another example, word rate can be used to adapt speech processing by having specific settings based on word rate settings in speech preprocessing or speech coding algorithms. This setting can, for example, trigger the use of different quantization vector codebooks in the encoder and decoder in a speech encoder.
[0205] In some embodiments, the speech processor 103 is arranged to determine the level of breathing sounds during inspiration / inhalation from the respiratory waveform signal and to adapt the speech processing according to the level of speech breathing sounds during inspiration / inhalation.
[0206] The breathing sound level may be determined, for example, by calculating the root mean square amplitude level of the inspiratory speech signal (the audio signal during detected inspiration / inhalation time intervals), optionally applying weighting to account for relative loudness as perceived by the human ear.
[0207] For example, in a recording, the sound of inhaling can distract the listener from what is being said or sung. In this case, the level of the detected inhaling sound can measure how much the inhaling sound needs to be attenuated to become inaudible.
[0208] In some embodiments, the speech processor 103 may be arranged to take into account other data than just the audio signal (and the respiratory waveform signal derived therefrom).
[0209] This approach can provide particularly advantageous and synergistic effects, for example, for embodiments in which speech processing is adapted based on parameters of the respiratory waveform signal that may not directly indicate, for example, physical exercise of the person, but may also be due to other issues. For example, it can distinguish between dyspnea caused by strenuous exercise and dyspnea caused by medical respiratory problems that may also occur during a relaxed state.
[0210] As a specific example, the speech processor 103 can make a prediction of the breathing rate or word rate based on the detected heart rate and use this estimate to set the parameters of the speech preprocessing or encoding algorithm as described above. This is based on the consideration that heart rate and breathing rate are generally correlated. This is advantageous for obtaining optimal performance of the system from the speaker's first utterance before the model used to estimate the breathing wave has access to a sufficiently long speech recording to be able to make a first estimate. This is also advantageous in situations where the speaker only uses very short utterances (e.g., voice commands for controlling a music playback application in a wearable audio player).
[0211] In some embodiments, the speech processor 103 may include a human breathing model that can be used to determine the speaker's breathing properties. This breathing model can be adapted based on the received heart rate and / or activity. In some embodiments, the speech processor 103 can estimate a computational model that generates an estimate of breathing characteristics from activity characteristics corresponding to current physical activity and vital signs measured using a wearable device.
[0212] For example, a human respiratory model can model specific biophysical processes in body organs, or it can be a functional high-level metabolic model. For example, metabolic equivalents (METs) relate to the oxygen uptake rate of a body for a given activity as a multiple of resting VO2. Based on the known METs for the activity, the effective lung volume VO2, and the oxygen concentration in the air, we can obtain an indirect estimate of the respiratory rate required to meet the subject's oxygen demand. This can be adapted based on heart rate by a more fine-grained model that contains specific MET values for different intensities of specific physical activities (e.g., from jogging to full-speed uphill running) measured based on heart rate and activity intensity.
[0213] This method can be applied to different languages. In many embodiments, the same method can be used for different languages, and for example, the same trained neural network can be used. Thus, in some embodiments, the processing can be language-agnostic. This can reflect that many speech parameters and their relationship to respiration can be less sensitive to changes in language, as many language properties are common across different languages. For example, speech does not occur during inhalation / inhalation, but occurs during exhalation / exhalation, which is common to all languages.
[0214] In fact, experiments performed by the applicants have shown that the effects of breathing on speech have similar temporal statistics in most languages, and the methods described in many embodiments and scenarios will be applicable to speech in many (if not all) languages. In many cases, no language adaptation or limitation is employed or envisioned accordingly.
[0215] However, in some embodiments, the process can be specific to a particular language. For example, using only a particular language to train a neural network will optimize / train it for that particular language. Thus, in this case, improved performance can be achieved for speech in that language.
[0216] In some embodiments, the method may be arranged to first detect the language and then apply settings specific to that language, such as loading the neural network with coefficients and parameters for that particular language. In other embodiments, the neural network may have been trained using a different language, for example, and the neural network may be automatically adapted to the different language.
[0217] The audio device may be embodied in one or more appropriately programmed processors. For example, an artificial neural network may be implemented in one or more such appropriately programmed processors. Different functional blocks (particularly the artificial neural network) may be implemented in separate processors and / or may be implemented, for example, in the same processor. Examples of suitable processors are provided below.
[0218] Figure 8 8 is a block diagram illustrating an example processor 800 according to an embodiment of the present disclosure. The processor 800 may be used to implement one or more processors for implementing the apparatus or elements thereof as previously described (including, in particular, one or more artificial neural networks). The processor 800 may be any suitable processor type, including but not limited to a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA) (where the FPGA has been programmed to form a processor), a graphics processing unit (GPU), an application specific integrated circuit (ASIC) (where the ASIC has been designed to form a processor), or a combination thereof.
[0219] Processor 800 may include one or more cores 802. Core 802 may include one or more arithmetic logic units (ALUs) 804. In some embodiments, core 802 may include a floating point logic unit (FPLU) 806 and / or a digital signal processing unit (DSPU) 808 in addition to or in place of ALU 804.
[0220] The processor 800 may include one or more registers 812 communicatively coupled to the core 802. The registers 812 may be implemented using dedicated logic gates (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 812 may be implemented using static memory. The registers may provide data, instructions, and addresses to the core 802.
[0221] In some embodiments, the processor 800 may include one or more levels of cache memory 810 communicatively coupled to the core 802. The cache memory 810 may provide computer-readable instructions to the core 802 for execution. The cache memory 810 may provide data for processing by the core 802. In some embodiments, the computer-readable instructions may have been provided to the cache memory 810 from a local memory (e.g., a local memory attached to the external bus 816). The cache memory 810 may be implemented using any suitable cache memory type (e.g., metal oxide semiconductor (MOS) memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology).
[0222] The processor 800 may include a controller 814 that can control inputs to the processor 800 from other processors and / or components included in the system and / or outputs from the processor 800 to other processors and / or components included in the system. The controller 814 can control data paths in the ALU 804, FPLU 806, and / or DSPU 808. The controller 814 can be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of the controller 814 can be implemented as independent gates, FPGAs, ASICs, or any other suitable technology.
[0223] Registers 812 and cache 810 may communicate with controller 814 and core 802 via internal connections 820A, 820B, 820C, and 820D. The internal connections may be implemented as a bus, a multiplexer, a crossbar switch, and / or any other suitable connection technology.
[0224] Input and output of the processor 800 may be provided via a bus 816, which may include one or more conductors. The bus 816 may be communicatively coupled to one or more components of the processor 800, such as the controller 814, the cache 810, and / or the registers 812. The bus 816 may be coupled to one or more components of the system.
[0225] Bus 816 can be coupled to one or more external memories. External memory can include read-only memory (ROM) 832. ROM 832 can be a shielded ROM, an electronically programmable read-only memory (EPROM), or any other suitable technology. External memory can include random access memory (RAM) 833. RAM 833 can be static RAM, backup static RAM, dynamic RAM (DRAM), or any other suitable technology. External memory can include electrically erasable programmable read-only memory (EEPROM) 835. External memory can include flash memory 834. External memory can include a magnetic storage device, such as a disk 836. In some embodiments, external memory can be included in the system.
[0226] The foregoing description (and claims / invention) relates to generating a respiratory waveform signal and adapting speech processing based on the respiratory waveform signal. However, it should be understood that the generation of a respiratory waveform signal is not limited to the present application. The generated respiratory waveform signal can be fed to a determiner, which is arranged to determine breathing / respiratory parameters / attributes from the respiratory waveform signal. The breathing / respiratory parameters / attributes can indicate the respiratory properties of the speaker. The breathing parameters can be, for example, breathing rate, breathing depth, etc. As mentioned above, such breathing parameters can be used for speech processing. However, it can also be used for many other purposes and can indeed be provided to a user via a suitable user interface. The user can, for example, be a health professional who can make use of the information provided (for example for diagnostic purposes).
[0227] Thus, although the described method comprises a speech processor arranged to apply speech processing to an audio signal (where the speech processing is dependent on the respiration waveform signal), the respiration waveform signal may be used for other purposes.
[0228] The present invention can be implemented in any suitable form, including any combination of hardware, software, firmware or these items. The present invention can optionally be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components of the embodiments of the present invention can be implemented physically, functionally and logically in any suitable manner. In fact, the function can be implemented in a single unit, in multiple units or as a part of other functional units. Like this, the present invention can be implemented in a single unit, or can be distributed between different units, circuits and processors physically and functionally.
[0229] Although the present invention has been described in conjunction with certain embodiments, it is not intended to be limited to the specific forms set forth herein. On the contrary, the scope of the present invention is limited only by the appended claims. In addition, although features may appear to be described in conjunction with specific embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0230] Furthermore, although listed separately, multiple devices, elements, circuits or method steps may be implemented by, for example, a single circuit, unit or processor. In addition, although individual features may be included in different claims, these features may be advantageously combined, and inclusion in different claims does not imply that the combination of features is not feasible and / or advantageous. Furthermore, the inclusion of a feature in one claim category does not imply a limitation to that category, but rather indicates that the feature is equally applicable to other claim categories. Furthermore, the order of features in the claims does not imply that the features must work in any particular order, and in particular, the order of the individual steps in a method claim does not imply that the steps must be performed in this order. On the contrary, the steps may be performed in any suitable order. Furthermore, singular references do not exclude pluralities. Thus, references to "a", "an", "first", "second", etc. do not exclude pluralities. The figure marks in the claims are provided merely as clarifying examples and should not be construed as limiting the scope of the claims in any way.
[0231] In general, the following embodiments indicate examples of apparatuses and methods for speech processing.
[0232] Example:
[0233] Embodiment 1: A device for speech processing of an audio signal, comprising:
[0234] an input (101) arranged to receive the audio signal, the audio signal comprising a speech audio component for a speaker;
[0235] a segmenter (105) arranged to generate segments of the audio signal, each segment having a segment duration and consecutive segments having an inter-segment duration, the inter-segment duration being the time between the consecutive segments, the segment duration exceeding the inter-segment duration such that the consecutive segments overlap;
[0236] a first generator (107) arranged to generate a segment of a breathing waveform signal for each segment of the audio signal;
[0237] a second generator (109) arranged to generate the respiratory waveform signal by combining the segments of the respiratory waveform signal; and
[0238] A speech processor (103) is arranged to apply speech processing to the audio signal, the speech processing being dependent on the respiration waveform signal.
[0239] Embodiment 2. The apparatus of embodiment 1, wherein at least some segments have a segment duration that exceeds the inter-segment duration.
[0240] Embodiment 3. The apparatus of embodiment 2, wherein the segment duration is equal to the sub-segment duration.
[0241] Embodiment 4. An apparatus according to embodiment 2 or 3, wherein the second generator (109) is arranged to combine the fragments by applying a weighted combination to samples of different fragments for the same moment, wherein the weights of the weighted combination are determined according to a fragment window function.
[0242] Embodiment 5: The apparatus according to any preceding embodiment, wherein the segment duration exceeds the inter-segment duration by no less than 50%.
[0243] Embodiment 6. An apparatus according to any preceding embodiment, wherein the segmentation of the audio signal comprises at least a first portion of the audio signal also included in a previous segment of the audio signal and a second portion of the audio signal also included in a subsequent segment of the audio signal.
[0244] Embodiment 7: The apparatus according to any of the preceding embodiments, wherein a sampling rate of the respiratory waveform signal is not less than 10 times lower than a sampling rate of the audio signal.
[0245] Embodiment 8. The apparatus according to any preceding embodiment, wherein the first generator (107) is arranged to generate the fragments to have a lower sampling rate than the segmentation of the audio signal.
[0246] Embodiment 9. The apparatus according to any preceding embodiment, wherein the segmenter (105) is arranged to generate the segments to have a lower sampling rate than the audio signal.
[0247] Embodiment 10. An apparatus according to any preceding embodiment, wherein the first generator (107) comprises a trained artificial neural network having input nodes for receiving samples of a segment of the audio signal and output nodes arranged to provide samples of fragments of the respiratory waveform signal for the segment of the audio signal.
[0248] Embodiment 11. An apparatus according to any preceding embodiment, wherein the first generator (107) comprises a model for human breathing, and wherein the first generator is arranged to generate the segments from the model and to adapt parameters of the model in response to the segmentation.
[0249] Embodiment 12. An apparatus according to any preceding embodiment, wherein the segmenter (105) is arranged to apply a frequency transform to the audio signal and generate each segment as a matrix of time-frequency samples, the matrix of time-frequency samples comprising at least samples for two different time intervals and for two different frequency ranges.
[0250] Embodiment 13. The apparatus according to any preceding embodiment further comprises an adapter (113), wherein the adapter (113) is arranged to adapt at least one of the inter-segment duration and the segment duration in response to at least one of the properties of the audio signal and the properties of the respiratory waveform signal.
[0251] Embodiment 14: A method for processing an audio signal, the method comprising:
[0252] receiving the audio signal, the audio signal including a speech audio component for a speaker;
[0253] generating segments of the audio signal, each segment having a segment duration and consecutive segments having an inter-segment duration, the inter-segment duration being the time between the consecutive segments, the segment duration exceeding the inter-segment duration such that the consecutive segments overlap;
[0254] generating a segment of a breathing waveform signal for each segment of the audio signal;
[0255] generating the respiratory waveform signal by combining the segments of the respiratory waveform signal; and
[0256] Speech processing is applied to the audio signal, the speech processing being dependent on the respiration waveform signal.
Claims
1. A device for speech processing of an audio signal, the device comprising: an input (101) arranged to receive the audio signal, the audio signal comprising a speech audio component for a speaker; a segmenter (105) arranged to generate segments of the audio signal, each segment having a segment duration and consecutive segments having an inter-segment duration, the inter-segment duration being the time between the consecutive segments, the segment duration exceeding the inter-segment duration such that consecutive segments overlap; a first generator (107) arranged to generate, for each segment of the audio signal, segments of the respiratory waveform signal from each segment of the audio signal, at least some of the segments having a segment duration exceeding the inter-segment duration; a second generator (109) arranged to generate the respiratory waveform signal by combining the segments of the respiratory waveform signal, the second generator (109) being arranged to combine segments of the at least some segments by applying a weighted combination to samples of different segments for the same time instant, wherein weights of the weighted combination are determined according to a segment window function; and A speech processor (103) is arranged to apply speech processing to the audio signal, the speech processing being dependent on the respiration waveform signal.
2. A device according to any preceding claim, wherein The segment duration exceeds the inter-segment duration by no less than 50%.
3. An apparatus according to any preceding claim, wherein The segmentation of the audio signal comprises at least a first portion of the audio signal also included in a previous segment of the audio signal and a second portion of the audio signal also included in a subsequent segment of the audio signal.
4. An apparatus according to any preceding claim, wherein The sampling rate of the respiratory waveform signal is not less than 10 times lower than the sampling rate of the audio signal.
5. An apparatus according to any preceding claim, wherein The first generator (107) is arranged to generate the segments to have a lower sampling rate than the segments of the audio signal.
6. An apparatus according to any preceding claim, wherein The segmenter (105) is arranged to generate the segments to have a lower sampling rate than the audio signal.
7. An apparatus according to any preceding claim, wherein The first generator (107) comprises a trained artificial neural network having input nodes for receiving samples of a segment of the audio signal and output nodes arranged to provide samples of segments of the respiratory waveform signal for the segment of the audio signal.
8. An apparatus according to any preceding claim, wherein The first generator (107) comprises a model for human breathing, and wherein the first generator is arranged to generate the segments from the model and to adapt parameters of the model in response to the segmentation.
9. An apparatus according to any preceding claim, wherein The segmenter (105) is arranged to apply a frequency transform to the audio signal and generate each segment as a matrix of time-frequency samples comprising at least samples for two different time intervals and for two different frequency ranges.
10. An apparatus according to any preceding claim, further comprising an adapter (113) arranged to adapt at least one of the inter-segment duration and the segment duration in response to at least one of a property of the audio signal and a property of the respiratory waveform signal.
11. A method for processing an audio signal, the method comprising: receiving an audio signal, the audio signal including a speech audio component for a speaker; generating segments of the audio signal, each segment having a segment duration and consecutive segments having an inter-segment duration, the inter-segment duration being the time between the consecutive segments, the segment duration exceeding the inter-segment duration such that consecutive segments overlap; generating a segment of a respiratory waveform signal for each segment of the audio signal according to each segment of the audio signal; generating the respiratory waveform signal by combining the segments of the respiratory waveform signal, the combining comprising combining segments of the at least some segments by applying weighted combination to samples of different segments for the same time instant, wherein weights of the weighted combination are determined according to a segment window function; and Speech processing is applied to the audio signal, the speech processing being dependent on the respiration waveform signal.
12. A computer program product comprising computer program code means adapted to perform all the steps of claim 11 when said program is run on a computer.
Citation Information
Patent Citations
Virtual Trainer
US20090098981A1
Automatic speech recognition system for people with speech-affecting disabilities
US7529670B1