Speech processing of audio signal
By generating breathing waveform signals and using artificial neural networks to adjust speech processing, the problem of insufficient performance of speech processing algorithms in the prior art under different conditions is solved, and more efficient speech enhancement and recognition is achieved.
Patent Information
- Application Number
- CN202380085420.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-15
- Filing Date
- 2023-12-11
- Publication Date
- 2025-07-22
AI Technical Summary
Existing speech processing algorithms cannot provide optimal performance when the actual situation is different from the expected situation, and often require more resources and complexity, making it difficult to adapt to the speaker's changes and environmental flexibility.
By generating a breath waveform signal, adjust speech processing, use an artificial neural network to determine breathing rate and other breathing parameters, and optimize the operation of the speech processor to adapt to different scenarios and user behavior.
Improve the adaptability and flexibility of speech processing, reduce complexity and resource use, improve the quality of speech enhancement and recognition, and provide more accurate word detection and user experience.
Smart Images

Figure CN120359568A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to performing speech processing on an audio signal and, in particular but not exclusively, to automatic speech recognition for capturing an audio signal of a speaker (e.g., performing an activity such as an exercise activity). Background Art
[0002] Speech processing of audio signals is widely applied in a variety of practical applications and is becoming increasingly important and prominent for many daily activities and devices.
[0003] For example, speech enhancement can be widely applied to provide improved clarity and reproducibility for speech captured by an audio signal (e.g., sounds captured in a real-life environment). Speech encoding is also often performed on the captured audio. For example, speech encoding of the captured audio signal is a component of a smart phone. Another speech processing application that has become increasingly frequent in recent years is the application of speech recognition (e.g., to provide a user interface to a device).
[0004] For example, the speech interface of a personal assistant or a home speaker enables simple control of media, navigation features, tracking, obtaining various information services, etc. by using a voice interface and, at the same time, performing other activities (e.g., during exercise).
[0005] Speech interfaces based on automatic speech recognition (ASR) have also become very important in controlling various devices such as wearable, accessible, and IoT consumer, health, and fitness devices. Speech is the most natural interface for interacting with a device interface during activities involving significant user movement and change. As an example, US20090098981A1 discloses a fitness device having a speech interface for a user to communicate with a virtual fitness coach.
[0006] However, although much effort has been invested in developing and optimizing speech processing algorithms for various applications and many very advantageous and effective methods have been developed, they are often not optimal in all cases. For example, many speech processing operations have been developed for specific nominal conditions (e.g., nominal speaker attributes, nominal acoustic environment, etc.). In scenarios where the actual conditions are different from the expected nominal conditions, many speech processing applications may not provide optimal performance. Many speech processing algorithms may also be more complex or resource-intensive than desired. Adjusting speech processing to attempt to compensate for varying attributes often results in less flexible, more complex, and / or resource-intensive implementations, which often do not provide optimal performance.
[0007] Accordingly, an improved method for speech processing would be advantageous. In particular, a method that allows for increased flexibility, increased adaptability, improved performance, improved (e.g., speech enhancement, encoding, and / or recognition) quality, reduced complexity and / or resource usage, improved remote control of audio processing, improved adaptation to changes in speaker attributes and / or activities, reduced computational load, improved user experience, facilitating convenient implementation and / or improved spatial audio experience would be advantageous. SUMMARY OF THE INVENTION
[0008] Accordingly, the present invention seeks to alleviate, mitigate, or eliminate one or more of the above disadvantages, preferably individually or in any combination.
[0009] According to one aspect of the present invention, there is provided an apparatus for speech processing of an audio signal, the apparatus comprising: an input arranged to receive an audio signal comprising a speech audio component from a speaker; a determiner arranged to generate a respiratory waveform signal from the audio signal, the respiratory waveform signal indicating the volume of air in the lungs of the speaker as a function of time; and a speech processor arranged to perform speech processing of the audio signal, the speech processing being dependent on the respiratory waveform signal; and wherein the speech processor is arranged to: determine a respiratory rate estimation result from the respiratory waveform signal, and adjust the speech processing based on the respiratory rate estimation result.
[0010] The method can provide improved speech processing in many embodiments and scenarios. For many signals and scenarios, the method can provide speech processing that more closely reflects the changes in the attributes and characteristics of the speaker's speech. For example, the method can be adjusted for the speaker's level of fatigue or, for example, difficulty in breathing. The method can provide facilitating adjustments and can in particular provide improved adjusted speech processing without relying on or requiring additional inputs from other devices that measure the speaker's attributes. The method can provide an effective implementation and can allow for reduced complexity and / or resource usage in many embodiments.
[0011] According to an optional feature of the present invention, the provided determiner comprises a trained artificial neural network having input nodes for receiving samples of the audio signal and output nodes arranged to provide samples of the respiratory waveform signal.
[0012] This can provide particularly advantageous operation and performance in many scenarios and applications.
[0013] The method can provide a particularly advantageous arrangement which, in many embodiments and scenarios, can allow the facilitation and / or improvement possibilities of an artificial neural network to be utilized when adjusting and optimizing speech processing, typically including speech enhancement and / or recognition).
[0014] In many embodiments, the method can allow the generation of an improved speech signal by allowing the system to adjust / modify the processing of the audio signal.
[0015] The method can provide an effective implementation and, in many embodiments, can allow for reduced complexity and / or resource usage.
[0016] The artificial neural network is a trained artificial neural network.
[0017] The artificial neural network can be a trained artificial neural network trained with training data that includes training speech audio signals and training respiration waveform signals generated based on measurements of respiration waveforms; the training employs a cost function that compares the training respiration waveform signals with the respiration waveform signals generated by the artificial neural network for the training speech audio signals. The artificial neural network can be a trained (one or more) artificial neural network trained with training data that includes training speech audio signals representing a range of different states of a range of relevant speakers performing different activities.
[0018] The artificial neural network can be a trained artificial neural network that is trained with training data having training input data that includes training speech audio signals and using a cost function that includes a contribution indicating the difference between a measured training respiration waveform signal and the respiration waveform signal generated by the artificial neural network in response to the training speech audio signals.
[0019] According to an optional feature of the present invention, the provided speech processing includes speech enhancement processing.
[0020] The method can provide improved speech enhancement and, in many scenarios, can allow the generation of an improved speech signal.
[0021] According to an optional feature of the present invention, the provided speech processing includes speech recognition processing.
[0022] The method can provide improved speech recognition. The method can allow for more accurate detection of words, terms, and sentences in many scenarios and can particularly allow for improved speech recognition of a speaker in various situations, states, and activities. For example, the method can allow for improved speech detection of a person who is exercising.
[0023] According to an optional feature of the present invention, the speech processing includes speech recognition processing for generating a plurality of candidate recognized words based on the audio signal; and the speech processor is arranged to select among the candidate recognized words in response to the respiration waveform signal.
[0024] This can provide advantageous speech recognition in many scenarios and can allow implementation with low complexity and resource requirements. For example, it can allow adapting to the respiration waveform signal and taking the respiration waveform signal into account by implementing post-processing on the results of existing speech recognition algorithms. This method can provide improved backward compatibility.
[0025] According to an optional feature of the present invention, the speech processor is arranged to: determine a speech word rate based on the respiration waveform signal, and adjust the speech processing according to the speech word rate.
[0026] This can provide improved operation and performance in many embodiments.
[0027] According to an optional feature of the present invention, the speech processor is arranged to: determine timing attributes of an inhalation interval based on the respiration waveform signal, and adjust the speech processing according to the timing attributes of the inhalation interval.
[0028] This can provide improved operation and performance in many embodiments. The timing attributes can generally be the duration, repetition time / frequency, and / or timing of the inhalation interval.
[0029] According to an optional feature of the present invention, the speech processor is arranged to: determine an inhalation time period of the audio signal based on the timing attributes, and attenuate the audio signal during the inhalation time period.
[0030] This can provide improved operation and performance in many embodiments. The attenuation can be partial attenuation or complete attenuation (mute) of the audio signal.
[0031] According to an optional feature of the present invention, the speech processor is arranged to: determine an inhalation time period of the audio signal based on the timing attributes, and exclude the inhalation time period from the speech recognition processing applied to the audio signal.
[0032] This can provide improved operation and performance in many embodiments.
[0033] According to an optional feature of the present invention, the speech processor is arranged to: determine a respiration sound level during inhalation based on the respiration waveform signal, and adjust the speech processing according to the respiration sound level during inhalation.
[0034] This can provide improved operation and performance in many embodiments.
[0035] According to an optional feature of the present invention, the apparatus further comprises: a receiver arranged to receive speaker data indicative of at least one of a heart rate of the speaker and an activity of the speaker; and the speech processor is arranged to adjust the speech processing based on the speaker data.
[0036] This can provide improved operation and / or performance in many embodiments. It can allow for more accurate adjustment of speech processing for the current condition of the speaker.
[0037] In some embodiments, the apparatus further comprises: a receiver arranged to receive heart rate data indicative of the heart rate of the speaker; and the speech processor is arranged to adjust the speech processing based on the heart rate.
[0038] In some embodiments, the apparatus further comprises: a receiver arranged to receive activity data indicative of an activity of the speaker; and the speech processor is arranged to adjust the speech processing based on the activity.
[0039] According to an optional feature of the present invention, the speech processor includes a human respiration model for determining a respiration attribute of the speaker, and the speech processor is further arranged to: adjust the speech processing in response to the respiration attribute; and adjust the human respiration model based on the speaker data.
[0040] This can provide improved operation and / or performance in many embodiments.
[0041] According to an optional feature of the present invention, the speech processor is arranged to set initial operation parameters for the speech processing in response to the speaker data.
[0042] This can provide improved operation and / or performance in many embodiments. The initial operation parameter values can be values used before determining the respiration waveform signal.
[0043] According to one aspect of the present invention, there is provided a method comprising: receiving an audio signal comprising a speech audio component from a speaker; generating a respiration waveform signal from the audio signal, the respiration waveform signal indicative of the volume of air in the lungs of the speaker as a function of time; and performing speech processing on the audio signal, the speech processing depending on the respiration waveform signal; and wherein the speech processing comprises: determining a respiration rate estimation result from the respiration waveform signal, and adjusting the speech processing based on the respiration rate estimation result.
[0044] These and other aspects, features, and advantages of the present invention will be apparent and elucidated with reference to the embodiments (one or more) described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Embodiments of the present invention will be described by way of example only with reference to the accompanying drawings, in which
[0046] Figure 1 some elements of an example of an audio device according to some embodiments of the present invention are illustrated;
[0047] Figure 2 an example of the structure of an artificial neural network is illustrated;
[0048] Figure 3 an example of a node of an artificial neural network is illustrated;
[0049] Figure 4 some elements of an example of a training setup for an artificial neural network are illustrated; and
[0050] Figure 5 some elements of a possible arrangement of a processor for elements of an implementation device according to some embodiments of the present invention are illustrated. DETAILED DESCRIPTION
[0051] Speech processing algorithms are typically designed and optimized for nominal conditions (e.g., a speaker speaking in a quiet and relaxed environment at rest). However, in many scenarios, the environment and conditions may be significantly different from the nominal conditions, and thus speech processing may degrade from the optimal achievable performance and results.
[0052] Figure 1 Some elements of a speech processing device according to some embodiments of the present invention are illustrated. The speech processing device may be adapted to provide improved speech processing for a speaker in many different environments and scenarios (e.g., if the speaker is exercising).
[0053] The speech processing device includes a receiver 101, which is arranged to receive an audio signal including a speech audio component. The audio signal may specifically be a microphone signal representing audio captured by a microphone. The audio signal may specifically be a speech signal, and in many embodiments may be a speech signal captured by a microphone (e.g., a microphone of a headset, a smart phone, or a body-worn microphone) arranged to capture the speech of a single person. Thus, the audio signal will include a speech audio component (hereinafter also referred to as a speech signal) and possibly other sounds (e.g., background sounds from the environment).
[0054] The speech processing apparatus further includes a speech processing unit 103, which is arranged to apply speech processing to the received audio signal, and thus specifically to the speech audio component of the audio signal. The speech processing can specifically be speech enhancement processing, speech encoding, or a speech recognition process.
[0055] The speech processing apparatus further includes a determiner / determination circuit 105, which is arranged to generate a respiration waveform signal based on the audio signal. The respiration waveform signal indicates the volume of air in the speaker's lungs as a function of time. Thus, it reflects the speaker's respiration and the airflow into and out of the speaker's lungs.
[0056] The determiner 105 is arranged to process the audio signal in order to extract information indicating the respiration of the speaker captured by the audio signal. The determiner 105 can be arranged to: receive samples of the audio input signal and produce a time series or parameters representing the time series of the measurement / estimation results of the volume of air in the lungs.
[0057] The respiration waveform signal can thus be a time-varying signal reflecting the respiration changes of the speaker, and specifically can reflect the changes in lung volume caused by the speaker's respiration. In some embodiments, the respiration waveform signal can directly represent the current lung volume. In other embodiments, the respiration waveform signal can represent, for example, the change in the current (e.g., when the respiration waveform signal represents the airflow inhaled / exhaled) lung volume.
[0058] Under normal circumstances, a separate sensor device (e.g., a respiration belt worn when the subject is speaking) can be used to capture the respiration waveform signal. However, in the current method, the determiner can determine the respiration waveform signal by deriving the respiration waveform signal from the speech data, even when no such sensor data is available. As a low-complexity example, the determiner can use a speech activity detection (SAD) algorithm to segment the speech data into speech segments and non-speech segments. Based on the knowledge that inhalation occurs during speech pauses, the determiner 105 can use the duration and frequency of the pauses as proxies for the respiration behavior, and thus determine the respiration waveform signal accordingly. In many embodiments, more accurate determinations can be used, including using artificial neural networks, as will be described later. This can, for example, compensate for or take into account that there are more pauses in speech than real inhalation events (which may lead to errors in the estimation of the respiration rate, for example).
[0059] The determiner 105 is coupled to the speech processor 103, which is arranged to adjust the speech processing in response to the respiration waveform signal. Thus, Figure 1 the speech processing apparatus is arranged to adjust the speech processing to reflect the current respiration of the speaker, rather than performing a predetermined speech processing based only on nominal or expected attributes.
[0060] This method reflects the inventors' recognition that a speaker's breathing affects speech, and that not only can time-varying properties generated by breathing be determined based on the captured audio, but also that these time-varying properties can be used to adjust the speech processing of the same audio signal to provide improved performance, particularly performance that can be better adjusted for different scenarios and user behaviors.
[0061] The specific speech processing and adjustments performed can depend on the individual embodiment. However, in many embodiments, the method can be used to adjust the operation to reflect different breathing patterns that a person may exhibit, for example, due to performing different activities. For example, when a person is exercising, the breathing may become strained, which can lead to frequent interruptions and slowed speech. Figure 1 The speech processing device can be adapted to optimize for slow speech with long silent intervals.
[0062] This method can provide significantly improved speech processing in many scenarios and can, for example, particularly provide favorable operation for speakers engaged in strenuous activities (such as exercise) or having breathing or speaking difficulties.
[0063] This method can be particularly advantageous for automatic speech recognition (such as providing a speech interface). For example, the speech interface of a personal assistant enables control of media, navigation features, motion tracking, and various other information services during exercise. However, for normal speech at rest, speech recognition is typically optimized and may be personalized. During and immediately after strenuous exercise, the demand for oxygen increases significantly, causing altered breathing that affects speech. This in turn increases the word error rate (WER), thus improving the usability and user experience provided by the speech interface. The speech processing device can detect the user's breathing pattern and adjust the speech recognition to account for the altered speech attributes.
[0064] Speech recognition algorithms have been optimized for user groups. Often, the voice prompts for starting an interaction are also optimized for individual users.
[0065] More specifically, when a user is under physical stress (such as during exercise or physically demanding work) and requires more breathing effort, it will tend to affect the user's speech. During strenuous exercise (such as during interval training), speaking is severely affected. For example, typically, a user may need to inhale after every 1 - 3 words, or may often swallow word endings. The condition of shortness of breath is called dyspnea, and speech in this condition will be referred to as dyspneic speech hereinafter. Dyspneic speech patterns have been observed under mild physical stress, but human listeners usually subconsciously adjust for this and do not have problems understanding such speech.
[0066] However, automatic speech recognition algorithms that have been trained on speech from resting humans have severe difficulties in accurately understanding dyspneic speech and may in fact even have difficulty accurately detecting a predefined speech cue or trigger term. During voice commands and conversations with an automatic speech recognition-based interface, increased breathing rate, rapid and audible inspirations, and swallowing of word endings can significantly degrade performance. This can make the speech interface difficult to use and the usage frustrating.
[0067] However, in Figure 1 the method of a speech processing device of
[0068] As a low-complexity example, the speech processor 103 can be arranged to select between a plurality of different speech processing algorithms / processes. Each of these different speech processing algorithms / processes can be optimized for a given breathing scenario. For example, one process may have been optimized / trained for a resting breathing pattern. Another process may have been optimized / trained for a more stressed but still controlled breathing pattern. Yet another process may have been optimized / trained for a maximum distress breathing pattern. Based on the respiratory waveform signal, the signal processor 103 can select the speech processing algorithm / process that most closely matches the current breathing pattern indicated by the respiratory waveform signal and then apply that processing to the received audio signal.
[0069] Thus, in some embodiments, the speech processor 103 can be arranged to switch between different speech processing algorithms / functions according to the respiratory waveform signal. The speech processor 103 can analyze the respiratory waveform signal to determine parameter values and specifically determine the breathing rate. Each processing algorithm / function can be associated with a range of breathing rates, and the speech processor 103 can select the algorithm that includes the rate determined from the respiratory waveform signal and apply it to the audio signal.
[0070] In some embodiments, the individual algorithms / functions from which to select can include, for example, methods that use substantially the same signal processing but have parameters that have been optimized for a given condition indicated by the respiratory waveform signal. For example, the same speech processing can be used, but the operating parameters are different for different breathing rates. The operating parameters can be determined, for example, by training or by individual manual optimization based on testing of captured audio (e.g., for which the breathing attributes of a person have been measured by other means).
[0071] In some embodiments, the individual algorithms / functions from which to choose can include methods that are fundamentally different and that use, for example, completely different methods and principles. For example, different speech recognition methods can be used for a person at rest breathing (and thus speaking quickly and continuously) and a person breathing in extreme distress (able to utter mainly individual words interrupted by a large amount of breathing noise).
[0072] As an example, in many embodiments, the speech processor 103 can be arranged to determine timing attributes of the inhalation intervals based on the respiratory waveform signal.
[0073] Respiration includes a series of alternating inhalations (drawing air into the lungs) and exhalations (expelling air from the lungs), and the speech processor 103 can be arranged to evaluate the respiratory waveform signal to determine the timing attributes of the intervals during which the speaker is inhaling. Specifically, the speech processor 103 can be arranged to evaluate the audio signal / respiratory waveform signal to determine the frequency of the inhalation intervals, the duration of the inhalation intervals, and / or the timing of when these intervals begin and / or end.
[0074] Such parameters can be determined, for example, by a peak picking algorithm that locates local minima and maxima in the respiratory waveform signal (which correspond to the start of the inhalation phase and the start of the exhalation phase). The respiratory rate can be determined based on the time between two successive inhalations or exhalations.
[0075] In some embodiments, the speech processor 103 can, for example, select between different algorithms or adjust the parameters of speech processing based on such timing indications. For example, if the duration of an inhalation (e.g., relative to exhalation) increases, it can indicate that the person is taking a deep breath, which can reflect a different speech pattern than during normal breathing.
[0076] However, in some embodiments, the speech processor 103 can be specifically arranged to modify the speech processing for the inhalation time intervals. In particular, in many embodiments, the speech processor 103 can be arranged to prohibit or stop speech processing during the inhalation time intervals. For example, speech processing in the form of speech recognition can simply ignore all inhalation time intervals and be applied only to non-inhalation time intervals.
[0077] The speech processor 103 can be arranged to determine the inhalation time periods of the audio signal. Thus, these time periods can reflect the time intervals during which the speaker is estimated to be breathing in the audio signal. In some embodiments, the speech processor 103 can be arranged to attenuate the audio signal during the inhalation time periods. For example, a speech enhancement algorithm arranged to reduce noise by emphasizing the speech component of the audio signal can attenuate the audio signal (including completely attenuating the signal in many embodiments) during the inhalation time intervals.
[0078] Such methods can tend to provide significantly improved performance and can allow speech processing to be adapted and specifically applied to the portions of the audio signal where speech is present, while allowing different methods to be adapted and specifically attenuate portions of the audio signal where speech is unlikely to be present. These methods can, for example, allow significant reduction of breathing sounds by eliminating the sound of heavy breathing during the inhalation phase of the breathing wave signal. This can significantly improve the perceived clarity of speech and can accordingly provide perceived speech. In some embodiments, it can be allowed to remove (estimated based on the breathing wave) the portions of the speech data that occur during inhalation prior to automatic speech recognition, thereby facilitating this and reducing the risk of misdetecting words. In speech coding embodiments, it can improve, for example, the coding efficiency, since speech coding can be not performed during the inhalation time intervals. For example, an indication of a silent or non-speech time interval can simply be inserted into the coded data stream to indicate that no speech data is provided for that interval.
[0079] In many embodiments, the determiner 105 can include a trained artificial neural network that is arranged to determine samples of the breathing waveform signal based on samples of the audio signal. The artificial neural network can include input nodes that receive samples of the audio signal and output nodes that generate samples of the breathing waveform signal. The artificial neural network can, for example, have been trained using speech data and sensor data representing lung air volume measurements during speech.
[0080] The artificial neural network used in the described functionality can be a network of nodes arranged in multiple layers, and each node holds a node value. Figure 2 An example of a portion of the artificial neural network is illustrated.
[0081] The node value of a given node can be calculated to include contributions from some or typically all of the nodes in the previous layer of the artificial neural network. Specifically, the node value of a node can be calculated as a weighted sum of the node values output by all the nodes in the previous layer. Typically, a bias can be added and the result can be subjected to an activation function. The activation function provides the necessary part of each neuron by typically providing non-linearity. This non-linearity and the activation function provide a significant impact in the learning and adaptation process of the neural network. Thus, the node value is generated based on the node values of the previous layer.
[0082] The artificial neural network can specifically include an input layer 201 that includes a plurality of nodes that receive input data values for the artificial neural network. Thus, the node values of the nodes in the input layer can typically be directly the input data values for the artificial neural network and thus can not be calculated from other node values.
[0083] The artificial neural network may also include zero, one, or more hidden layers 203 or processing layers. For each such layer, the node values are typically generated as a function of the node values of the nodes in the previous layer, and specifically, a weighted combination and an added bias may be applied, and then an activation function (e.g., sigmoid, ReLU, or Tanh function) may be applied.
[0084] Specifically, as Figure 3 shown, each node (which may also be referred to as a neuron) may receive input values (from the nodes in the previous layer) and compute a node value as a function of these values. Typically, this includes first generating a value that is a linear combination of the input values, where each of these values is weighted by a weight:
[0085]
[0086] where w refers to the weight, x refers to the nodes in the previous layer, and n refers to the index of the different nodes in the previous layer.
[0087] Then the activation function may be applied to the resulting combination. For example, the node value l may be determined as:
[0088] l = f(k)
[0089] where the function may be, for example, the rectified linear unit function as described by Xavier Glorot, Antoine Bordes, and Yoshua Bengio in "Proceedings of the Fourteenth International Conference on Artificial Intelligence" (PMLR 15:315 - 323, 2011):
[0090] f(k) = ReLU(k) = max(0, k)
[0091] Other commonly used functions include the sigmoid function or the tanh function. In many embodiments, multiple functions may be used to compute the node output or value. For example, the ReLU function and the Sigmoid function may be combined using an activation function such as the following:
[0092] f(k) = ReLU(k) + σ(k)
[0093] Such operations may be performed by each node (usually except for the input nodes) of the artificial neural network.
[0094] The artificial neural network further includes an output layer 205 that provides an output from the artificial neural network, i.e., the output data of the artificial neural network is the node value of the output layer. Regarding the hidden / processing layer, the output node value is generated by a function of the node values of the previous layer. However, compared to the hidden / processing layer where the node values are typically not accessible or further used, the node values of the output layer are accessible and provide the result of the operation of the artificial neural network.
[0095] Many different network architectures and toolboxes have been developed for artificial neural networks, and in many embodiments, the artificial neural network can be based on adjusting and customizing such networks. An example of a network architecture that can be applied to the above applications is the long short-term memory (LSTM) of Sepp Hochsreiter [described by Hochreiter, Sepp and Jürgen Schmidhuber in "Long Short-Term Memory" (Neural Computation 9.8, 1997, pp. 1735 - 1780)].
[0096] LSTM is an architecture for classifying and regressing time-domain signals using recursive causal or bidirectional evaluation and has been successfully applied to audio signals. For
[0097]
[0098] where * denotes matrix multiplication, o denotes the Hadamard product, x is the input vector, h t-1 denotes the output vector of the previous time step, W, V, U are the network weights, and b is the bias vector.
[0099] In theory, a classical (or "vanilla") artificial neural network is capable of tracking any long-term dependencies in an input sequence. The problem with a vanilla artificial neural network is essentially computational (or practical): when using backpropagation to train a vanilla artificial neural network, the long-term gradients of backpropagation can "vanish" (i.e., they can tend to zero) or "explode" (i.e., they can tend to infinity) because the process involves computations using finite-precision numbers. An artificial neural network using LSTM cells partially solves the vanishing gradient problem because LSTM cells allow the gradients to also remain constant. However, LSTM networks can still be troubled by the gradient explosion problem.
[0100] In some cases, an artificial neural network can also be arranged to include an additional contribution that allows for dynamically adjusting or customizing the artificial neural network for a particular desired property or characteristic of the generated output. For example, a set of values can be provided to adjust the artificial neural network. These values can be included by providing a contribution to some nodes of the artificial neural network. These nodes can be specific input nodes, but can generally be nodes of a hidden or processing layer. Such adaptive values can be weighted, for example, and added as a contribution to the weighted sum / correlation value of a given node.
[0101] The above description relates to neural network methods that can be applicable to many embodiments and implementations. However, it should be understood that many other types and architectures of neural networks can be used. In fact, many different methods for generating neural networks have been developed and are being developed, including neural networks using complex architectures and processes different from those described above. The method is not limited to any particular neural network method and any suitable method can be used without departing from the present invention.
[0102] An artificial neural network is adapted to a particular purpose through a training process that is used to adjust / tune / modify the weights and other parameters (e.g., biases) of the artificial neural network. It should be understood that many different training processes and algorithms are known for training artificial neural networks. Generally, the training is based on a large training set, where a large number of input data examples are provided to the network. Additionally, the output of the artificial neural network is typically compared (directly or indirectly) with the expected or ideal result. A cost function can be generated to reflect the desired result of the training process. In a typical scenario known as supervised learning, the cost function often represents the distance between the predicted result of a particular input data and the ground truth. Based on the cost function, the weights can be changed, and by repeating the process for the modified weights, the artificial neural network can be adjusted towards a state that minimizes the cost function.
[0103] More specifically, during the training step, the neural network can have two different information flows from input to output (forward pass) and from output to input (backward pass). In the forward pass, the data is processed by the neural network as described above, while in the backward pass, the weights are updated to minimize the cost function. Generally, this backpropagation follows the gradient direction of the cost function landscape. In other words, by comparing the predicted output and the ground truth for a batch of data inputs, it is possible to estimate the direction in which the cost function is minimized and propagate it backward by updating the weights accordingly. Other methods known for training artificial neural networks include, for example, the Levenberg - Marquardt algorithm, conjugate gradient methods, and Newton methods, etc.
[0104] In the current scenario, the training can specifically include a training set that includes a potentially large number of data pairs of captured audio signals and measured respiratory waveform signals. The measured respiratory waveform signals can be generated based on the sensor signals of a sensor that is arranged to measure an attribute that depends on the volume of air in the lungs. Thus, the training set can be used to perform the training of an artificial neural network, which includes linked audio / speech data / signals and sensor data / signals representing the measurement results of the volume of air in the lungs during speech.
[0105] In some embodiments, the training data can be the audio signals in a time period corresponding to the processing time interval of the artificial neural network being trained. For example, the number of samples in the training audio signals can correspond to the number of samples corresponding to the input nodes of the artificial neural network(s) being trained. Thus, each training example can correspond to one operation of the artificial neural network(s) being trained. However, typically, a batch of training samples is considered for each step to accelerate the training process. In addition, many upgrades to gradient descent can also accelerate convergence or avoid local minima in the cost function landscape.
[0106] Figure 4 An example of how training data can be generated through a dedicated test is illustrated. The speech audio signal 401 can be captured by the microphone 403 during the time interval when a person is speaking. The resulting captured audio signal can be fed as input data to the artificial neural network 405 (suitable processing (such as amplification, digitization, filtering) can be performed before generating the test audio signal fed to the artificial neural network). Additionally, the respiratory waveform signal 407 is determined from the sensor signals of a suitable lung volume sensor 409 located on the user, for example. For example, the sensor 409 can be a respiration-induced plethysmography (rip) sensor or a nasal cannula flow sensor.
[0107] A large number of such measurements can be made to generate a large number of data pairs of training audio signals and respiratory waveform signals. Then these signals can be applied and used to train the artificial neural network, where the cost function is determined as the difference between the respiratory waveform signal generated by the artificial neural network for the training audio signals and the measured training respiratory waveform signals. Then the artificial neural network can be adjusted based on such cost functions known in the art. For example, the cost value for each training audio signal and / or the cost value for a combined set of training downmixed audio signals can be determined (for example, determining the average cost value for the training set). Generally, the cost function will include at least one component that reflects the proximity of the generated signal to the reference signal, that is, the so-called reconstruction error. In some embodiments, the cost function will include at least one component that reflects the proximity of the generated signal to the reference signal from a perceptual perspective. InFigure 4 In the example of, the reference signal can specifically be the measured respiratory waveform signal.
[0108] As previously mentioned, the method can provide improved performance for a range of speech processing applications.
[0109] Especially for automatic speech recognition, the method can provide significantly improved performance by adjusting the process according to the current attributes of the speaker.
[0110] For example, as previously mentioned, the speech processor 103 can be arranged to determine the inhalation time period in the audio signal corresponding to the time when the speaker is breathing and thus not speaking. Then, the speech processor 103 can be arranged to exclude these inhalation time periods from the speech recognition processing applied to the audio signal. Therefore, it is possible to exclude from speech recognition the times in the respiratory waveform signal that may indicate times when speech is unlikely to be produced, thereby reducing the risk of erroneously detecting words when there is no speech. It can also improve accurate detection at other times because it is possible to estimate when words are likely to be spoken.
[0111] As another example, the respiratory waveform signal can be used to estimate the word rate, for example, according to the level of effort and distress in the respiratory pattern estimated from the respiratory waveform signal, and then adjust speech recognition.
[0112] In some embodiments, the speech processor 103 can be arranged to (first) perform speech recognition operations without considering the respiratory waveform signal or any parameters derived therefrom. However, instead of only providing an estimated term or sentence, speech recognition can generate multiple different candidates for estimating what may have been said. Then, the speech processor 103 can be arranged to select between the different candidates based on the respiratory waveform signal. For example, the active waveform for the estimated sentence can be compared with the respiratory waveform signal, and this selection can be based on the closeness of these waveform matches. For example, a sentence aligned with the inspiration interval can be selected instead of a sentence not aligned with the inspiration interval (and which, for example, includes speech during the estimated inspiration interval).
[0113] As another example, the selection between candidates can be based on the level of distress of the respiration. For example, when the user is relaxed, the word rate of speech is generally much higher than when the user is highly distressed and experiencing dyspnea (e.g., due to strenuous physical exercise). Therefore, the speech processor 103 can be arranged to select between the candidates based on the word rate of these candidates. For example, if the respiratory waveform signal indicates that the speaker is relaxed, then the candidate with the highest word rate can be selected, while if the respiratory waveform signal indicates that the speaker is experiencing dyspnea, then the candidate with the lowest word rate can continue to be selected.
[0114] Of course, in many scenarios, the respiratory waveform signal can be just one factor in choosing between candidates, and other parameters can also be considered. For example, in addition to candidate terms, speech recognition can also estimate a confidence or reliability value for each candidate / estimated result, and such a selection can further take into account such a confidence level.
[0115] A particular advantage of such an approach is the ability to achieve improved detection by further considering the speaker's respiratory attributes, but still being able to use a conventional speech recognition module. Thus, it is possible to use post-processing of the results from existing speech recognition algorithms to provide improved performance, while allowing backward compatibility and allowing the use of existing recognition algorithms.
[0116] In many embodiments, speech processing can be speech encoding, where, for example, the captured speech is encoded for efficient transmission or distribution over a suitable communication channel. As previously mentioned, more efficient encoding can typically be achieved simply by not encoding any signal during the inhalation period, or by allocating fewer bits to the signal in the entropy encoding during the inhalation period.
[0117] As another example, information about the respiratory wave can be used to control speech manipulation (e.g., hiding the state of the speaker's fatigue). For example, the system can detect a state of dyspnea by detecting an increase in the speaker's respiratory rate. Next, a speech synthesizer such as a Wavenet model can be used to resynthesize a new speech signal that is similar to the speech of the same speaker, but has a different respiratory wave pattern corresponding to a speaker with a lower respiratory rate.
[0118] In some embodiments, speech processing can be speech enhancement. Such speech enhancement can be used, for example, to provide a clearer speech signal, where, for example, other audio sources and noise in the audio signal are reduced. For example, as previously described, the audio signal can be attenuated during the inhalation period. As another example, different frequencies can be attenuated according to the respiratory waveform signal, thereby specifically attenuating the frequencies that tend to be relatively dominant for some respiratory noises. As another example, the respiratory waveform signal can indicate transitions or phases in the respiratory cycle that tend to be associated with specific sounds, and the speech processing can be arranged to generate corresponding signals and apply them in antiphase to provide an audio cancellation effect for such sounds.
[0119] In many embodiments, the speech processor 103 can be arranged to evaluate the respiratory waveform signal to determine one or more parameters of the speaker's respiration. It can also be arranged to adjust the speech processing based on the determined parameter values.
[0120] The speech processor 103 is arranged to determine a respiratory rate estimation result based on a respiratory waveform signal and to adjust speech processing based on the respiratory rate.
[0121] Since respiration is inherently a repetitive / periodic process, the respiratory waveform signal will also be a periodic signal (or at least have a strong periodic component), and the respiratory rate can be determined by determining the repetition rate / duration of the periodic component of the respiratory waveform signal. For example, the respiratory waveform signal can be correlated with itself, and the duration between correlation peaks can be determined and used as a measure of the respiratory rate. It should be understood that many different techniques and algorithms are known for detecting periodic components in a signal, and any method can be used without departing from the present invention.
[0122] The respiratory rate can indicate a plurality of parameters that can affect a person's speech. For example, the respiratory rate can be a strong indicator of the degree of relaxation or distress / fatigue of the speaker, and as previously mentioned, this can affect speech in different ways (including, for example, word rate, amount of respiratory noise, etc.). As previously mentioned, speech processing can be adapted to reflect these parameters.
[0123] The respiratory rate can also provide information related to the timing of the inhalation time intervals / segments. In particular, the frequency and time between inhalation time intervals can be directly given by the respiratory rate.
[0124] As another example, the respiratory rate can be used to adjust speech processing by having a specific speech preprocessing method selected based on the speaker's respiratory rate. For example, different parameters of speech activity detection or speech enhancement algorithms are used when the speaker has a lower respiratory rate than an elevated or higher respiratory rate.
[0125] The respiratory rate provides a particularly accurate adjustment of speech processing in many scenarios and for many situations and speech signals and applications.
[0126] In some embodiments, the speech processor 103 is arranged to determine a speech word rate based on a respiratory waveform signal and to adjust speech processing based on the speech word rate.
[0127] For example, the speech word rate can be determined by the output of an automatic speech recognition method.
[0128] As described, the speaker's word rate can vary significantly based on the speaker's level of relaxation or stress. The respiratory waveform signal can be used to determine its level and thus the degree of dyspnea the speaker is currently experiencing. A word rate for a given dyspnea level can be determined (e.g., as a predetermined value for a given level of respiratory distress). Speech processing can then be performed based on the determined word rate. For example, as previously described, the speech processor 103 can perform speech recognition that provides multiple possible candidates, and can select the candidate having a word rate that most closely matches the word rate determined from the respiratory waveform signal. In some embodiments, the speech recognition can thus provide several candidates for a speech utterance, and based on the expected word rate determined from the respiratory waveform signal, the speech processor 103 can select, in the case of high dyspnea, the utterance corresponding to the lowest word rate and having a long speech pause duration.
[0129] In fact, many speech recognition systems have difficulty in the case of low word rates and pause in dyspneic speech. However, in this method, the parameters of the speech recognition can be adjusted to accommodate this, or the output of the speech recognition system can be processed to accommodate this condition.
[0130] As another example, the word rate can be used to adjust speech processing by having specific settings based on the word rate in the speech preprocessing or speech coding algorithm. The settings can, for example, trigger the use of different quantization vector codebooks in the encoder and decoder in the speech encoder.
[0131] In some embodiments, the speech processor 103 is arranged to: determine the level of respiratory sound during inspiration based on the respiratory waveform signal, and adjust speech processing based on the level of speech respiratory sound during inspiration.
[0132] For example, the level of respiratory sound can be determined by calculating the root mean square amplitude level of the inspiratory signal, and optionally, weighting can be applied to account for the relative loudness perceived by the human ear.
[0133] For example, in an audio recording, the inspiratory signal may distract the listener from the content being spoken or sung. In this case, the sound level of the detected inspiratory signal can be a measure of how much the inspiratory signal needs to be attenuated to become inaudible.
[0134] In some embodiments, the speech processor 103 can be arranged to consider other data, rather than just the audio signal (and the respiratory waveform signal derived therefrom).
[0135] In some embodiments, the apparatus may include a second receiver 107 arranged to receive data indicative of the heart rate of the speaker. The heart rate data may be received, for example, from a body-worn heart rate monitor, a pulse meter, or any other device arranged to measure a person's pulse or heart rate.
[0136] The heart rate may provide further information about the user's state and the likelihood of dyspnea speech, and the second position parameter may be arranged to further take this into account when adjusting speech processing.
[0137] For example, when storing multiple different speech processes or different parameters for a given speech process, these different processes or parameters may be divided into different ranges of parameters determined from the respiratory waveform signal, but these ranges are further subdivided into multiple ranges of heart rate. Thus, a more accurate and / or more fine-grained adjustment of speech processing can be achieved.
[0138] For example, the method may provide particularly advantageous and synergistic effects for embodiments that adjust speech processing based on parameters of the respiratory waveform signal, which may not directly indicate, for example, a person's physical exercise, but may also be due to other problems. For example, it may allow a distinction to be made between dyspnea caused by strenuous exercise and dyspnea caused by medical respiratory problems (and thus also occurring in a relaxed state).
[0139] As a specific example, the speech processor 103 may predict the respiratory rate or word rate based on the detected heart rate and use the estimated result to set parameters of the speech preprocessing or encoding algorithm, as discussed above. This is based on the consideration that the heart rate and respiratory rate are generally correlated. This is beneficial for obtaining the best performance of the system from the speaker's first utterance before a model for estimating the respiratory wave can access a long enough speech record to give a first estimated result. This is also advantageous in cases where the speaker uses only very short utterances (e.g., voice commands) to control a music playback application in a wearable audio player.
[0140] In some embodiments, the second receiver 107 is arranged to receive data indicative of the speaker's activity.
[0141] In some embodiments, such data may be provided by the user directly providing user input (e.g., manually selecting between a set of predetermined activities). As another example, the activity may be detected by detecting the use of an external device. For example, the speech processing device may be coupled to exercise equipment, and if the exercise equipment is in use, the exercise equipment may transmit data to the speech processing device to indicate that the exercise equipment is in use, and the speech processing device may be arranged to adjust speech processing according to the activity.
[0142] As another example, the activity data may represent detected physical activities (e.g., sitting, walking, running, cycling, or swimming). It may also represent specific situations (e.g., walking up stairs, sitting in a meeting, giving a presentation in front of an audience, or driving a car).
[0143] As an example, similar to the heart rate case above, the speech processor 103 may be arranged to adjust speech processing based on the activity data to trigger speech processing or speech encoding patterns optimized for a breathing rate that is typically associated with or has been observed in the past in the context of the detected activity. This is also beneficial at the start of a speech session and in the case of very short speech utterances such as voice commands.
[0144] In some embodiments, the speech processor 103 may include a human breathing model that can be used to determine the breathing attributes of the speaker. Such a breathing model may be adjusted based on the received heart rate and / or activity. In some embodiments, the speech processor 103 may estimate a computational model that produces an estimate of the breathing characteristics based on activity characteristics corresponding to the current physical activity and vital signs measured using the wearable device.
[0145] For example, the human breathing model may model specific biophysical processes in the body organs, or it may be a functional high-level metabolic model. For example, metabolic equivalents (METs) relate to the rate of oxygen uptake of the body for a given activity as a multiple of the resting VO2. Based on known activity METs, effective lung capacity, VO2, and the oxygen concentration in the air, we can obtain an indirect estimate of the breathing rate required to meet the oxygen demand of the subject. This can be adjusted based on the heart rate through a more granular model that includes specific MET values (based on heart rate and activity intensity measurements) for different intensities of a specific physical activity (e.g., from a light jog to a full-speed uphill run).
[0146] In some embodiments, the heart rate and / or activity data may be specifically used to set initial operating parameters for speech processing. In many embodiments, the operating parameters may subsequently be adjusted and changed based on the breathing waveform signal.
[0147] In fact, in many embodiments, determining the breathing waveform signal from the audio signal and determining the breathing parameters from the breathing waveform signal may require some processing, analysis, and a sufficient duration of the audio signal. Accordingly, there may be a significant delay in adjusting speech processing for a specific breathing condition. The heart rate and / or activity data can be used to provide a more accurate guess / estimate of the appropriate operating parameters.
[0148] For example, when starting up, the second receiver 107 can receive heart rate data from a chest-worn heart rate monitor. Based on the current heart rate / pulse, the speech processor 103 can set a plurality of operating parameters (e.g., by extracting appropriate values from a look-up table). Then, it can start speech processing based on these parameters, while starting to determine the respiratory waveform signal and determine the respiratory parameters based on the respiratory waveform signal. When a sufficiently accurate respiratory waveform signal / parameter has been generated, it can continue to determine the appropriate operating parameters therefor (again, e.g., by extracting values from a previously populated look-up table (e.g., based on manual testing / optimization)). Then, the speech processor 103 can gradually transition the operating values from their initial value set to the expected values determined based on the respiratory waveform signal based on the heart rate.
[0149] (One or more) audio devices can be specifically implemented in one or more appropriately programmed processors. For example, the artificial neural network of the determiner can be implemented in one or more such appropriately programmed processors. Different functional blocks (especially in the artificial neural network) can be implemented in separate processors and / or can be implemented, for example, in the same processor. Examples of suitable processors are provided below.
[0150] Figure 5 is a block diagram illustrating an example processor 500 according to an embodiment of the present disclosure. The processor 500 can be used to implement one or more processors that implement the device or its elements as described above (specifically including one or more artificial neural networks). The processor 500 can be any suitable type of processor, including but not limited to a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable array (FPGA) (wherein the FPGA has been programmed to form a processor), a graphics processing unit (GPU), an application specific integrated circuit (ASIC) (wherein the ASIC has been designed to form a processor), or a combination thereof.
[0151] The processor 500 can include one or more cores 502. The cores 502 can include one or more arithmetic logic units (ALUs) 504. In some embodiments, in addition to or instead of the ALU 504, the cores 502 can include a floating point logic unit (FPLU) 506 and / or a digital signal processing unit (DSPU) 508.
[0152] The processor 500 can include one or more registers 512 communicatively coupled to the cores 502. The registers 512 can be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 512 can be implemented using static memory. The registers can provide data, instructions, and addresses to the cores 502.
[0153] In some embodiments, the processor 500 may include one or more levels of cache memory 510 communicatively coupled to the core 502. The cache memory 510 may provide computer-readable instructions for execution to the core 502. The cache memory 510 may provide data for processing by the core 502. In some embodiments, the computer-readable instructions may have been provided to the cache memory 510 by local memory (e.g., local memory attached to the external bus 516). The cache memory 510 may be implemented using any suitable cache memory type (e.g., metal-oxide semiconductor (MOS) memory such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology).
[0154] The processor 500 may include a controller 514 that may control the input to the processor 500 from other processors and / or components included in the system and / or the output from the processor 500 to other processors and / or components included in the system. The controller 514 may control the data paths in the ALU 504, FPLU 506, and / or DSPU 508. The controller 514 may be implemented as one or more state machines, data paths, and / or dedicated control logic units. The gates of the controller 514 may be implemented as discrete gates, FPGAs, ASICs, or any other suitable technology.
[0155] The registers 512 and the cache memory 510 may communicate with the controller 514 and the core 502 via internal connections 520A, 520B, 520C, and 520D. The internal connections may be implemented as buses, multiplexers, crossbars, and / or any other suitable connection technology.
[0156] Input and output for the processor 500 may be provided via the bus 516, which may include one or more wires. The bus 516 may be communicatively coupled to one or more components of the processor 500 (e.g., the controller 514, the cache memory 510, and / or the registers 512). The bus 516 may be coupled to one or more components of the system.
[0157] The bus 516 can be coupled to one or more external memories. The external memories can include a read-only memory (ROM) 532. The ROM 532 can be a mask ROM, an electrically programmable read-only memory (EPROM), or any other suitable technology. The external memories can include a random access memory (RAM) 533. The RAM 533 can be a static RAM, a battery-backed static RAM, a dynamic RAM (DRAM), or any other suitable technology. The external memories can include an electrically erasable programmable read-only memory (EEPROM) 535. The external memories can include a flash memory 534. The external memories can include a magnetic storage device (e.g., a disk 536). In some embodiments, the external memories can be included in the system.
[0158] The present invention can be implemented in any suitable form, including hardware, software, firmware, or any combination of these items. The present invention can optionally be implemented, at least in part, as computer software running on one or more data processors and / or digital signal processors. The elements and components of the embodiments of the present invention can be physically, functionally, and logically implemented in any suitable manner. In fact, the functions can be implemented in a single unit, in multiple units, or as part of other functional units. For this reason, the present invention can be implemented in a single unit, or can be physically and functionally distributed among different units, circuits, and processors.
[0159] Although the present invention has been described in connection with some embodiments, the present invention is not intended to be limited to the specific forms set forth herein. Instead, the scope of the present invention is limited only by the appended claims. Additionally, although it may seem that features are described in connection with specific embodiments, those skilled in the art will recognize that the various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0160] In addition, although listed individually, multiple devices, components, circuits, or method steps may also be implemented by, for example, a single circuit, unit, or processor. Additionally, although individual features may be included in different claims, these features may be advantageously combined, and the inclusion of these features in different claims does not mean that the combination of these features is not feasible and / or advantageous. Moreover, including a feature in one type of claim does not imply a limitation to that type, but rather indicates that the feature equivalently applies to other types of claims where appropriate. Furthermore, the order of features in a claim does not mean that the features must be in any particular order in which they operate, and in particular, the order of individual steps in a method claim does not mean that the steps must be performed in that order. Instead, these steps may be performed in any suitable order. Additionally, a singular reference does not exclude a plurality. Thus, references to "a", "an", "first", "second", etc. do not exclude a plurality. The reference signs in the claims are provided only as clarifying examples and should not be construed as limiting the scope of the claims in any way.
[0161] Generally, examples of the method are indicated by the following embodiments.
[0162] Example:
[0163] 1. An apparatus for speech processing of an audio signal, the apparatus comprising:
[0164] An input unit (101) arranged to receive an audio signal comprising a speech audio component of a speaker;
[0165] A determiner (105) arranged to generate a respiration waveform signal based on the audio signal, the respiration waveform signal indicating the volume of air in the lungs of the speaker as a function of time; and
[0166] A speech processor (103) arranged to perform speech processing on the audio signal, the speech processing being dependent on the respiration waveform signal.
[0167] 2. The apparatus according to embodiment 1, wherein the determiner (105) comprises:
[0168] A trained artificial neural network having input nodes for receiving samples of the audio signal and output nodes arranged to provide samples of the respiration waveform signal.
[0169] 3. The apparatus according to any one of embodiments 1 to 2, wherein the speech processing comprises speech enhancement processing.
[0170] 4. The apparatus according to any one of embodiments 1 to 3, wherein the speech processing comprises speech recognition processing.
[0171] 5. The apparatus according to embodiment 4, wherein the speech processing includes a speech recognition process of generating a plurality of candidate recognized words based on the audio signal; and the speech processor (103) is arranged to select among the candidate recognized words in response to the respiration waveform signal.
[0172] 6. The apparatus according to any one of embodiments 1 to 5, wherein the speech processor (103) is arranged to: determine a respiration rate estimation result based on the respiration waveform signal, and adjust the speech processing based on the respiration rate estimation result.
[0173] 7. The apparatus according to any one of embodiments 1 to 6, wherein the speech processor (103) is arranged to: determine a speech word rate based on the respiration waveform signal, and adjust the speech processing based on the speech word rate.
[0174] 8. The apparatus according to any one of embodiments 1 to 7, wherein the speech processor (103) is arranged to: determine a timing attribute of an inhalation interval based on the respiration waveform signal, and adjust the speech processing based on the timing attribute of the inhalation interval.
[0175] 9. The apparatus according to embodiment 8, wherein the speech processor (103) is arranged to: determine an inhalation time period of the audio signal based on the timing attribute, and attenuate the audio signal during the inhalation time period.
[0176] 10. The apparatus according to embodiment 8 or 9, wherein the speech processor (103) is arranged to: determine an inhalation time period of the audio signal based on the timing attribute, and exclude the inhalation time period from the speech recognition process applied to the audio signal.
[0177] 11. The apparatus according to any one of embodiments 1 to 2, wherein the speech processor (103) is arranged to: determine a respiration sound level during inhalation based on the respiration waveform signal, and adjust the speech processing based on the respiration sound level during inhalation.
[0178] 12. The apparatus according to any one of embodiments 1 to 11, further comprising: a receiver (107) arranged to receive speaker data indicating at least one of the heart rate of the speaker and the activity of the speaker; and wherein the speech processor (103) is arranged to adjust the speech processing based on the speaker data.
[0179] 13. The apparatus according to embodiment 12, wherein the speech processor (103) includes a human respiration model for determining a respiration attribute of the speaker, and the speech processor (103) is further arranged to: adjust the speech processing in response to the respiration attribute; and adjust the human respiration model according to the speaker data.
[0180] 14. The apparatus according to embodiment 12 or 13, wherein the speech processor (103) is arranged to set initial operation parameters for the speech processing in response to the speaker data.
[0181] 15. A method for speech processing of an audio signal, the method comprising:
[0182] Receiving an audio signal including a speech audio component of a speaker;
[0183] Generating a respiration waveform signal based on the audio signal, the respiration waveform signal indicating a volume of air in the lungs of the speaker as a function of time; and
[0184] Performing speech processing on the audio signal, the speech processing being dependent on the respiration waveform signal.
[0185] More specifically, the invention is defined by the appended claims.
Claims
1. An apparatus for speech processing of an audio signal, the apparatus comprising: An input unit (101) arranged to receive an audio signal including a speech audio component of a speaker; A determiner (105) arranged to generate a respiration waveform signal based on the audio signal, the respiration waveform signal indicating the volume of the speaker's pulmonary air as a function of time; And A speech processor (103) arranged to perform speech processing of the audio signal, the speech processing being dependent on the respiration waveform signal; And wherein the speech processor (103) is arranged to: determine a respiration rate estimation result based on the respiration waveform signal, and adjust the speech processing based on the respiration rate estimation result.
2. The device according to claim 1, wherein, The determiner (105) comprises: A trained artificial neural network having input nodes for receiving samples of the audio signal and output nodes arranged to provide samples of the respiration waveform signal.
3. The device according to any one of the preceding claims, wherein, The speech processing includes speech enhancement processing.
4. The apparatus according to any one of the preceding claims, wherein, The speech processing includes speech recognition processing.
5. The device according to claim 4, wherein, The speech processing includes speech recognition processing for generating a plurality of candidate recognized words based on the audio signal; and the speech processor (103) is arranged to select between the candidate recognized words in response to the respiration waveform signal.
6. The apparatus according to any one of the preceding claims, wherein, The speech processor (103) is arranged to: determine a speech word rate based on the respiration waveform signal, and adjust the speech processing based on the speech word rate.
7. The apparatus according to any one of the preceding claims, wherein, The speech processor (103) is arranged to: determine a timing attribute of an inhalation interval based on the respiration waveform signal, and adjust the speech processing based on the timing attribute of the inhalation interval.
8. The apparatus according to claim 7, wherein The speech processor (103) is arranged to: determine an inhalation time period of the audio signal based on the timing attribute, and attenuate the audio signal during the inhalation time period.
9. The apparatus according to claim 7 or 8, wherein, The speech processor (103) is arranged to: determine an inhalation time period of the audio signal based on the timing attribute, and exclude the inhalation time period from speech recognition processing applied to the audio signal.
10. The apparatus according to any one of the preceding claims, wherein, The speech processor (103) is arranged to: determine a respiration sound level during inhalation based on the respiration waveform signal, and adjust the speech processing based on the respiration sound level during inhalation.
11. The device according to any one of the preceding claims, further comprising: A receiver (107) arranged to receive speaker data indicating at least one of the speaker's heart rate and the speaker's activity; And wherein the speech processor (103) is arranged to adjust the speech processing based on the speaker data.
12. The device according to claim 11, wherein The speech processor (103) includes a human respiration model for determining the respiration attributes of the speaker, and the speech processor (103) is further arranged to: adjust the speech processing in response to the respiration attributes; and adjust the human respiration model based on the speaker data.
13. The device according to claim 11 or 12, wherein, The speech processor (103) is arranged to set initial operation parameters for the speech processing in response to the speaker data.
14. A method for speech processing of an audio signal, the method comprising: Receiving an audio signal including a speech audio component of a speaker; Generating a respiration waveform signal based on the audio signal, the respiration waveform signal indicating the volume of air in the lungs of the speaker as a function of time; And Performing speech processing on the audio signal, the speech processing being dependent on the respiration waveform signal; And wherein the speech processing includes: determining a respiration rate estimation result based on the respiration waveform signal, and adjusting the speech processing based on the respiration rate estimation result.
15. A computer program product comprising a computer program code unit adapted to perform all the steps according to claim 14 when the program is run on a computer.
Citation Information
Patent Citations
Virtual Trainer
US20090098981A1