Audio signal processing
By segmenting audio signals based on a speaker's respiration waveform, the apparatus adapts speech processing to improve flexibility and efficiency, addressing suboptimal performance in varying conditions.
Patent Information
- Application Number
- JP2025538228
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-13
- Filing Date
- 2024-01-02
- Publication Date
- 2026-01-23
AI Technical Summary
Existing speech processing technologies are not optimal in scenarios where actual conditions differ from expected nominal conditions, leading to suboptimal performance, complexity, and resource inefficiency, particularly in adapting to variations in speaker properties and activities.
An apparatus that segments audio signals based on a speaker's respiration waveform signal to generate speech segments, using an artificial neural network trained with respiratory waveform signals to adapt speech processing, allowing for improved flexibility and reduced complexity.
This approach provides adaptive speech processing that closely correlates with speaker variations, enhancing speech quality and recognition, reducing computational burden, and improving user experience.
Smart Images

Figure 2026502451000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to performing speech processing of audio signals, and in particular, but not exclusively, to automatic speech recognition or speech enhancement processing of audio signals to capture speaker information. [Background technology]
[0002] Speech processing of audio signals is widely applied in a variety of practical applications and is becoming increasingly important and meaningful for many everyday activities and devices.
[0003] For example, speech enhancement is widely applied to improve the intelligibility and reproducibility of speech captured by audio signals, e.g., sounds captured in real-world environments. Speech coding is also frequently performed from captured audio. For example, speech coding of captured audio signals is an integral part of smartphones. Another speech processing application that has become increasingly frequent in recent years is the application of speech recognition, e.g., to provide a user interface for devices.
[0004] For example, a voice interface on a personal assistant or home speaker may allow one to control media, navigate functions, track, obtain various information services, etc., simply by using the voice interface and while performing other activities, such as exercising.
[0005] Voice interfaces based on Automatic Speech Recognition (ASR) have also become very important in controlling various devices, such as wearable devices, wearable devices, IoT consumer devices, health devices, and fitness devices. Voice is the most natural interface for interacting with devices during activities involving large user movements and fluctuations. As an example, US20090098981A1 discloses a fitness device with a voice interface used for communication between a user and a virtual fitness coach.
[0006] However, while significant effort has been expended in developing and optimizing speech processing algorithms for various applications, resulting in many highly advantageous and efficient approaches, these may not be optimal in all situations. For example, many speech processing operations are developed for specific nominal conditions, such as nominal speaker properties or nominal acoustic environments. Many speech processing applications may not provide optimal performance in scenarios where actual conditions differ from the expected nominal conditions. Additionally, some speech processing algorithms may be unnecessarily complex or resource-intensive. Adapting speech processing to compensate for various properties tends to result in less flexible, more complex, and more resource-intensive implementations, often resulting in suboptimal performance. Summary of the Invention [Problem to be solved by the invention]
[0007] Therefore, improved approaches to speech processing would be advantageous, particularly approaches that allow for greater flexibility, greater adaptability, improved performance, improved quality such as speech enhancement, encoding, and / or recognition, reduced complexity and / or resource usage, improved remote control of audio processing, improved adaptation to variations in speaker properties and / or activity, reduced computational burden, improved user experience, easier implementation, and / or an improved spatial audio experience. [Means for solving the problem]
[0008] Accordingly, the Invention seeks to preferably mitigate, reduce or eliminate one or more of the above mentioned disadvantages singly or in any combination.
[0009] According to one aspect of the present invention, there is provided an apparatus for speech processing of an audio signal, the apparatus comprising: an input configured to receive an audio signal comprising a speech audio component of a speaker; a determiner configured to generate a respiration waveform signal indicative of the speaker's lung air volume as a function of time; a segmenter configured to segment the audio signal in response to the respiration waveform signal to generate speech segments of the audio signal; and a speech processing circuit configured to perform speech processing of the audio signal, the speech processing being segment-based processing applied to the speech segments.
[0010] This approach provides improved speech processing in many embodiments and scenarios. For many signals and scenarios, this approach can provide speech processing that more closely reflects variations in the speaker's speech properties and characteristics. For example, it can adapt to a speaker's level of fatigue, breathing difficulties, etc. This approach can provide accelerated adaptation and, in particular, improved adaptive speech processing. This approach provides efficient implementation and, in many embodiments, allows for reduced complexity and / or resource usage.
[0011] Adapting segmentation based on respiration waveform signals, in particular, allows for closer correlation and adaptation of speech processing to a speaker's actual, current speech. Adapting segmentation based on respiration waveform signals typically allows for closer correlation to sentence structure, words, etc. For example, in some cases, segmentation is more likely to match the perceived content of the speech or the speaker's particular speaking patterns.
[0012] Segment-based processing can be sequential processing, in which speech segments are processed sequentially. Segment-based processing can involve processing in which an output speech signal segment for a time interval of an audio signal is generated based on the speech segments of the audio signal for that time interval. In some embodiments, the speech processing of a speech segment can be independent of other parts of the audio signal other than the speech segment. In some embodiments, the speech processing of a speech segment does not include samples of the audio signal other than samples belonging to the speech segment.
[0013] This approach can sense the temporal breathing patterns of a speaker (or multiple speakers), for example, using an audio signal and / or dedicated sensor input. The temporal breathing patterns can be represented by a breathing waveform signal. Segmentation of the speech can be performed according to respiratory events, such as inhalations or speech and silent breathing. The approach can then process the speech in segments determined by the respiratory events. The breathing waveform signal can represent / reflect / be a measure of the speaker's temporal breathing patterns (particularly throughout a breathing cycle). The breathing waveform signal can represent / reflect / be a measure of the speaker's temporal breathing patterns. The breathing waveform signal can represent / reflect / be a measure of the variation (as a function of time) of lung air volume during (and typically throughout) a breathing cycle.
[0014] In many embodiments, the respiration waveform signal can be independent of the level of the audio signal. In many embodiments, the respiration waveform signal is independent of the audio signal.
[0015] According to an optional feature of the invention, the determiner is configured to determine a respiration waveform signal from the audio signal.
[0016] This approach provides accelerated adaptation, and in particular can provide improved adaptive speech processing without relying on or requiring other input from other devices that measure speaker characteristics. This approach provides efficient implementation and, in many embodiments, allows for reduced complexity and / or resource usage.
[0017] According to an optional feature of the invention, the determiner includes a trained artificial neural network having an input node for receiving the sample of the audio signal and an output node configured to provide the sample of the respiration waveform signal.
[0018] This results in particularly advantageous behavior and performance in many scenarios and applications.
[0019] This approach can provide a particularly advantageous configuration that, in many embodiments and scenarios, enables enhanced and / or improved utilization of artificial neural networks in adapting and optimizing audio processing, typically including audio enhancement and / or recognition.
[0020] An artificial neural network is a trained artificial neural network.
[0021] The artificial neural network can be an artificial neural network trained with training data including training speech audio signals and training respiratory waveform signals generated from measurements of respiratory waveforms, the training using a cost function that compares the training respiratory waveform signals to the respiratory waveform signals generated by the artificial neural network for the training speech audio signals. The artificial neural network is a trained artificial neural network that has been trained with training data including training speech audio signals representing various relevant speakers in various different states and performing different activities.
[0022] The artificial neural network can be an artificial neural network trained with training data having training input data including a training speech audio signal, and trained with a cost function including a contribution indicative of the difference between the measured training respiration waveform signal and the respiration waveform signal generated by the artificial neural network in response to the training speech audio signal.
[0023] According to an optional feature of the invention, the determiner is configured to detect extrema in the respiration waveform signal and determine the speech segments in response to the extrema.
[0024] According to an optional feature of the invention, the determiner is configured to detect local maxima in the respiratory waveform signal and determine the speech segments in response to the local maxima.
[0025] This can provide particularly advantageous speech segmentation, which in many embodiments allows for improved speech segmentation that adapts closely to the content of the utterance as reflected in the speaker's speech behavior and / or sentence structure, etc.
[0026] In many embodiments, the determiner may be configured to determine the start of the audio segment depending on the timing of the local maximum, and in particular to determine the start of the audio segment as having a fixed time offset relative to the time of the local maximum (this offset may be zero).
[0027] According to an optional feature of the invention, the determiner is configured to detect minima in the respiratory waveform signal and determine the speech segments in response to the minima.
[0028] This can provide particularly advantageous speech segmentation, which in many embodiments allows for improved speech segmentation that adapts closely to the content of the utterance as reflected in the speaker's speech behavior and / or sentence structure, etc.
[0029] In many embodiments, the determiner may be configured to determine the end of the audio segment depending on the timing of the local minimum, and in particular to determine the end of the audio segment as having a fixed time offset relative to the time of the local minimum (the offset may be zero, or often negative).
[0030] According to an optional feature of the invention, the determiner is configured to determine the at least one audio segment as a segment of the audio signal between a time point of a maximum in the respiratory waveform signal and a time point of a (shorter than) minimum in the respiratory waveform signal.
[0031] This can provide particularly advantageous speech segmentation, which in many embodiments allows for improved speech segmentation that adapts closely to the content of the utterance as reflected in the speaker's speech behavior and / or sentence structure, etc.
[0032] The determiner may be specifically configured to determine the start of the audio segment as the time of the local maximum and the end of the audio segment as the time of the immediately following local minimum.
[0033] In accordance with an optional feature of the invention, the segmenter is configured to divide the audio signal into voiced and non-voiced segments.
[0034] According to an optional feature of the invention, the segmenter is configured to determine an expiratory time interval and an inhalation time interval from the respiratory waveform signal, and is configured to determine speech segments as segments of the audio signal during the expiratory time interval and to determine non-speech segments as segments of the audio signal during the inhalation time interval.
[0035] In accordance with an optional feature of the invention, the speech processing circuitry is configured to select speech segments containing speech from the speaker rather than speech from other speakers, the selection being dependent on the respiration waveform signal.
[0036] This approach allows for speech segmentation that reflects the currently active speaker in a scenario with multiple potential speakers, which leads to improved speech segmentation and, consequently, improved speech processing.
[0037] According to an optional feature of the invention, the apparatus further comprises a sensor input configured to receive a chest sensor signal, the determiner being configured to determine a respiration waveform signal from the chest sensor signal.
[0038] This allows for improved and / or easier operation in many embodiments and scenarios, often allowing for improved speech segmentation and resulting speech processing.
[0039] According to an optional feature of the invention, the apparatus further comprises a video input configured to receive a video signal comprising a video image of the speaker, and the determiner configured to determine a respiration waveform signal from the video image.
[0040] This allows for improved and / or easier operation in many embodiments and scenarios, often allowing for improved speech segmentation and resulting speech processing.
[0041] In accordance with an optional feature of the invention, the speech processing includes speech enhancement processing.
[0042] This approach provides improved speech enhancement and allows for the generation of improved speech signals in many scenarios.
[0043] In accordance with an optional feature of the invention, the speech processing includes speech recognition processing.
[0044] This approach improves speech recognition, allowing for more accurate detection of words, terms and sentences in many scenarios, and in particular improving speech recognition of speakers in a variety of situations, conditions and activities.
[0045] One aspect of the present invention provides a method comprising the steps of receiving an audio signal including a speech audio component of a speaker; generating a respiration waveform signal indicative of the speaker's lung air volume as a function of time; segmenting the audio signal in response to the respiration waveform signal to generate speech segments of the audio signal; and performing speech processing of the speech signal, wherein the speech processing is segment-based processing applied to the speech segments.
[0046] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief explanation of the drawings]
[0047] Embodiments of the present invention will now be described, by way of example only, with reference to the drawings in which: [Figure 1] 1 illustrates some elements of an example audio device according to some embodiments of the present invention. [Figure 2] FIG. 1 shows an example of the structure of an artificial neural network. [Figure 3] FIG. 1 shows an example of a node in an artificial neural network. [Figure 4] Diagram showing some elements of an example training setup for an artificial neural network. [Figure 5] FIG. 10 is a diagram showing an example of a respiration waveform signal. [Figure 6] FIG. 1 illustrates some elements of a possible configuration of a processor for implementing elements of an apparatus according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0048] 1 illustrates some elements of a speech processing apparatus according to some embodiments of the present invention, which is suitable for providing improved speech processing for speakers in many different environments and scenarios.
[0049] The audio processing device comprises a receiver 101 configured to receive an audio signal including a speech audio component. The audio signal may in particular be a microphone signal representing audio captured by a microphone. The audio signal may in particular be a voice signal, and in many embodiments a voice signal captured by a microphone positioned to capture the voice of a single person, such as a headset microphone, a smartphone, or a body-worn microphone. Thus, the audio signal may include not only the speech audio component (hereinafter also referred to as the speech signal) but also other sounds, such as background sounds from the environment.
[0050] The audio processing device further comprises an audio processing circuit 103 configured to apply audio processing to the received audio signal, and thus in particular to the audio component of the audio signal. The audio processing can in particular be a speech enhancement process, a speech coding process or a speech recognition process.
[0051] In many practical applications, speech processing is performed segmentally, where an audio signal is divided into segments and each segment is processed independently. In some applications, an audio / voice signal is fixedly divided into consecutive segments, which are processed individually. For example, in many applications, speech processing algorithms are based on dividing an input audio signal into periodic time intervals of fixed duration and processing these segments individually. However, in the speech processing device of FIG. 5, speech processing is not performed based on predetermined fixed segments / time intervals, but rather on segmenting the audio signal. In particular, the speech processing device includes a determiner 105 configured to determine a respiration waveform signal indicative of the air volume in a speaker's lungs as a function of time, and a segmenter 107 configured to segment the audio signal to generate speech segments of the audio signal in response to the respiration waveform signal. The speech processing circuit 103 is then configured to process the audio signal based on segment-based processing applied to the speech segments generated by the segmenter 107.
[0052] The respiration waveform signal indicates / is a measure of / reflects / represents the air volume in a speaker's lungs as a function of time, specifically the variation in air volume in the lungs during a breathing cycle. It thus reflects the speaker's breathing and the flow of air into and out of the speaker's lungs. The respiration waveform signal can thus be a time-varying signal that reflects the variation in the speaker's breathing, and in particular the changes in lung volume due to the speaker's breathing during and throughout the breathing cycle. In some embodiments, the respiration waveform signal represents current lung volume. In other embodiments, the respiration waveform signal can represent changes in current lung volume, for example, when the respiration waveform signal represents inhaled / exhaled airflow.
[0053] In many embodiments, the audio processing device may have a sensor input 109 configured to receive a sensor signal from which a respiration waveform signal is determined.
[0054] The sensor signal can be a chest sensor signal, particularly a signal from a chest sensor that is part of a chest belt worn around a person's chest / body region. Such a sensor can directly generate a sensor signal that reflects the circumference of the chest, or can generate a sensor signal that directly reflects more localized chest expansion and contraction. Such a sensor signal can directly reflect the time variation of lung air volume and, in many embodiments, can be used directly as a respiration waveform signal. In many embodiments, some low-pass filtering, noise suppression, and amplification can be applied to the input sensor signal to generate the respiration waveform signal.
[0055] The sensor signal can be, in particular, from a breathing belt sensor, for example, or from other sensors that sense chest or abdominal movement of the speaker (or each speaker). Breathing sensors can also be incorporated into other devices, such as car safety belts.
[0056] In some embodiments, the sensor signal does not directly reflect chest contraction and expansion, but can be, for example, from an optical sensor such as a video camera. The sensor signal can be, for example, a video signal from a camera capturing a speaker. The determiner 105 can be configured to analyze the video image to determine the movement of the speaker's chest and, therefrom, determine a respiratory waveform signal that reflects the expansion / contraction of the speaker's chest and, therefore, the variation in lung volume. It should be understood that various approaches for evaluating images to determine the variation in lung volume are known and will not be further described herein for the sake of brevity. For example, US9301710B2 discloses a method for estimating a subject's respiratory rate based on changes in illumination patterns in a video signal showing the subject's chest region.
[0057] As another example, the sensor signal can be from a thermal or flow-based sensor that monitors airflow through the mouth and nose.
[0058] In many embodiments, the determiner / decision circuitry 105 can be specifically configured to generate the respiration waveform signal at least in part from the audio signal. The determiner 105 can be configured to process the audio signal to extract information indicative of the speaker's respiration captured by the audio signal. The determiner 105 can be configured to receive samples of the audio input signal and generate time series or parameters representing a time series of measurements / estimations of lung air volume.
[0059] The following discussion focuses on approaches that allow a determiner to determine a respiration waveform signal even when specific sensor data is unavailable by deriving the respiration waveform signal from a captured audio signal, particularly from speech data / its components. As a low-complexity example, the determiner can use a voice activity detection (SAD) algorithm to segment the speech data into speech and non-speech segments. Based on the knowledge that inspiration occurs during pauses in speech, the determiner 105 can use the duration and frequency of the pauses as a proxy for breathing activity and determine the respiration waveform signal accordingly. In many embodiments, more accurate determinations can be used, including the use of artificial neural networks. This can, for example, correct for or account for more pauses in speech than true inspiration events, which can introduce errors into respiration rate estimates.
[0060] In many embodiments, the determiner 105 can include a trained artificial network configured to determine samples of the respiration waveform signal based on samples of the audio signal. The artificial neural network can include input nodes that receive the samples of the audio signal and output nodes that generate the samples of the respiration waveform signal. The artificial neural network can be trained using, for example, speech data and sensor data representing measurements of lung air volume during speech.
[0061] The artificial neural network used in the described functions can be a network of nodes organized as layers, with each node holding a node value. Figure 2 shows an example of a section of an artificial neural network.
[0062] The node value of a given node can be calculated to include contributions from some, or often all, nodes in previous layers of the artificial neural network. Specifically, the node value of a node can be calculated as a weighted sum of the node values of all node outputs in the previous layer. Typically, a bias is added, and the result is applied to an activation function. The activation function typically provides nonlinearity to account for the key role of each neuron. Such nonlinearity and activation function have a significant impact on the learning and adaptation process of the neural network. Thus, the node value is generated as a function of the node values in the previous layer.
[0063] The artificial neural network may specifically include an input layer 201 that includes a plurality of nodes that receive input data values to the artificial neural network. Thus, the node values of the nodes in the input layer may typically be direct input data values to the artificial neural network and therefore may not be calculated from other node values.
[0064] An artificial neural network may further include zero, one, or more hidden layers 203 or processing layers. For each such layer, node values are typically generated as a function of the node values of the nodes in the previous layer; in particular, an activation function (such as a sigmoid, ReLU, or Tanh function) may be applied, followed by a weighted combination and an added bias.
[0065] Specifically, as shown in Figure 3, each node, sometimes called a neuron, can receive input values (from nodes in the previous layer) and then calculate the node value as a function of these values. Often, this involves first generating a value as a linear combination of the input values, each weighted by a weight:
number
[0066] Here, w refers to the weight, x refers to the node in the previous layer, and n is an index that refers to the respective node in the previous layer.
[0067] An activation function can then be applied to the resulting combination. For example, the node value l can be l=f(k) can be determined as:
[0068] This function can be, for example, a modified linear identity function, as described in Xavier Glorot, Antoine Bordes, Yoshua Bengio Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, PMLR 15:315-323, 2011. f(k) = ReLU(k) = max(0, k)
[0069] Other commonly used functions include the sigmoid function or the tanh function. In many embodiments, a node output or value can be calculated using multiple functions. For example, both the ReLU and the Sigmoid functions can be used in f(k)=ReLU(k)+σ(k) They can be combined using an activation function like this:
[0070] Such operations can be performed by each node of the artificial neural network (typically except for the input node).
[0071] The artificial neural network further comprises an output layer 205 that provides output from the artificial neural network, i.e., the output data of the artificial neural network are the node values of the output layer. As for the hidden / processing layers, the output node values are generated by functions of the node values of the previous layer. However, in contrast to the hidden / processing layers, where the node values are typically not accessible or further used, the node values of the output layer are accessible and provide the results of the operation of the artificial neural network.
[0072] Several different network structures and toolboxes for artificial neural networks have been developed, and in many embodiments, artificial neural networks can be based on the adaptation and customization of such networks. An example of a network architecture suitable for the above applications is the Long Short-Term Memory (LSTM) by Sepp Hochsreiter, described in Hochreiter, Sepp, and Jurgen Schmidhuber. "Long Short-Term Memory." Neural Computation 9.8 (1997): 1735-1780.
[0073] LSTM is an architecture used for classification and regression of time-domain signals using iterative causal or bidirectional evaluation, and has been successfully applied to audio signals. For example,
number
[0074] where * is matrix multiplication, o is Hadamard product, x is input vector, and h t-1 is the output vector of the previous time step, W, V, U are the network weights, and b is the bias vector.
[0075] In theory, classical (or "vanilla") artificial neural networks can track arbitrary long-term dependencies in input sequences. The problem with vanilla artificial neural networks is computational (or practical) in nature. When training vanilla artificial neural networks using backpropagation, the long-term gradients that are backpropagated can "vanish" (i.e., tend to zero) or "explode" (i.e., tend to infinity) due to the calculations involved in the process, which use finite-precision numbers. Artificial neural networks that use LSTM units partially solve the vanishing gradient problem because the LSTM units allow the gradients to flow unchanged. However, LSTM networks can still suffer from the exploding gradient problem.
[0076] In some cases, an artificial neural network can be further configured to include additional contributions that allow the artificial neural network to be dynamically adapted or customized to specific desired characteristics or features of the generated output. For example, a set of values can be provided to adapt the artificial neural network. These values can be included by providing contributions to several nodes of the artificial neural network. These nodes can specifically be input nodes, but typically can be nodes in hidden or processing layers. Such adaptation values can be weighted and summed, for example, as contributions to a weighted sum / correlation value for a given node.
[0077] The above description relates to a neural network approach that may be suitable for many embodiments and implementations. However, it will be understood that many other types and structures of neural networks can be used. Indeed, many different approaches for generating neural networks have been developed, including neural networks that use complex structures and processes different from those described above. This approach is not limited to any particular neural network approach, and any suitable approach can be used without detracting from the invention.
[0078] Artificial neural networks are adapted for specific purposes through a training process used to adapt / tune / modify the weights and other parameters (e.g., biases) of the artificial neural network. It will be appreciated that many different training process algorithms for training artificial neural networks are known. Typically, training is based on a large training set, in which a large number of examples of input data are provided to the network. Furthermore, the output of the artificial neural network is typically compared (directly or indirectly) to expected or ideal results. A cost function can be generated to reflect the desired outcome of the training process. In a typical scenario known as supervised learning, the cost function often represents the distance between the prediction for particular input data and the ground truth. Based on the cost function, the weights can be modified, and by repeating the process with the modified weights, the artificial neural network can be adapted to a state where the cost function is minimized.
[0079] More specifically, during the training step, a neural network may have two distinct information flows: from input to output (forward pass) and from output to input (backward pass). In the forward pass, data is processed by the neural network as described above, and in the backward pass, weights are updated to minimize the cost function. Typically, such backward propagation follows the gradient direction of the cost function landscape. In other words, by comparing the predicted output for a batch of data input with the ground truth, the direction in which the cost function is minimized can be estimated, and backward propagation can be performed by updating the weights accordingly. Other known approaches for training artificial neural networks include, for example, the Levenberg-Marquardt algorithm, the conjugate gradient method, and the Newton method.
[0080] In this case, training can include, inter alia, a training set including a potentially large number of pairs of captured audio signals and measured respiratory waveform signals. The measured respiratory waveform signals can be generated from sensor signals of a sensor configured to measure lung air volume-dependent characteristics. Thus, training of the artificial neural network can be performed using a training set comprising linked audio / speech data / signals and sensor data / signals representing lung air volume measurements during speech.
[0081] In some embodiments, the training data is an audio signal in a time segment corresponding to the processing interval of the artificial neural network being trained; for example, the number of samples in the training audio signal may correspond to the number of samples corresponding to the input nodes of the artificial neural network being trained. Thus, each training example may correspond to one operation of the artificial neural network being trained. However, typically, to speed up the training process, a batch of training samples is considered for each step. Furthermore, many upgrades to gradient descent are possible to speed up convergence or avoid local minima in the cost function landscape.
[0082] FIG. 4 shows an example of how training data is generated by dedicated testing. A speech audio signal 401 can be captured by a microphone 403 during a time interval in which a person is speaking. The resulting captured audio signal is provided as input data to an artificial neural network 405 (appropriate processing such as amplification, digitization, filtering, etc. can be performed before generating the test audio signal provided to the artificial neural network). Additionally, a respiratory waveform signal 407 is determined as a sensor signal, for example, from a suitable lung volume sensor 409 placed on the user. For example, the sensor 409 can be a respiratory inductive plethysmography (RIP) sensor or a nasal cannula flow sensor.
[0083] A large number of such measurements can be performed to generate a large number of pairs of training audio signals and respiration waveform signals. These signals are then applied to train an artificial neural network, where a cost function is determined as the difference between the respiration waveform signal generated by the artificial neural network for the training audio signals and the measured training respiration waveform signal. The artificial neural network is then adapted based on the cost function, as is known in the art. For example, a cost value can be determined for each combination set of training audio signals and / or training downmix audio signals (e.g., an average cost value for the training set is determined). Typically, the cost function includes at least one component that reflects how close the generated signal is to a reference signal, i.e., the so-called reconstruction error. In some embodiments, the cost function includes at least one component that reflects how close the generated signal is to the reference signal from a perceptual point of view. In the example of FIG. 4, the reference signal can specifically be the measured respiration waveform signal.
[0084] The segmenter can determine speech segments of the audio signal based on the respiration waveform signal, an example of which is shown in FIG.
[0085] As illustrated in FIG. 5, a respiration waveform signal typically has periodic components that reflect chest expansion and contraction caused by the speaker's breathing. The repetition rate, duration between peaks and valleys, etc., vary, indicating that breathing is not perfectly periodic. The inventors have recognized that a speaker's breathing affects their speaking style and that it is not only possible to determine the time-varying characteristics caused by breathing from captured audio, but that the time-varying characteristics can be used to adapt speech segmentation for speech processing of the same audio signal to improve performance, particularly to provide performance that can better adapt to different scenarios and user behaviors.
[0086] The particular segmentation approach used to generate the audio segments may vary in different embodiments and applications.
[0087] In some embodiments, the segmenter 107 can be configured to perform signal analysis of the respiratory waveform signal to extract characteristics of the periodic component of the signal. Different algorithms are known for extracting periodic components from a signal, and it will be understood that any suitable algorithm can be used. Based on this analysis, periodic characteristics such as frequency / repetition rate and timing of the periodic component can be determined. In some embodiments, these characteristics can be used to generate an audio segment. For example, an audio segment can be generated as a time interval that follows the maximum value of the periodic component of the respiratory waveform signal and is a given percentage (e.g., 90%) of the total period. Typically, the maximum value of the respiratory waveform signal corresponds to maximum chest expansion and air volume, and therefore corresponds to the point of transition from the inhalation interval to the exhalation interval.
[0088] Thus, a voice segment can be determined to correspond to a period of time that is most likely to correspond to speaking time. This time interval can be set to a large percentage of the period to reflect that speech / exhalation intervals are typically significantly longer than non-speech / inhalation intervals. This approach can provide a more reliable estimate for situations where the speaker's breathing is very rhythmic and constant. This allows voice segments of the same duration and containing the same number of samples to be generated. This can facilitate computation and implementation in many embodiments, for example, allowing for the implementation of constant segment / block processing. For example, the same number can be processed for each successive processing.
[0089] In many embodiments, the determiner is configured to detect local maxima in the respiratory waveform signal and determine the audio segments according to the timing of the extrema, and often the determiner 105 is configured to determine the start and / or end points of the audio segments according to the timing of one or more of the extrema. The extrema can often be local minima, local maxima, or both local minima and maxima.
[0090] In many embodiments, determiner 105 can be configured to determine the timing of extrema. For example, determiner 105 can be configured to determine the time points of peaks in the respiratory waveform signal, such as the peaks in the respiratory waveform signal of FIG. 5. In some embodiments, the determined timing can be fitted to the periodic component. For example, the timing can be determined so that the least squares error between the periodic timing time points and the time points at which the minima / peaks are measured is minimized. The periodic time points can then be used to generate audio segments that result in periodic time intervals / audio segments.
[0091] In many embodiments, segmenter 107 can be configured to generate segments of different durations / lengths / sizes. Specifically, segmenter 107 can be configured to determine audio segments to reflect dynamic changes in breathing rhythm and changes in the duration of inhalation and exhalation intervals. In particular, in some embodiments, each audio segment can be determined based on local characteristics of the respiratory waveform signal rather than average characteristics over a larger number of repetitions of the periodic component of the respiratory waveform signal.
[0092] In many embodiments, the timing of an audio segment is based on the timing of a local maximum, a local minimum, or a pair of a local maximum and an adjacent local minimum.
[0093] For example, in some embodiments, determiner 105 can be configured to detect a local maximum. Once a local maximum is detected, an audio segment can be determined as the audio signal for the time interval following the time of the local maximum. In many embodiments, determiner 105 can be configured to determine the start of the audio segment as a function of the time of the local maximum in the respiratory waveform signal, often as the time of the local maximum in the respiratory waveform signal.
[0094] The expiration time interval will typically follow the maximum lung volume corresponding to the peak of the respiration waveform signal. Therefore, determining the time interval of the speech segment as following the maximum value of the respiration waveform signal will tend to generate the speech segment as the expiration time interval. This approach can provide an appropriate determination of the time interval of the speech segment during which the speaker is active.
[0095] In some embodiments, determiner 105 can additionally or alternatively be configured to detect local minima. Once a local minima is detected, an audio segment can be determined as the time interval of the audio signal preceding the time point of the local minima. In many embodiments, determiner 105 can be configured to determine the end of the audio segment as a function of, and often as, the time point of, the local minima in the respiratory waveform signal.
[0096] The expiration time interval will typically precede the minimum lung volume corresponding to the minimum value of the respiration waveform signal. Therefore, determining the time interval of the speech segment as preceding the minimum value of the respiration waveform signal tends to generate the speech segment as the expiration time interval. This approach can provide an appropriate determination of the time interval of the speech segment during which the speaker is active.
[0097] The duration of the time interval may be predetermined in some cases or may be dynamically determined, for example, based on different parameters. For example, in some cases, the duration of the time interval may depend on available computational resources. For example, if there are no computational resource limitations (e.g., if there are not many other processes currently active), the time interval may be set long, but if computational resources are currently limited (e.g., due to other complex processes currently active), the length of the time interval may be shortened. Thus, the time interval and length of the speech segment may vary in some embodiments depending on different parameters. In certain instances, this may affect the proportion of the speech signal actually processed by the speech processing. Of course, this may not be suitable for some processes, but may be acceptable for others. For example, in the case of speech recognition, acceptable performance may be possible based on recognizing only the beginning of a sentence, which usually coincides with the start of the exhalation interval, often a maximum value. Therefore, for some applications, speech recognition may be acceptable even if it is performed only for a short time at the start of the exhalation time interval and improves over longer time intervals, and the audio device may be configured to adapt the considered duration, and therefore the quality of speech recognition, based on the available computational resources.
[0098] In some embodiments, the duration of the time intervals can depend on the respiratory waveform signal. For example, the average period of the periodic component of the respiratory waveform signal can be determined, and the duration of the time intervals for the audio segments can be set as a given percentage of the periodic duration. In other embodiments, the duration of the time intervals can vary dynamically and can be different for successive audio segments.
[0099] In many embodiments, an audio segment can be determined as a segment of the audio signal between a time point of a maximum in the respiratory waveform signal and a time point of a minimum in the respiratory waveform signal. The start of the time segment can be determined as the time point of the maximum, and the end of the time segment can be determined as the time point of the subsequent minimum. Thus, the time interval can be the interval from the maximum to the minimum.
[0100] The maximum value of the respiratory waveform signal typically corresponds to the current maximum lung volume and the start of the expiratory interval, and the minimum value of the respiratory waveform signal typically corresponds to the current minimum lung volume, the start of the inspiration time interval, and the end of the expiratory time interval.
[0101] This approach can therefore provide an efficient approach for determining exhalation / expiration and inhalation time intervals from a respiratory waveform signal, whereby speech segments can be determined as segments of the audio signal during the exhalation / expiration time intervals, and non-speech segments can be determined as segments of the audio signal during the inhalation time intervals.
[0102] Breathing involves a series of alternating inhalations (drawing air into the lungs) and exhalations (exhaling air from the lungs), and the speech processor 103 can be configured to evaluate the breathing waveform signal to determine timing characteristics of the speaker's inhalation / exhalation intervals. Specifically, the speech processor 103 can be configured to evaluate the timing of the inhalation / exhalation intervals, particularly when these intervals begin and / or end.
[0103] Such parameters can be determined, for example, by a peak detection algorithm that identifies local minima and maxima in the respiratory waveform signal, which correspond to the onset of the inspiration / inhalation and expiration / exhalation phases.
[0104] In some approaches, the audio signal can be divided into speech and non-speech segments, particularly speech segments corresponding to inhalation intervals and non-speech segments corresponding to inhalation intervals. Speech processing can then be configured to process each speech segment differently. For example, in some embodiments, attenuation can be increased for non-speech segments compared to speech segments, or speech recognition can be performed only during speech segments.
[0105] In this approach, the respiration waveform signal can represent the speaker's (or multiple speakers') breathing pattern over time, which can be used to segment the speech according to respiratory events such as inspiration or speech and silent breaths. Speech processing is performed based on these segments reflecting respiratory events.
[0106] This approach can exploit the insight that speech is typically produced subconsciously, with breathing synchronizing with the grammatical structure of the utterance, for example, at the beginning of a new sentence or clause. Therefore, if speech audio segmentation is performed according to breathing events, the segmented speech will likely represent the intended grammatical structure underlying the speech content. Therefore, adapting speech processing to breathing patterns has the potential to improve speech processing in many scenarios.
[0107] The audio device is configured to adapt its speech processing in response to the respiration waveform signal by adapting the segmentation of the audio signal into speech segments based on the respiration waveform signal. Thus, rather than simply performing a predetermined speech segmentation based on nominal or expected characteristics, the speech processing device of FIG. 1 is configured to adapt this speech segmentation to reflect the speaker's current respiration.
[0108] This approach reflects the inventors' recognition that a speaker's breathing affects speech and that it is not only possible to determine the time-varying characteristics due to breathing, but that these time-varying characteristics can further be used to adapt speech segmentation for speech processing of the same audio signal, providing improved performance, particularly performance that is better adapted to different scenarios and user behavior.
[0109] The particular speech segmentation, processing, and adaptation performed can depend on the particular implementation. However, in many embodiments, this approach can be used to adapt operations to reflect different breathing patterns a person may exhibit, for example, due to performing different activities. For example, when exercising, breathing may become labored and speech may become distorted or slow. The speech processing device of FIG. 1 can automatically adapt to optimize for slow speech with longer silence intervals.
[0110] This approach can provide significantly improved speech processing in many scenarios, and can provide benefits, for example, particularly for speakers engaged in strenuous activities such as exercise, or speakers who have breathing or speech difficulties.
[0111] The audio processing circuit 103 is configured to perform segment-based audio processing of the audio signal.
[0112] In some embodiments, the audio processor 103 can be configured to specifically modify audio processing during inspiration time intervals. In particular, the audio processor 103 can be configured in many embodiments to suppress or stop audio processing during inspiration time intervals. For example, audio processing in the form of voice recognition can simply ignore all inspiration time intervals and be applied only to expiration time intervals.
[0113] The speech processor 103 can be configured to determine inspiration time segments of the audio signal. These time segments can therefore reflect time intervals of the audio signal during which the speaker is presumed to be inhaling and can therefore be considered non-speech segments. In some embodiments, the speech processor 103 can be configured to attenuate the audio signal during inspiration time segments. For example, a speech enhancement algorithm configured to reduce noise by enhancing the speech component of the audio signal can attenuate the audio signal during inspiration time intervals, and in many embodiments, can attenuate the signal completely.
[0114] Such an approach tends to provide significantly improved performance, allowing speech processing to be adapted and applied specifically to portions of the audio signal where speech is present, while enabling a different approach, particularly attenuation, to portions of the audio signal where speech is less likely to be present. This approach can, for example, allow for significantly reducing breathing sounds by removing loud breathing sounds during the inspiration phase of the respiration waveform signal. This can significantly improve the perceived clarity of speech and provide a correspondingly better perceived speech. In some embodiments, portions of speech data occurring during inspiration (as estimated from the respiration waveform) can be removed prior to automatic speech recognition, thereby facilitating automatic speech recognition and reducing the risk of false word detection. In speech coding embodiments, for example, coding efficiency can be increased because speech coding is not performed during inspiration intervals. For example, an indicator of a silent or non-speech interval can simply be inserted into the coded data stream to indicate that no speech data is provided for this interval.
[0115] In many embodiments, the speech processing circuitry 103 is specifically configured to perform segment-based speech enhancement processing.
[0116] In some embodiments, the audio processing can be audio enhancement. Such audio enhancement can be used, for example, to provide a clearer audio signal with reduced noise from other audio sources and in the audio signal. For example, as described above, the audio signal can be attenuated during inspiration time segments. As another example, a respiratory waveform signal can indicate transitions or phases in the respiratory cycle that tend to be associated with particular sounds, and audio processing can be configured to generate corresponding signals and apply them in antiphase to provide an audio canceling effect for such sounds. Segmentation can then be used to focus particular sounds into specific audio segments.
[0117] As another example, speech processing may include estimating acoustic models of speech and background (out-of-segment data) for, e.g., noise reduction, dereverberation, speech enhancement, echo cancellation, multi-microphone beamforming, or other processing tasks aimed at improving speech quality or separating multiple speakers from one of multiple speech signals.
[0118] More specifically, in some embodiments, the segmented speech enhancement process can be used to separate and enhance the speech of individual speakers, for example, in teleconferencing or diarization applications. Given a continuous speech recording, a determiner can estimate a respiratory waveform signal and determine its corresponding minima and maxima. The continuous audio signal is then divided into speech segments s corresponding to the time range between the i-th maximum and the next minimum of the respiratory waveform. i and the background / inhalation segment b corresponding to the time range between the ith minimum and the next maximum. i In some embodiments, each audio segment s i It can be assumed that the segment contains the speech of one speaker and the background segment represents a static noise background. Based on this formulation, several example speech enhancement methods can be implemented.
[0119] In noise reduction applications, b i The sequence of segments is used to generate the background noise spectrum model W n (ω), which allows us to estimate s i It can be adaptively used to suppress background noise in the input audio signal within a segment, thus reducing the effect of background noise on further processing and transmission of the audio content.
[0120] In the case of a close-talking microphone, breathing sounds may become disturbingly loud. i The segments are amplitude suppressed to reduce the influence of breath sounds. In one embodiment, this type of enhancement is i It can be turned on based on the detection of high levels of breath sounds in the segment.
[0121] Adaptive beamforming for multi-microphone voice capture is based on multi-channel signal processing model A. k The goal is to estimate the parameters of (ω) to capture clean speech for a target (or angular direction) k around a multi-microphone device. A speaker identification algorithm is used to identify two consecutive speech segments s i and s i+1 It is possible to detect whether two segments contain audio from the same speaker or from different speakers. If two segments are determined to be from the same speaker k, then both segments can be used to generate A k The parameters of (ω) are updated. This process can be performed in parallel for multiple targets (or angular directions) k to construct a multi-beam speech capture algorithm. This method effectively exploits the insight that turn-taking in conversation often occurs during inspiration, i.e., waiting for the end of a sentence and starting to speak when the other person is inhaling.
[0122] The same principle of segmentation based on extrema of the respiratory waveform signal, combined with speaker identification in successive speech segments, can also be used to control the adaptation of the acoustic echo cancellation (AEC) filter coefficients. In one embodiment, the algorithm parameters are i and s i+1 The filter can be kept static if the identity of the speaker detected in remains the same, and can be adapted if the speaker changes.
[0123] In many embodiments, the speech processing circuitry 103 is specifically configured to perform segment-based speech recognition processing.
[0124] In automatic speech recognition, speech is divided into processing frames, and individual words are recognized in the context of adjacent utterances. While the underlying segmentation problem does not have a unique solution, the recognized utterance with the highest likelihood score by the model is usually selected. However, it is well known that segmenting speech in a manner that differs from the intended sentence uttered by the speaker often results in significant errors in the recognized speech. Speech segmentation based on consideration of respiratory waveform signals can provide more accurate and appropriate segmentation in many scenarios and for many signals. In particular, it can often enable the segmentation of speech at the input to automatic speech recognition to follow the actual sentence and clause structure of the speech. This can significantly improve the performance of automatic speech recognition of continuous speech.
[0125] As mentioned above, the impact of different sentence structures on speech recognition processing can be significant. For example, sentence structure can result in very different meanings, as shown for example in the following text differences: “Dear John: I want a man who knows what love is all about. You are generous, kind, thoughtful. People who are not like you admit to being useless and inferior. You have ruined me for other men. I yearn for you. I have no feelings whatsoever when we're apart. I can be forever happy-will you let me be yours? Jane” And the following text: “Dear John, I want a man who knows what love is. All about you are generous, kind, thoughtful people, who are not like you. Admit to being useless and inferior. You have ruined me. For other men, I yearn. For you, I have no feelings whatsoever. When we're apart, I can be forever happy. Will you let me be? Yours, Jane”
[0126] By applying segmentation based on the respiratory waveform signal, the segmentation can reflect the structure of the sentence, thereby enabling a speech recognition algorithm to distinguish, for example, the meaning of the above example.
[0127] In particular, for automatic speech recognition, this approach can provide significantly improved performance by adapting the segmentation and process to the speaker's current breathing pattern.
[0128] For example, as described above, the speech processor 103 can be configured to determine inspiration time segments of the audio signal corresponding to times when the speaker is inhaling and therefore not speaking. The speech processor 103 can then be configured to exclude these inspiration time segments from the speech recognition processing applied to the speech signal. Thus, times when the respiration waveform signal may indicate very little likelihood of speech can be excluded from speech recognition, thereby reducing the risk of falsely detecting words when no speech is present. This can also improve the accuracy of detection at other times, since it allows for estimation of when words are likely to be spoken. In particular, this allows segmentation to follow sentence structure, resulting in significantly improved speech recognition accuracy.
[0129] This approach is particularly advantageous for automatic speech recognition, e.g., for providing voice interfaces. For example, a personal assistant's voice interface allows for control of media, navigation functions, sports tracking, and various other information services while exercising. However, voice recognition is typically optimized for stationary, normal speech, and sometimes personalized. During and immediately after intense exercise, oxygen needs increase significantly, resulting in changes in breathing and affecting speech. Specifically, breathing rate increases, which can result in shorter speech intervals and sentences, and the speech recognition performed by the audio device can reflect this.
[0130] More specifically, in some embodiments, a segmented speech recognition process can be applied to select the utterance that best matches the determined respiratory waveform signal. Automatic speech recognition (ASR) algorithms typically generate multiple candidate word sequences that correspond to the best matches with the input speech sequence. In traditional ASR systems, the candidate sequences are evaluated based on the well-known Viterbi algorithm. In the above example, the ASR algorithm may generate several equally likely candidates, such as: “I yearn. For you, I have no feelings whatsoever! When we're apart?” “I yearn for you! I have no feelings whatsoever when we're apart..” “None for you I have, no feelings whatsoever. When we're apart.”
[0131] The timing of respiratory events can be combined with the sequence of tokens from the ASR to calculate a score for how well each sequence matches a breathing pattern. In one embodiment, a score can be calculated for how often text symbols, such as full stops, commas, exclamation points, or question marks, match the minimum or maximum values of the respiratory waveform signal. The sequence with the highest score is selected.
[0132] In many embodiments, the speech processing can be, for example, speech encoding, where captured speech is encoded for efficient transmission or distribution over an appropriate communication channel. More efficient encoding can typically be achieved by not encoding the signal in non-speech segments and only encoding the audio signal in speech segments. In some embodiments, the encoding in speech segments can be further based on assumptions, such as, for example, that speech segments reflect complete sentences. For example, a speech processor can process each speech segment s iThe speech signal may be coded in speech segments by treating each speech segment s as a separate coding entity, so that a specific bit budget is allocated to each speech segment s. i and a smaller bit budget is allocated to the background b i assigned to a segment.
[0133] In some embodiments, the speech processing circuitry is configured to select speech segments as including speech from a particular speaker (the speaker for whom the breath waveform signal is provided) rather than from other speakers. For example, the audio signal may include audio from multiple speakers, and the segmenter may be configured to generate speech segments that may include speech from different speakers. For example, the segmenter may generate speech segments based solely on the audio signal without considering the breath waveform signal. As an example, speech segments may be generated when the instantaneous level of the audio signal exceeds a threshold, and non-speech / silence segments may be generated when the instantaneous level is below the threshold. Thus, the segmenter may generate speech segments that reflect speech with a high probability but without considering the current speaker, i.e., without considering the identity of the active speaker.
[0134] The segmenter 107 can proceed to select audio segments for a given speaker based on the given speaker's respiration waveform signal. As a low-complexity example, the segmenter 107 may simply select audio segments during exhalation periods determined from the respiration waveform signal. In some embodiments, the respiration waveform signal can reflect different breathing patterns when a speaker is speaking rather than silent, and the segmenter 107 can be configured to identify such periods and select the audio segments to be for the given speaker.
[0135] The audio device can then use certain (subset) of the speech segments for further speaker-specific processing, for example, adapting beamforming based only on such segments.
[0136] In multi-speaker applications, such as multi-beamforming in voice capture and speech diarization between multiple conversation partners, the problem of speech segmentation and active speaker identification is a challenging problem. This is a very difficult task using traditional acoustic signal processing techniques. Speakers often interrupt, overlap, and even sound similar. However, it has been observed that in multi-speaker scenarios, humans are good at finding the right moment to take a turn. Often, this occurs at (or just before) the end of the previous speaker's sentence. Therefore, considering the breathing waveform signals of one or more speakers enables the detection of different speakers taking turns.
[0137] As a specific example, audio equipment can be i Based on the segments, a machine learning model for speaker identification can be configured to determine turn-taking changes by comparing successive audio segments.
[0138] The audio device may in particular be implemented as one or more suitably programmed processors. For example, the artificial neural network of the determiner may be implemented as one or more such suitably programmed processors. Different functional blocks, in particular the artificial neural network, may be implemented in separate processors and / or may, for example, be implemented in the same processor. Examples of suitable processors are given below:
[0139] 6 is a block diagram illustrating an exemplary processor 600 according to an embodiment of the disclosure. The processor 600 can be used to realize one or more processors that implement the apparatus or elements thereof as described above (including, in particular, one or more artificial neural networks). The processor 600 can be any suitable processor type, including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA) (wherein the FPGA is programmed to form a processor), a graphics processing unit (GPU), an application specific integrated circuit (ASIC) (wherein the ASIC is designed to form a processor), or a combination thereof.
[0140] The processor 600 may include one or more cores 602. The cores 602 may include one or more arithmetic logic units (ALUs) 604. In some embodiments, the cores 602 may include a floating point logic unit (FPLU) 606 and / or a digital signal processing unit (DSPU) 608 in addition to or instead of the ALUs 604.
[0141] The processor 600 may include one or more registers 612 communicatively coupled to the cores 602. The registers 612 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 612 may be realized using static memory. The registers may provide data, instructions, and addresses to the cores 602.
[0142] In some embodiments, processor 600 may include one or more levels of cache memory 610 communicatively coupled to cores 602. Cache memory 610 may provide computer-readable instructions to cores 602 for execution. Cache memory 610 may provide data for processing by cores 602. In some embodiments, computer-readable instructions may be provided to cache memory 610 by local memory, for example, local memory attached to external bus 616. Cache memory 610 may be implemented using any suitable cache memory type, such as, for example, static random access memory, dynamic random access memory, and / or any other suitable memory technology.
[0143] Processor 600 may include a controller 614, which may control input to processor 600 from other processors and / or components in the system and / or control output from processor 600 to other processors and / or components in the system. Controller 614 may control data paths within ALU 604, FPLU 606, and / or DSPU 608. Controller 614 may be implemented as one or more state machines, data paths, and / or dedicated control logic. Gates in controller 614 may be implemented as standalone gates, FPGAs, ASICs, or any other suitable technology.
[0144] Registers 612 and cache 610 may communicate with controller 614 and core 602 via internal connections 620A, 620B, 620C, and 620D. The internal connections may be implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection technology.
[0145] Inputs and outputs of processor 600 are provided via bus 616, which may include one or more conductive lines. Bus 616 may be communicatively coupled to one or more components of processor 600, such as controller 614, cache memory 610, and / or registers 612. Bus 616 may be coupled to one or more components of the system.
[0146] The bus 616 may be coupled to one or more external memories. The external memory may include read-only memory 632. The ROM 632 may be masked ROM, electrically programmable read-only memory (EPROM), or any other suitable technology. The external memory may include random access memory 633. The RAM 633 may be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory may include electrically erasable programmable read-only memory (EEPROM) 635. The external memory may include flash memory 634. The external memory may include a magnetic storage device such as a disk 636. In some embodiments, external memory may be included in the system.
[0147] The present invention may be implemented in any suitable form including hardware, software, firmware or any combination of these. The present invention may optionally be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the present invention may be physically, functionally, and logically implemented in any suitable way. Indeed, functionality may be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the present invention may be implemented in a single unit, or may be physically and functionally distributed between different units, circuits, and processors.
[0148] Although the present invention has been described in connection with several embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, while certain features may appear to be described in connection with particular embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0149] Furthermore, although individually listed, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Furthermore, although individual features may be included in different claims, these may be advantageously combined in some cases, and their inclusion in different claims does not imply that a combination of features is not feasible and / or advantageous. Furthermore, the inclusion of a feature in one category of claims does not imply limitation to this category, but rather indicates that the feature is equally applicable to other claim categories, where appropriate. Furthermore, the order of features in the claims does not imply a particular order in which the features must operate, and in particular the order of individual steps in method claims does not imply that the steps must be performed in this order. Rather, steps may be performed in any suitable order. Furthermore, a reference to the singular does not exclude a plurality; thus, references to "a," "an," "first," "second," etc. do not exclude a plurality. Reference signs in the claims are provided merely as a clarifying example and should not be construed as limiting the scope of the claims in any way.
Claims
1. 1. An apparatus for sound processing of an audio signal, comprising: an input configured to receive the audio signal comprising a speaker's voice audio component; a determiner configured to generate a respiratory waveform signal indicative of the speaker's lung air volume as a function of time; a segmenter configured to segment the audio signal in response to the respiration waveform signal to generate speech segments of the audio signal; an audio processing circuit configured to perform audio processing of the audio signal; wherein the audio processing is segment-based processing applied to the audio segments.
2. The apparatus of claim 1 , wherein the determiner is configured to determine the respiration waveform signal from the audio signal.
3. 3. The apparatus of claim 2, comprising a trained artificial neural network having the determiner, an input node for receiving samples of the audio signal, and an output node configured to provide samples of the respiration waveform signal.
4. 4. The apparatus of claim 1, wherein the determiner is configured to detect local maxima in the respiratory waveform signal and to determine the speech segments in response to timing of the maxima.
5. 5. The apparatus of claim 1, wherein the determiner is configured to detect minima in the respiratory waveform signal and to determine the speech segments in response to timing of the minima.
6. 6. The apparatus of claim 1, wherein the determiner is configured to determine at least one speech segment as a segment of the audio signal between a time point of a maximum value of the respiratory waveform signal and a time point of a minimum value of the respiratory waveform signal.
7. 7. The apparatus of claim 1, wherein the segmenter is configured to divide the audio signal into speech segments and non-speech segments.
8. 8. The apparatus of claim 1, wherein the segmenter is configured to determine an expiration time interval and an inspiration time interval from the respiratory waveform signal, determine the speech segment as a segment of the audio signal during the expiration time interval, and determine a non-speech segment as a segment of the audio signal during the inspiration time interval.
9. 9. The apparatus of claim 1, wherein the speech processing circuitry is configured to select speech segments containing speech from the speaker rather than speech from other speakers, the selection being dependent on the respiration waveform signal.
10. 10. The apparatus of claim 1, further comprising a sensor input configured to receive a chest sensor signal, and wherein the determiner is configured to determine the respiration waveform signal from the chest sensor signal.
11. 11. The apparatus of claim 1, further comprising a video input arranged to receive a video signal comprising a video image of the speaker, and wherein the determiner is configured to determine the respiration waveform signal from the video image.
12. The device of claim 1 , wherein the speech processing comprises speech enhancement processing.
13. The device of claim 1 , wherein the speech processing comprises speech recognition processing.
14. 1. A method for sound processing of an audio signal, comprising: receiving the audio signal containing the speaker's voice audio component; generating a respiratory waveform signal indicative of the speaker's lung air volume as a function of time; segmenting the audio signal in response to the respiration waveform signal to generate voice segments of the audio signal; performing sound processing of the audio signal; wherein the audio processing is segment-based processing applied to the audio segments.
15. A computer program product which, when executed by a computer, causes the computer to carry out the method of claim 14.