Speech processing of audio signals

By generating breathing waveform signal segmented audio signals and using artificial neural networks to process voice fragments, the problem of poor performance of speech processing algorithms in the prior art under actual conditions is solved, and more efficient and flexible speech processing is achieved.

CN120513481APending Publication Date: 2025-08-19KONINKLIJKE PHILIPS NV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202480007318.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-13
Filing Date
2024-01-02
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Existing voice processing algorithms cannot provide optimal performance when actual conditions do not match the expected conditions, and are often complex or have high resource requirements, lack flexibility and adaptability.

Method used

By generating a breath waveform signal to segment the audio signal, the artificial neural network adaptively process voice fragments, including speech enhancement and recognition, reducing dependence on other sensor data.

Benefits of technology

Improve the adaptability and flexibility of speech processing, reduce complexity and resource use, and improve the accuracy of user experience and speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120513481A_ABST
    Figure CN120513481A_ABST
Patent Text Reader

Abstract

An apparatus for speech processing comprises an input (101) arranged to receive an audio signal comprising a speech audio component. A determiner (105) generates a breathing waveform signal indicative of the amount of lung air of the speaker as a function of time. A sectionalizer (107) segments the audio signal in response to the respiratory waveform signal to generate speech segments of the audio signal. A speech processing circuit (103) performs speech processing on the audio signal by segment-based processing applied to the speech segment. In some cases, a breathing waveform signal may be generated from an audio signal, for example, using an artificial neural network. The method may provide improved speech processing that more closely reflects current speech attributes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to performing speech processing on audio signals, in particular but not limited to automatic speech recognition or speech enhancement processing on audio signals capturing a speaker. Background Art

[0002] Speech processing of audio signals is widely used in various practical applications and is becoming increasingly important and significant for many daily activities and devices.

[0003] For example, speech enhancement is widely used to improve the clarity and reproducibility of speech captured by an audio signal (e.g., sounds captured in real-life environments). Speech encoding is also often performed on the captured audio. For example, speech encoding of captured audio signals is an integral part of smartphones. Another speech processing application that has become increasingly frequent in recent years is the use of speech recognition, for example, to provide user interfaces for devices.

[0004] For example, a voice interface of a personal assistant or home speaker makes it possible to control media, navigation functions, track tracking, obtain various information services, etc., simply by using the voice interface and while performing other activities (e.g., during exercise).

[0005] Voice interfaces based on automatic speech recognition (ASR) are also becoming increasingly important in controlling various devices, such as wearables, proximity and IoT consumer, health, and fitness devices. Voice is the most natural interface for interacting with devices during activities that involve a lot of user movement and change. As an example, US20090098981A1 discloses a fitness device with a voice interface for users to communicate with a virtual fitness trainer.

[0006] However, although a great deal of effort has been put into developing and optimizing speech processing algorithms for various applications, and many very advantageous and efficient methods have been developed, they are often not optimal in all cases. For example, many speech processing operations are developed for specific nominal conditions (such as nominal speaker attributes, nominal acoustic environment, etc.). When actual conditions differ from the expected nominal conditions, many speech processing applications may not provide optimal performance. Many speech processing algorithms may also be more complex or more resource-intensive than expected. Adapting speech processing to attempt to compensate for different attributes often results in less flexible, more complex and / or more resource-intensive implementation methods that generally do not provide optimal performance.

[0007] Therefore, an improved method for speech processing would be advantageous. In particular, a method that allows for increased flexibility, improved adaptability, improved performance, improved quality of, for example, speech enhancement, encoding, and / or recognition, reduced complexity and / or resource usage, improved remote control of audio processing, improved adaptation to changes in speaker attributes and / or activity, reduced computational load, improved user experience, simplified implementation, and / or improved spatial audio experience would be advantageous. Summary of the Invention

[0008] Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination.

[0009] According to one aspect of the present invention, there is provided an apparatus for performing speech processing on an audio signal, the apparatus comprising: an input terminal arranged to receive the audio signal, the audio signal comprising a speech audio component for a speaker; a determiner arranged to generate a breathing waveform signal, the breathing waveform signal indicating the amount of lung air of the speaker as a function of time; a segmenter arranged to segment the audio signal to generate speech segments of the audio signal, the segmentation being responsive to the breathing waveform signal; and a speech processing circuit arranged to perform speech processing on the audio signal, the speech processing being segment-based processing applied to the speech segments.

[0010] This method can provide improved speech processing in many embodiments and scenarios. For many signals and scenarios, it can provide speech processing that more closely reflects variations in the speaker's speech properties and characteristics. For example, it can adapt to the speaker's level of fatigue or, for example, dyspnea. The method can provide convenient adaptation and, in particular, can provide improved adaptive speech processing. The method can provide efficient implementation and, in many embodiments, can allow for reduced complexity and / or resource usage.

[0011] Adaptation based on the segmentation of the respiratory waveform signal can specifically allow speech processing to be more closely correlated and adaptive to the actual current speech of the speaker. Adaptation based on the segmentation of the respiratory waveform signal can generally allow for closer correlation with sentence structure, words, etc. For example, in some cases, it can provide a higher probability that the segmentation is consistent with the cognitive content of the speech and the speaker's specific speaking patterns.

[0012] Segment-based processing can be sequential processing, wherein speech segments are processed sequentially. Segment-based processing can perform processing wherein output speech signal segments within a time interval of the audio signal are generated based on the speech segments of the audio signal within the time interval. In some embodiments, speech processing of a speech segment can be independent of portions of the audio signal other than the speech segment. In some embodiments, speech processing of a speech segment does not include samples in the audio signal other than those belonging to the speech segment.

[0013] The method can, for example, use an audio signal and / or a dedicated sensor input to sense the temporal breathing pattern of a speaker (or multiple speakers). The temporal breathing pattern can be represented by a respiratory waveform signal. Segmentation of the speech can be performed based on respiratory events (such as inspiration or speaking and silent breathing). The method can then process the speech in segments determined by the respiratory events. The respiratory waveform signal can represent / reflect / be a measure of the temporal breathing pattern of the speaker (particularly over the entire respiratory cycle). The respiratory waveform signal can represent / reflect / be a measure of the temporal breathing pattern of the speaker. The respiratory waveform signal can represent / reflect / be a measure of the change in lung air volume (as a function of time) during a respiratory cycle (typically over the entire respiratory cycle).

[0014] In many embodiments, the breathing waveform signal can be independent of the level of the audio signal. In many embodiments, the breathing waveform signal can be independent of the audio signal.

[0015] According to an optional feature of the invention, the determiner is arranged to determine the respiration waveform signal from the audio signal.

[0016] The method may provide convenient adaptation and may in particular provide improved adaptive speech processing without relying on or requiring further input from other devices measuring properties of the speaker. The method may provide efficient implementation and, in many embodiments, may allow for reduced complexity and / or resource usage.

[0017] According to an optional feature of the invention, the determiner comprises a trained artificial neural network having input nodes for receiving samples of the audio signal and output nodes arranged to provide samples of the respiration waveform signal.

[0018] This can provide particularly advantageous operation and performance in many scenarios and applications.

[0019] The approach may provide a particularly advantageous arrangement which, in many embodiments and scenarios, may allow for leveraging the facilities and / or improved possibilities of artificial neural networks when adapting and optimizing speech processing (typically including speech enhancement and / or recognition).

[0020] An ANN is a trained artificial neural network.

[0021] The artificial neural network may be a trained artificial neural network trained using training data, the training data including a training speech audio signal and a training breathing waveform signal generated from measurements of a breathing waveform; the training employing a cost function that compares the training breathing waveform signal with a breathing waveform signal generated by the artificial neural network for the training speech audio signal. The artificial neural network may be one or more trained artificial neural networks trained using training data including training speech audio signals representing a series of related speakers in a series of different states and performing different activities.

[0022] The artificial neural network may be a trained artificial neural network that is trained by having training data including training input data comprising a training speech audio signal and using a cost function that includes a contribution indicative of a difference between a measured training breathing waveform signal and a breathing waveform signal generated by the artificial neural network in response to the training speech audio signal.

[0023] According to an optional feature of the invention, the determiner is arranged to detect a local extrema of the respiratory waveform signal and to determine the speech segment in response to the local extrema.

[0024] According to an optional feature of the invention, the determiner is arranged to detect a local maximum of the respiratory waveform signal and to determine the speech segment in response to the local maximum.

[0025] This can provide particularly advantageous speech segmentation. In many embodiments, it can allow the improved speech segmentation to be closely adapted to the speaker's speech behavior and / or speech content as reflected in sentence structure, etc.

[0026] In many embodiments, the determiner can be arranged to determine the start time of the speech segment in response to the timing of the local maximum, and specifically determine the start time of the speech segment as a start time having a fixed time offset from the time of the local maximum (the offset can be zero).

[0027] According to an optional feature of the invention, the determiner is arranged to detect a local minimum of the respiratory waveform signal and to determine the speech segment in response to the local minimum.

[0028] This can provide particularly advantageous speech segmentation. In many embodiments, it can allow the improved speech segmentation to be closely adapted to the speaker's speech behavior and / or speech content as reflected in sentence structure, etc.

[0029] In many embodiments, the determiner can be arranged to determine the end time of the speech segment in response to the timing of the local minimum, and specifically determine the end time of the speech segment as an end time having a fixed time offset from the time of the local minimum (the offset can be zero or typically negative).

[0030] According to an optional feature of the invention, the determiner is arranged to determine at least one speech segment as a segment of the audio signal between a time of a local maximum of the respiratory waveform signal and a time of a (next) local minimum of the respiratory waveform signal.

[0031] This can provide particularly advantageous speech segmentation. In many embodiments, it can allow the improved speech segmentation to be closely adapted to the speaker's speech behavior and / or speech content as reflected in sentence structure, etc.

[0032] The determiner may be specifically arranged to determine the start time of the speech segment as the time of the local maximum and the end time of the speech segment as the time of the next local minimum.

[0033] According to an optional feature of the invention, the segmenter is arranged to divide the audio signal into speech segments and non-speech segments.

[0034] According to an optional feature of the invention, the segmenter is arranged to determine a respiratory time interval and an inspiratory time interval based on the respiratory waveform signal; and determine the speech segment as a segment of the audio signal during the respiratory time interval, and determine the non-speech segment as a segment of the audio signal during the inspiratory time interval.

[0035] According to an optional feature of the invention, the speech processing circuitry is arranged to select the speech segment as comprising speech from the speaker and not from another speaker, the selection being dependent on the respiration waveform signal.

[0036] The method may allow speech segmentation reflecting the current active speaker in a scene with multiple potential speakers.The method may allow improved speech segmentation and result in improved speech processing.

[0037] According to an optional feature of the invention, the apparatus further comprises a sensor input arranged to receive a chest sensor signal, and the determiner is arranged to determine the respiration waveform signal based on the chest sensor signal.

[0038] This may allow improved and / or facilitated operation in many embodiments and scenarios. It may generally allow improved speech segmentation and resulting speech processing.

[0039] According to an optional feature of the invention, the apparatus further comprises a video input arranged to receive a video signal comprising a video image of the speaker, and the determiner is arranged to determine the respiration waveform signal based on the video image.

[0040] This may allow improved and / or facilitated operation in many embodiments and scenarios. It may generally allow improved speech segmentation and resulting speech processing.

[0041] According to an optional feature of the invention, the speech processing comprises speech enhancement processing.

[0042] The approach may provide improved speech enhancement and, in many scenarios, may allow the generation of improved speech signals.

[0043] According to an optional feature of the invention, the speech processing comprises speech recognition processing.

[0044] The method can provide improved speech recognition. In many scenarios, it can allow more accurate detection of words, terms, and sentences, and in particular can allow improved speech recognition for speakers in a variety of situations, states, and activities.

[0045] According to one aspect of the present invention, a method is provided, comprising: receiving an audio signal including a speech audio component for a speaker; generating a respiration waveform signal, the respiration waveform signal indicating a lung air volume of the speaker as a function of time; segmenting the audio signal to generate speech segments of the audio signal, the segmentation being responsive to the respiration waveform signal; and performing speech processing on the audio signal, the speech processing being segment-based processing applied to the speech segments.

[0046] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Embodiments of the present invention will now be described, by way of example only, with reference to the accompanying drawings, in which

[0048] Figure 1 shows some elements of an example of an audio device according to some embodiments of the present invention;

[0049] Figure 2 An example of the structure of an artificial neural network is shown;

[0050] Figure 3 An example of a node of an artificial neural network is shown;

[0051] Figure 4 Some elements of an example of a training setup for an artificial neural network are shown;

[0052] Figure 5 shows an example of a respiratory waveform signal; and

[0053] Figure 6 Some elements of a possible arrangement of a processor for implementing elements of an apparatus according to some embodiments of the present invention are shown. DETAILED DESCRIPTION

[0054] Figure 1 Some elements of a speech processing apparatus according to some embodiments of the present invention are shown.The speech processing apparatus may be adapted to provide improved speech processing for a speaker in many different environments and scenarios.

[0055] The speech processing device comprises a receiver 101, which is arranged to receive an audio signal comprising a speech audio component. The audio signal may specifically be a microphone signal representing audio captured by a microphone. The audio signal may specifically be a speech signal, and in many embodiments may be a speech signal captured by a microphone arranged to capture the speech of a single person (such as a microphone of a headset, a smartphone, or a body-worn microphone). Therefore, the audio signal will include a speech audio component (hereinafter also referred to as a speech signal) and possibly other sounds, such as background sounds from the environment.

[0056] The speech processing device further comprises a speech processing circuit 103 arranged to apply speech processing to the received audio signal, and thus in particular to the speech audio component of the audio signal.The speech processing may in particular be a speech enhancement process, speech coding or a speech recognition process.

[0057] In many practical applications, speech processing is performed in segments, where the audio signal is divided into segments and each segment is processed by itself. In some applications, the audio / speech signal is fixedly divided into consecutive segments, which are then processed individually. For example, in many applications, speech processing algorithms can be based on dividing the input audio signal into periodic time intervals of fixed duration and processing these time intervals individually. However, in Figure 5 In the speech processing apparatus of the present invention, speech processing is not performed in predetermined and fixed segments / time intervals, but is instead based on segmentation of the audio signal. In particular, the speech processing apparatus comprises: a determiner 105 arranged to determine a respiration waveform signal indicating the volume of air in the speaker's lungs as a function of time; and a segmenter 107 arranged to segment the audio signal in response to the respiration waveform signal to generate speech segments of the audio signal. The speech processing circuit 103 is then arranged to process the audio signal based on the segment-based processing applied to the speech segments generated by the segmenter 107.

[0058] The respiratory waveform signal indicates / measures / reflects / represents the speaker's lung air volume as a function of time, in particular, changes in lung air volume during a respiratory cycle. Therefore, it reflects the speaker's breathing and the airflow in and out of the speaker's lungs. Therefore, the respiratory waveform signal can be a time-varying signal that reflects changes in the speaker's breathing, and in particular, can reflect changes in lung capacity caused by the speaker's breathing, in particular during and throughout a respiratory cycle. In some embodiments, the respiratory waveform signal can represent the current lung capacity. In other embodiments, the respiratory waveform signal can, for example, represent changes in the current lung capacity, such as when the respiratory waveform signal represents inhaled / exhaled airflow. In many embodiments, the speech processing device can include a sensor input terminal 109, which is arranged to receive a sensor signal and determine the respiratory waveform signal based on the sensor signal.

[0059] The sensor signal may specifically be a chest sensor signal, such as a signal from a chest sensor that is part of a chest strap worn around a person's chest / body region. Such a sensor may directly generate a sensor signal reflecting the periphery of the chest, or may directly reflect a more localized chest expansion and contraction. Such a sensor signal may directly reflect temporal changes in lung air volume and, in many embodiments, may be used directly as a respiratory waveform signal. In many embodiments, some low-pass filtering, noise suppression, and amplification may be applied to the input sensor signal to generate the respiratory waveform signal.

[0060] The sensor signal may specifically come from, for example, a breathing belt sensor, or from other sensors that sense chest or abdominal movements of the speaker (or each speaker). Breathing sensors may also be integrated into other devices, such as seat belts in cars.

[0061] In some embodiments, the sensor signal may not directly reflect chest contraction and expansion, but may, for example, come from an optical sensor such as a camera. The sensor signal may, for example, be a video signal from a camera capturing the speaker. The determiner 105 may be arranged to analyze the video image to determine the speaker's chest movement, and thereby determine a respiratory waveform signal reflecting the speaker's chest expansion / contraction and, therefore, changes in lung capacity. It will be appreciated that various methods for evaluating images to determine changes in lung capacity are known and, for the sake of brevity, will not be described further herein. For example, US9301710B2 discloses a method for estimating a subject's respiratory rate based on changes in an illumination pattern in a video signal showing a chest region of the subject.

[0062] As another example, the sensor signal may be from a thermal or flow based sensor that monitors airflow through the mouth and nose.

[0063] In many embodiments, the determiner / determination circuit 105 may be specifically arranged to generate a breathing waveform signal based at least in part on an audio signal. The determiner 105 may be arranged to process the audio signal to extract information indicative of the speaker's breathing captured by the audio signal. The determiner 105 may be arranged to receive samples of the audio input signal and generate a time series or parameter representing a time series of lung air volume measurements / estimates.

[0064] The following description will focus on a method in which a determiner can determine a respiratory waveform signal by deriving it from a captured audio signal and, in particular, from its speech data / component, even if no specific sensor data is available. As a low complexity example, the determiner can use a speech activity detection (SAD) algorithm to segment the speech data into speech and non-speech segments. Based on the knowledge that inhalation occurs during pauses in speech, the determiner 105 can use the duration and frequency of the pauses as a proxy for respiratory behavior, thereby determining the respiratory waveform signal accordingly. In many embodiments, a more accurate determination can be used, including the use of an artificial neural network. This can, for example, compensate for or take into account the presence of more pauses in speech than true inhalation events, which can lead to errors in the estimation of respiratory rate, for example.

[0065] In many embodiments, the determiner 105 may include a trained artificial neural network that is arranged to determine samples of the respiratory waveform signal based on samples of the audio signal. The artificial neural network may include an input node that receives samples of the audio signal and an output node that generates samples of the respiratory waveform signal. The artificial neural network may, for example, have been trained using speech data and sensor data representing lung air volume measurements during speech.

[0066] The artificial neural network used in the described functions may be a network of nodes arranged in layers, with each node holding a node value. Figure 2 An example of a portion of an artificial neural network is shown.

[0067] The node value of a given node can be calculated as including the contributions of some or usually all nodes from the previous layer of the artificial neural network. Specifically, the node value of a node can be calculated as the weighted sum of the node values output by all nodes in the previous layer. Typically, a bias can be added and the result can be processed by an activation function. The activation function provides the necessary part of each neuron by generally providing nonlinearity. This nonlinearity and activation function have a significant impact on the learning and self-adaptation process of the neural network. Therefore, the node value is generated based on the node values of the previous layer.

[0068] The artificial neural network may specifically include an input layer 201, which includes a plurality of nodes that receive input data values of the artificial neural network. Therefore, the node value of the input layer node can usually be directly the input data value of the artificial neural network, and therefore may not be calculated from other node values.

[0069] The artificial neural network may also include none, one or more hidden layers 203 or processing layers. For each such layer, the node value is generally generated as a function of the node value of the nodes in the previous layer, and in particular, a weighted combination and an added bias may be applied, followed by an activation function (such as sigmoid, ReLU or Tanh function).

[0070] Specifically, if Figure 3 As shown, each node (also called a neuron) can receive input values (from the nodes in the previous layer) and calculate the node value as a function of these values. Typically, this involves first generating a value that is a linear combination of the input values, where each of these input values is weighted by a weight:

[0071]

[0072] Where w refers to the weight, x refers to the nodes in the previous layer, and n refers to the index of different nodes in the previous layer.

[0073] Then, an activation function can be applied to the resulting combination. For example, the node value l can be determined as:

[0074] l=f(k)

[0075] The function may be, for example, a modified linear unit function, as described in Xavier Glorot, Antoine Bordes, Yoshua Bengio (Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, PMLR 15: 315-323, 2011):

[0076] f(k)=ReLU(k)=max(0,k)

[0077] Other commonly used functions include sigmoid function or tanh function. In many embodiments, multiple functions can be used to calculate node output or value. For example, ReLU and Sigmoid functions can be combined using activation functions such as the following:

[0078] f(k)=ReLU(k)+σ(k)

[0079] Such operations can be performed by every node of the artificial neural network (usually except for the input nodes).

[0080] The artificial neural network also includes an output layer 205, which provides the output from the artificial neural network. That is, the output data of the artificial neural network are the node values of the output layer. For hidden / processing layers, the output node values are generated by a function of the node values of the previous layer. However, in contrast to hidden / processing layers, where the node values are generally not accessible or further used, the node values of the output layer are accessible and provide the results of the operation of the artificial neural network.

[0081] Many different network structures and toolkits have been developed for artificial neural networks, and in many embodiments, artificial neural networks can be based on adapting and customizing such networks. An example of a network architecture that can be suitable for the above applications is the long short-term memory (LSTM) [Sep Hochreiter, Sepp, and Jürgen Schmidhuber, "Long Short-Term Memory," Neural Computation 9.8 (1997): 1735-1780].

[0082] LSTM is an architecture for classification and regression of time-domain signals using recursive causal or bidirectional evaluation and has been successfully applied to audio signals.

[0083]

[0084] Where * represents matrix multiplication, o represents Hadamard product, x is the input vector, h t-1 represents the output vector of the previous time step, W, V, U are the network weights, and b is the bias vector.

[0085] In theory, classical (or "vanilla") artificial neural networks can track arbitrary long-term dependencies in input sequences. The problem with vanilla artificial neural networks is computational (or practical) in nature: when backpropagation is used to train a vanilla artificial neural network, the backpropagated long-term gradients can "vanish" (i.e., they can tend to zero) or "explode" (i.e., they can tend to infinity) because the calculations involved in the process use finite precision numbers. Artificial neural networks using LSTM units partially solve the vanishing gradient problem because LSTM units allow the gradients to also remain constant. However, LSTM networks can still suffer from the exploding gradient problem.

[0086] In some cases, the artificial neural network can also be configured to include additional contributions that allow the artificial neural network to be dynamically adapted or customized for specific desired properties or characteristics of the generated output. For example, a set of values can be provided to adapt the artificial neural network. These values can be included by providing contributions to certain nodes of the artificial neural network. These nodes can be specific input nodes, but can generally be nodes of hidden or processing layers. For example, such adaptive values can be weighted and added as contributions to a weighted sum / correlation value for a given node.

[0087] The above description relates to a neural network approach that may be applicable to many embodiments and implementations. However, it should be understood that many other types and structures of neural networks may be used. Indeed, many different methods have been and are being developed for generating neural networks, including neural networks that use complex structures and processes different from those described above. The method is not limited to any particular neural network approach, and any suitable approach may be used without departing from the present invention.

[0088] Artificial neural networks are adapted to specific purposes through a training process that is used to adapt / adjust / modify the weights and other parameters (e.g., bias) of the artificial neural network. It should be understood that many different training processes and algorithms are known for training artificial neural networks. Typically, training is based on a large training set, in which a large number of input data examples are provided to the network. In addition, the output of the artificial neural network is typically (directly or indirectly) compared with an expected or ideal result. A cost function can be generated to reflect the expected results of the training process. In a typical scenario known as supervised learning, a cost function typically represents the distance between the prediction of specific input data and a reference true value. Based on the cost function, the weights can be changed, and by repeating the process for the modified weights, the artificial neural network can be adapted to a state in which the cost function is minimized.

[0089] In more detail, during the training step, the neural network can have two different information flows, from input to output (forward pass) and from output to input (backward pass). In the forward pass, as described above, the data is processed by the neural network, and in the backward pass, the weights are updated to minimize the cost function. Typically, this backpropagation follows the gradient direction of the cost function loss landscape. In other words, by comparing the predicted output with the true value of the reference of a batch of data inputs, the direction in which the cost function is minimized and propagated backward can be estimated by updating the weights accordingly. Other methods known for training artificial neural networks include, for example, the Levenberg-Marquardt algorithm, the conjugate gradient method, and the Newton method.

[0090] In the present case, training may specifically include a training set comprising a potentially large number of pairs of captured audio signals and measured respiratory waveform signals. The measured respiratory waveform signals may be generated from sensor signals of a sensor arranged to measure a characteristic that depends on lung air volume. Thus, training of an artificial neural network may be performed using a training set comprising linked audio / speech data / signals and sensor data / signals representing lung air volume measurements during speech.

[0091] In some embodiments, the training data can be an audio signal in a time period corresponding to the processing time interval of the trained artificial neural network, for example, the number of samples in the training audio signal can correspond to the number of samples corresponding to the input nodes of the trained (one or more) artificial neural network. Thus, each training example can correspond to one operation of the trained (one or more) artificial neural network. However, typically, a batch of training samples is considered for each step to speed up the training process. In addition, many upgrades to gradient descent can also speed up convergence or avoid local minima in the cost function loss landscape.

[0092] Figure 4 An example of how training data can be generated through a dedicated test is shown. A speech audio signal 401 can be captured by a microphone 403 during a time interval when a person is speaking. The resulting captured audio signal can be fed as input data to an artificial neural network 405 (suitable processing, such as amplification, digitization, and filtering, can be performed before generating a test audio signal that is fed to the artificial neural network). In addition, a respiratory waveform signal 407 is determined as a sensor signal from a suitable lung capacity sensor 409, for example, located on the user. For example, sensor 409 can be a respiratory induction plethysmography (RIP) sensor or a nasal cannula flow rate sensor.

[0093] A large number of such measurements can be made to generate a large number of pairs of training audio signals and breathing waveform signals. These signals can then be applied and used to train an artificial neural network, where the cost function is determined as the difference between the breathing waveform signal generated by the artificial neural network for the training audio signal and the measured training breathing waveform signal. The artificial neural network can then be adapted based on such a cost function as is known in the art. For example, a cost value can be determined for each combined set of training audio signals and / or training downmix audio signals (e.g., determining the average cost value for the training set). Typically, the cost function will include at least one component that reflects how close the generated signal is to the reference signal, the so-called reconstruction error. In some embodiments, the cost function will include at least one component that reflects how close the generated signal is to the reference signal from a perceptual perspective. In Figure 4 In the example of , the reference signal may specifically be a measured respiratory waveform signal.

[0094] The segmenter may determine speech segments of the audio signal based on the breathing waveform signal. Figure 5 An example of a respiratory waveform signal is shown in .

[0095] like Figure 5 As shown, a respiration waveform signal typically has a periodic component reflecting the expansion and contraction of the chest caused by the speaker's breathing. The repetition rate, duration between peaks or troughs, and other factors can vary, reflecting that respiration is not completely periodic. The inventors have recognized that a speaker's respiration affects speech and that not only can the time-varying properties caused by respiration be determined from captured audio, but these time-varying properties can also be used to adapt the speech segments used for speech processing of the same audio signal, providing improved performance, particularly better adaptability to different scenarios and user behaviors.

[0096] The specific segmentation method used to generate the speech segments may vary in different embodiments and applications.

[0097] In some embodiments, the segmenter 107 can be arranged to perform signal analysis of the respiratory waveform signal to extract properties of the periodic component of the signal. It should be understood that different algorithms for extracting periodic components from the signal are known, and any suitable algorithm can be used. Based on the analysis, periodic properties such as frequency / repetition rate, timing of the periodic component, etc. can be determined. In some embodiments, these properties can be used to generate speech segments. For example, a speech segment can be generated as a time interval that is a given percentage (e.g., 90%) of a complete cycle following the maximum value of the periodic component of the respiratory waveform signal. Typically, the maximum value of the respiratory waveform signal corresponds to the maximum chest expansion and air volume, and therefore corresponds to the time when the inhalation (inhalation) interval switches to the exhalation (expiration) interval.

[0098] Thus, it is possible to determine that the speech segment corresponds to the time period that is most likely to correspond to the time of speech. The time interval can be arranged to be the majority of the time period to reflect that speech / expiratory intervals are generally substantially longer than non-speech inspiratory intervals. This method can also provide reliable estimates for situations where the speaker's breathing is very rhythmic and constant. It can allow the generation of speech segments of the same duration and containing the same number of samples. In many embodiments, this can facilitate operation and implementation and can, for example, allow for fixed segment / block processing. For example, the same number of samples can be processed for each sequential processing.

[0099] In many embodiments, the determiner is arranged to detect local maxima of the respiratory waveform signal and determine a speech segment in response to the timing of the local extrema, and typically the determiner 105 is arranged to determine the start time and / or end time of the speech segment in response to the timing of one or more local extrema. The local extrema may typically be a local minimum, a local maximum, or both a local minimum and a local maximum.

[0100] In many embodiments, the determiner 105 may be arranged to determine the timing of the local extrema. For example, the determiner 105 may be arranged to determine the moment of the peak of the respiratory waveform signal, e.g. Figure 5 The peak value of the respiratory waveform signal. In some embodiments, the determined timing can be fitted to the periodic component. For example, the timing can be determined so that the least square error between the periodic timing moment and the measurement moment of the minimum / peak value is minimized. The periodic timing can then be used to generate speech segments, thereby generating periodic time intervals / speech segments.

[0101] In many embodiments, segmenter 107 can be arranged to generate segments of varying durations / lengths / sizes. Specifically, segmenter 107 can be arranged to determine speech segments to reflect dynamic changes in respiratory rhythm and changes in the duration of inspiratory and expiratory time intervals. In particular, in some embodiments, each speech segment can be determined based on local properties of the respiratory waveform signal rather than based on average properties of a larger number of repetitions of the periodic component of the respiratory waveform signal.

[0102] In many embodiments, the timing of a speech segment is based on the timing of a local maximum, the timing of a local minimum, and / or on a pair of a local maximum and an adjacent local minimum.

[0103] For example, in some embodiments, the determiner 105 may be arranged to detect a local maximum. When a local maximum is detected, the speech segment may be determined as the audio signal in the time interval following the time of the local maximum. In many embodiments, the determiner 105 may be arranged to determine the start time of the speech segment in response to, and typically as a result of, the time of the local maximum of the respiratory waveform signal.

[0104] The expiration time interval may typically follow the maximum lung volume corresponding to a peak in the respiratory waveform signal. Therefore, determining the time interval at which a speech segment follows a maximum in the respiratory waveform signal will tend to result in speech segments being generated for the expiration time interval. This method can provide suitable time interval determination for speech segments of speaker activity.

[0105] In some embodiments, the determiner 105 may additionally or alternatively be arranged to detect a local minimum. When a local minimum is detected, the speech segment may be determined as the audio signal in the time interval preceding the local minimum. In many embodiments, the determiner 105 may be arranged to determine the end time of the speech segment in response to, and typically as a result of, the time of the local minimum of the respiratory waveform signal.

[0106] The expiratory time interval may typically precede the minimum lung volume corresponding to a minimum in the respiratory waveform signal. Therefore, determining the time interval of a speech segment preceding a minimum in the respiratory waveform signal will tend to result in generating speech segments for the expiratory time interval. This method can provide suitable time interval determination for speech segments of speaker activity.

[0107] In some cases, the duration of the time interval can be predetermined or dynamically determined based on, for example, various parameters. For example, in some cases, the duration of the time interval can depend on available computing resources. For example, when computing resources are not limited (e.g., because not many other processes are currently active), the time interval can be set to be very long, while if computing resources are currently limited (e.g., because other complex processes are currently active), the length of the time interval can be reduced. Thus, in some embodiments, the length of the time interval and the length of the speech segment can vary based on various parameters. In certain instances, this can affect the proportion of the speech signal that is actually processed by the speech processing. Of course, this may not be suitable for some processes, but may be acceptable in others. For example, for speech recognition, it is possible to recognize only the beginning of a sentence, which is typically aligned with the beginning of the exhalation interval and, therefore, with a maximum value. Therefore, for some applications, speech recognition may be acceptable even if it is performed only for a period of time at the beginning of the exhalation interval and improves over longer time intervals. The audio device can be configured to adapt the duration considered, and therefore the quality of the speech recognition, based on the available computing resources.

[0108] In some embodiments, the duration of the time interval can be dependent on the respiratory waveform signal. For example, the average cycle time of the periodic component of the respiratory waveform signal can be determined, and the duration of the time interval for the speech segment can be set to a given percentage of the cycle duration. In other embodiments, the duration of the time interval can be dynamically variable and can be different for consecutive speech segments.

[0109] In many embodiments, a speech segment can be determined as a segment of the audio signal between the time of a local maximum of the respiratory waveform signal and the time of a local minimum of the respiratory waveform signal. The start time of the segment can be determined as the time of the local maximum, and the end time of the segment can be determined as the time of the next local minimum. Thus, the time interval can be the interval from the local maximum to the local minimum.

[0110] The maximum value of the respiratory waveform signal generally corresponds to the current maximum lung volume and the beginning of the expiratory interval, and the minimum value of the respiratory waveform signal generally corresponds to the current minimum lung volume, the beginning of the inspiratory interval, and the end of the expiratory interval.

[0111] Therefore, the method can provide an efficient method for determining the respiratory / expiratory time interval and the inspiratory time interval based on the respiratory waveform signal. This may result in speech segments being determined as segments of the audio signal during the respiratory / expiratory time interval, and non-speech segments being determined as segments of the audio signal during the inspiratory time interval.

[0112] Respiration comprises a series of alternating inhalations (inhalation of air into the lungs) and exhalations / expirations (exhalation of air from the lungs), and the speech processor 103 may be arranged to evaluate the respiratory waveform signal to determine timing properties of the speaker's inhalation / exhalation intervals. In particular, the speech processor 103 may be arranged to evaluate the timing of the inhalation / exhalation intervals, specifically when these intervals begin and / or cease.

[0113] Such parameters may be determined, for example, by a peak picking algorithm that locates local minima and maxima in the respiratory waveform signal, which correspond to the start of the inspiration / inhalation and expiration / exhalation phases.

[0114] In some methods, the audio signal can be specifically divided into speech segments and non-speech segments, specifically, speech segments corresponding to exhalation intervals and non-speech segments corresponding to inhalation intervals. The speech processing can then be arranged to treat different speech segments differently. For example, in some embodiments, the attenuation of non-speech segments can be increased relative to that of speech segments, or speech recognition can be performed only during speech segments.

[0115] In this method, the respiratory waveform signal can represent the temporal breathing pattern of the speaker (or speakers), which can be used to segment the speech according to respiratory events (such as inspiration or speech and non-speech breaths). Speech processing is performed based on these segments reflecting the respiratory events.

[0116] This approach can leverage the insight that speech is often performed subconsciously, synchronizing breathing with the grammatical organization of spoken language, such as at the beginning of a new sentence or clause. Therefore, if speech audio is segmented based on breathing events, the segmented speech may represent the underlying intended grammatical structure of the speech content. Therefore, adapting speech processing to breathing patterns can allow for improved speech processing in many scenarios.

[0117] The audio device is arranged to adjust speech processing in response to the breathing waveform signal by segmenting the audio signal into speech segments based on the breathing waveform signal. Figure 1 The speech processing means is arranged to adapt the speech segmentation to reflect the current breathing of the speaker, rather than performing a predetermined speech segmentation based solely on nominal or expected properties.

[0118] The method reflects the inventors' recognition that a speaker's breathing affects speech, and that not only can time-varying characteristics resulting from breathing be determined, but these time-varying characteristics can also be used to adapt speech segmentation to speech processing of the same audio signal to provide improved performance, in particular performance that can better adapt to different scenarios and user behaviors.

[0119] The specific speech segmentation, processing, and adaptation performed may depend on the individual embodiment. However, in many embodiments, the method can be used to adjust operation to reflect the different breathing patterns that a person may exhibit, for example, due to performing different activities. For example, when a person exercises, breathing may become strained, resulting in very garbled and slow speech. Figure 1 The speech processing device can automatically adapt to optimize for slow speech with longer silence intervals.

[0120] The approach may provide significantly improved speech processing in many scenarios, and may, for example, provide particularly advantageous operation for speakers who engage in strenuous activities (such as exercise) or who have breathing or speaking difficulties.

[0121] The speech processing circuit 103 is configured to perform segment-based speech processing on the audio signal.

[0122] In some embodiments, the speech processor 103 can be specifically configured to modify speech processing for inspiratory intervals. Specifically, in many embodiments, the speech processor 103 can be configured to inhibit or stop speech processing during inspiratory intervals. For example, speech processing in the form of speech recognition can simply ignore all inspiratory intervals and only apply to expiratory intervals.

[0123] The speech processor 103 may be configured to determine inhalation periods of the audio signal. These periods may therefore reflect time intervals in the audio signal during which the speaker is estimated to be inhaling, and thus may be considered non-speech segments. In some embodiments, the speech processor 103 may be configured to attenuate the audio signal during the inhalation periods. For example, a speech enhancement algorithm configured to reduce noise by emphasizing the speech component of the audio signal may attenuate the audio signal during the inhalation periods, including, in many embodiments, completely attenuating the signal.

[0124] Such methods can tend to provide significantly improved performance and can allow speech processing to be adapted and specifically applied to portions of the audio signal where speech is present, while allowing different approaches to be taken and specifically attenuated for portions of the audio signal where speech is not possible. These methods can, for example, allow breathing sounds to be significantly reduced by eliminating heavy breathing sounds during the inspiratory phase of the respiratory wave signal. This can significantly improve the clarity of perceived speech and can provide perceived speech accordingly. In some embodiments, it can allow portions of the speech data that occur during inspiration (estimated from the respiratory waves) to be removed prior to automatic speech recognition, thereby facilitating this process and reducing the risk of false word detection. In speech coding embodiments, it can improve, for example, coding efficiency because speech coding can not be performed during the inspiratory time interval. For example, an indication of a silent or non-speech time interval can simply be inserted into the coded data stream to indicate that no speech data is provided for that interval.

[0125] In many embodiments, the speech processing circuitry 103 is specifically arranged to perform segment-based speech enhancement processing.

[0126] In some embodiments, the speech processing may be speech enhancement. Such speech enhancement may, for example, be used to provide a clearer speech signal while reducing noise from, for example, other audio sources and the audio signal. For example, as previously described, the audio signal may decay during the inhalation period. As another example, the respiratory waveform signal may indicate a transition or phase in the respiratory cycle that tends to be associated with a particular sound, and the speech processing may be arranged to generate corresponding signals and apply them inversely to provide an audio cancellation effect for such sounds. In this case, segmentation may allow a particular sound to be concentrated in a particular speech segment.

[0127] As another example, speech processing may include, for example, estimating acoustic models of speech and background (out-of-segment data) for use in noise reduction, dereverberation, speech enhancement, echo cancellation, multi-microphone beamforming, or other processing tasks aimed at improving speech quality or segmenting multiple speakers from one of multiple audio signals.

[0128] In more detail, in some embodiments, the segmented speech enhancement process can be used to separate and enhance the speech of individual speakers, for example, in teleconferencing or logging applications. Given a continuous speech recording, a determiner can estimate a respiratory waveform signal and determine its corresponding local minima and maxima. Next, the continuous audio signal can be segmented into speech segments s corresponding to the time range between the i-th maximum and the next minimum of the respiratory waveform. i , and the background / inspiratory segment b corresponding to the time range between the i-th local minimum and the next local maximum i In some embodiments, it can be assumed that each speech segment s i Contains the speech of a speaker, and the background segment represents the static noise background. Based on this formula, we can implement several examples of speech enhancement methods:

[0129] In noise reduction applications, b i The sequence of segments can be used to estimate the background noise spectrum model W n (ω), the background noise spectrum model can be adaptively used to suppress s i The background noise in the incoming speech signal in each segment is reduced in this way, reducing the influence of the background noise in the further processing and transmission of the speech content.

[0130] In the case of close-talk microphones, breathing sounds may be disturbingly high. In one embodiment, b i The amplitude of each segment is suppressed to reduce the influence of breathing sound. i This type of enhancement is enabled by the detection of high-level breathing sounds in a clip.

[0131] Adaptive beamforming in multi-microphone speech capture aims to estimate the multi-channel signal processing model A k The parameters of (ω) are used to capture the clean speech of the object (or angular direction) k around the multi-microphone device. We can use the speaker recognition algorithm to detect two consecutive speech segments s i and s i+1 Whether the two segments contain speech from the same or different speakers. If it is determined that the two segments come from the same speaker k, then the two segments are used to update A k The parameter of (ω) is used. This process can be performed in parallel with multiple objects (or angular directions) k to build a multi-beam speech capture algorithm. The method effectively uses the following insight: turn transitions in a conversation usually occur during inhalation, that is, we wait for the end of a sentence and start speaking when the other speaker inhales.

[0132] The same segmentation principle based on the extrema of the respiratory waveform signal, combined with speaker recognition in continuous speech segments, can also be used to control the adaptation of filter coefficients in acoustic echo cancellation (AEC). i and s i+1 The parameters of the algorithm can remain unchanged while the identity of the speaker detected in the image remains the same, and the filters can be adapted if the speaker has changed.

[0133] In many embodiments, the speech processing circuitry 103 is specifically arranged to perform segment-based speech recognition processing.

[0134] In automatic speech recognition, speech can be segmented into processing frames, where individual words are identified in the context of adjacent utterances. There is no single solution to the basic segmentation problem, but the utterance with the maximum likelihood score recognized by the model is usually selected. However, it is well known that segmenting speech in a way that is different from the expected sentence spoken by the speaker often leads to serious errors in the recognized speech. Speech segmentation based on taking into account the respiratory waveform signal can provide more accurate and appropriate segmentation in many scenarios and for many signals. In particular, in many cases, it can allow the speech to be segmented at the input of the automatic speech recognition to follow the actual sentence and clause structure of the speech. This can allow the performance of automatic speech recognition of continuous speech to be significantly improved.

[0135] As mentioned above, the impact of different sentence structures on speech recognition processing can be very significant. As an example, sentence structure can lead to very different meanings, such as indicated by the difference between the following text:

[0136] “Dear John,

[0137] I want someone who knows what love is. You are generous, kind, and thoughtful. People who are not like you admit that they are useless and inferior.

[0138] You ruined me for other men. I longed for you. When we were apart, I felt nothing. I could be happy forever - would you make me yours?

[0139] simple"

[0140] and the following text:

[0141] "Dear John,

[0142] I want someone who knows what love is. Surround yourself with generous, kind, thoughtful people who are not like you.

[0143] Admit your worthlessness and inferiority. You've ruined me. I long for other men. I have no feelings for you.

[0144] I will be forever happy when we are apart. Will you let me go?

[0145] your,

[0146] simple"

[0147] By applying segmentation based on the respiratory waveform signal, the segmentation can reflect the sentence structure, allowing speech recognition algorithms to, for example, distinguish meaning in the above example.

[0148] Especially for automatic speech recognition, the approach can provide significantly improved performance by adapting the process and segmentation to the speaker's current breathing pattern.

[0149] For example, as previously described, the speech processor 103 can be arranged to determine inhalation periods of the audio signal, corresponding to times when the speaker is inhaling and therefore not speaking. The speech processor 103 can then be arranged to exclude these inhalation periods from the speech recognition processing applied to the audio signal. Thus, times when the respiratory waveform signal may indicate that speech is unlikely can be excluded from speech recognition, thereby reducing the risk of falsely detecting words when no speech is present. It can also improve accuracy detection at other times because it can estimate when words may be spoken. In particular, it can allow segmentation to follow sentence structure, which can lead to a significant improvement in the accuracy of speech recognition.

[0150] The method may be particularly advantageous for automatic speech recognition (e.g., providing a voice interface). For example, a personal assistant's voice interface can control media, navigation features, exercise tracking, and various other information services during exercise. However, speech recognition is typically optimized for normal resting speech and may be personalized. During and after intense exercise, the demand for oxygen increases significantly, resulting in changes in breathing that affect speech. Specifically, it may increase the breathing rate, which may result in shorter speech intervals and sentences, and the speech recognition performed by the audio device may reflect this.

[0151] More specifically, in some embodiments, a segmented speech recognition process may be applied to select the utterance that best matches the determined respiratory waveform signal. Automatic speech recognition (ASR) algorithms typically generate multiple candidate word sequences corresponding to the highest matches to the input speech sequence. In conventional ASR systems, candidate sequences are ranked based on the well-known Viterbi algorithm. For the above example, the ASR algorithm may generate several equally likely candidates, such as

[0152] "I long for it. I have no feelings for you at all! When will we part?"

[0153] "I long for you! I felt nothing when we were apart."

[0154] "I didn't feel anything for you. When we were apart."

[0155] When the timing of respiratory events is combined with the ASR signature sequence, we can calculate a score for how well each sequence matches the respiratory pattern. In one embodiment, we can calculate a score for how often text symbols such as periods, commas, exclamation points, or question marks coincide with the minimum or maximum values of the respiratory waveform signal. The sequence with the highest score is selected.

[0156] In many embodiments, speech processing may be speech coding, wherein the captured speech is, for example, encoded for efficient transmission or distribution over a suitable communication channel. More efficient coding can often be achieved simply by not encoding any signal during non-speech segments and encoding the audio signal only during speech segments. In some embodiments, the coding during speech segments may be further based on assumptions about the speech segments, such as reflecting complete sentences. For example, the speech processor may be arranged to encode the audio signal in the speech segments by processing each speech segment s_i as a separate coding entity, such that a specific bit budget is allocated to each speech segment s_i, and a smaller bit budget is allocated to background b_i segments.

[0157] In some embodiments, the speech processing circuit is arranged to select speech segments to include speech from a particular speaker (the speaker for whom the breathing waveform signal is provided) rather than from another speaker. For example, the audio signal may include audio from multiple speakers, and the segmenter may be arranged to generate speech segments that may include speech from different speakers. For example, the segmenter may generate speech segments based solely on the audio signal without considering the breathing waveform signal. As an example, a speech segment may be generated when the instantaneous level of the audio signal exceeds a threshold, while a non-speech / silence segment may be generated when the instantaneous level is below the threshold. Thus, the segmenter may generate speech segments that reflect a high probability of speech, but without considering the current speaker, that is, without considering the identity of the active speaker.

[0158] Segmenter 107 may continue to select speech segments for a given speaker based on the given speaker's breathing waveform signal. As a low-complexity example, segmenter 107 may simply select speech segments during exhalation times determined from the breathing waveform signal. In some embodiments, the breathing waveform signal may reflect different breathing patterns when the speaker is speaking versus when the speaker is silent, and segmenter 107 may be configured to identify such periods and select speech segments for the given speaker.

[0159] The audio device may then use (a subset of) the specific speech segments in further speaker-specific processing, eg adapting the beamforming based on such segments only.

[0160] In multi-speaker applications, such as multi-beam beamforming in voice capture or dualization of conversations with multiple conversation partners, the problem is the segmentation of speech, while the identification of the active speaker is a challenging problem. Traditional acoustic signal processing techniques have difficulty performing this task. Speakers often interrupt and talk back to each other, and their speech may be similar. However, in multi-speaker scenarios, it has been observed that humans are good at detecting when to take a turn. This is usually done at (or before) the end of the previous speaker's sentence. Therefore, considering the breathing waveform signal of one or more speakers can allow the detection of different speakers taking turns.

[0161] As a specific example, the audio device may be configured such that segment clustering is based on s i segments and determines changes in speaking turns by comparing consecutive speech segments using a machine learning model for speaker identification.

[0162] The audio device can be implemented in one or more appropriately programmed processors. For example, the artificial neural network of the determiner can be implemented in one or more such appropriately programmed processors. Different functional blocks, in particular the artificial neural network, can be implemented in separate processors and / or can be implemented in the same processor, for example. Examples of suitable processors are provided below.

[0163] Figure 6 6 is a block diagram illustrating an example processor 600 according to an embodiment of the present disclosure. The processor 600 may be used to implement one or more processors that implement the apparatus or elements thereof as described above (including, in particular, one or more artificial neural networks). The processor 600 may be any suitable processor type, including but not limited to a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA) (where the FPGA has been programmed to form a processor), a graphics processing unit (GPU), an application specific integrated circuit (ASIC) (where the ASIC has been designed to form a processor), or a combination thereof.

[0164] Processor 600 may include one or more cores 602. Core 602 may include one or more arithmetic logic units (ALUs) 604. In some embodiments, core 602 may include a floating point logic unit (FPLU) 606 and / or a digital signal processing unit (DSPU) 608 in addition to or in place of ALU 604.

[0165] Processor 600 may include one or more registers 612 communicatively coupled to core 602. Registers 612 may be implemented using dedicated logic gates (e.g., flip-flops) and / or any memory technology. In some embodiments, registers 612 may be implemented using static memory. Registers may provide data, instructions, and addresses to core 602.

[0166] In some embodiments, processor 600 may include one or more levels of cache memory 610 communicatively coupled to core 602. Cache memory 610 may provide computer-readable instructions to core 602 for execution. Cache memory 610 may provide data for processing by core 602. In some embodiments, computer-readable instructions may have been provided to cache memory 610 by local memory (e.g., local memory attached to external bus 616). Cache memory 610 may be implemented using any suitable cache memory type, such as metal oxide semiconductor (MOS) memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.

[0167] Processor 600 may include a controller 614 that can control inputs to processor 600 from other processors and / or components included in the system and / or outputs from processor 600 to other processors and / or components included in the system. Controller 614 can control data paths in ALU 604, FPLU 606, and / or DSPU 608. Controller 614 can be implemented as one or more state machines, data paths, and / or dedicated control logic. Gates of controller 614 can be implemented as independent gates, FPGAs, ASICs, or any other suitable technology.

[0168] Registers 612 and cache 610 may communicate with controller 614 and core 602 via internal connections 620A, 620B, 620C, and 620D. The internal connections may be implemented as a bus, multiplexer, crossbar switch, and / or any other suitable connection technology.

[0169] Input and output of processor 600 may be provided via bus 616, which may include one or more conductors. Bus 616 may be communicatively coupled to one or more components of processor 600, such as controller 614, cache 610, and / or registers 612. Bus 616 may be coupled to one or more components of the system.

[0170] Bus 616 can be coupled to one or more external memories. External memory can include read-only memory (ROM) 632. ROM 632 can be a mask ROM, an electronically programmable read-only memory (EPROM), or any other suitable technology. External memory can include random access memory (RAM) 633. RAM 633 can be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. External memory can include electrically erasable programmable read-only memory (EEPROM) 635. External memory can include flash memory 634. External memory can include a magnetic storage device such as a disk 636. In some embodiments, external memory can be included in the system.

[0171] The present invention can be implemented in any suitable form, including hardware, software, firmware or any combination thereof. The present invention can optionally be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components of the embodiments of the present invention can be implemented physically, functionally and logically in any suitable manner. In fact, the functions can be implemented in a single unit, in multiple units or as part of other functional units. Like this, the present invention can be implemented in a single unit, or can be physically and functionally distributed between different units, circuits and processors.

[0172] Although the present invention has been described in conjunction with certain embodiments, it is not intended that the present invention be limited to the specific forms set forth herein. Rather, the scope of the present invention is limited solely by the claims. In addition, although features may appear to be described in conjunction with specific embodiments, those skilled in the art will recognize that the various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0173] Furthermore, although listed separately, multiple devices, elements, circuits or method steps may be implemented by, for example, a single circuit, unit or processor. In addition, although individual features may be included in different claims, these features may be advantageously combined together, and inclusion in different claims does not mean that the combination of these features is not feasible and / or advantageous. Furthermore, the inclusion of a feature in one category of claims does not imply a limitation to that category, but rather indicates that the feature is equally applicable to claims in other categories. Furthermore, the order of features in the claims does not imply that the features must work in any particular order, in particular, the order of the individual steps in a method claim does not imply that the steps must be performed in that order. On the contrary, the steps may be performed in any suitable order. Furthermore, singular references do not exclude a plurality. Thus, references to "a", "an", "first", "second", etc. do not exclude a plurality. The reference numerals in the claims are provided merely as clarifying examples and should not be construed as limiting the scope of the claims in any way.

Claims

1. A device for performing speech processing on an audio signal, the device comprising: an input terminal (101) arranged to receive the audio signal, the audio signal comprising a speech audio component for a speaker; a determiner (105) arranged to generate a breathing waveform signal indicative of a lung air volume of the speaker as a function of time; a segmenter (107) arranged to segment the audio signal to generate speech segments of the audio signal, the segmentation being responsive to the breathing waveform signal; as well as Speech processing circuitry (103) is arranged to perform speech processing on the audio signal, the speech processing being segment-based processing applied to the speech segments.

2. The device according to claim 1, wherein The determiner (105) is arranged to determine the breathing waveform signal based on the audio signal.

3. The device according to claim 2, wherein The determiner (105) comprises a trained artificial neural network having an input node for receiving samples of the audio signal and an output node arranged to provide samples of the respiration waveform signal.

4. An apparatus according to any preceding claim, wherein The determiner (105) is arranged to detect a local maximum of the respiratory waveform signal and determine the speech segment in response to the timing of the local maximum.

5. An apparatus according to any preceding claim, wherein The determiner (105) is arranged to detect a local minimum of the respiratory waveform signal and determine the speech segment in response to the timing of the local minimum.

6. An apparatus according to any preceding claim, wherein The determiner (105) is arranged to determine at least one speech segment as a segment of the audio signal between a time of a local maximum value of the respiratory waveform signal and a time of a local minimum value of the respiratory waveform signal.

7. An apparatus according to any preceding claim, wherein The segmenter (107) is arranged to divide the audio signal into the speech segments and the non-speech segments.

8. An apparatus according to any preceding claim, wherein The segmenter (107) is arranged to determine a respiratory time interval and an inspiratory time interval based on the respiratory waveform signal; and determine the speech segment as a segment of the audio signal during the respiratory time interval, and determine the non-speech segment as a segment of the audio signal during the inspiratory time interval.

9. An apparatus according to any preceding claim, wherein The speech processing circuit (103) is arranged to select a speech segment as comprising speech from the speaker rather than from another speaker, the selection being dependent on the respiration waveform signal.

10. The apparatus of any preceding claim, further comprising a sensor input (109) arranged to receive a chest sensor signal, and wherein The determiner (105) is arranged to determine the respiratory waveform signal from the chest sensor signal.

11. An apparatus according to any preceding claim, further comprising a video input (109) arranged to receive a video signal comprising a video image of the speaker, and wherein The determiner (105) is arranged to determine the respiratory waveform signal based on the video image.

12. An apparatus according to any preceding claim, wherein The speech processing includes speech enhancement processing.

13. An apparatus according to any preceding claim, wherein The speech processing includes speech recognition processing.

14. A method for performing speech processing on an audio signal, the method comprising: receiving the audio signal, the audio signal including a speech audio component for a speaker; generating a breathing waveform signal indicative of a volume of lung air of the speaker as a function of time; segmenting the audio signal to generate speech segments of the audio signal, the segmenting being responsive to the breathing waveform signal; and Speech processing is performed on the audio signal, the speech processing being a segment-based processing applied to the speech segments.

15. A computer program product comprising computer program code means adapted to perform all the steps of claim 14 when said program is run on a computer.

Citation Information

Patent Citations

  • Virtual Trainer

    US20090098981A1

  • Processing a video for respiration rate estimation

    US9301710B2