Detecting accidental activation of speech interface
A machine-learning based replay detector in speech interfaces differentiates between human and non-human voice inputs, reducing accidental activations and improving speech recognition reliability.
Patent Information
- Application Number
- PCT/US2024/060713
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-24
- Filing Date
- 2024-12-18
- Publication Date
- 2025-07-31
AI Technical Summary
Existing speech interfaces are prone to accidental activation by ambient speech, particularly when the wake-word is set to a common term, leading to unintended processing of non-human voice inputs from devices like loudspeakers.
A replay detector utilizing machine-learning techniques distinguishes between anomalous and non-anomalous audio inputs by analyzing acoustic signals, preventing speech interfaces from acting on technically-produced audio and allowing processing of human-generated audio.
Reduces the likelihood of accidental speech interface activation by effectively differentiating between human and non-human voice inputs, enhancing the reliability of speech recognition systems.
Smart Images

Figure US2024060713_31072025_PF_FP_ABST
Abstract
Description
DETECTING ACCIDENTAL ACTIVATION OF SPEECH INTERFACERelated Applications
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 624,325, filed on January 24, 2024, the contents of which are incorporated herein by reference.Background
[0002] As natural-language processing has become more sophisticated, devices with speech interfaces have become more ubiquitous. These devices typically understand speech and also emit speech. As a result, it is no longer uncommon for a first device to receive and process speech that originated from a second device, rather than from a human being. The second device, in such cases, could be another device equipped with a speech interface. But it could also be a loudspeaker connected to a media source, such as a music source or audiovisual content.
[0003] A useful safeguard is to use a wake-word. In such cases, a device will ignore speech that does not begin with the wake-word. However, the effectiveness of this safeguard depends on how uncommon the wake- word happens to be in ambient speech. Because wake-words can be set by a user, it is quite possible for a user to set the wakeword to something that commonly occurs in normal speech.
[0004] For example, since one is ultimately addressing a computer, it is not uncommon for a user to reset the default wake-word to simply be “computer.” This promotes a more natural interaction with the device. Because this word is common in ambient speech, there exists a not insignificant risk of accidentally triggering the computer’s speech processing function.Summary
[0005] A technical problem to be addressed by the invention is that of distinguishing between a voice from a loudspeaker and voice emitted by a human being. Embodiments include those in which the voice emitted by a loudspeaker is recording of a human being’s voice. This often occurs when an attempt is being made to spoof the speech interface. Embodiments also include those in which the voice emitted by a loudspeaker is that of an automotive assistant or other digital assistant. In either case, the subject matterdisclosed herein reduces the likelihood that a voice emitted by a loudspeaker will trigger processing by a speech interface.
[0006] In one aspect, the invention includes an automotive assistant that executes in an infotainment system of a vehicle having a microphone that provides an audio input and a loudspeaker that provides an audio output, a speech interface that is in communication with the automotive assistant, and a replay detector that prevents the speech interface from acting on the audio input when the audio input is from a first class of audio inputs and allows the speech interface to act on the audio input when the audio input is from a second class of audio inputs. In this context, audio inputs from the first class result from technically-produced acoustic signals that have been produced in an environment of said microphone and audio signals from the second class result from acoustic signals produced by one or more persons in the environment.
[0007] Without loss of generality, and for greater ease of exposition, audio inputs from the first and second classes will be referred to herein as “anomalous” and “non- anomalous” audio inputs. The distinction between an anomalous and non-anomalous audio input can be understood by recognizing that the acoustic signal received by the microphone is a sound wave whose existence depends on the concurrent existence of a human being that is producing the sound wave, whereby termination of that human being’s existence results in termination of the acoustic signal at the microphone following a time interval that depends on the speed of sound and the distance between the microphone and the location formerly occupied by the human being.
[0008] In contrast, a non-anomalous audio input results from an acoustic signal whose existence does not depend on the concurrent existence of a human being. Such an acoustic signal can be present even if no human being is present. Examples of such an acoustic signal include a recording of a human being a voice output of a digital assistant, or the sound of a radio program.
[0009] The distinction between an anomalous and non-anomalous audio input can also be understood by recognizing that the acoustic signal at the microphone can originate in two ways: as a result of a conversion of electrical energy into mechanical energy and as a result of a conversion of chemical energy into mechanical energy. Examples of the former include playback from a loudspeaker. Examples of the latter include the human voice as it is emitted from the human using chemical energy derived from the human’s metabolic processes and converted into mechanical energy that carries out mechanical work on the human’s diaphragm to provide fluid velocity sufficient to setthe human’s vocal cords into vibration and to cause time-varying tension on the vocal cords using muscles that likewise convert chemical energy into mechanical energy to perform mechanical work on the vocal cords. The former is considered a “technically- produced acoustic signal” whereas the latter is not.
[0010] An example of a technically-produced acoustic signal is one that results from replaying a recorded voice of a human being. In such a case, the acoustic signal that impinges upon the microphone is a result of conversion of chemical energy into mechanical energy by a human being, the conversion of that mechanical energy into electrical energy for storage in a recording medium, the conversion of electrical energy back into mechanical energy upon being replayed via the loudspeaker, and the conversion of that mechanical energy back into electrical energy at the microphone.[Oil] Another example of a technically-produced acoustic signal is one that results from a spoken utterance by an automotive assistant or digital assistant. In that case, electrical energy associated with an electrical signal representing speech is converted into mechanical energy by the loudspeaker, which then is converted back into electrical energy by the microphone.
[0012] An example of a non-technically-produced acoustic signal, which would in turn result in a non-anomalous audio signal, is that produced by a human being in the environment of the microphone. In this case, chemical energy from human metabolism is converted into mechanical energy, which then impinges onto the microphone to be converted into electrical energy.
[0013] The invention relies on the discovery that there exist telltale features in the acoustic signal that can ultimately be used by a machine-learning system to inspect audio signals to distinguish these two types of acoustic signal.
[0014] As used herein, an “acoustic signal” is a sound wave whereas an “audio signal” is a signal that occupies a frequency band that is audible to human beings.
[0015] Among the embodiments are those in which the replay detector has been trained using machine-learning techniques. Among those are those in which the replay detector has been trained to distinguish between an anomalous audio input and a non- anomalous audio input, those in which it has been trained to distinguish between a spectrogram that represents an anomalous audio input and a spectrogram that represents a non-anomalous audio input, those in which it has been trained to distinguish between an output of a generic signal processing method that represents an anomalous audio inputand an output of a generic signal processing method that represents a non-anomalous audio input, those in which it has been trained based on training audio signals obtained using acoustic transducers in the vehicle, among which are the loudspeaker and the microphone, and those in which it has been trained using the impulse response of the vehicle.
[0016] Still other embodiments include a signal conditioner that conditions the audio input from the microphone before the audio input reaches the speech interface. Among these are those in which the replay detector receives the audio input after the audio input has been conditioned, those in which it receives it before it has been conditioned, and those in which it receives it directly from the microphone.
[0017] Still other embodiments include a switch that switches between providing the replay detector with a conditioned audio input and providing the replay detector with an unconditioned audio input.
[0018] Other embodiments include a switch having a state that prevents the speech interface from receiving the audio input from the microphone wherein the replay detector causes the switch to be in the state when the microphone is receiving the anomalous speech and out of the state when the microphone is receiving the natural speech. Among these are embodiments in which the switch is set to permit the speech interface to receive the audio input until the replay detector has determined that the audio input is anomalous, at which point it disconnects the switch, and those in which the switch is set to block the speech interface from receiving the audio input until the replay detector has determined that the audio input is non-anomalous.
[0019] In one aspect, the invention includes a digital assistant, a speech interface, and a replay detector that prevents the speech interface from acting on the audio input when the acoustic input is an anomalous speech and allows the speech interface to act on the audio input when the audio input is a non-anomalous audio input.
[0020] In yet another aspect, the invention features a method that includes detecting that a first audio input obtained from a microphone is an anomalous audio input, preventing a speech interface from acting upon the first acoustic input, detecting that a second audio input obtained from the microphone is a non-anomalous audio input, and allowing the speech interface to act upon the second acoustic input.
[0021] Among the practices of the foregoing method are those in which the microphone is in a vehicle.
[0022] These and other features of the invention will be apparent from the following detailed description and the accompanying figure, in which:Description of Drawings
[0023] FIG. 1 shows an exemplary architecture for a system that reduces accidental triggering of a speech interface.Detailed Description
[0024] FIG. 1 shows a vehicle 10 having an infotainment system 12 that receives audio input 14 from a microphone 16 and provides audio output 18 via a loudspeaker 20.
[0025] The infotainment system 12 hosts an automotive assistant 22. A speech interface 24 between the microphone 16 and the automotive assistant 22 establishes two- way speech communication with an occupant in the vehicle 10.
[0026] A first switch 26 interrupts the signal path between the microphone 16 and the speech interface 24. Whether the switch 24 is open or closed determines whether the speech interface 24 receives the audio input 14 ever reaches the speech interface 24.
[0027] A replay detector 28 controls the state of the first switch 26. It does so by classifying the audio input 14 as being anomalous or non-anomalous.
[0028] In some embodiments, the first switch 26 remains open by default, as shown in FIG. 1. Upon receiving an audio signal, the replay detector 28 classifies that audio signal as being anomalous or non-anomalous. Upon having classified the audio signal as an anomalous audio signal, the replay detector 28 does nothing to change the state of the first switch 26. In contrast, upon having classified the audio signal as a non-anomalous audio signal, the replay detector 28 sends a first control signal 30 to close the first switch 26. The first switch 26 remains closed until the speech interface 24 detects completion of speech in the audio signal, at which point the first switch 26 is reopened.
[0029] A disadvantage of the foregoing embodiment is that there is inevitably some delay in classification. In an alternative embodiment, the first switch 26 instead remains closed by default and is only opened if and when the replay detector 28 has classified the audio signal as being an anomalous audio signal.
[0030] In either case, the speech interface 24 then provides the speech to the automotive assistant 22 in a form suitable for the automotive assistant 22 to process. Ifnecessary, the automotive assistant 22 provides data to the speech interface 24, which the speech interface 24 then uses to formulate suitable speech to deliver via the loudspeaker 20.
[0031] In some embodiments, a signal conditioner 32 disposed to receive the audio input 14 from the microphone 16 conditions the audio input 14 it prior to sending it to the speech interface 24. Examples of a signal conditioner 32 arc those that remove background noise, e.g., road noise from the audio input 14 and those that enhances various psychoacoustic metrics in the audio input 14. For convenience in exposition, the signal that leaves the signal conditioner 32 shall also be referred to as the “audio input 14” even though certain of its features may have changed in subtle ways. When necessary, reference will be made to the signal “upstream” of and “downstream” from the signal conditioner 32.
[0032] To carry out the function of classifying the audio input 14, it is useful for the replay detector 28 to receive the audio input 14. In some cases, a lower error rate results when the replay detector 28 receives the raw audio input 14 prior to processing by the signal conditioner 32. In other cases, the converse is true. For example, if the signal conditioner 32, in the course of carrying out its function, happens to suppress artifacts that the replay detector 28 relies upon for classification, a lower error rate results from taking the audio input 14 upstream of the signal conditioner 34. On the other hand, if the signal conditioner 32, in the course of carrying out its function, happens to make the artifacts that the replay detector 28 relies upon for classification more prominent, then a lower error rate results from taking the audio input 14 from downstream of the signal conditioner 32. To provide additional flexibility for the replay detector 28, a second switch 34 allows connects the replay detector 28 either upstream or downstream of the signal conditioner 32 in response to a second control signal 36.
[0033] The replay detector 28 implements a machine-learning system 38 that has been trained using training signals that include anomalous speech from multiple combinations of loudspeakers and microphones as well as natural speech from multiple combinations of human beings and microphones. Based on this training data, the machine-learning system 38 grows increasingly adept at identifying subtle artifacts that are present in anomalous speech and absent from natural speech. This results in a machine-learning system 38 that classifies between natural and anomalous speech with a low error rate.
[0034] In a preferred embodiment, the training signals are converted into spectrograms. This is particularly useful because it permits leveraging conventional image classification algorithms for use in classifying speech signals instead of conventional images. The resulting spectrograms in the training data are either non- anomalous spectrograms, which represent non-anomalous audio inputs and anomalous spectrograms, which represent anomalous audio inputs. With this being the case, it becomes possible to develop signal-processing methods to enhance those features that distinguish between these two types of spectrograms to then apply such signal-processing techniques to incoming audio inputs.
[0035] An example of such a signal-processing technique is “mean spectral subtraction,” in which one determines a mean based only on training data in the first class. The resulting mean spectrogram is then subtracted from all other spectrograms, including training and testing spectrograms for both anomalous and non-anomalous spectrograms. The resulting classifier is thus tuned to detect anomalous audio input and to flag anything that is not non-anomalous audio input as being “anomalous.” This type of classifier is thus robust to unanticipated spoof attacks. After all, it does not rely on knowing the features of such spoof attacks in advance.
[0036] In other embodiments, the replay detector 28 implements a different method for rejecting anomalous speech. These include methods that require no machine learning, such as extracting calculated features from audio input.
[0037] Embodiments of the replay detector 28 that rely on machine learning include those that rely on deep learning, such as that carried out by deep neural networks, as well as those that rely on classical machine-learning methods, such as Markov models, support vector machines, and gaussian mixture models.
[0038] In the context of a vehicle 10, the locations of loudspeakers 20 and microphones 16 as well as their respective properties are known in advance. In addition, the vehicle’s impulse response is also known in advance. In such cases, it is particularly useful for the training signals to incorporate signals obtained from speech within the vehicle 10. In those cases that include a signal conditioner 32, it is useful to provide training signals that incorporate the effect of the signal conditioner 32.
[0039] Having described the invention and a preferred embodiment thereof, what is claimed as new and secured by letters patent is:
Claims
CLAIMS1. An apparatus comprising an automotive assistant (22) that executes in an infotainment system (12) of a vehicle (10) having a microphone (16) that provides an audio input (14) and a loudspeaker (20) that provides an audio output (18), a speech interface (24), and a replay detector (28) that prevents said speech interface from acting on said audio input when said audio input is from a first class of audio inputs and permits said speech interface to act on said audio input when said audio input is from a second class of audio inputs, wherein said first class of audio inputs comprises audio inputs that result from technically-produced acoustic signals that have been produced in an environment of said microphone and wherein said second class of audio inputs comprises audio inputs that result from acoustic speech signals that have been produced by at least one person in said environment, wherein said at least one person comprises a speech organ, and wherein said speech organ produces said acoustic speech signals while said at last one person is in said environment.
2. The apparatus of claim 1, wherein said replay detector has undergone machine learning to be trained to distinguish between audio inputs from said first class of audio inputs and audio inputs from said second class of audio inputs.
3. The apparatus of claim 1, wherein said replay detector comprises a deep neural network that has been trained to distinguish between audio inputs from said first class of audio inputs and audio inputs from said second class of audio inputs4. The apparatus of claim 1, wherein said replay detector comprises a classifier that has been trained to distinguish between a first spectrogram and a second spectrogram, wherein said first spectrogram is a spectrogram from a first class of spectrograms, wherein said second spectrogram is a spectrogram from a second class of spectrograms, wherein said first class of spectrograms comprises spectrograms from audio inputs from said first class of audio inputs, and wherein said second class of spectrograms comprises spectrograms from audio inputs from said second class of audio inputs.
5. The apparatus of claim 1 , wherein said vehicle comprises acoustic transducers, among which arc said loudspeaker and said microphone, wherein said replay detector comprises a classifier that has been trained based on training audio signals obtained using said acoustic transducers in said vehicle.
6. The apparatus of claim 1, wherein said replay detector comprises a classifier that has been trained to distinguish between audio inputs from said first class and audio inputs from said second class using an impulse response of a cabin in said vehicle.
7. The apparatus of claim 1, wherein said replay detector compares acoustic features obtained from said audio input to features calculated from a training set.
8. The apparatus of claim 1, wherein said replay detector comprises a classifier that has been trained based on acoustic characteristics of a cabin in said vehicle.
9. The apparatus of claim 1, further comprising a signal conditioner (32) that conditions the audio input before said audio input reaches said speech interface.
10. The apparatus of claim 1 , wherein said replay detector receives said audio input after said audio input has been conditioned.
11. The apparatus of claim 1, wherein said replay detector receives said audio input directly from said microphone.
12. The apparatus of claim 1, further comprising a switch (34) that switches between providing said replay detector with a conditioned audio input and providing said replay detector with an unconditioned audio input.
13. The apparatus of claim 1, further comprising a switch (26) having a state that prevents said speech interface from receiving said audio input from said microphone, wherein said replay detector causes said switch to be in said state in response to receiving an audio input from said first class and out of said state in response to receiving an audio input from said second class.
4. A method comprising receiving, from a microphone, a first audio input, determining that said first audio input is from a first class of audio inputs, preventing a speech interface from acting on said first audio input, receiving, from said microphone, a second audio input, determining that said second audio input is from a second class of audio inputs, and allowing said speech interface to act upon said second acoustic input, wherein said first class comprises audio inputs that result from technically-produced acoustic signals that have been produced in an environment of said microphone and wherein said second class comprises audio inputs that result from acoustic speech signals that have been produced by at least one person in said environment.
Citation Information
Patent Citations
Method, device, mobile user apparatus and computer program for controlling an audio system of a vehicle
EP3661797B1
System and method for processing audio data of aircraft cabin environment
US20230373654A1
Method of operating an audio device system and an audio device system
WO2023110836A1