METHOD AND VOICE-CONTROLLED NODE FOR AUDIO PATTERN RECOGNITION

The voice-controlled node uses advanced noise reduction techniques to isolate voice commands from speaker tones, addressing interference challenges in audio pattern recognition, improving accuracy and efficiency.

DE112018002871B4Active Publication Date: 2026-04-23INFINEON TECHNOLOGIES AMERICAS CORP
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
INFINEON TECHNOLOGIES AMERICAS CORP
Filing Date
2018-05-10
Publication Date
2026-04-23

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A procedure (600) that includes the following: Providing (602) an input signal (232, 332) based on sound waves (205, 207, 209) received by one or more microphones (220), wherein the input signal (232, 332) includes a voice command component and one or more interference components; Intermediate storage (604) of the input signal (232, 332); detection (606) of an indication to receive audio data (261, 361); receipt (608) of the audio data (261, 361) via one or more computer networks (114), wherein the audio data (261, 361) corresponds to one or more interference components, the receipt of the audio data (261, 361) is a response to the detection of the indication, and the intermediate storage of the input signal (232, 332) takes place at least until the audio data (261) is received; by using (612) the audio data (261, 361), removing at least one portion of the one or more interference components from the input signal (232, 332) to generate an output signal (282, 382); and Providing (616) the output signal (282, 382) as a calculation of the voice command component for speech recognition (290).
Need to check novelty before this filing date? Find Prior Art

Description

RELATED REGISTRATIONS

[0001] This application is an international application of US patent application No. 15 / 716,173, filed on September 26, 2017, claiming the priority benefit of provisional US patent application No. 62 / 515,712, filed on June 6, 2017, all of which are hereby incorporated by reference in their entirety. AREA OF INVENTION

[0002] The subject matter relates to the field of connectivity solutions. More specifically, but not limited to, the subject matter discloses techniques for enabling audio pattern recognition. In this sense, the present invention relates to a method according to claim 1 and a voice-controlled node according to claim 10. Advantageous embodiments may include features of dependent claims. STATE OF THE ART

[0003] Audio pattern recognition typically involves an audio processing device receiving a predefined audio pattern (e.g., via a microphone) and performing audio pattern recognition, either locally or remotely, to assign a corresponding meaning to the predefined audio pattern. The environment of the audio processing device may include not only the source of the predefined audio pattern but also sources of interfering audio input. If the interfering audio input is received as sound waves at the microphone of the audio processing device, it mixes with the sound waves of the predefined audio pattern, making pattern recognition a technical challenge. Furthermore, the disclosures in US 9,472,203 B1, US 9,653,060 B1, and DE 11 2014 000 709 T5 may be helpful for understanding the present invention.US patent 9,472,203 B1 discloses an acoustic echo cancellation (AEC) system that detects and compensates for differences in sampling rates between the AEC system and a range of wireless loudspeakers using a search-based trial-and-error technique. The system individually determines a frequency offset for each microphone-speaker pair using an iterative process, calculating an echo return loss enhancement (ERLE) value for each offset tested and selecting the frequency offset with the highest ERLE value.

[0004] US Patent 9,653,060 B1 relates to an echo cancellation system that uses a combined reference signal consisting of a playback reference signal and an adaptive reference signal. The playback reference signal is generated from a playback signal sent to a loudspeaker, and the adaptive reference signal is generated using beamforming at microphone inputs corresponding to the audio received by the loudspeaker. The system applies a low-pass filter to the playback reference signal and a high-pass filter to the adaptive reference signal to create the combined reference signal. The system can remove the combined reference signal from the target signals connected to the microphone inputs to isolate speech contained within the target signals.

[0005] Finally, DE 11 2014 000 709 T5 relates to a method for operating a speech trigger. In some embodiments, the method is carried out on an electronic device comprising one or more processors and a memory that stores instructions to be executed by the one or more processors. The method involves receiving a sound input. The sound input can correspond to a spoken word or sentence, or a part thereof. The method includes determining whether at least part of the sound input corresponds to a predefined type of sound, such as a human voice. Upon determining that at least part of the sound input corresponds to the predefined type, the method includes determining whether the sound input contains predefined content, such as a predefined trigger word or a predefined trigger phrase.The procedure also includes, upon determining that the sound input contains the specified content, initiating a voice-based service, such as a voice-based digital assistant. BRIEF DESCRIPTION OF THE FIGURES

[0006] Some embodiments are illustrated in the figures of the accompanying drawings as examples and without limitation, wherein: Fig. 1 is a block diagram illustrating an audio processing device that is communicatively coupled to other devices via one or more networks, according to various embodiments; Fig. 2 is a block diagram illustrating interactions of an audio processing device according to embodiments; Fig. 3 is a block diagram illustrating aspects of a reference generator, according to one embodiment; Fig. 4 is a block diagram illustrating the interactions of a loudspeaker system, according to embodiments; Fig. 5 a timing diagram illustrating the operating modes of a device according to embodiments; Fig. 6 is a flowchart illustrating a method for enabling audio pattern recognition, according to embodiments; Fig. 7 is a flowchart illustrating a method for providing a sound output and corresponding audio data, according to embodiments; Fig. 8 is a block diagram illustrating an IoT system comprising a voice-controlled hub (VCH), according to embodiments; and Fig. 9 is a block diagram illustrating an electronic device according to embodiments. DETAILED DESCRIPTION

[0007] This document describes systems and methods for audio pattern recognition. Numerous examples and embodiments are presented in the following description to provide a thorough understanding of the claimed subject matter. It will be obvious to a person skilled in the art that the claimed subject matter can be implemented in other embodiments. Some embodiments are briefly introduced and then discussed in more detail together with other embodiments, beginning with Fig. 1.

[0008] A connected device capable of perceiving audio elements can be used as an audio processing device to enable local and / or remote audio pattern recognition, according to the embodiments described herein. For example, a smartphone has interfaces for connecting to remote computing devices (e.g., via personal area networks (PANs), local area networks (LANs), the internet, and / or other network types) and has a microphone for perceiving audio elements (e.g., voice commands). Another example of an audio processing device is a voice-controlled node (VCH) that can control devices over a network (e.g., Bluetooth or Wi-Fi) in response to voice commands.The VCH can be configured to control devices that include, without limitation, household appliances, thermostats, lighting, media devices, and / or applications running on a computing device. In some embodiments, the VCH receives voice commands from a user via microphone, enables voice recognition, and instructs a connected device to perform a corresponding action (e.g., turning on a light switch, playing music, changing a television channel, or performing an internet search). Enabling voice recognition can involve either local speech recognition or speech recognition by a remote speech recognition application. As such, the VCH can free up resources on connected devices by eliminating the need for them to perform speech recognition.

[0009] For audio processing devices that enable local or remote audio pattern recognition, interfering audio elements from sources in the environment (e.g., televisions, speakers, or other ambient noise) are typically not helpful when they mix with the audio pattern to be recognized. Indeed, operations by a speech recognition application (e.g., automatic speech recognition (ASR)) to resolve the audio pattern of a mixed signal can require more processing power and time, and may be less consistent and less accurate, compared to the case where the speech recognition application is presented with only the audio pattern.

[0010] In some embodiments, interfering audio elements emitted by an ambient audio source to the microphone of an audio processing device are also provided to the audio processing device as audio data via a network (e.g., by the ambient audio source). The audio processing device then uses this shared audio data to reduce the interfering audio elements that are mixed with the desired audio pattern presented to the microphone. As a result, the audio processing device can provide the speech recognition application with a calculation of the desired audio pattern.

[0011] Returning to the VCH example, a microphone array of the VCH can receive both a voice command from a human user and speaker tones generated by a remote speaker system. The microphone array then provides an audio input signal that includes a voice command component and a speaker tone component. The VCH also uses its RF transceiver to receive stream data corresponding to the speaker tones from the remote speaker system via one or more RF channels. The processing system can use a packet-loss masker to conceal packets of the stream data that are unusable.

[0012] In embodiments, a VCH processing system uses noise reduction to distinguish between the voice command component of the audio input signal and the loudspeaker component of the audio input signal. In one embodiment, an adaptive filter can calculate the loudspeaker component by using the corresponding stream data and compare the calculated stream data with the audio input signal (e.g., mathematically) to generate a calculated voice command component.

[0013] In embodiments, the processing system uses a VCH storage system to buffer the audio input signal, allowing the VCH to synchronize the timing of the calculated stream data with the timing of the audio input signal. This synchronization enables the adaptive filter to accurately compare the calculated stream data with the audio input signal when generating the calculated voice command component. In embodiments, the VCH can control when to receive the stream data, depending on various conditions, including but not limited to voice activity, network congestion, power consumption, and / or transmission time. For example, the processing system can first detect an indication of VCH activity (e.g., voice activation) and, in response, request and receive the stream data from the loudspeaker system.After the processing system detects an indication of inactivity (e.g., a lack of voice input), the VCH can request that the loudspeaker system stop transmitting the stream data.

[0014] The embodiments described herein can enhance the user experience by accelerating the voice recognition process, increasing recognition rates (e.g., understanding a larger percentage of voice commands), and improving recognition accuracy, while still allowing users to continue using their connected media devices. Compared with previous techniques, these embodiments can reduce and / or offload the power and resource consumption associated with speech recognition by removing interfering audio signals from the input to the speech recognition application. These and / or similar performance improvements extend to any connected device that uses the systems and methods described herein to enable audio pattern recognition. For example, any device (e.g.,Detect media devices connected in a peer-to-peer (P2P) network and / or request the transmission of audio elements from other devices that are currently playing interfering audio elements. These and other embodiments are described in more detail herein.

[0015] The detailed description below includes cross-references to the accompanying drawings, which form part of the detailed description. The drawings show illustrations of embodiments. These embodiments, also referred to herein as "examples," are described in sufficient detail to allow those skilled in the art to implement embodiments of the claimed subject matter. The embodiments can be combined, other embodiments can be used, or structural, logical, and electrical modifications can be made without departing from the claimed scope. The following detailed description is therefore not to be interpreted in a limiting sense, and its scope is defined by the attached claims and their equivalents.

[0016] Fig. Figure 1 is a block diagram 100 illustrating an audio processing device 102 that is communicatively coupled to other devices via one or more networks 114, according to various embodiments. The audio processing device 102 serves to enable audio pattern recognition and can control a device or application based on a recognized audio pattern. As shown, the audio processing device 102 receives sound waves 105 from the audio pattern source 104 and sound waves 107 and 109 from the audio interference sources 106 and 108, respectively. The audio processing device can itself also emit audio interference (not shown), for example, via loudspeakers.

[0017] As also shown, the audio processing device 102 interacts with one or more networks 114 via one or more communication links. To enable pattern recognition, the audio processing device 102 provides noise reduction to remove some or all of the audio interference by using appropriate audio data that it receives from the audio interference sources 106 and 108 via the network(s) 114 or that is generated internally. In one embodiment, the noise reduction can be implemented by using ICA (Independent Component Analysis), in which incoming signals (e.g., from a microphone) are separated according to source (e.g., signals from the audio pattern source and the audio interference sources), and then the audio data is compared with the separated signals to determine which should be removed to obtain a calculated audio pattern.In other embodiments, the noise suppression can employ adaptive filters, neural networks, or other techniques known in the prior art that can be used to attenuate non-target components of a signal. In some embodiments, the audio processing device 102 can be coupled to a controlled device 103 (e.g., a local device or application) that it can control based on detected audio patterns.

[0018] The audio pattern source 104 serves to provide the sound waves 105 corresponding to a recognizable audio pattern. In some embodiments, the audio pattern source 104 can interact with the network(s) 114 via one or more communication links. In embodiments, an audio pattern is a predefined audio pattern and / or an audio pattern that can be recognized by a pattern recognition application associated with the audio processing device 102. The audio pattern source 104 can be an animate entity (e.g., a human being) or an inanimate object (e.g., a machine).

[0019] The audio interference sources 106 and 108 are sources of sound waves 107 and 109, respectively, which interfere with the detection of the audio patterns corresponding to the sound waves 105. As shown, the audio interference sources 106 and 108 interact with the network(s) 114 via one or more communication links. The audio interference sources 106 and 108 can provide the audio processing device 102 with audio data corresponding to the audio interference via the network(s) 114. Audio interference sources can include loudspeakers, televisions, video games, industrial noise sources, or any other noise source whose audio output is digitized or can be digitized and which is provided to the audio processing device 102 via the network(s) 114.

[0020] As shown, the controlled device 110 is coupled to the network(s) 114 via the link(s). The controlled devices 110 and 103 can comprise any device with a function that can be initiated in response to audio pattern recognition enabled by the audio processing device 102. Example controlled devices include household appliances, thermostats, lighting, automated blinds, automated door locks, automotive controls, windows, industrial controls, and actuators. As used herein, controlled devices can comprise any logic, firmware, or software application running on the controlled device 110.

[0021] A pattern recognition application 112 is operated to recognize audio patterns and associate the recognized audio patterns with a corresponding meaning. The pattern recognition application can reside on one or more computing devices that are coupled to the network(s) 114 via the link(s) and use or be implemented using processors, memory, circuits, arithmetic logic, software, algorithms, and data structures to organize and process attributes of an audible sound, including pitch, volume, tone, repetitive or rhythmic sounds, and / or spoken sounds such as words, sentences, and the like. In one embodiment, the pattern recognition application 112 comprises an ASR that identifies predefined audio patterns and associates them with each other (e.g., by using a data structure) and / or with a corresponding meaning.The pattern recognition application can, for example, enable music recognition, song recognition, voice recognition and speech recognition by recognizing 112 audio patterns, but this is not limited to.

[0022] The network(s) 114 can comprise one or more types of wired and / or wireless networks to connect the network nodes from Fig. 1. to communicatively couple with one another. For example, and without limitation, the network(s) may include a wireless local area network (WLAN) (e.g., Wi-Fi, compliant with 802.11), PANs (e.g., Bluetooth SIG standard or Zigbee, compliant with IEEE 802.15.4), and the Internet. In one embodiment, the audio processing device 102 is communicatively coupled with the pattern recognition application 112 via Wi-Fi and the Internet. The audio processing device 102 may be communicatively coupled with the audio interference sources 106 and 108 and with the controlled device 110 via Bluetooth and / or Wi-Fi.

[0023] Fig. Figure 2 is a block diagram illustrating the interactions of an audio processing device 202 according to embodiments. As shown, the audio processing device 202 comprises various interactive functional blocks. Each functional block can be implemented using hardware (e.g., circuits), instructions (e.g., software and / or firmware), or a combination of hardware and instructions.

[0024] The microphone array 220 is designed to receive the sound waves 205 of a voice command from a human speaker 204 and / or the interfering sound waves 207 and 209 emitted by the loudspeaker systems 206 and 208, respectively. Each microphone of the microphone array 220 includes a mechanism (e.g., comprising a diaphragm) to convert the energy of sound waves into an electronic signal. When the sound waves 205, 207, and 209 are received during a common period, the electronic signal includes components corresponding to the voice command and the audio interference. In some embodiments, one or more microphones of the array may be a digital microphone.

[0025] An audio input processor 230 comprises circuitry for processing and analyzing the electronic audio signal received by the microphone array 220. In embodiments, the audio input processor 230 provides analog-to-digital conversion to digitize the electronic audio signals. After digitization, the audio input processor 230 can provide signal processing (e.g., demodulation, mixing, filtering) to analyze or manipulate attributes of the audio inputs (e.g., phase, wavelength, frequency). The audio input processor 230 can isolate audio components (e.g., by using beamforming) or determine distance and / or position information associated with one or more audio sources (e.g., by using techniques such as TDOA (Time Distance Of Arrival) and / or AoA (Angle of Arrival)).In some embodiments, the audio input processor 230 can provide noise suppression (e.g., of background noise) and / or calculation and suppression of noise components by using one or more adaptive filters. After such additional processing, if present, is completed, the audio input processor 230 can provide a resulting input signal 232 (e.g., an input signal for each microphone of the microphone array 220) to the combiner 280 (e.g., via the buffer 240).

[0026] An activation detection 250 serves to detect an indication to initiate pattern recognition (e.g., an active mode) and / or to detect an indication to stop pattern recognition (e.g., an inactive mode). Enabling pattern recognition may involve receiving audio data from an interference source. In some embodiments, an activity, inactivity, behavior, or state of a person, device, or network may provide the indication. For example, a user may provide the indication via touch input at an interface (not shown) (e.g., touchscreen or mechanical button) or by making a predetermined sound, gesture, gaze, or eye contact.The activation detection 250 can detect the indication by recognizing a keyword spoken by a user or by recognizing the user's voice. In some embodiments, the indication can be a timeout indicating the absence of voice activity, user activity, or device activity. The indication can be provided by hardware or software based on a timer, power consumption, network states (e.g., congestion), or any other internal or external device or network states. In embodiments, the activation detection 250 can report the detection to other components and / or remote devices and / or initiate a mode change in response to the indication detection.

[0027] An RF / packet processor 260 serves to provide one or more network interfaces, process network packets, implement codecs, (de)compress audio elements, and / or provide any other analog or digital packet processing. In one embodiment, the RF / packet processor 260 includes an RF transceiver for wirelessly transmitting and receiving packets. In other embodiments, the RF / packet processor 260 communicates with the loudspeaker system 206, the loudspeaker system 208, and a controlled device 210, respectively, via links 215, 216, and 217.

[0028] For example, the RF / packet processing unit can receive audio data from the loudspeaker systems 206 and 208 corresponding to the interfering sound waves 207 and 209 received by the microphone array 220. In embodiments, the RF / packet processing unit 260 can provide packet loss masking to calculate and mask portions of missing audio data due to lost or corrupted packets. The RF / packet processing unit 260 can also include circuitry to analyze and process attributes of RF signals (e.g., phase, frequency, amplitude, signal strength) to calculate the location and / or distance of transmitting devices relative to the audio processing device 202. The RF / packet processing unit 260 can examine packet contents to determine properties, status, and capabilities associated with transmitting devices, such as timestamps, location, loudspeaker specifications, audio quality, and volume.In embodiments, the RF / packet processing 260 provides the reference generator 270 with received audio data 261 with any associated information for processing.

[0029] The audio processing device from Fig. Figure 2 is an embodiment that provides noise suppression by using the reference generator 270 and the combiner 280. Other noise suppression techniques known in the prior art can be used without departing from the inventive subject matter. In this embodiment, noise is suppressed at the input signal 232 (e.g., a single channel) generated by the audio input processor 230, which comprises an input combination from each microphone in the array. In other embodiments, the noise suppression can be provided prior to the generation of the input signal 232 for each microphone signal between the microphone array 220 and the audio input processor 230.

[0030] The reference generator 270 serves to generate a reference signal 272 based on the audio data 261, which is a calculation of one or more interference components of the input signal 232. In one embodiment, the reference signal 272 is a combination of audio data 261 from different interference sources. To provide a calculation of interference components, the reference generator 270 can use information provided by the audio input processor 230 and / or the RF / packet processor 260. For example, the reference generator 270 can use attributes of the input signal, the distance and / or angle of an interference source relative to the audio processing device 202, and / or attributes of the audio data 261 to calculate the interference components resulting from the interfering sound waves 207 and 209 at the microphone array 220.In embodiments, the reference generator 270 can use adaptive filtering, which is implemented in firmware on a digital signal processor (DSP), an application processor, in dedicated hardware, or a combination of both (e.g., hardware accelerators). One embodiment of the reference generator 270 is described with reference to... Fig. 3 described.

[0031] The combiner 280 serves to generate an output signal 282 based on a comparison of the reference signal 272 with the input signal 232. In embodiments, the buffer 240, which can temporarily store the input signal 232, is used to synchronize the timing of the input signal 232 with the timing of the reference signal 272 before combining. An exemplary buffer timing is described with reference to Fig. 5 described in more detail.

[0032] In one embodiment, the combiner 280 removes all or part of the interference components from the input signal 232 to produce an output signal 282 that represents a calculation of the voice command component of the input signal 232. For example, the output signal 282 can be a difference between the input signal 232 and the reference signal 272, calculated by the combiner 280. Alternatively or additionally, the combiner 280 can use addition, multiplication, division, or other mathematical operations and / or algorithms, alone or in combination, to generate the output signal 282.

[0033] The combiner 280 can provide the output signal 282 to the pattern recognition application 290. The pattern recognition application 290 does not need to be located within the audio processing device 202. In other embodiments, all or part of the output signal 282 can be transmitted by the RF / packet processor 260 to a remote audio processing application (located, for example, on one or more computing devices connected to the internet) for voice recognition purposes. The audio processing device 202 can then transmit the voice command via the link 217 to the controlled device 210 to complete the corresponding action.

[0034] Fig. Figure 3 is a block diagram illustrating aspects of a reference generator 300 according to one embodiment. In one embodiment, the adder 369 receives audio data 318 from a number, I, of audio interference sources n1(t), n2(t)...n1(t) and combines the audio data 318 to generate the audio data 361, N(t). The adder 369 can be configured in the RF / packet processing 260 from Fig. 2 and / or in the reference generator 270 from Fig. 2. The adaptive filter 371 receives the audio data 361, N(t), and outputs the reference signal 372 to the subtractor 383. The input signal 332, y(t) is also provided to the subtractor 383. The subtractor 383 can be located in the reference generator 270 and / or the combiner 280. Fig. 2. The input signal 332, y(t), can include a desired audio pattern component, x(t), and / or the audio interference components, N(t)', after it has undergone modification during sound wave propagation from the audio interference sources. Thus, the input signal 332, y(t), can be expressed as follows: y(t)=x(t)+N(t)'

[0035] In one embodiment, the audio data 361, N(t), can represent the audio interference components without having undergone any modification during sound wave propagation. The adaptive filter 371 calculates the modified audio interference components N(t)' by performing a convolution operation (e.g., represented by “*”) on the audio data 361, N(t), and a calculated impulse response h(t). For example, the adaptive filter 371 can derive the calculated impulse response, h(t), from the propagation through a loudspeaker, through a room, and to the microphone. The modified audio interference components N(t)' can be expressed as follows: N(t)'=N(t)*h_(t)

[0036] Thus, the output signal 282, x(t), from the subtractor 383, is a calculation of the desired audio pattern component, x(t), which is given by: x_(t)=x(t)+N(t)'−N(t)*h_(t)

[0037] In some embodiments, a processing logic of the audio processing device 202 (not shown) can evaluate the output signal by using any prior art voice activation detection algorithm (e.g., by using neural network techniques) to determine whether the input signal includes a useful component from a target source of the specified audio pattern. Depending on the determination, the processing logic can initiate or suppress pattern recognition. For example, if the computation of the desired audio pattern, x(t), is less than a specified threshold, the processing logic can determine that the input signal does not include a useful component from the target source.If the calculated desired audio pattern, x(t), reaches or exceeds the predefined threshold, the processing logic can determine that the input signal contains a useful component from the target source. In some embodiments, the predefined threshold can be selected and stored in memory during a manufacturing process and / or dynamically determined and stored in memory during runtime based on operating conditions.

[0038] In one embodiment, operations of the adder 369, the adaptive filter 371, and / or the subtractor 383 are implemented in the frequency domain. Any adaptive filter technique known in the prior art, including the NLMS algorithm (NLMS: Normalized Least Mean Squares), can be used.

[0039] Fig. Figure 4 is a block diagram illustrating the interactions of a loudspeaker system 206 according to embodiments. As shown, the loudspeaker system 206 comprises various interactive functional blocks. Each functional block can be implemented using hardware (e.g., circuits), instructions (e.g., software and / or firmware), or a combination of hardware and instructions.

[0040] In one embodiment, the loudspeaker system 206 provides audio playback via the loudspeaker 437 and supplies corresponding audio data 436 to the audio processing device 202 via the RF / packet processing unit 460 and the link 217. The RF / packet processing unit 460 provides a communication interface to one or more networks and can be the same as that described in relation to Fig. The audio data generator 435, as described in Section 2, or a similar device, forwards the audio data 436 for output by the loudspeaker 437 and for transmission to the audio processing device 202. The audio data generator 435 may be located in an audio forwarding block of the loudspeaker system 206. In one embodiment, the loudspeaker system 206 transmits the audio data 436 to the audio processing device 202 whenever the loudspeaker system 206 reproduces audio elements through the loudspeaker 437. In some embodiments, the loudspeaker system 206 transmits the audio data 436 based on a request or other indication from the audio processing device 202, which operates in a mode to enable pattern recognition.

[0041] The activation detection 450 serves to detect indications of a mode in which the audio data 436 should be transmitted, or of a mode in which the transmission of the audio data 436 to the audio processing device 202 should be stopped. In various embodiments, the indications can be the same as those relating to the activation detection 250. Fig. 2 described or similar. The buffer 240 is used to store the audio data 436 in case the audio processing device 202 requests the audio data 436 corresponding to the audio interference already played from the loudspeaker 437. Exemplary buffer requirements associated with operating modes are given below with reference to Fig. 5 described.

[0042] Fig. Figure 5 is a timing diagram 500 illustrating the operating modes of a device according to embodiments. At T1 502, the audio processing device 202 receives sound waves at its microphones. The audio processing device 202 begins to process the input signal (e.g., the input signal 232 from Fig. 2) to generate and buffer the input signal so that it can be used to enable pattern recognition when the audio processing device 202 enters an active mode. Also at T1 502, the loudspeaker system 206 outputs audio elements via its loudspeaker 437 and begins to process the corresponding audio data (e.g., the audio data 436 from Fig. 4) to temporarily store.

[0043] At T2 504, the audio processing device 202 detects an indication to activate before requesting audio data from the interference source at T3 506 in response. In one embodiment, the audio processing device 202 can specify a start time in the request from which the requested audio data should begin. At T4 508, the loudspeaker system 206 begins transmitting the audio data it has cached. For example, although the loudspeaker system 206 may have cached data since T1 502, it can only transmit the audio data cached since the activation detection at T2 504. According to T4 508, in some embodiments, the loudspeaker system 206 can stream the audio data without caching.In the T5 510, the audio processing device 202 receives the audio data and can continue to buffer the input signal until the reference signal has been generated and is ready to be combined with the buffered input signal.

[0044] In one embodiment, the input signal, which has been buffered from approximately T2 504 (e.g., activation) to approximately T5 510 (audio data received), is used for the transition to the active mode. After using this initial buffering, the audio processing device 202 can continue buffering as needed (not shown) to assist in synchronizing the input signal with the reference signal.

[0045] At T6 512, the audio processing device 202 is deactivated before, in response, it requests at T7 514 that the interference source stop transmitting the audio data. In preparation for reactivation, the audio processing device 202 at T6 512 can resume buffering, and the loudspeaker system 206 at T8 516 can stop transmitting audio data and resume buffering.

[0046] Fig. Figure 6 is a flowchart illustrating a method 600 for enabling audio pattern recognition according to embodiments. The method 600 can be performed by processing logic that includes hardware (circuits, dedicated logic, etc.), software (such as that running on a general-purpose computer or a dedicated machine), firmware (embedded software), or a combination thereof. In various embodiments, the method 600 can be performed by the audio processing device consisting of Fig. 2 will be carried out.

[0047] For example, the audio input processor 230 provides an input signal 232 to the buffer 240 in block 602. This input signal includes a predefined or otherwise identifiable audio pattern (e.g., a voice command component) and audio interference (e.g., an interference component). The buffer 240 temporarily stores the input signal 232 before making it available to the combiner 280 in block 612 for use. The buffer 240 can process the input signal 232 as described above with respect to Fig. 5 described storage. In some embodiments, the buffer stores the input signal 232 for a period of time long enough to support synchronization of the timing of the input signal with the timing of the reference signal, such that at least a portion of the buffered input signal 232 is used when generating the output signal 282.

[0048] In some embodiments, the audio processing device 202, or another device, can control whether the audio processing device 202 receives audio data from interference sources. An interference source, or another device, can also control whether the interference source transmits audio data to the audio processing device 202. The audio processing device 202 can receive audio data in an active mode and not receive audio data in an inactive mode. For example, the transmission of audio data from interference sources can be controlled based on levels of network congestion and / or network activity. Some embodiments include blocks 606 and 618 to detect a mode change, and blocks 608 and 620 to start or stop receiving audio data in response to the detection of a mode change.

[0049] For example, the activation detection 250 in block 606 can determine whether the audio processing device 202 has received an indication to enter a mode of active preparation and provision of the output signal 282 for audio pattern recognition (e.g., an active mode) or to remain in a current operating mode (e.g., an inactive mode). If the activation detection 250 determines that operation should continue in an inactive mode, the buffer 240 in block 604 continues to buffer the input signal 232. If the activation detection 250 determines that operation should switch to the active mode, the activation detection 250 in block 608 causes the RF / packet processor 260 to request audio data corresponding to the audio interference.

[0050] In block 610, the RF / packet processor 260 receives the audio data corresponding to the audio interference and transmits it to the reference generator 270. In embodiments, the audio data can be requested and received from any source of audio interference components that is communicatively coupled to the audio processing device 202.

[0051] In block 612, the reference generator 270 and the combiner 280 use the audio data to remove interference components from the input signal and generate an output signal. For example, based on the audio data 261, the reference generator generates a reference signal 272 and provides the combiner 280 with the reference signal 272. As with reference to Fig. As described in section 3, the adder 369 can combine audio data from multiple interference sources. The combiner 280 from Fig. 2 generates an output signal 282 based on the reference signal 272 and the input signal 232. For example, the subtractor 383 generates from Fig. 3. The output signal 382 is generated by subtracting the reference signal 372 from the input signal 332. In block 616, the combiner 280 provides the output signal 282 for audio pattern recognition.

[0052] In some embodiments, the processing logic of the audio processing device 202 can provide voice activity detection to determine, based on the value of the resulting output signal 282, whether the input signal 232 includes an audio pattern component to be detected. For example, the processing logic can distinguish whether the input signal includes a voice command or just audio input from a connected device by using any suitable voice activity algorithm known in the prior art (e.g., using neural network techniques).For example, if a value of the output signal 282 is below a predefined threshold, the processing logic can determine that the values ​​of the input signal 232 are the same as or similar to the values ​​of the reference signal 272, and / or that the input signal 232 contains little or no voice command component from a user, and therefore, there is no need to proceed to a voice recognition application. If the output signal 382 reaches or exceeds the predefined threshold, the processing logic can determine that the values ​​of the input signal 232 are not the same as or similar to the values ​​of the reference signal 272, and / or that the input signal 232 contains a sufficient amount of voice command component from the user, and therefore, a proceed to the voice recognition application is necessary.

[0053] If the activation detection 250 in block 618 determines that the operation should switch from active mode to inactive mode, the activation detection 250 in block 620 causes the audio processing device 202 to stop receiving the audio data corresponding to the audio interference. In some embodiments, the activation detection 250 can cause the RF / packet processor 260 to request that the loudspeaker systems 206 and 208 stop transmitting audio data. Alternatively or additionally, the activation detection 250 can cause the RF / packet processor 260 to block or reject any incoming audio data.

[0054] Fig. Figure 7 is a flowchart illustrating a method 700 for providing sound output and corresponding audio data according to one embodiment. The method 700 can be performed by processing logic that includes hardware (circuits, dedicated logic, etc.), software (such as that running on a general-purpose computer or a dedicated machine), firmware (embedded software), or a combination thereof. In various embodiments, the method 700 can be performed by the audio processing device consisting of Fig. 4 will be carried out.

[0055] In block 702, the audio data generator generates the audio data 436 and makes it available to the loudspeaker 437 for playback in block 704. Simultaneously, but depending on the current operating mode, the audio data generator 435 can either make the audio data 436 available to the buffer 440 in block 706 (e.g., in an inactive mode) or bypass the buffer and make the audio data 436 available to the audio processing device 202 in block 710 (e.g., in an active mode).

[0056] If the audio data is buffered in block 706, the process can proceed to block 708, where the activation detection 450 determines whether the loudspeaker system 206 has received an indication that the audio processing device 202 has entered an active mode or that it is maintaining operation in an inactive mode. In some embodiments, the operating mode change indicator is a request or instruction from the audio processing device 202. If the activation detection 450 detects an inactive mode, the buffer 440 in block 606 continues to buffer the audio data. If the activation detection 450 detects a change from the inactive mode to the active mode, the process proceeds to block 710, where the loudspeaker system 206 provides the audio data 436 to the audio processing device 202.

[0057] In block 712, the activation detection 450 monitors whether a change from an active to an inactive mode has occurred. If no mode change has occurred, the loudspeaker system 206 continues to provide the audio data as described in block 710. If the mode changes to inactive, the loudspeaker system 206 stops providing the audio data 436 to the audio processing device 202 in block 710 and instead makes it available to the buffer 440 for temporary storage, as described in block 706.

[0058] Fig. Figure 8 is a block diagram illustrating an IoT system 800 comprising a VCH 802, according to embodiments. In one embodiment, smart devices in a room or zone are controlled by the VCH 802, which is connected to the Cloud-ASR 812. The VCH 802 is positioned in a zone to listen to everything in that zone, to detect and interpret commands, and to perform the requested actions. For example, the VCH 802 can control a connected light switch 811, thermostat 810, television 808, and speaker 806 within the zone.

[0059] By centrally positioning voice control in the VCH 802, connected devices can be controlled via a control link through the VCH 802. This can eliminate the need to implement a voice interface on each connected device and can result in significant cost savings for device manufacturers in terms of both hardware capabilities and development effort. The VCH 802 has internet access and can perform cloud-based ASR instead of local ASR, providing access to advanced ASR capabilities and processing power, and improving performance and user experience.

[0060] An embodiment implementing the audio output from the television 808 and the loudspeakers 806, which is received by the VCH 802, in order to avoid the interference with voice control operation, the reduction in speech recognition rates, and disadvantages to the user experience suffered by previous technologies. For example, suppose the human target speaker 804 utters the command "VCH, what time is it?" to the VCH 802. At the same time, the loudspeakers 806 are playing a song with the lyrics "Come as you are," while the television 808 is showing the Super Bowl with the audio element "Touchdown Patriots." The microphone 802.1 in the VCH 802 captures all audio signals and passes a mixed audio signal to the acoustic echo cancellation (AEC) block 802.2.

[0061] The AEC 802.2 can prevent or remove echoes or any unwanted signal from an audio input signal. Echo suppression can involve detecting the reappearance of an originally transmitted signal, which reappears with a delay in the transmitted or received signal. Once the echo is detected, it can be removed by subtracting it from the transmitted or received signal. This technique is typically implemented digitally using a digital signal processor or software, although it can also be implemented in analog circuits.

[0062] The speakers 806 also transmit "Come as you are" to the VCH 802 via Bluetooth or Wi-Fi audio links, and the television 808 transmits "Touchdown Patriots" to the VCH 802 via Bluetooth or Wi-Fi audio links. The transfer of audio elements over wireless links will lead to bit errors and packet loss; therefore, a packet loss cover (PLC) algorithm 802.4 can be applied when the signals are received at the VCH 802 and before they are used as a reference by the AEC 802.2. In one embodiment, the audio signal driving the VCH speaker 802.5 can also be used as a reference by the AEC 802.2.

[0063] These signals are used as reference input for the AEC 802.2, which subtracts their calculation at 802.3 from the mixed signal captured by the microphone 802.1. The result is the interference-free target signal "VCH, what time is it?" This signal is passed to the Cloud-ASR 812 and allows for accurate recognition, resulting in the response from the VCH speaker 802.5, "It is 6:30."

[0064] In some embodiments, connected devices transmit audio elements to the VCH 802 whenever audio is played, regardless of whether a user is issuing commands. Depending on the wireless environment and the VCH 802's wireless capabilities, this can cause network congestion. If the continuous transmission of audio elements via the audio links is a concern due to power consumption, congestion, or other reasons, the VCH 802 can signal to the connected devices via a control link when to transmit, based on the detection of a speech signal or keyword phrase.

[0065] The connected devices can store a certain amount of played audio in a buffer. When the VCH 802 detects active speech from a user, it uses the control link to inform the connected devices to begin audio transmission. Alternatively or additionally, the VCH 802 can use the control link to mute or reduce the volume on the connected devices playing audio when the VCH 802 detects a voice and / or keyword phrase. The VCH 802 can also use a buffer of the audio received from the 802.3 microphone, allowing it to properly synchronize the AEC 802.2 with the detected audio.In one embodiment, the requirement for intermediate storage is based on the delay in voice activity detection and the time required for signaling to the remote devices and for the start of receiving the reference signals.

[0066] The embodiments described herein can be applied to a peer-to-peer (P2P) connected scenario where device A wants to interact with a user and perform voice recognition, but device B is playing audio. Device A can issue a request on the control channel to determine if any connected devices are currently playing audio. In response to the request, an audio link is established, and device B sends its audio to device A. Device A can then suppress the audio from device B, which it detects as interference, when the user attempts to use voice recognition.

[0067] Fig. Figure 9 is a block diagram illustrating an electronic device 900 according to embodiments. The electronic device 900 can be the embodiments of the audio processing device 102, the audio pattern source 104, the audio interference sources 106 and 108, the controlled devices 103 and 110, and / or the pattern recognition application 112. Fig.1. comprise and / or operate the electronic device 900, in whole or in part. The electronic device 900 may be in the form of a computer system within which sets of instructions can be executed to cause the electronic device 900 to perform any one or more of the methodologies discussed herein. The electronic device 900 may be operated as a standalone device or be connected (e.g., networked) to other machines. In networked use, the electronic device 900 may be operated in the function of a server or a client machine in a server-client network environment, or as a peer machine in a P2P (or distributed) network environment.

[0068] The electronic device 900 can be an IoT device (iT = Internet of Things), a server computer, a client computer, a personal computer (PC), a tablet, a set-top box (STB), a VCH, a personal digital assistant (PDA), a mobile phone, a web appliance, a network router, a network switch or network bridge, a television, a speaker, a remote control, a screen, a portable multimedia device, a portable video player, a portable gaming device or a control panel, or any other machine capable of executing (sequentially or otherwise) a set of instructions specifying actions to be taken by that machine.Furthermore, although only a single electronic device 900 is illustrated, the term “device” is also to be understood as encompassing any collection of machines which, individually or collectively, execute a set (or several sets) of instructions to carry out any one or more of the methodologies discussed herein.

[0069] As shown, the electronic device 900 comprises one or more processors 902. In embodiments, the electronic device 900 and / or the one or more processors 902 may comprise one or more processing devices 905, such as a system-on-a-chip processing device developed by Cypress Semiconductor Corporation, San Jose, California. Alternatively, the electronic device 900 may comprise one or more other processing devices known to those skilled in the art, such as a microprocessor or central processing unit, an application processor, a host controller, a controller, a special-purpose processor, a DSP, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or the like.A 901 bus system can include a communication block (not shown) to communicate with an internal or external component, such as an embedded controller or an application processor, via one or more communication interfaces 909 and / or the 901 bus system.

[0070] Components of the electronic device 900 can be located on a common substrate, such as a die substrate with an integrated circuit (IC), a multi-chip module substrate, or the like. Alternatively, components of the electronic device 900 can be one or more separate integrated circuits and / or discrete components.

[0071] The memory system 904 can include volatile and / or non-volatile memory that can communicate with each other via the bus system 901. For example, the memory system 904 can include main memory (RAM) and program flash. The RAM can be static RAM (SRAM) and the program flash can be non-volatile memory, which can be used to store firmware (e.g., control algorithms that can be executed by the processor(s) 902 to implement operations described herein). The memory system 904 can include instructions 903 that, when executed, perform the procedures described herein. Sections of the memory system 904 can be dynamically allocated to provide caching, buffering, and / or other memory-based functionalities.

[0072] The storage system 904 may include a drive unit providing a machine-readable medium on which one or more sets of instructions 903 (e.g., software) embodying any one or more of the methodologies or functions described herein may be stored. During their execution by the electronic device 900, which in some embodiments constitutes machine-readable media, the instructions 903 may also reside wholly or at least partially within the other storage devices of the storage system 904 and / or within the processor(s) 902. Furthermore, the instructions 903 may be transmitted or received over a network via the communication interface(s) 909.

[0073] While in some embodiments a machine-readable medium is a single medium, the term "machine-readable medium" is to be understood as encompassing a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) that store the one or more sets of instructions. The term "machine-readable medium" is also to be understood as encompassing any medium capable of storing or encoding a set of instructions for execution by the machine, causing the machine to perform any one or more of the exemplary operations described herein. Accordingly, the term "machine-readable medium" is to be understood as encompassing, but not limited to, solid-state storage and optical and magnetic media.

[0074] As further shown, the electronic device 900 comprises one or more display interfaces 906 (e.g., a liquid crystal display (LCD), a touchscreen, a cathode ray tube (CRT), and software and hardware support for display technologies), one or more audio interfaces 908 (e.g., microphones, loudspeakers, and software and hardware support for microphone input / output and loudspeaker input / output). As also shown, the electronic device 900 comprises one or more user interfaces 910 (e.g., keyboard, buttons, switches, touchpad, touchscreens, and software and hardware support for user interfaces).

[0075] The above description is to be understood as illustrative and not limiting. For example, the embodiments described above (or one or more aspects thereof) may be used in combination with one another. Other embodiments will be clear to those skilled in the art after reviewing the above description. In this document, as is customary in patent documents, the terms "a" or "an" are used to encompass one or more than one. In this document, the term "or" is used to mean a non-exclusive "or," such that "A or B" includes "A but not B," "B but not A," and "A and B," unless otherwise specified.In the event of inconsistent usage between this document and the documents incorporated by reference, the usage in the incorporated reference(s) should be considered as complementary to that of this document; in the event of incompatible usage, the usage in this document shall prevail over the usage in any incorporated reference(s).

[0076] Although the claimed subject matter has been described with reference to specific embodiments, it is obvious that various modifications and changes to these embodiments can be made without departing from the broader spirit and scope of the claimed subject matter. Accordingly, the patent description and the drawings should be considered in an illustrative rather than a limiting sense. The scope of the claims should be determined by reference to the appended claims together with the full scope of equivalents to which such claims are entitled. In the appended claims, the terms "comprising" and "in which" are used as the plain language equivalents of the corresponding terms "including" and "whereby".Likewise, in the following claims, the terms "comprising" and "including" are extensible; if a system, device, article, or process includes elements in addition to those listed behind such a term in a claim, they are still to be considered as belonging to the scope of that claim. Furthermore, in the following claims, the terms "first," "second," and "third," etc., are used merely as distinguishing features and are not intended to impose any numerical requirements on their objects.

Claims

[1] A method (600) comprising the following: Providing (602) an input signal (232, 332) based on sound waves (205, 207, 209) received by one or more microphones (220), wherein the input signal (232, 332) includes a voice command component and one or more interference components; Intermediate storage (604) of the input signal (232, 332); detection (606) of an indication to receive audio data (261, 361); receipt (608) of the audio data (261, 361) via one or more computer networks (114), wherein the audio data (261, 361) corresponds to one or more interference components, the receipt of the audio data (261, 361) is a response to the detection of the indication, and the intermediate storage of the input signal (232, 332) takes place at least until the audio data (261) is received; by using (612) the audio data (261, 361), removing at least one portion of the one or more interference components from the input signal (232, 332) to generate an output signal (282, 382); and Providing (616) the output signal (282, 382) as a calculation of the voice command component for speech recognition (290). [2] Method (600) according to claim 1, wherein the use (612) of the audio data (261, 361) comprises generating a reference signal (272, 372) by combining first audio data of the audio data with second audio data of the audio data, wherein the first audio data correspond to a first interference component of one or more interference components and the second audio data correspond to a second interference component of one or more interference components, wherein removing at least the portion of one or more interference components from the input signal (232, 332) comprises subtracting the reference signal (272, 372) from the input signal (232, 332). [3] Method (600) according to claim 2, wherein removing at least the portion of one or more interference components from the input signal (232, 332) comprises combining at least one portion of the audio data (261, 361) with at least one portion of the input signal (232, 332). [4] Method / 600) according to claim 1, wherein removing at least the portion of one or more interference components from the input signal (232, 332) comprises comparing at least one portion of the audio data (261, 361) with at least one portion of the input signal (232, 332). [5] Method (600) according to claim 1, further comprising using at least a portion of the cached input signal (232, 332) to synchronize a timing of the audio data (261, 361) with a timing of the input signal (232, 332). [6] Method (600) according to claim 1, further comprising detecting (618) an indication to stop receiving audio data (261, 361), and in response thereto stopping (620) the reception of the audio data (261, 361), and temporarily storing (604) the input signal. [7] Method (600) according to claim 1, wherein receiving (610) the audio data (261, 361) via the one or more computer networks (114) comprises wirelessly receiving first audio data of the audio data via one or more radio frequency channels. [8] Method according to claim 1, wherein receiving (610) the audio data (261, 361) includes receiving first audio data of the audio data from a source of a first interference component of one or more interference components and receiving second audio data of the audio data from a source of a second interference component of one or more interference components. [9] Method (600) according to claim 1, wherein providing the output signal (282, 382) for speech recognition (290) includes transmitting the output signal to a speech recognition application which is communicatively coupled to a first computer network of one or more computer networks (114). [10] A voice-controlled node (802) that includes the following: a microphone array (801.2) configured to receive a voice command from a human user (804) and a loudspeaker tone generated by a remote loudspeaker system (806, 808) and to provide an audio input signal that includes a voice command component and a loudspeaker tone component; a high-frequency transceiver configured to receive stream data corresponding to the loudspeaker sound from the remote loudspeaker system (806, 808); and a processing system configured to use a noise suppressor (802.2) to distinguish between the voice command component of the audio input signal and the loudspeaker sound component of the audio input signal, wherein the noise suppressor (802.2) is configured to use the stream data to remove at least a portion of the loudspeaker sound component from the audio input signal to generate a computed voice command component, wherein The node (802) further includes a storage system, wherein the processing system is configured to use the storage system to buffer the audio input signal at least until the radio frequency transceiver receives the stream data from the remote speaker system (806, 808), wherein the processing system is configured to detect a voice activation of the voice-controlled node (802) and, in response, to use the high-frequency transceiver to request the stream data from the remote loudspeaker system (806, 808). [11] Voice-controlled node (802) according to claim 10, further comprising a packet loss cover (802.4), wherein the processing system is configured to use the packet loss cover (802.4) to cover packets of the stream data that are not usable for the noise suppressor (802.2). [12] Voice-controlled node (802) according to claim 10, wherein the processing system is configured to use the cached audio input signal to synchronize a timing of the stream data with a timing of the audio input signal. [13] Voice-controlled node (802) according to claim 10, wherein the high-frequency transceiver is configured to transmit the calculated voice command component to a remote automatic speech recognition application (812) and to receive voice command data corresponding to the calculated voice command, and wherein the processing system is configured to initiate a device function based on the voice command data.

Citation Information

Patent Citations

  • Audio Enhancement Architecture for the Smart Home

    US62515712P0

  • METHOD AND DEVICE FOR OPERATION OF A VOICE TRIGGER FOR A DIGITAL ASSISTANT

    DE112014000709T5

  • Clock synchronization for multichannel system

    US9472203B1

  • Hybrid reference signal for acoustic echo cancellation

    US9653060B1

  • US15716173B2