Enhanced noise reduction in voice activated device
Patent Information
- Application Number
- JP2022163746
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-08
- Filing Date
- 2022-10-12
- Publication Date
- 2025-10-16
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] This application claims priority and benefit under 35 U.S.C. § 119(e) to U.S. Provisional Application No. 63 / 262,630, filed Oct. 17, 2021, which is hereby incorporated by reference in its entirety.
[0002] This implementation generally relates to voice-activated devices, and more particularly, to systems and methods for noise reduction for voice-activated devices.
Background Art
[0003] Voice-activated devices provide hands-free operation by listening for and responding to a user's voice. For example, a user may query a voice-activated device for information (such as a recipe, instructions, directions, etc.) to play media content (such as music, videos, audiobooks, etc.), or to control various devices in the user's home or office environment (such as lighting, thermostats, garage doors, and other home automation devices). Some voice-activated devices may interpret a user's query and communicate with one or more network (such as cloud computing) resources to generate a response to the query. Additionally, some voice-activated devices may first listen for a predefined "trigger word" or "wake word" before generating a query to be sent to network resources.
Summary of the Invention
[0004] This summary is provided to introduce, in a simplified form, a selection of concepts that are further described below in the "Detailed Description." This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
[0005] Noise reduction of audio signals received by a voice-activated device is supported using one or more motion sensors. The motion sensors provide indications of movement or motion information, such as linear or rotational displacement of the voice-activated device. A noise reduction unit within the voice-activated device, which is in standby mode to conserve power, may be activated in response to indications of movement. When activated, the noise reduction unit may adapt to ambient noise from the new location or direction before returning to standby mode. Subsequently, if speech is detected in the audio signal, the noise reduction unit may be activated accordingly, and noise in the audio signal may be suppressed with a short delay or without delay. In addition to, or instead of, motion information may be provided to the noise reduction unit and used to quickly adapt to noise in the audio signal.
[0006] In one embodiment, a method for processing an audio signal in a voice-activated device includes detecting the movement of the voice-activated device, switching a noise reduction unit within the voice-activated device from an inactive mode to an active mode after detecting the movement, and performing noise reduction on the audio signal received after detecting the movement.
[0007] In one embodiment, the controller for the voice-activated device includes a processing system comprising one or more processors coupled to at least one memory. The processing system is configured to detect the movement of the voice-activated device, switch the noise reduction within the voice-activated device from an inactive mode to an active mode at least in part based on the detection of the movement, and perform noise reduction on the voice signal received after the movement has been detected.
[0008] In one embodiment, the voice-activated device comprises one or more motion sensors configured to detect the movement of the voice-activated device, and a noise reduction unit configured to switch from an inactive mode to an active mode at least partially based on the detected movement, and to reduce noise in the audio signal received after the movement has been detected. [Brief explanation of the drawing]
[0009] This implementation is illustrated as an example and is not intended to be limited by the form shown in the attached drawings.
[0010] [Figure 1] Figure 1 illustrates an example of a voice-activated device.
[0011] [Figure 2] Figure 2 shows a timing diagram for the audio input signal and illustrates the operation of the audio activity detector.
[0012] [Figure 3] Figure 3 is a timing diagram for the audio input signal, illustrating the noise in the input signal after the operation of the voice-activated device.
[0013] [Figure 4] Figure 4 shows an exemplary voice-activated device configured to detect motion, used to enhance noise reduction.
[0014] [Figure 5] Figure 5 shows a timing diagram for the audio input signal, illustrating the noise reduction of the input signal in response to the detection of movement of the voice-activated device.
[0015] [Figure 6] Figure 6 illustrates a voice-activated device that is moved relative to a sound source and adapts to changes in the direction of the sound source based on detected motion information.
[0016] [Figure 7] Figure 7 shows a block diagram of an exemplary voice-activated device with several implementations.
[0017] [Figure 8] Figure 8 shows an exemplary flowchart illustrating the typical operation of a voice-activated device in several implementations. [Modes for carrying out the invention]
[0018] The following description includes many specific details, such as examples of specific components, circuits, and processes, to provide a deeper understanding of the disclosure. The term “combined” as used in this application means directly connected or connected via one or more intervening components or circuits. The terms “electronic system” and “electronic device” may be used synonymously to refer to any system capable of electronically processing information. Furthermore, specific nomenclature is specified in the following description for illustrative purposes to provide a deeper understanding of the aspects of the disclosure. However, it will be apparent to those skilled in the art that these specific details may not be necessary to carry out exemplary embodiments. In other examples, well-known circuits and devices are shown in the form of block diagrams to avoid obscuring the disclosure. Some parts of the following detailed description are presented in the form of other symbolic representations of procedures, logic blocks, processes, and operations on data bits within computer memory.
[0019] These descriptions and expressions are intended to be the means used by those skilled in the art of data processing technology to communicate the content of their work to others skilled in the art in the most efficient manner. In this disclosure, procedures, logical blocks, processes, etc., are considered to be a consistent sequence of steps or instructions that lead to a desired result. Such steps require the physical manipulation of physical quantities. Although not required, these quantities usually take the form of electrical or magnetic signals that can be stored, transmitted, combined, or otherwise manipulated in a computer system. However, it should be noted that all of these and similar terms should be associated with the appropriate physical quantities and are merely convenient labels applicable to those quantities.
[0020] As is apparent from the following discussion, unless otherwise specified, throughout this application, discussions using terms such as "access," "receive," "transmit," "use," "select," "judge," "normalize," "multiply," "average," "monitor," "compare," "apply," "update," "measure," "derive," etc. refer to the operations and processes of a computer system or similar electronic computing device that manipulates data represented as physical (electronic) quantities in the registers and memories of the computer system and converts it into other data represented as physical quantities in the memories or registers of the computer system or other such information storage devices, transmission devices, or display devices.
[0021] In the figures, a single block may sometimes be described as performing one or more functions. However, in actual implementation, one or more functions performed by the block may be performed in a single component, may be distributed across multiple components, and / or may be performed using hardware, using software, or using a combination of hardware and software. To clearly illustrate such hardware-software compatibility, various exemplary components, blocks, modules, circuits, and processes are hereinafter generally described in terms of their functions. Whether such functions are implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functions in various ways according to each specific application, but such implementation choices should not be construed as deviating from the scope of this disclosure. Also, exemplary input devices may include components different from those shown, including well-known components such as processors, memories, etc.
[0022] The technology described in this application may be implemented in hardware, software, firmware, or any combination thereof, unless specifically described as being implemented in a particular way. Also, any mechanism described as a module or component may be implemented together in an integrated logic device or separately in discrete but cooperative logic devices. When implemented in software, the technology may be at least partially realized by a non-transitory processor-readable storage medium that contains instructions to perform the functions or methods described as being executed. The non-transitory processor-readable data storage medium may form part of a computer program product that may include packaging material.
[0023] The non-transitory processor-readable storage medium may include random access memory (RAM), such as synchronous dynamic random access memory (SRAM), read only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read only memory (EEPROM), flash memory, and other known storage media. Additionally or alternatively, the technology may be at least partially realized by a processor-readable communication medium that transmits or communicates code in the form of instructions or data structures and is accessible, readable, and / or executable by a computer or other processor.
[0024] Various exemplary logic blocks, modules, circuits, and instructions described in connection with the embodiments disclosed in this application may be executed by one or more processors (or processing systems). The term “processor” as used in this application means any general-purpose processor, dedicated processor, conventional processor, controller, microcontroller, and / or state machine capable of executing scripts or instructions of one or more software programs stored in memory. The term “voice-activated device” or “voice-enabled device” as used in this application may mean any device capable of performing voice-search operations and / or responding to voice inquiries. Examples of voice-activated devices include, but are not limited to, smart speakers, home automation devices, voice command devices, virtual assistants, personal computing devices (e.g., desktop computers, laptop computers, tablets, web browsers, personal digital assistants (PDAs)), data input devices (e.g., remote controls and mice), data output devices (e.g., display screens and printers), remote terminals, kiosks, game consoles (e.g., game consoles, portable game consoles, etc.), communication devices (e.g., mobile phones such as smartphones), and media devices (e.g., recorders, editors, playback devices such as televisions, set-top boxes, music players, digital photo frames, digital cameras).
[0025] Voice-activated devices provide hands-free operation by listening to and responding to user voice commands. Many voice-activated devices are always on so that they can receive and respond to voice commands at any time. Therefore, average power consumption is subject to strict requirements to maintain battery power for a reasonable amount of time. To meet these strict power requirements, voice-activated devices may include a voice activity detector (VAD) used to detect the presence or absence of speech in the received voice signal. When there is no speech, power consumption may be reduced by, for example, setting other components of the voice-activated device to standby mode. When speech is detected by the VAD, the other components are switched from standby mode to active mode. Once activated, noise reduction components would ideally suppress noise in the received voice signal so that speech can be easily identified so that the voice-activated device can respond to the user's voice, for example, by detecting keywords spoken by the user, receiving and analyzing queries, etc.
[0026] Aspects of this disclosure recognize problems related to noise reduction after a change in the position or orientation of a voice-activated device. For example, noise suppression of a speech signal may depend on the relative position or orientation of the voice-activated device with respect to a sound source, e.g., a noise source or a speech source. However, the voice-activated device may change its position, orientation, or both while the noise reduction component is in standby mode. For example, when the noise reduction component is activated in response to speech detection by a VAD, the noise reduction component may attempt to reduce noise in the speech signal based on its previous position or orientation. This may not be applicable to the current position or orientation. As a result, the noise reduction component may not be able to adequately reduce noise in the speech signal immediately afterward (or, equivalently, highlight the desired speech source), and may be required to adapt to the new position or orientation of the sound source before the speech (or other signals) in the speech signal can be accurately identified, resulting in delays and the possibility of missing keywords spoken by the user.
[0027] Various aspects generally relate to the suppression of noise in speech signals by speech-activated devices, and in particular to adaptation to noise in the environment after the speech-activated device has moved. In some implementations, a noise reduction unit switches from an inactive mode to an active mode in response to the detection of movement of the speech-activated device. The noise reduction unit may adapt to the ambient noise in the speech signal from the new position or orientation before returning to the inactive mode. Then, when speech is detected, the noise reduction unit switches to the active mode and may accurately suppress the ambient noise in the speech signal with a small delay or no delay at all. In some implementations, the noise reduction unit may use motion information determined from the detection of movement of the speech-activated device. Motion information may include, for example, the amount of relative change in the position or orientation of the speech-activated device. Then, when speech is detected, the noise detection unit switches to the active mode and may use the motion information to accurately suppress the ambient noise in the speech signal in order to quickly adapt to the new position or orientation. For example, motion information may be used to change or manipulate the direction of beamforming used for noise suppression of the speech signal.
[0028] For example, Figure 1 illustrates an example of a voice-activated device 100 that does not detect motion and therefore cannot dynamically adjust noise suppression in response to motion detection. The voice-activated device 100 is illustrated as including a microphone 110, a switch 120, a voice activity detector (VAD) 130, a noise reduction unit 140, and a wake word engine 150. The voice-activated device 100 may include additional components not shown, such as a speech analysis unit, an application processor, a communication unit, etc.
[0029] The microphone 110 illustrated in Figure 1 may be, for example, a single microphone or a microphone array. The microphone 110 receives speech 101 generated by one or more sound sources, including a human voice and / or an ambient noise source, and provides an audio signal 112. The VAD 130 receives the audio signal 112 and determines whether speech or other target speech is present in the audio signal 112. The VAD 130 may be implemented in hardware and / or software, may be included in the microphone 110, for example, in the VM3011 microphone from Vesper Technologies, or may be part of a codec chip or any component in the audio flow.
[0030] When there is no speech or other target sound in the voice signal 112, the power consumption of the voice activation device 100 may be reduced by setting other units, such as the noise reduction unit 140 and the wake word engine 150, to standby mode. Figure 1 illustrates, as an example, the enabling of the noise reduction unit 140 and the wake word engine 150 by switch 120 in response to the detection of speech or other target sound in the voice signal 112. Switch 120 should be understood as being illustrated simply as an example of power management. For example, VAD 130 may control the power supply to one or more components based on the presence or absence of speech or other target sound from the voice signal 112. For example, in some implementations, components such as the noise reduction unit 140 and the wake word engine 150 are continuously connected to the microphone 110, but the VAD 130 may switch from standby mode to active mode when it detects speech or other target sound from the audio signal 112, and the VAD 130 may switch from active mode to standby mode when it detects that there is no speech or other target sound from the audio signal 112.
[0031] Figure 2 illustrates the timing diagram 200, which includes a simulation of the operation of the audio input signal 202 and the VAD130. In Figure 2, the X axis represents time, and the Y axis represents the amplitude of the audio signal.
[0032] As illustrated, the input signal 202 may contain some noise and may also contain intermittent periods in which utterance 206 (or other target speech) is present. When the VAD 130 detects utterance 206 in the input signal 202, the VAD 130 enables active mode 208. During active mode 208, other components (e.g., noise reduction unit 140 and wake word engine 150) may process the input signal 202. For example, as illustrated in Figure 2, utterance 206 may be used by the VAD 130 to enable active mode 208 for other components, and wake word 210 may be used by the wake word engine 150 to trigger the activation of other components, such as the utterance analysis unit, application processor, and communication unit, in order to analyze query 212. After a predetermined length of time has elapsed, for example, after the utterance 206 is no longer detected in the input signal 202, the active mode 208 is disabled, thereby setting other components (e.g., noise reduction unit 140, wake word engine 150, etc.) into standby mode.
[0033] As illustrated in Figure 1, once enabled (i.e., set to active mode), the noise reduction unit 140 receives a noisy signal. The noise reduction unit 140 processes the noisy signal (for example, adapting to noise in the speech signal) and provides the enhanced signal to other components, such as the wake word engine 150. Figure 1 illustrates, for example, the enhanced signal to the wake word engine 150. When the wake word engine 150 detects a wake word, it may trigger the activation of other components such as the speech analysis unit, application processor, and communication unit.
[0034] The noise reduction unit 140 may apply one or more noise reduction techniques. For example, the noise reduction unit 140 may apply one or more speech enhancement, signal-to-noise ratio (SNR) enhancement, such as noise reduction or suppression, dynamic beamforming, dynamic interference cancellation, dynamic noise cancellation, etc.
[0035] For example, in an implementation where microphone 110 is a single microphone, the noise reduction technique used by noise reduction unit 140 may be highly dependent on the noise energy level, for instance, only temporal information may be considered during the filtering of the audio signal 112. In such an implementation, sudden changes in the noise level that may occur due to the movement of the voice-activated device while in "sleep mode" may be mistakenly classified as speech or other target sounds in the audio signal 112.
[0036] In implementations where microphone 110 is a microphone array, the noise reduction techniques used by noise reduction unit 140 may, in addition to or instead rely on space. For example, dynamic beamforming may be performed by a beamforming unit (not shown) to track the direction of noise sources and / or speech sources, and noise reduction unit 140 may apply spatial filtering for speech enhancement or to increase the SNR of the output signal from beamforming. Speech enhancement generally includes, for example, reducing the amount of distortion in the speech signal and increasing the SNR. To effectively track the direction of noise using beamforming, a “noise-only” signal frame (i.e., a period in which the speech signal contains only noise and not speech or other target speech) may be used, and adaptations may be applied across the “noise-only” signal frame. If a “noise-only” signal frame is not available, beamforming may not converge to the correct direction of noise, especially when a dynamic environment is not considered, resulting in sub-best performance, which may even inadvertently suppress the speech signal in the speech signal 112.
[0037] Therefore, the voice-activated device 100 may apply one or more noise reduction techniques that depend on position, orientation, or both. However, if the position and / or orientation of the voice-activated device 100 is changed, the position and / or orientation-dependent techniques for noise reduction may not function correctly until the noise reduction unit 140 can adapt to the noise in the correct position and / or orientation, which may take some time. Therefore, if the voice-activated device 100 is in sleep mode, for example, when components such as the noise reduction unit 140 are in standby mode, and the position and / or orientation of the voice-activated device 100 is changed, i.e., the voice-activated device 100 is moved, even if the VAD 130 activates the noise reduction unit 140 in response to speech or other target speech, the noise reduction performed by the noise reduction unit 140 may not function properly for a period of time. As a result, the noise reduction unit 140 may not properly suppress the noise, components such as the wake word engine 150 may not be able to distinguish between speech and noise, and may miss wake words or other queries.
[0038] Figure 3 illustrates the timing diagram 300 of the audio input signal 302, showing the noise in the input signal after the voice-activated device has moved. In Figure 3, the X axis represents time, and the Y axis represents the amplitude of the input signal 302. Figure 3 shows a series of events illustrating how the noise reduction unit attempts to reduce noise in the input signal 302 after the position and / or orientation of the voice-activated device has changed.
[0039] As illustrated by arrow 304 in Figure 3, the noise reduction unit may initially reduce ambient noise in the input signal 302, for example, after being initially adapted to ambient noise. If speech is present in box 306 in the input signal 302, it is clear and easily distinguishable from ambient noise.
[0040] After the position and / or orientation of the voice-activated device is changed at arrow 308, the noise reduction unit can no longer suppress the ambient noise in the input signal 302 in box 310 based on its initial adaptation to ambient noise. Box 312 illustrates speech in the noisy input signal 302 after the voice-activated device has been moved and the noise reduction unit has been switched to active mode. Speech in box 312 may be difficult to distinguish from noise by, for example, the wake word engine or other components, which can result in wake words or other information being missed.
[0041] The noise reduction unit, as illustrated in box 314, adapts to the ambient noise over time until the ambient noise is adequately suppressed so that speech can be clearly distinguished from the noise, as illustrated in box 316.
[0042] Figure 4 illustrates an example of a voice-activated device 400 configured to detect the movement of the voice-activated device 400. This example may be used to enhance noise reduction in response to such movement. The voice-activated device 400 is illustrated to include a microphone 410, a switch 420, a voice activity detector (VAD) 430, a noise reduction unit 440, and a wake word engine 450, which may be the same as the microphone 110, switch 120, voice activity detector (VAD) 130, noise reduction unit 140, and wake word engine 150 discussed with reference to Figure 1. The voice-activated device 400 further includes a motion sensor 435, which is illustrated to control the switch 420 and / or provide motion information to the noise reduction unit 440. The voice-activated device 400 may include additional components not shown, such as a speech analysis unit, an application processor, and a communication unit.
[0043] The microphone 410 illustrated in Figure 4 may be, for example, a single microphone or a microphone array. The microphone 410 receives speech 401 generated by one or more sound sources, including a human voice and / or an ambient noise source, and provides a speech signal 412. The VAD 430 receives the speech signal 412 and determines whether speech or other target speech is present in the speech signal 412. The VAD 430 may be implemented in hardware and / or software, may be included in the microphone 410, for example, in the VM3011 microphone from Vesper Technologies, or may be part of a codec chip or any component in the speech flow.
[0044] Switch 420 is illustrated as an example of power management, allowing components such as the noise reduction unit 440 and the wake word engine 450 to be set to standby mode until the VAD 430 detects the presence of speech or other target sound in the audio signal 412. When the VAD 430 detects the presence of speech or other target sound, components such as the noise reduction unit 440 and the wake word engine 450 may be switched to active mode (illustrated using switch 420), as discussed in Figures 1 and 2, for example.
[0045] The voice-activated device 400 further includes a motion sensor 435 capable of detecting linear motion, rotational motion, or a combination thereof. The motion sensor 435 may include, for example, one or more accelerometers, one or more gyroscopes, a magnetometer, a digital compass, or any combination thereof. In some implementations, the motion sensor 435 may detect the occurrence of linear motion or rotational motion and generate a control signal (to the switch 420) when motion is detected. In some implementations, the motion sensor 435 may, in addition to or instead, measure motion, for example, linear displacement and / or rotational displacement, and supply motion information to the noise reduction unit 440.
[0046] The motion sensor 435, like the VAD 430, may be always or almost always active, and may detect (and / or measure) motion while other components such as the noise reduction unit 440 and the wake word engine 450 are in standby mode.
[0047] In one implementation, when motion is detected by the motion sensor 435, the motion sensor 435 may provide a control signal to switch one or more components, such as the noise reduction unit 440 or the wake word engine 450, from standby mode to active mode. The motion sensor 435 may operate independently of the VAD 430; that is, it should be understood that the VAD 430 does not need to detect speech (or other target speech) in the audio signal, and that components may transition from standby mode to active mode based on the motion detected by the motion sensor 435. For example, Figure 4 illustrates the motion sensor 435 providing a control signal to the switch 420 to switch other components to active mode, but any power management technique may be used. For example, the motion sensor 435 may control the power supply to one or more components based on the detected motion of the voice-activated device 400. For example, in some implementations, components such as a noise reduction unit 440 and a wake word engine 450 are continuously connected to the microphone 410, but they may be switched from standby mode to active mode in response to a motion sensor 435 detecting movement of the voice-activated device 400.
[0048] The noise reduction unit 440 may apply one or more noise reduction techniques similar to those discussed above for the noise reduction unit 140. For example, the noise reduction unit 140 may apply one or more speech enhancement, signal-to-noise ratio (SNR) enhancement, such as noise reduction or suppression, dynamic beamforming, dynamic interference cancellation, or dynamic noise cancellation. The one or more noise reduction techniques applied by the noise reduction unit 440 may be position-dependent, orientation-dependent, or both position-dependent.
[0049] By switching the noise reduction unit 440 to active mode in response to motion detection (without requiring speech detection by VAD430), the noise reduction unit 440 may adapt to ambient noise at a new location and / or orientation even if there is no speech in the speech signal 412. Thus, the noise reduction unit 440 can receive a "noise only" signal frame and adapt to any new noise characteristics, such as the direction and energy level of the sound source, as soon as the location and / or orientation of the speech-activated device 400 changes. In some implementations, the noise reduction unit 440 may switch to active mode when motion of the speech-activated device 400 is detected, may begin to adapt to changes in location and / or orientation while the speech-activated device is moving, or the noise reduction unit 440 may switch to active mode after the motion detected by the motion sensor 435 has completed.
[0050] Figure 5 illustrates the timing diagram 500 for the audio input signal 502, illustrating noise reduction in the input signal in response to the detection of movement of the voice-activated device. In Figure 5, the X axis represents time, and the Y axis represents the amplitude of the input signal 502. Figure 5 shows a series of events illustrating noise reduction in the input signal 502 by the noise reduction unit 440 in response to the motion sensor 435 detecting a change in the position and / or orientation of the voice-activated device 400.
[0051] As illustrated by arrow 504 in Figure 5, the noise reduction unit 440 may initially reduce the ambient noise in the input signal 502 after being initially adapted to, for example, ambient noise. If speech is present in the box 506 in the input signal 502, it is clear and easily distinguishable from the ambient noise.
[0052] The movement of the voice-activated device 400 is detected by the motion sensor 435 at arrow 508, and the noise reduction unit 540 switches to active mode accordingly. As illustrated by box 510, the input signal 502 received by the noise reduction unit 540 contains ambient noise but no speech. By receiving a signal consisting only of noise, the noise reduction unit 540 may adapt to the ambient noise at the new position and / or orientation of the voice-activated device 400. After a predetermined length of time, or in response to an instruction from the noise reduction unit 540 that the ambient noise has been adequately reduced, the noise reduction unit 540 may return to standby mode, for example, at the end of box 510. Thus, when speech is detected in the input signal 502 (after the voice-activated device 400 has moved and adapted to the noise from the new position or orientation), the speech is clear and easily distinguishable from the ambient noise, for example, as illustrated in box 512.
[0053] In additional or alternative implementations, the movement of the voice-activated device 400 may be measured by a motion sensor 435, and motion information, such as the displacement and / or rotation of the voice-activated device 400, may be provided to a noise reduction unit 440. The noise reduction unit 440 may use this motion information to perform noise reduction in the voice signal 412.
[0054] In one implementation, motion information may be used by the noise reduction unit 440 to adapt to ambient noise with a relatively short adaptation time, or without any adaptation time at all. For example, the noise reduction unit 440 may be switched to active mode based on motion detected from the motion sensor 435, or the noise reduction unit 440 may receive motion information from the motion sensor 435. This motion information may be used to adapt more quickly to ambient noise when speech is not present, for example (as shown in box 510 in Figure 5). In another example, the noise reduction unit 440 may receive motion information from the motion sensor 435, but otherwise remain in standby mode (for example, the motion information may be stored in a buffer and provided to the noise reduction unit 440 when it enters active mode). If the VAD 430 detects speech (or other target speech) in the audio signal 412, the noise reduction unit 440 may be switched to active mode and quickly adapt to ambient noise using motion information from the motion sensor 435.
[0055] Figure 6 illustrates an environment including a voice-activated device 600 and a sound source 620, which may be a noise source or a speech source. The voice-activated device 600 may be an example of the voice-activated device 400 in Figure 4. The voice-activated device 600 is illustrated as moving from a first position and orientation at a first time (t1) relative to the sound source 620 to a second position at a second time (t2) relative to the sound source 620 (as indicated by arrows 612 and 614).
[0056] The voice-activated device 600 includes a microphone 602, shown as a microphone array. The microphone 602 receives sound from the sound source 620 (shown by arrow 622) at a first energy level and angle α1 using beamforming. The voice-activated device 600 further includes a motion sensor 604, which may include, for example, one or more accelerometers and / or gyroscopes, a compass, etc. The motion sensor 604 measures the linear and / or rotational displacement of the voice-activated device 600, shown by arrows 612 and 614, when the voice-activated device 600 moves from its first position and orientation at a first time (t1) to its second position and orientation at a second time (t2). The motion sensor 604 provides motion information to a noise reduction unit 606. The noise reduction unit 606 uses the measurement information to determine the current direction of the sound source 620 (for example, at time t2) in order to quickly adapt to ambient noise. For example, the noise reduction unit 606 may estimate a new direction (e.g., angle α2) based on the previous direction (angle α1) and energy level of the sound source from a first time t1 (before the voice activation device 600 moves) and the measured linear displacement 612 and rotational displacement 614 as measured by the motion sensor 604, and estimate the second energy level of the sound (indicated by arrow 624) from the sound source 620 at a second time (t2) (after the voice activation device 600 moves).
[0057] Thus, the noise reduction unit 606 may use motion information to make adjustments based on changes in the measured position and / or orientation. For example, the newly estimated orientation of the sound source 620 may be used for a new steering direction for beamforming using the microphone 602 to receive (or suppress) sound from the sound source 620.
[0058] Figure 7 illustrates a block diagram of an example of a voice-activated device 700 in several implementations. More specifically, as discussed in this application, the voice-activated device 700 is configured to detect motion and enhance noise reduction of the audio signal in response to that motion. In some implementations, the voice-activated device 700 may be an example of the voice-activated device 400 in Figure 4 or the voice-activated device 600 in Figure 6. The voice-activated device 700 or a part thereof may be a controller for enhancing noise reduction in response to motion. The voice-activated device 700 is illustrated as including a device interface 710, a network interface 716, one or more motion sensors 718, a VAD 719, a processing system 720, and a memory 730. It should be understood that additional components may be included in the voice-activated device 700.
[0059] The device interface 710 is configured to communicate with one or more components of the voice-activated system. In some implementations, the device interface 710 may include a microphone interface (I / F) 712, a media output interface 714, and a network interface 716. The microphone interface 712 may communicate with the microphone of the voice-activated device 700 (e.g., microphone 410 in Figure 4 and / or microphone 602 in Figure 6). For example, the microphone interface 712 may receive voice signals from the microphone, and in some implementations, it may provide control signals to the microphone, for example, to control beamforming.
[0060] The media output interface 714 may be used to communicate with one or more media output components of the voice-activated device 700. For example, the media output interface 714 may send information and / or media content to a media output component (e.g., a speaker and / or display) to generate a response to a user's voice input or inquiry.
[0061] The network interface 716 may be used to communicate with external network resources of the voice-activated device 700. For example, the network interface 716 may send voice queries to network resources and receive results from those network resources.
[0062] One or more motion sensors 718 may include one or more accelerometers, one or more gyroscopes, magnetometers, digital compasses, or any combination thereof. In some implementations, one or more motion sensors 718 may detect the occurrence of linear or rotational motion and generate control signals when motion is detected. In some implementations, one or more motion sensors 718 may generate motion information, such as measured linear and / or rotational displacement. It should be understood that a processing system 720 (or other processing system) may cooperate with the one or more motion sensors 718 to generate motion information based on the raw signals generated by the one or more motion sensors 718.
[0063] VAD719 is a voice activity detector that detects the presence or absence of speech (or other trigger sounds) in the voice signal received via the microphone interface 712. Although VAD719 is illustrated as a separate component in Figure 7, it should be understood that VAD719 can be implemented in hardware and / or software. Furthermore, VAD719 can be coupled to receive voice signals directly from the microphone, from the microphone interface 712, or from the processing system 720. Moreover, VAD719 may be included in the microphone itself, or as part of a codec chip, or within any component in the voice flow.
[0064] The processing system 720 may include one or more suitable processors capable of executing scripts or instructions of one or more software programs stored in the voice-activated device 700 (for example, in memory 730). The processing system 720 may be implemented using a combination of hardware, firmware, and software. In some embodiments, the processing system 720 may represent one or more circuits that can be configured to perform at least a portion of data signal calculation procedures or processes related to the operation of the voice-activated device 700.
[0065] Memory 730 may contain one or more software (SW) modules containing executable code or software instructions that, when executed by the processing system 720, cause one or more processors in the processing system 720 to operate as a dedicated computer programmed to perform the technology disclosed herein (in particular, including one or more non-volatile memory elements such as EPROM, EEPROM, flash memory, or hard drive). Although a component or module is illustrated as software in memory 730 executable by one or more processors in the processing system 720, it should be understood that such component or module may be stored in memory 730, located in one or more processors in the processing system 720, or in dedicated hardware separate from the processors. The set of contents of memory 730 as illustrated in the voice-activated device 700 is merely illustrative, and therefore, the functionality of modules and / or data structures may be combined, separated, and / or constructed in different ways depending on the implementation of the voice-activated device 700.
[0066] The memory 730 may include a standby / start SW module 731 that, when executed by the processing system 720, receives control signals from the VAD 719 and, in some implementations, from one or more motion sensors 718, and, in response to such control signals, configures one or more processors to switch one or more components of the voice-activated device 700, including a noise reduction unit, between standby and active modes. In some implementations, one or more processors may be configured to receive signals from other components in the voice-activated device 700 indicating that there is no longer any speech in the voice signal, so that the component may be switched from active to standby mode.
[0067] Memory 730 may include a noise reduction SW module 732 that configures one or more processors to reduce noise in the received audio signal when noise reduction is in active mode, when executed by the processing system 720. The noise reduction SW module 732 may include, for example, one or more submodules for noise reduction. For example, a speech enhancement SW module 734 may configure one or more processors in the processing system 720 to enhance speech in the audio signal by, for example, one or more temporal or frequency filters. A spatial filtering SW module 736 may configure one or more processors in the processing system 720 to spatially filter the audio signal by, for example, beamforming. A beamforming SW module 738 may configure one or more processors in the processing system 720 to perform dynamic beamforming, for example, to adjust or manipulate the beam of a microphone array to direct the received beam towards a desired sound source or away from a noise source. An interference cancellation SW module 740 may configure one or more processors in the processing system 720 to perform dynamic interference cancellation. The noise cancellation SW module 742 may be configured to perform dynamic noise cancellation by one or more processors in the processing system 720.
[0068] The memory 730 may include a wake word SW module 744 that, when executed by the processing system 720, configures one or more processors to identify wake words (or other target noise) in audio signals received while in active mode.
[0069] Each software module contains instructions that cause the voice-activated device 700 to perform a corresponding function when executed by one or more processors of the processing system 720. The non-temporary computer-readable medium of memory 730 therefore contains instructions for performing all or part of the operations described below with respect to Figure 8.
[0070] Figure 8 illustrates an exemplary flowchart depicting an exemplary operation 800 for processing an audio signal according to the implementation described in this application. In some implementations, the exemplary operation 800 may be performed by an audio-activated device, such as the audio-activated devices 400, 600, or 700 shown in Figures 4, 6, and 7, respectively.
[0071] As illustrated, the voice-activated device may detect its movement, as discussed with reference to, for example, Figures 4, 5, 6, and 7 (810). For example, the controller may include a processing system configured to detect the movement of the voice-activated device, such as shown in Figure 7. The movement of the voice-activated device may be detected using, for example, a motion sensor 435, a motion sensor 604, or one or more motion sensors 718, as illustrated in Figures 4, 6, and 7, respectively, and a processing system 720 which consists of dedicated hardware or executes executable code or software instructions in memory 730.
[0072] The voice-activated device may switch a noise reduction unit within the voice-activated device from inactive mode to active mode based at least partially on the detection of motion, as discussed with reference to, for example, Figures 4, 5, 6, and 7 (820). For example, the controller may include a processing system configured to switch the noise reduction within the voice-activated device from inactive mode to active mode based at least partially on the detection of motion, as illustrated, for example, in Figure 7. The noise reduction unit may be configured to switch from inactive mode to active mode based at least partially on the detection of motion, for example, using a switch 420 as illustrated in Figures 4 and 7, or using a processing system 720 that implements executable code or software instructions in memory, such as a standby / start SW module 731, which is composed of dedicated hardware.
[0073] The voice-activated device may perform noise reduction of the received voice signal after motion detection by a noise reduction unit (830), as discussed with reference to, for example, Figures 4, 5, 6, and 7. In some embodiments, the noise reduction of the voice signal may be one or more of the following: speech enhancement, signal-to-noise ratio (SNR) enhancement, spatial filtering, beamforming, interference cancellation, noise cancellation, or any combination thereof. For example, the controller may include a processing system configured to perform noise reduction of the received voice signal after motion detection, as illustrated, for example, in Figure 7. The noise reduction unit may be configured to perform noise reduction of the received voice signal after motion detection using, for example, noise reduction unit 440 or noise reduction unit 606, as illustrated in Figures 4, 6, and 7, respectively, or using a processing system 720 that executes executable code or software instructions in memory 730, such as a noise reduction SW module 732 (and optionally one or more submodules), which is composed of dedicated hardware.
[0074] In some embodiments, the switching of the noise reduction unit from an inactive mode to an active mode may be in response to motion detection, and performing noise reduction on the received audio signal after motion detection may include adapting the audio signal to ambient noise before returning from the active mode to the inactive mode.
[0075] For example, in some embodiments, after adapting to ambient noise in the audio signal, the voice-activated device may further detect speech in the audio signal, as discussed with reference to, for example, Figures 2, 4, and 5. Speech in the audio signal may be detected using a VAD430 or VAD719, as illustrated in Figures 4 and 7, respectively, and a processing system 720 that consists of dedicated hardware or executes executable code or software instructions in memory 730. The noise reduction unit may be switched from an inactive mode to an active mode in response to the detection of speech. Here, the noise reduction unit is adapted to ambient noise in the audio signal, as discussed with reference to Figures 4 and 5. For example, the switching of the noise reduction unit from an inactive mode to an active mode in response to the detection of speech may use a switch 420, as shown in Figures 4 and 7, respectively, or a processing system 720 that consists of dedicated hardware or executes executable code or software instructions in memory 730, such as a standby / activation SW module 731.
[0076] In some embodiments, the voice-activated device may further generate motion information from motion detection. Here, the motion information is used in the noise reduction operation after motion detection, as discussed with reference to, for example, Figures 4 and 6. For example, the motion information may be generated from motion detected using, for example, motion sensors 435, 604, or one or more motion sensors 718, as illustrated in Figures 4, 6, and 7, and a processing system 720 which consists of dedicated hardware or executes executable code or software instructions in memory 730.
[0077] For example, in some embodiments, speech may be detected after the voice-activated device has detected motion. Here, the switching of the noise reduction unit from inactive mode to active mode may be in response to the detection of speech, as discussed with reference to, for example, Figures 4 and 6. Speech may be detected after motion detection using a VAD430 or VAD719, as illustrated in Figures 4 and 7, and a processing system 720 that is composed of dedicated hardware or executes executable code or software instructions in memory 730.
[0078] For example, in some embodiments, the voice-activated device may further determine a steering direction for beamforming to receive the voice signal based on the steering state and motion information prior to motion detection. Here, the noise reduction after motion detection uses the steering direction, as discussed with reference to, for example, Figures 4 and 6. For example, the steering direction may be determined for beamforming to receive the voice signal based on the steering state and motion information prior to motion detection. Here, the noise reduction may be performed based on the steering direction after motion detection, consisting of, for example, a noise reduction unit 440, a noise reduction unit 606, as illustrated in Figures 4, 6, and 7, respectively, or dedicated hardware, or a processing system 720 that executes executable code or software instructions in memory 730, such as a noise reduction SW module 732 (and optionally one or more submodules such as a beamforming SW module 738).
[0079] Those skilled in the art will understand that information and signals can be represented using any of the various different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols and chips that may have been referenced throughout the above description may be represented by voltage, electric current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.
[0080] Furthermore, those skilled in the art will understand that various exemplary logic blocks, modules, circuits, and algorithmic steps described in relation to the embodiments disclosed in this application may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this hardware-software compatibility, various exemplary components, blocks, modules, circuits, and steps have been described above in general terms of their function. Whether such functions are implemented in hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functions in various ways to suit their specific applications, but such implementation choices should not be interpreted as resulting in a deviation from the scope of this disclosure.
[0081] The methods, sequences, or algorithms described in relation to the embodiments disclosed in this application may be implemented directly in hardware, in software modules executed by a processor, or in a combination of the two. The software modules may be RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art. The exemplary storage medium is coupled to the processor so that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integrated with the processor.
[0082] In the above specification, embodiments have been described with reference to specific examples. However, it will be apparent that various modifications and changes can be made to the embodiments without deviating from the broader scope of disclosure presented in the appended claims. Therefore, the specification and drawings should be evaluated in an illustrative rather than restrictive sense.
Claims
1. 1. A method for processing an audio signal in a voice-activated device, comprising: Detecting movement of the voice-activated device; switching a noise reduction unit in the voice-activated device from an inactive mode to an active mode based at least in part on detecting the motion; performing noise reduction on the received audio signal after detecting the motion with the noise reduction unit; Contains method.
2. Performing noise reduction of the audio signal includes one or more of speech enhancement, signal-to-noise ratio (SNR) enhancement, spatial filtering, beamforming, interference cancellation, noise cancellation, or any combination thereof. The method of claim 1.
3. Switching the noise reduction unit from the inactive mode to the active mode is responsive to detecting the motion, and performing the noise reduction of the received audio signal after detecting the motion includes adapting to environmental noise of the audio signal before returning from the active mode to the inactive mode. The method of claim 1.
4. After adapting to the environmental noise of the audio signal, the method further comprises: Detecting speech in an audio signal; switching the noise reduction unit from the inactive mode to the active mode in response to detecting the speech; Including, The noise reduction unit is adapted to the environmental noise of the audio signal. The method of claim 3.
5. Furthermore, generating motion information from the motion detection; performing the noise reduction after the motion detection uses the motion information. The method of claim 1.
6. further comprising detecting speech after detecting the movement; Switching the noise reduction unit from the inactive mode to the active mode is in response to detecting the speech. The method of claim 5.
7. Further, the method includes determining a steering direction for beamforming for receiving the audio signal based on a steering state before detecting the motion and the motion information, Performing the noise reduction after detecting the movement uses the steering direction. The method of claim 5.
8. 1. A controller for a voice-activated device, comprising: at least one memory; a processing system comprising one or more processors coupled to the at least one memory; Equipped with the processing system comprising: Detecting movement of the voice-activated device; switching a noise reduction unit in the voice-activated device from an inactive mode to an active mode based at least in part on detecting the motion; configured to perform noise reduction on the received audio signal after detecting the motion. controller.
9. the processing system is configured to switch the noise reduction from the inactive mode to the active mode in response to detecting the motion; The processing system is configured to perform the noise reduction of the received audio signal after the motion is detected by adapting to environmental noise of the audio signal before returning from the active mode to the inactive mode. The controller of claim 8 .
10. 1. A voice-activated device, comprising: one or more sensors configured to detect movement of the voice-activated device; a noise reduction unit configured to switch from an inactive mode to an active mode based at least in part on the detected motion and to perform noise reduction on the received audio signal after the motion is detected; Equipped with Voice-activated devices.