Acoustic zoning with distributed microphones

By coordinating multiple asynchronous microphones and classifier models, the inefficiency of user location determination and audio output in intelligent audio device systems is solved, enabling accurate positioning and optimized audio interaction in complex environments, and improving the efficiency of wake word detection and audio interaction.

CN114402385BActive Publication Date: 2025-11-21DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080064826.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-18
Filing Date
2020-07-28
Publication Date
2025-11-21
Estimated Expiration
2040-07-28

AI Technical Summary

Technical Problem

Existing smart audio device systems struggle to effectively coordinate multiple microphones to determine user location and optimize audio output, especially in noisy and complex environments, resulting in inefficient wake word detection and audio interaction.

Method used

By coordinating multiple asynchronous microphones and using a classifier model to estimate the user's location in different user zones, combined with light signals and speaker control, accurate user location and optimized audio output can be achieved.

Benefits of technology

It improves the accuracy of wake word detection and the efficiency of audio interaction, and can accurately identify the user's location and optimize audio output in complex environments, thereby enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114402385B_ABST
    Figure CN114402385B_ABST
Patent Text Reader

Abstract

A method for estimating a location of a user in an environment can involve receiving an output signal from each microphone of a plurality of microphones in the environment. At least two microphones of the plurality of microphones can be included in separate devices at separate locations in the environment, and the output signals can correspond to a current utterance of a user. The method can involve determining a plurality of current acoustic features from the output signal of each microphone, and applying a classifier to the plurality of current acoustic features. Applying the classifier can involve applying a model trained on previously determined acoustic features resulting from a plurality of previous utterances spoken by the user in a plurality of user zones in the environment. The method can involve determining an estimate of the user zone in which the user is currently located based at least in part on an output from the classifier.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-citation of related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 950,004, filed December 18, 2019; U.S. Provisional Patent Application No. 62 / 880,113, filed July 30, 2019; and EP Patent Application No. 19212391.7, filed November 29, 2019, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to systems and methods for coordinating and implementing intelligent audio devices. Background Technology

[0004] Smart audio devices have been widely deployed and are becoming a common feature in many homes. While existing systems and methods for implementing smart audio devices offer advantages, improved systems and methods are desirable.

[0005] Notation and nomenclature

[0006] In this document, we use the term "intelligent audio device" to refer to a smart device, which is a single-purpose audio device or a virtual assistant (e.g., a connected virtual assistant). A single-purpose audio device is a device that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker) and is generally or primarily designed to perform a single purpose (e.g., a television (TV) or a mobile phone). While televisions can generally play (and are considered capable of playing) audio from program material, in most cases, modern televisions run some kind of operating system on which applications run natively, including applications for watching television. Similarly, audio input and output in a mobile phone may do many things, but these are all served by applications running on the phone. In this sense, a single-purpose audio device with a speaker and microphone is typically configured to run local applications and / or services to directly utilize the speaker and microphone. Some single-purpose audio devices may be configured to be grouped together to enable audio playback in a zone or user-configured area.

[0007] In this document, a “virtual assistant” (e.g., a connected virtual assistant) is a device that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker) and provides the ability to use multiple devices (unlike the virtual assistant) for some form of cloud-enabled or otherwise not implemented in or on the virtual assistant itself (e.g., a smart speaker, a smart display, or a voice assistant integrated device). Virtual assistants may sometimes work together, for example, in a very discrete and conditionally defined manner. For example, two or more virtual assistants may work together in the sense that one of them, the one most confident that it has heard the wakeword, responds to the wakeword. Connected devices may form a constellation that can be managed by a master application, which may be (or include or implement) the virtual assistant.

[0008] In this document, "wake word" is used broadly to refer to any sound (e.g., a human-spoken word or some other sound) in which a smart audio device is configured to wake up in response to the detection ("hearing") of a sound (using at least one microphone contained in or coupled to the smart audio device, or at least one other microphone). In this context, "wake up" means that the device enters a state of waiting (i.e., listening) for a sound command.

[0009] In this paper, the term "wake word detector" refers to a device (or software containing instructions for configuring the device) configured to continuously search for alignments between real-time sound (e.g., speech) features and a trained model. Typically, a wake word event is triggered whenever the probability of a wake word being detected by the wake word detector exceeds a predefined threshold. For example, the threshold may be a predetermined threshold tuned to provide a good trade-off between false acceptance and false rejection rates. Following a wake word event, the device may enter a state in which it listens for commands and passes the received commands to a larger, more computationally intensive recognizer (which may be referred to as a "wake-up" state or a "focused" state).

[0010] Throughout this disclosure, the terms "speaker" and "loudspeaker" are used synonymously to refer to any acoustic transducer (or group of transducers) driven by a single loudspeaker feed. A typical set of headphones includes two loudspeakers. A loudspeaker may be implemented to include multiple transducers (e.g., a woofer and a tweeter), all of which are driven by a single common loudspeaker feed. In some examples, the loudspeaker feed undergoes different processing in different branches of the circuitry coupled to the different transducers.

[0011] In this document, the expression "microphone location" refers to the location of one or more microphones. In some instances, a single microphone location may correspond to a microphone array residing in a single audio device. For example, a microphone location may be a single location corresponding to an entire audio device containing one or more microphones. In some such instances, a microphone location may be a single location corresponding to the centroid of the microphone array of a single audio device. However, in some examples, a microphone location may be the location of a single microphone. In some such instances, the audio device may have only a single microphone.

[0012] Throughout this disclosure, the expression “to” a signal or data (e.g., to filter, scale, transform, or apply gain to the signal or data) is used broadly to mean performing an operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a signal version that has undergone preliminary filtering or preprocessing before the operation is performed on it).

[0013] Throughout this disclosure, and as included in the claims, the term "system" is used broadly to refer to an apparatus, system, or subsystem. For example, a subsystem implementing a decoder may be referred to as a decoder system, and a system containing such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, wherein the subsystem generates M inputs and the other XM inputs are received from an external source) may also be referred to as a decoder system.

[0014] Throughout this disclosure, and as included in the claims, the term "processor" is used broadly to refer to a system or apparatus that is programmable or otherwise configurable (e.g., using software or firmware) to perform operations on data (e.g., audio or video or other image data). Examples of processors include field-programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipelined processing of audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets. Summary of the Invention

[0015] A coordinating system comprising multiple smart audio devices may require some knowledge of the user's location in order to at least: (a) select the optimal microphone for voice pickup; and (b) transmit audio from a reasonable location. Existing techniques include selecting a single microphone with high wake-word confidence and sound source localization algorithms that use multiple synchronized microphones to estimate the user's coordinates relative to the device. As used herein, a microphone may be referred to as "synchronized" if it uses the same sampling clock or synchronized sampling clock to digitally sample the sound detected by the microphone. For example, a first microphone of a plurality of microphones in an environment may sample audio data according to a first sampling clock, and a second microphone of the plurality of microphones may sample audio data according to a second sampling clock.

[0016] Some embodiments consider a system for coordinating (or managing) smart audio devices, where each of the smart audio devices is (or includes) a wake-word detector. Multiple microphones (e.g., asynchronous microphones) may be available, each of which is included in or configured to communicate with a device implementing a classifier; in some instances, the classifier may be one of the smart audio devices. (The classifier may also be referred to herein as a “region classifier” or a “classifier model.”) In other instances, the classifier may be implemented by another type of device (e.g., a smart device not configured to provide audio) configured to communicate with the microphones. For example, at least some of the microphones may be discrete microphones (e.g., in a home appliance) that are not included in any smart audio device but are configured to communicate with the device implementing a classifier configured to estimate the user’s region based on multiple acoustic features derived from at least some of the microphones. In some such embodiments, the goal is not to estimate the user’s precise geometric location, but rather to form a robust estimate of the discrete region (e.g., in the presence of heavy noise and residual echoes). As used herein, the “geometric location” of an object or user in an environment refers to a location based on a coordinate system, whether the coordinate system is referenced to GPS coordinates, to the entire environment (e.g., according to a Cartesian or polar coordinate system, having its origin at a certain location within the environment), or to a specific device within the environment (e.g., according to a Cartesian or polar coordinate system, with the device as its origin), such as a smart audio device. According to some examples, an estimate of a user’s location in an environment can be determined without referencing the geometric locations of multiple microphones.

[0017] The user, smart audio device, and microphone may be located in an environment where sound can propagate from the user to the microphone (e.g., the user's residence or business premises), and the environment may include predetermined user areas. For example, the environment may include one or more of the following user areas: a food preparation area; a dining area; an open area of ​​living space; a television area (including a television and sofa) of living space; a bedroom separate from the open living space area; and so on. During the operation of the system according to some examples, it can be assumed that if the user is in the environment, then the user is physically located in or near one of the user areas at any given time, and the user area in which the user is currently located may change over time.

[0018] The microphone may be “asynchronous.” As used herein, a microphone can be referred to as “asynchronous” if different sampling clocks are used to digitally sample the sound detected by the microphone. For example, a first microphone among a plurality of microphones in the environment may sample audio data according to a first sampling clock, and a second microphone among a plurality of microphones may sample audio data according to a second sampling clock. In some examples, the microphones in the environment may be randomly located, or at least distributed in an irregular and / or asymmetrical manner within the environment. In some instances, the user’s area can be estimated via a data-driven method involving multiple high-level acoustic features at least partially derived from at least one of the wake word detectors. In some implementations, these acoustic features (which may include wake word confidence and / or reception levels) may consume very little bandwidth and may be asynchronously transmitted to the means of implementing the classifier with very little network load. Data regarding the geometric location of the microphones may or may not be provided to the classifier, depending on the specific implementation. As mentioned elsewhere herein, in some instances, an estimate of the user’s location in the environment can be determined without reference to the geometric locations of multiple microphones.

[0019] Some aspects of the embodiments relate to implementing and / or coordinating intelligent audio devices.

[0020] According to some embodiments, in the system, multiple smart audio devices respond in a coordinated manner (e.g., by emitting light signals) to (e.g., to indicate focus or availability) the system's determination of a common operating point (or operating state). For example, the operating point could be a focused state entered in response to a wake word from a user, wherein all devices have an estimate of the user's location (e.g., with at least one uncertainty), and wherein the devices emit light of different colors depending on their estimated distance from the user.

[0021] Some publicly available methods involve estimating a user's location within an environment. Some of these methods may involve receiving output signals from each of a plurality of microphones in the environment. Each of the plurality of microphones may reside at a microphone location within the environment. In some instances, the output signals may correspond to the user's current utterance.

[0022] In some instances, at least two of the plurality of microphones are contained in separate devices at separate locations within the environment.

[0023] Some such methods may involve determining multiple current acoustic features from the output signal of each microphone and applying a classifier to the multiple current acoustic features. Applying the classifier may involve applying a model trained on previously determined acoustic features derived from multiple previous utterances made by the user in multiple user zones within the environment. Some such methods may involve determining an estimate of the user zone in which the user is currently located, at least in part based on the output of the classifier. For example, the user zone may include a sink area, a food preparation area, a refrigerator area, a dining area, a sofa area, a television area, and / or a porch area.

[0024] In some instances, a first microphone of the plurality of microphones may sample audio data according to a first sampling clock, and a second microphone of the plurality of microphones may sample audio data according to a second sampling clock. In some instances, at least one of the plurality of microphones may be included in a smart audio device or configured to communicate with a smart audio device. According to some instances, the plurality of user areas may involve multiple predetermined user areas.

[0025] In some instances, the estimation can be determined without reference to the geometric positions of the plurality of microphones. In some instances, the plurality of current acoustic features can be determined asynchronously.

[0026] In some instances, the current utterance and / or the previous utterance may contain a wake-up word utterance. In some instances, the user region may be estimated as the category with the highest posterior probability.

[0027] According to some implementations, the model is trained using training data labeled with user areas. In some examples, the classifier may involve applying a model trained using unlabeled training data without user areas. In some instances, applying the classifier may involve applying a Gaussian mixture model trained on one or more of normalized wake-word confidence, normalized average received level, or maximum received level.

[0028] In some instances, training of the model can continue during the application of the classifier. For example, the training can be based on explicit feedback from the user, i.e., feedback provided by the user. Alternatively or additionally, training can be based on implicit feedback, i.e., automatically provided feedback, such as implicit feedback regarding the success (or failure) of beamforming or microphone selection based on the estimated user area. In some instances, the implicit feedback can include a determination that the user has anomalously terminated the voice assistant's response. According to some embodiments, the implicit feedback can include a command recognizer returning a low-confidence result. In some examples, the implicit feedback can include a second-pass wake word detector returning a low-confidence result for the spoken wake word.

[0029] According to some implementations, the method may involve selecting at least one speaker based on the estimated user area and controlling the at least one speaker to provide sound to the estimated user area. Alternatively or further, the method may involve selecting at least one microphone based on the estimated user area and providing a signal output by the at least one microphone to a smart audio device.

[0030] Some publicly available training methods may involve prompting a user to say a training utterance at least once in each of multiple locations within a first user area of ​​the environment. According to some examples, the training utterance may be a wake-up word. Some such methods may involve receiving a first output signal from each of multiple microphones in the environment. Each of the multiple microphones may reside in a microphone location within the environment. In some examples, the first output signal may correspond to an instance of a detected training utterance received from the first user area.

[0031] Some such methods may involve determining a first acoustic feature from each of the first output signals and training a classifier model to correlate the first user area with the first acoustic feature. In some examples, the classifier model may be trained without reference to the geometric positions of the plurality of microphones. In some examples, the first acoustic feature may include normalized wake-word confidence, normalized average received level, or maximum received level.

[0032] Some such methods may involve prompting a user to speak the training utterance at each of multiple locations within a second to a Kth user area of ​​the environment. Some such methods may involve receiving second to Hth output signals from each of multiple microphones in the environment. The second to Hth output signals may each correspond to an instance of the detected training utterance received from the second to the Kth user areas. Some such methods may involve determining second to Gth acoustic features from each of the second to the Hth output signals, and training the classifier model to associate the second to the Kth user areas with the second to the Gth acoustic features, respectively.

[0033] In some examples, a first microphone of the plurality of microphones may sample audio data according to a first sampling clock, and a second microphone of the plurality of microphones may sample audio data according to a second sampling clock.

[0034] Some or all of the operations, functions, and / or methods described herein can be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include, for example, the memory devices described herein, including (but not limited to) random access memory (RAM) devices, read-only memory (ROM) devices, etc. Therefore, some innovative aspects of the subject matter described herein can be implemented in one or more non-transitory media having software stored thereon. For example, the software may include instructions for controlling one or more devices to at least partially perform one or more of the disclosed methods.

[0035] One innovative aspect of the subject matter described in this disclosure can be implemented in an apparatus. The apparatus may include an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof. According to some examples, the control system may be configured to at least partially perform one or more of the disclosed methods.

[0036] This disclosure includes a system configured (e.g., programmed) to perform any embodiment of the disclosed methods or steps thereof, and one or more tangible, non-transitory, computer-readable media (e.g., disks or other tangible storage media) implementing non-transitory storage of data, storing code for performing (e.g., executable code to perform) any embodiment of the disclosed methods or steps thereof. For example, embodiments of the disclosed system may be or include a programmable general-purpose processor, digital signal processor, or microprocessor, programmed with software or firmware and / or otherwise configured to perform various operations on data, including embodiments of the disclosed methods or steps thereof. Such a general-purpose processor may be or include a computer system including input devices, memory, and processing subsystems programmed (and / or otherwise configured) to perform embodiments of the disclosed methods (or steps thereof) in response to data asserted thereto. Attached Figure Description

[0037] Figure 1A This indicates the environment based on an instance.

[0038] Figure 1B This indicates the environment based on another instance.

[0039] Figure 1C This is a block diagram illustrating examples of components of a device capable of implementing various aspects of this disclosure.

[0040] Figure 1D An overview can be provided by, for example Figure 1C The flowchart shown is an example of a method executed by the device.

[0041] Figure 2 This is a block diagram of elements of an example of an embodiment configured to implement a region classifier.

[0042] Figure 3 An overview can be provided by, for example Figure 1C A flowchart of an instance of a method executed by the device.

[0043] Figure 4 An overview can be provided by, for example Figure 1C The flowchart is another instance of the method executed by the device.

[0044] Figure 5 An overview can be provided by, for example Figure 1C The flowchart is another instance of the method executed by the device.

[0045] Similar reference numerals and names in the various figures indicate similar elements. Detailed Implementation

[0046] Many embodiments of this disclosure are technically possible. How to implement the embodiments described herein will be apparent to those skilled in the art. Some embodiments of the systems and methods are described herein.

[0047] A coordinating system consisting of multiple smart audio devices can be configured to determine when a "wake word" (as defined above) from a user is detected. At least some devices in such a system can be configured to listen for commands from the user.

[0048] Next, referring to Figures 1 to 3, we describe some examples of embodiments of the present disclosure. Figure 1A This is a schematic diagram of an environment (living space) that includes a system comprising a set of intelligent audio devices (devices 1.1) for audio interaction, a speaker (1.3) for audio output, a microphone 1.5, and a controllable light (1.2). In some examples, one or more of the microphones 1.5 may be part of, or associated with, one of the devices 1.1, the light 1.2, or the speaker 1.3. Alternatively or additionally, one or more of the microphones 1.5 may be attached to another part of the environment, such as a wall, ceiling, furniture, appliance, or another device in the environment. In one instance, each of the intelligent audio devices 1.1 includes at least one microphone 1.5 (and / or is configured to communicate with at least one microphone 1.5). Figure 1A The system can be configured to implement embodiments of this disclosure. Various methods can be used to obtain... Figure 1A Information is collectively acquired in the microphone 1.5 and provided to a device implementing a classifier configured to provide a location estimate of the user who read the wake word.

[0049] In living spaces (e.g., Figure 1A Within a living space (as defined in the text), there exists a set of natural activity zones where people will perform tasks or activities, or cross thresholds. In some instances, these zones (which may be referred to herein as user zones) can be defined by the user without specifying geometric coordinates or other markings. Figure 1A In the example shown, the user area may contain:

[0050] 1. Kitchen sink and food preparation area (in the upper left area of ​​the living space);

[0051] 2. Refrigerator door (right side of the sink and food preparation area);

[0052] 3. Dining area (located in the lower left area of ​​the living space);

[0053] 4. The open areas of the living space (to the right of the sink, food preparation area, and dining area);

[0054] 5. TV sofa (on the right side of the open area);

[0055] 6. The television itself;

[0056] 7. Table; and

[0057] 8. Doorway or entrance passage (in the upper right area of ​​the living space).

[0058] According to some embodiments, a system for estimating where a sound (e.g., a wake word or other signal used for focus) is generated or originates may have a certain level of confidence (or multiple assumptions) in the estimation. For example, if the user happens to be near the boundary between zones in the system environment, then an uncertain estimate of the user's location may include a certain level of confidence in the user's location in each of the zones. In some conventional implementations of voice interfaces, it is required that the voice assistant's voice can only be emitted from one location at a time, which forces a single selection of a single location (e.g., Figure 1A (One of the eight speaker positions 1.1 and 1.3). However, based on simple imaginative role-playing, it is obvious that (in such conventional implementations) the selected location of the assistant's voice source (e.g., the location of the speaker contained in the assistant or configured to communicate with the assistant) is unlikely to be a natural return response that expresses focus or attention.

[0059] Next, refer to Figure 1B We describe another environment 109 (acoustic space) that includes a user (101) speaking direct speech 102, and an example of a system including a set of intelligent audio devices (103 and 105), speakers for audio output, and a microphone. The system can be configured according to embodiments of this disclosure. The speech spoken by user 101 (sometimes referred to herein as the speaker) can be recognized as a wake word by elements of the system.

[0060] To be more specific, Figure 1B The system components include:

[0061] 102: Direct local voice (generated by user 101);

[0062] 103: Voice assistant device (coupled to one or more speakers). Device 103 is positioned closer to user 101 than device 105, and therefore device 103 is sometimes referred to as the "near" device, and device 105 is referred to as the "far" device;

[0063] 104: Multiple microphones in (or coupled to) the proximity device 103;

[0064] 105: Voice assistant device (coupled to one or more speakers);

[0065] 106: Multiple microphones in (or coupled to) remote device 105;

[0066] 107: Household appliances (e.g., lamps); and

[0067] 108: Multiple microphones in (or coupled to) household appliance 107. In some instances, each of the microphones 108 may be configured to communicate with a device configured to implement a classifier, which in some examples may be at least one of devices 103 or 105.

[0068] Figure 1B The system may also include at least one classifier (e.g., the one described below). Figure 2 (Classifier 207). For example, device 103 (or device 105) may include a classifier. Alternatively or additionally, the classifier may be implemented by another device configured to communicate with device 103 and / or 105. In some instances, the classifier may be implemented by another local device (e.g., a device within environment 109), while in other instances, the classifier may be implemented by a remote device located outside environment 109 (e.g., a server).

[0069] Figure 1C This is a block diagram illustrating examples of components of a device capable of implementing various aspects of this disclosure. According to some examples, device 100 may be or may include a smart audio device configured to perform at least some of the methods disclosed herein. In other embodiments, device 100 may be or may include another device configured to perform at least some of the methods disclosed herein. In some such embodiments, device 100 may be or may include a server.

[0070] In this example, device 110 includes an interface system 115 and a control system 120. In some embodiments, the interface system 115 may be configured to receive input from each of a plurality of microphones in the environment. The interface system 115 may include one or more network interfaces and / or one or more external device interfaces (e.g., one or more Universal Serial Bus (USB) interfaces). According to some embodiments, the interface system 115 may include one or more wireless interfaces. The interface system 115 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 115 may include a control system 120 and a memory system (e.g., Figure 1C The system 120 may include one or more interfaces between the optional memory systems 125 shown in the diagram. However, the control system 120 may also include a memory system.

[0071] For example, control system 120 may include a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components. In some embodiments, control system 120 may reside in more than one device. For example, a portion of control system 120 may reside in... Figure 1A and 1B The control system 120 may reside in a device within the environment depicted, while another part of the control system 120 may reside in a device outside the environment, such as a server, mobile device (e.g., a smartphone or tablet), etc. In some such instances, the interface system 115 may also reside in more than one device.

[0072] In some implementations, the control system 120 may be configured to perform the methods disclosed herein at least in part. According to some examples, the control system 120 may be configured to implement a classifier, for example, such as the classifier disclosed herein. In some such examples, the control system 120 may be configured to determine an estimate of the user's current user area based at least in part on the output from the classifier.

[0073] Some or all of the methods described herein can be executed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include, for example, the memory devices described herein, including (but not limited to) random access memory (RAM) devices, read-only memory (ROM) devices, etc. For example, one or more non-transitory media may reside on... Figure 1C The optional memory system 125 and / or control system 120 shown herein. Therefore, various innovative aspects of the subject matter described herein can be implemented in one or more non-transitory media having software stored thereon. For example, the software may contain instructions for controlling at least one device to process audio data. For example, the software may be executed by one or more components of the control system, such as... Figure 1C The control system 120.

[0074] Figure 1D An overview can be provided by, for example Figure 1C The flowchart illustrates an example of a method performed by the device shown herein. As with other methods described herein, the blocks of method 130 are not necessarily executed in the indicated order. Furthermore, such methods may contain more or fewer blocks than shown and / or described. In this embodiment, method 130 involves estimating the user's location in the environment.

[0075] In this example, box 135 relates to receiving an output signal from each of a plurality of microphones in the environment. In this example, each of the plurality of microphones resides in a microphone location in the environment. According to this example, the output signal corresponds to the user's current utterance. In some instances, the current utterance may be or may include a wake word utterance. For example, box 135 may relate to a control system (e.g., Figure 1C The control system 120) is connected via an interface system (e.g., Figure 1C The interface system 115) receives output signals from each of the multiple microphones in the environment.

[0076] In some instances, at least some of the microphones in the environment may provide an output signal asynchronous with respect to an output signal provided by one or more other microphones. For example, a first microphone of a plurality of microphones may sample audio data according to a first sampling clock, and a second microphone of a plurality of microphones may sample audio data according to a second sampling clock. In some examples, at least one of the microphones in the environment may be included in or configured to communicate with a smart audio device.

[0077] According to this example, block 140 relates to determining a plurality of current acoustic features from the output signal of each microphone. In this example, the “current acoustic features” are acoustic features derived from the “current utterance” in block 135. In some embodiments, block 140 may relate to receiving a plurality of current acoustic features from one or more other devices. For example, block 140 may relate to receiving at least some of the plurality of current acoustic features from one or more wake word detectors implemented by one or more other devices. Alternatively or additionally, in some embodiments, block 140 may relate to determining a plurality of current acoustic features from the output signal.

[0078] Whether acoustic features are determined by a single device or multiple devices, they can be determined asynchronously. If acoustic features are determined by multiple devices, they will typically be determined asynchronously unless the devices are configured to coordinate the process of determining the acoustic features. If acoustic features are determined by a single device, then in some implementations, the acoustic features can still be determined asynchronously because the single device can receive the output signal from each microphone at different times. In some instances, acoustic features can be determined asynchronously because at least some of the microphones in the environment can provide output signals asynchronously relative to the output signals provided by one or more of these microphones.

[0079] In some instances, acoustic features may include a wake-word confidence metric, a wake-word duration metric, and / or at least one receive level metric. The receive level metric may indicate the received level of sound detected by the microphone and may correspond to the level of the microphone's output signal.

[0080] Alternatively or additionally, acoustic features may include one or more of the following:

[0081] • The average state entropy (purity) of each wake word state aligned with the 1-best (Viterbi) of the acoustic model.

[0082] • CTC loss (connectionist temporal classification loss) for the acoustic model of the wake word detector.

[0083] In addition to wake-word confidence, the wake-word detector can also be trained to provide speaker-microphone distance estimates and / or reverberation time (RT60) estimates. Distance estimates and / or RT60 estimates can be acoustic features.

[0084] • In addition to or excluding the wideband receive level / power at the microphone, the acoustic characteristics can be the receive level in multiple log / Mel / Bark spaced frequency bands. The frequency bands can vary depending on the specific implementation (e.g., 2 bands, 5 bands, 20 bands, 50 bands, 1 octave band, or 1 / 3 octave band).

[0085] • The cepstral representation of the spectral information at the previous point is calculated by obtaining the logarithmic DCT (discrete cosine transform) of the band power.

[0086] • Band power within a frequency band weighted for human speech. For example, acoustic features might be based only on a specific frequency band (e.g., 400 Hz to 1.5 kHz). In this instance, higher and lower frequencies can be ignored.

[0087] • Confidence of the speech activity detector per frequency band or per-bin.

[0088] • Acoustic characteristics may be based at least in part on long-term noise estimates to ignore microphones with poor signal-to-noise ratios.

[0089] Kurtosis is a measure of the peak value of a speech pattern. Kurtosis can also be an indicator of the tailing effect of a long reverberation.

[0090] • Estimated wake word start time. Similar to duration, one would expect the start and duration to be equal within a frame or therefore across all microphones. Outliers can provide clues to unreliable estimates. This assumption, rather than some degree of synchronization over frames of tens of milliseconds, might be reasonable.

[0091] Based on this example, box 145 relates to applying a classifier to multiple current acoustic features. In some such instances, applying a classifier might involve applying a model trained on previously determined acoustic features derived from multiple previous utterances made by a user in multiple user zones within an environment. Various examples are provided herein.

[0092] In some instances, the user area may include a sink area, food preparation area, refrigerator area, dining area, sofa area, television area, bedroom area, and / or porch area. In some instances, one or more user areas may be pre-defined. In some such instances, one or more pre-defined user areas may have been selected by the user during the training process.

[0093] In some implementations, applying a classifier may involve applying a Gaussian mixture model trained on previous utterances. According to some such implementations, applying a classifier may involve applying a Gaussian mixture model trained on one or more of the normalized wake-word confidence, normalized average received level, or maximum received level of the previous utterance. However, in alternative implementations, the classifier may be based on a different model, such as one of the other models disclosed herein. In some examples, the model may be trained using training data labeled with user regions. However, in some instances, the classifier involves applying a model trained using unlabeled training data without labeled user regions.

[0094] In some instances, the preceding utterance may already be, or may already contain, the wake word utterance. According to some such instances, the preceding utterance and the current utterance may already be utterances of the same wake word.

[0095] In this example, box 150 relates to determining an estimate of the user's current location based at least in part on the output from the classifier. In some such instances, the estimate can be determined without reference to the geometric positions of multiple microphones. For example, the estimate can be determined without reference to the coordinates of individual microphones. In some instances, the estimate can be determined without estimating the user's geometric position.

[0096] Some embodiments of method 130 may involve selecting at least one speaker based on an estimated user area. Some such embodiments may involve controlling at least one selected speaker to provide sound to the estimated user area. Alternatively or additionally, some embodiments of method 130 may involve selecting at least one microphone based on the estimated user area. Some such embodiments may involve providing a signal output by at least one microphone to a smart audio device.

[0097] refer to Figure 2 We will now describe embodiments of this disclosure. Figure 2 This is a block diagram of elements of an example of an embodiment configured to implement a region classifier. According to this example, system 200 includes components distributed in the environment (for example, e.g., Figure 1A or Figure 1BThe system 200 comprises at least a portion of the environment described herein, including multiple speakers 204. In this example, the system 200 includes a multi-channel speaker renderer 201. According to this embodiment, the output of the multi-channel speaker renderer 201 is used both as a speaker drive signal (for driving the speaker feeds of the speakers 204) and as an echo reference. In this embodiment, the echo reference is provided to the echo management subsystem 203 via multiple speaker reference channels 202, which include at least some of the speaker feed signals output from the renderer 202.

[0098] In this embodiment, system 200 includes multiple echo management subsystems 203. According to this example, the echo management subsystems 203 are configured to implement one or more echo suppression processes and / or one or more echo cancellation processes. In this example, each of the echo management subsystems 203 provides a corresponding echo management output 203A to one of the wake-word detectors 206. The echo management output 203A has an attenuated echo relative to the input of the corresponding echo management subsystem 203.

[0099] According to this implementation scheme, system 200 includes components distributed in the environment (e.g., Figure 1A or Figure 1B The environment described herein contains at least N microphones 205 (N is an integer). The microphones may include array microphones and / or spot microphones. For example, one or more smart audio devices located in the environment may include a microphone array. In this example, the outputs of the microphones 205 are provided as inputs to the echo management subsystem 203. According to this embodiment, each of the echo management subsystems 203 captures the output of an individual microphone 205 or an individual group or subset of microphones 205.

[0100] In this example, system 200 includes multiple wake word detectors 206. According to this example, each of the wake word detectors 206 receives audio output from one of the echo management subsystems 203 and outputs multiple acoustic features 206A. The acoustic features 206A output from each echo management subsystem 203 may include (but are not limited to): wake word confidence, wake word duration, and a measure of received level. Although the three arrows depicting the three acoustic features 206A are shown as output from each echo management subsystem 203, in alternative embodiments, more or fewer acoustic features 206A may be output. Furthermore, although the three arrows more or less strike the classifier 207 along a vertical line, this does not indicate that the classifier 207 must receive acoustic features 206A from all the wake word detectors 206 simultaneously. As mentioned elsewhere herein, in some examples, acoustic features 206A may be determined and / or provided to the classifier asynchronously.

[0101] According to this implementation, system 200 includes a region classifier 207, which may also be referred to as classifier 207. In this example, the classifier receives multiple features 206A from multiple wake word detectors 206 used for multiple (e.g., all) microphones 205 in the environment. According to this example, the output 208 of the region classifier 207 corresponds to an estimate of the user's current user region. According to some such examples, the output 208 may correspond to one or more posterior probabilities. Based on Bayesian statistics, the estimate of the user's current user region may be or may correspond to the maximum posterior probability.

[0102] We will now describe an example implementation of the classifier, in which, in some instances, the classifier may correspond to... Figure 2 The region classifier is 207. Let x be... i (n) represents the i-th microphone signal in discrete time n, where i = {1…N} (i.e., microphone signal x). i (n) represents the output of N microphones 205). In the echo management subsystem 203, N signals x... i The processing of (n) produces a 'clean' microphone signal e. i (n), where i = {1…N}, each occurring in discrete time n. In this example, Figure 2 The clean signal e referred to as 203A in the middle i (n) is fed into the wake word detector 206. Here, each wake word detector 206 generates a feature w. i The vector of (j), in Figure 2 The term 206A is used to refer to the utterance of the j-th wake word, where j = {1…J} is the index corresponding to the j-th wake word utterance. In this example, classifier 207 obtains the total feature set. As input.

[0103] According to some implementation schemes, a set of area labels C k (For k = {1…K}) can correspond to the number K of different user areas in the environment. For example, user areas may include a sofa area, a kitchen area, a reading chair area, etc. Some instances may define more than one area within a kitchen or other room. For example, a kitchen area may include a sink area, a food preparation area, a refrigerator area, and a dining area. Similarly, a living room area may include a sofa area, a TV area, a reading chair area, one or more porch areas, etc. These areas can be labeled by the user, for example, during the training phase.

[0104] In some implementations, classifier 207 estimates the posterior probability p(C) of the feature set W(j) for example by using a Bayesian classifier. k |W(j)). Probability p(C) k |W(j)) indicates that the user is in zone C kThe probability of each of the following (for the "j"th utterance and the "k"th region, for region C) k Each of the elements in the text and each of the elements in the discourse, and is an instance of the output 208 of classifier 207.

[0105] Based on some examples, training data (e.g., for each user area) can be collected by prompting users to select or define areas (e.g., a sofa area). The training process may involve prompting the user to say a training phrase, such as a wake word, near the selected or defined area. In the sofa area example, the training process may involve prompting the user to say the training phrase at the center and end edges of the sofa. The training process may involve prompting the user to repeat the training phrase several times at each location within the user area. Then, the user may be prompted to move to another user area and continue until all designated user areas have been covered.

[0106] Figure 3 An overview can be provided by, for example Figure 1C A flowchart of an example of a method executed by device 110. As with other methods described herein, the boxes in method 300 are not necessarily executed in the indicated order. Furthermore, such methods may contain more or fewer boxes than shown and / or described. In this embodiment, method 300 involves training a classifier to estimate the user's position in the environment.

[0107] In this example, box 305 relates to prompting the user to say a training utterance at least once in each of several locations within a first user area of ​​the environment. In some instances, the training utterance may be one or more instances of a wake-up word utterance. According to some implementations, the first user area may be any user area selected and / or defined by the user. In some instances, the control system may create corresponding area labels (e.g., area label C described above). k (a corresponding example of one of them), and can associate the area label with the training data obtained for the first user area.

[0108] An automatic prompting system can be used to collect this training data. As mentioned above, the interface system 115 of device 110 may include one or more means for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. For example, device 110 may provide the user with the following prompts on the screen of the display system, or hear them spoken via one or more speakers during the training process:

[0109] • Move to the sofa.

[0110] • Say the wake word ten times while turning your head.

[0111] • Move to a position halfway between the sofa and the reading chair, and say the wake word ten times.

[0112] • “Stand in the kitchen as if you are cooking, and say the wake word ten times.”

[0113] In this example, block 310 relates to receiving a first output signal from each of a plurality of microphones in the environment. In some instances, block 310 may relate to receiving the first output signal from all active microphones in the environment, while in other instances, block 310 may relate to receiving the first output signal from a subset of all active microphones in the environment. In some instances, at least some of the microphones in the environment may provide an output signal asynchronous to an output signal provided by one or more other microphones. For example, a first microphone of the plurality of microphones may sample audio data according to a first sampling clock, and a second microphone of the plurality of microphones may sample audio data according to a second sampling clock.

[0114] In this example, each of the multiple microphones resides at a microphone location in the environment. In this example, the first output signal corresponds to an instance of the detected training utterance received from the first user area. Since box 305 relates to prompting the user to speak the training utterance at least once at each of the multiple locations within the first user area of ​​the environment, in this example, the term "first output signal" refers to a set of all output signals corresponding to the training utterance in the first user area. In other examples, the term "first output signal" may refer to a subset of all output signals corresponding to the training utterance in the first user area.

[0115] According to this example, block 315 relates to determining one or more first acoustic features from each of the first output signals. In some instances, the first acoustic features may include a wake-up word confidence metric and / or a received level metric. For example, the first acoustic features may include a normalized wake-up word confidence metric, an indication of a normalized average received level, and / or an indication of a maximum received level.

[0116] As mentioned above, since box 305 involves prompting the user to speak the training utterance at least once in each of the multiple locations within the first user area of ​​the environment, in this example, the term "first output signal" refers to a set of all output signals corresponding to the training utterance in the first user area. Therefore, in this example, the term "first acoustic feature" refers to a set of acoustic features derived from the set of all output signals corresponding to the training utterance in the first user area. Thus, in this example, the set of first acoustic features is at least as large as the set of first output signals. For example, if two acoustic features are determined from each of the output signals, then the set of first acoustic features will be twice as large as the set of first output signals.

[0117] In this example, box 320 relates to training a classifier model to correlate a first user area with a first acoustic feature. For example, the classifier model could be any of the models disclosed herein. According to this embodiment, the classifier model is trained without reference to the geometric positions of the multiple microphones. In other words, in this example, during training, no data regarding the geometric positions of the multiple microphones (e.g., microphone coordinate data) is provided to the classifier model.

[0118] Figure 4 An overview can be provided by, for example Figure 1C A flowchart of another example of the method performed by device 110. As with other methods described herein, the boxes of method 400 are not necessarily executed in the indicated order. For example, in some embodiments, at least a portion of the acoustic feature determination process for box 425 may be performed before box 415 or box 420. Furthermore, this method may include more or fewer boxes than shown and / or described. In this embodiment, method 400 relates to training a classifier for estimating the location of a user in an environment. Method 400 provides examples of extending method 300 to multiple user areas of the environment.

[0119] In this example, box 405 involves prompting the user to say the training utterance at least once within the user area of ​​the environment. In some examples, box 405 can be referenced above. Figure 3 The procedure described in box 305 is performed in any manner other than those described in box 405 concerning a single location within the user area. In some instances, the training utterance may be one or more instances of a wake-up word utterance. According to some implementations, the user area may be any user area selected and / or defined by the user. In some instances, the control system may create corresponding area labels (e.g., area label C described above). k (a corresponding example of one of them), and can associate the area label with training data obtained for the user area.

[0120] Based on this example, it is roughly as described in the reference above. Figure 3 Box 410 is generally performed as described in box 310. However, in this example, the procedure in box 410 is generalized to any user area, not necessarily to the first user area for which training data is acquired. Therefore, the output signal received in box 410 is "the output signal from each of a plurality of microphones in the environment, each of which resides at a microphone location in the environment, and the output signal corresponds to an instance of the detected training utterance received from the user area." In this example, the term "output signal" refers to a set of all output signals corresponding to one or more training utterances at a location in the user area. In other examples, the term "output signal" may refer to a subset of all output signals corresponding to one or more training utterances at a location in the user area.

[0121] According to this example, box 415 relates to determining whether sufficient training data has been acquired for the current user area. In some such examples, box 415 may relate to determining whether an output signal corresponding to a threshold number of training utterances has been acquired for the current user area. Alternatively or additionally, box 415 may relate to determining whether an output signal for training utterances at a location corresponding to a threshold number within the current user area has been acquired. If not, then in this example, method 400 returns to box 405 and prompts the user to speak at least one additional utterance at the location within the same user area.

[0122] However, if it is determined in box 415 that sufficient training data has been acquired for the current user area, then in this example, the process continues to box 420. According to this example, box 420 involves determining whether training data should be acquired for additional user areas. According to some examples, box 420 may involve determining whether training data has been acquired for each user area previously identified by the user. In other examples, box 420 may involve determining whether training data has been acquired for a minimum number of user areas. This minimum number may have been selected by the user. In other examples, the minimum number may be a suggested minimum number per environment, a suggested minimum number of rooms per environment, etc.

[0123] If it is determined in box 420 that training data should be obtained for an additional user area, then in this instance, the process continues to box 422, which involves prompting the user to move to another user area in the environment. In some instances, the user may be allowed to select the next user area. According to this instance, after the prompt in box 422, the process continues to box 405. In some such instances, after the prompt in box 422, the user may be prompted to confirm that they have arrived in the new user area. According to some such instances, the user may need to confirm that they have arrived in the new user area before the prompt in box 405 is provided.

[0124] If it is determined in box 420 that training data should not be obtained for additional user areas, then in this example, the process continues to box 425. In this example, method 400 involves acquiring training data for K user areas. In this embodiment, box 425 involves determining first to G acoustic features from first to H output signals corresponding to each of the first to Kth user areas for which training data has been obtained. In this example, the term "first output signal" refers to a set of all output signals corresponding to the training utterance for the first user area, and the term "Hth output signal" refers to a set of all output signals corresponding to the training utterance for the Kth user area. Similarly, the term "first acoustic feature" refers to a set of acoustic features determined from the first output signal, and the term "Gth acoustic feature" refers to a set of acoustic features determined from the Hth output signal.

[0125] Based on these examples, box 430 relates to training a classifier model to correlate the first through Kth user regions with the first through Kth acoustic features, respectively. For example, the classifier model can be any of the classifier models disclosed herein.

[0126] In the aforementioned example, the user area is labeled (e.g., according to the area label C described above). k (The corresponding example of one of them). However, depending on the specific implementation, the model can be trained based on labeled or unlabeled user regions. In the case of labeling, each training utterance can be paired with a label corresponding to a user region, for example, as follows:

[0127]

[0128] Training a classifier model may involve determining the best fit for labeled training data. Without loss of generality, appropriate classification methods for a classifier model may include:

[0129] • Bayesian classifiers, for example, have each class distribution described by a multivariate normal distribution, a full covariance Gaussian mixture model, or a diagonal covariance Gaussian mixture model;

[0130] • Vector quantization;

[0131] • Nearest neighbor (k-means);

[0132] • A neural network with a softmax output layer, where each output corresponds to a class;

[0133] • Support Vector Machine (SVM); and / or

[0134] • Propulsion technologies, such as gradient propulsion machines (GBM)

[0135] In one instance of implementing the unlabeled case, the data can be automatically partitioned into K clusters, where K can also be unknown. For example, unlabeled automatic partitioning can be performed using classic clustering techniques such as k-means or Gaussian mixture modeling.

[0136] To improve robustness, regularization can be applied to classifier model training, and the model parameters can be updated as new utterances are made over time.

[0137] We will now describe further aspects of the embodiments.

[0138] Instance acoustic feature sets (e.g., Figure 2The acoustic features (206A) may include the probability of the wake-up word confidence, the average received level over the estimated duration of the most confident wake-up word, and the maximum received level over the duration of the most confident wake-up word. Features can be normalized relative to their maximum value for each wake-up word utterance. Training data can be labeled, and a fully covariant Gaussian mixture model (GMM) can be trained to maximize the expected value of the training labels. The estimation region can be the class that maximizes the posterior probability.

[0139] The above description of some embodiments discusses learning an acoustic zone model from a set of training data collected during the cue collection process. In this model, training time (or configuration mode) and runtime (or normal mode) can be viewed as two different modes in which the microphone system can be placed. An extension of this approach is online learning, where some or all of the acoustic zone model is learned online or adaptively (i.e., during runtime or in normal mode). In other words, even after a classifier is applied to the "runtime" process to estimate the user's current user zone (e.g., based on...), the acoustic zone model can be learned online. Figure 1D In some implementations of Method 130, the process of training the classifier can continue.

[0140] Figure 5 An overview can be provided by, for example Figure 1C The flowchart shows another example of the method executed by device 110. As with other methods described herein, the boxes in method 500 are not necessarily executed in the indicated order. Furthermore, such methods may contain more or fewer boxes than shown and / or described. In this embodiment, method 500 involves continuous training of a classifier during a “runtime” process of estimating the user’s position in the environment. Method 500 is an example of the online learning paradigm referred to herein.

[0141] In this example, box 505 of method 500 corresponds to boxes 135 to 150 of method 130. Here, box 505 involves providing an estimate of the user's current user region based at least in part on the output from the classifier. According to this embodiment, box 510 involves obtaining implicit or explicit feedback on the estimate of box 505. In box 515, the classifier is updated based on the feedback received in box 505. For example, box 515 may involve one or more reinforcement learning methods. As implied by the dashed arrow from box 515 to box 505, in some embodiments, method 500 may involve returning to box 505. For example, method 500 may involve providing a future estimate of the user's user region at the future time based on an updated model.

[0142] Explicit techniques for obtaining feedback may include:

[0143] • Use a voice user interface (UI) to ask the user if the prediction is correct. For example, you could provide the user with a voice indicating the following: "I think you are on the sofa, please say 'yes' or 'no'."

[0144] • Use the voice UI to notify users at any time to correct incorrect predictions. (For example, you could provide the user with a voice indicating the following: "I am now able to predict where you are when you are talking to me. If I am wrong, I can say something like, 'Amanda, I am not on the sofa. I am in the reading chair.'")

[0145] • Use the voice UI to notify users that correct predictions can be rewarded. (For example, you could provide a voice prompt indicating something like, "I am now able to predict where you are when you are talking to me. If I am right, you can help improve my prediction by saying something like, 'Amanda, yes. I am on the sofa.'")

[0146] • Includes physical buttons or other UI elements that can be operated by the user to provide feedback (e.g., thumb up and / or thumb down buttons on a physical device or in a smartphone app).

[0147] The target for predicting the user's location within the user's acoustic zone can be a notification microphone selection or adaptive beamforming scheme, which attempts to more effectively pick up sound from the user's acoustic zone, for example, to better recognize commands following a wake word. In such scenarios, implicit techniques for obtaining feedback on the quality of the zone prediction may include:

[0148] • Penalties lead to misrecognition of the predicted command following the wake word. Agents that might indicate misrecognition could involve the user interrupting the voice assistant's response to a command, for example, by saying the opposite command, such as "Amanda, stop!";

[0149] • Penalty results in a low-confidence prediction that the speech recognizer has successfully identified the command. Many automatic speech recognition systems have the ability to return to the confidence level, and the results can be used for this purpose;

[0150] • The penalty causes the second-pass wake word detector to fail to retrospectively detect wake word predictions with high confidence; and / or

[0151] • Enhanced predictions that lead to highly reliable recognition of wake words and / or correct recognition of user commands.

[0152] The following is an example of a second-pass wake word detector failing to detect the wake word retrospectively with high confidence. Assume that after obtaining an output signal corresponding to the current utterance from a microphone in the environment, and after determining acoustic features based on the output signal (e.g., via multiple first-pass wake word detectors configured to communicate with the microphone), the acoustic features are provided to a classifier. In other words, assume the acoustic features correspond to the detected wake word utterance. Further assume the classifier determines that the person uttering the current utterance is most likely located in zone 3, which corresponds to the reading chair in this example. For example, there may be a known specific microphone or known combination of microphones best suited for listening to a person's voice when they are in zone 3, for example, to be sent to a cloud-based virtual assistant service for voice command recognition.

[0153] Further assuming that, after determining which microphone(s) will be used for speech recognition, but before the person's voice is actually sent to the virtual assistant service, a second wake-word detector operates on the microphone signals corresponding to the speech detected by the selected microphone in region 3, which you will be submitting for command recognition. If the second wake-word detector is inconsistent with the multiple first-pass detectors that actually utter the wake-word, it is likely because the classifier mispredicted the region. Therefore, the classifier should be penalized.

[0154] Techniques for posteriorly updating the region mapping model after one or more wake words have been spoken may include:

[0155] • Adaptive maximization of posterior probability (MAP) using Gaussian mixture models (GMMs) or nearest neighbor models; and / or

[0156] • For example, reinforcement learning in neural networks, such as determining new network weights by associating appropriate “one-hot” (in the case of correct prediction) or “one-cold” (in the case of incorrect prediction) ground truth labels with the softmax output and applying online backpropagation.

[0157] In this context, some instances of MAP adaptation might involve adjusting the means in the GMM each time a wake word is spoken. In this way, these means might become more like the acoustic features observed when subsequent words are spoken. Alternatively or additionally, such instances might involve adjusting the variance / covariance or mixed weight information in the GMM each time a wake word is spoken.

[0158] For example, the MAP adaptive scheme can be as follows:

[0159] μ i,new =μi i,old *α+x*(1-α)

[0160] In the aforementioned equation, μi,old Let represent the mean of the i-th Gaussian in the mixture, α represent a parameter controlling the positivity of the MAP adaptation (α may be in the range [0.9, 0.999]), and x represent the feature vector of the new wake-word utterance. The exponent "i" will correspond to the mixture element, which returns the highest prior probability containing the speaker's position at the wake-word time.

[0161] Alternatively, each of the elements in the mixture can be adjusted based on its prior probability of containing a wake word, for example, as follows:

[0162] M i,new =μ i,old *β i *x(1-β i )

[0163] In the aforementioned equation, βi=α*(1-P(i)), where P(i) represents the prior probability that observation x is caused by the mixture element i.

[0164] In a reinforcement learning instance, there might be three user regions. Suppose that for a given wake word, the model predicts the probabilities of the three user regions as [0.2, 0.1, 0.7]. If a second information source (e.g., a second wake word detector) confirms that the third region is correct, then the ground truth label could be [0, 0, 1] (“one-hot”). The posterior update of the region mapping model might involve backpropagating the error through the neural network, which essentially means that if the same input is shown again, the neural network will predict region 3 more strongly. Conversely, if the second information source shows that region 3 is an incorrect prediction, then in one instance, the ground truth label could be [0.5, 0.5, 0.0]. If the same input is shown in the future, then backpropagating the error through the neural network will make the model less likely to predict region 3.

[0165] Some aspects of the embodiments include one or more of the following:

[0166] Example 1. A method for estimating the location of a user in an environment (e.g., as a zone label), wherein the environment comprises a plurality of user zones and a plurality of microphones (e.g., each of the microphones is included in or configured to communicate with at least one smart audio device in the environment), the method comprising the steps of: (e.g., determining an estimate of one of the user zones in which the user is located, at least in part from the output signals of the microphones);

[0167] Example 2. The method according to Example 1, wherein the microphones are asynchronous and / or randomly distributed;

[0168] Example 3. According to the method of Example 1, wherein determining the estimation involves applying a classifier model that has been trained on acoustic features derived from multiple wake word detectors, the acoustic features being based on multiple wake word utterances in multiple locations;

[0169] Example 4. The method described in Example 3, wherein the user region is estimated as the class with the highest posterior probability;

[0170] Example 5. The method described in Example 3, wherein the classifier model is trained using training data labeled with a reference region;

[0171] Example 6. The method according to Example 1, wherein the classifier model is trained using unlabeled training data;

[0172] Example 7. The method according to Example 1, wherein the classifier model includes a Gaussian mixture model trained based on normalized wake word confidence, normalized average received level and maximum received level;

[0173] Example 8. The method described in any of the previous examples, wherein the adaptation of the acoustic region model is performed online;

[0174] Example 9. The method according to Example 8, wherein the adaptation is based on explicit feedback from the user;

[0175] Example 10. The method according to Example 8, wherein the adaptation is based on implicit feedback to the success of beamforming or microphone selection based on the predicted acoustic zone;

[0176] Example 11. The method according to Example 10, wherein the implicit feedback includes a response from the user prematurely terminating the voice assistant;

[0177] Example 12. The method according to Example 10, wherein the implicit feedback includes a command identifier returning a low-confidence result; and

[0178] Example 13. According to the method of Example 10, wherein the implicit feedback includes a second pass of the wake word detector returning a low confidence level of saying the wake word.

[0179] Some embodiments include systems or apparatus configured (e.g., programmed) to perform one or more of the disclosed methods, and tangible computer-readable media (e.g., optical discs) storing code for implementing one or more of the disclosed methods or steps thereof. For example, the system may be or may include a programmable general-purpose processor, digital signal processor, or microprocessor, any of which is programmed with software or firmware and / or otherwise configured to perform various operations on data, including embodiments of the disclosed methods or steps thereof. Such a general-purpose processor may be or may include a computer system comprising input devices, memory, and processing subsystems, which are programmed (and / or otherwise configured) to perform the disclosed methods (or steps thereof) in response to data asserted thereto.

[0180] Some embodiments of the disclosed system may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed and otherwise configured) to perform desired processing on audio signals, including the execution of embodiments of the disclosed methods. Alternatively, embodiments of the disclosed system (or elements thereof) may be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor that may include input devices and memory) programmed with software or firmware and / or otherwise configured to perform any of the various operations including embodiments of the disclosed methods. Alternatively, elements of some embodiments of the disclosed system may be implemented as a general-purpose processor or DSP configured (e.g., programmed) to execute embodiments of the disclosed methods, and the system may also include other elements (e.g., one or more speakers and / or one or more microphones). A general-purpose processor configured to execute embodiments of the disclosed methods may be coupled to an input device (e.g., a mouse and / or keyboard), memory, and in some instances, to a display device.

[0181] Another aspect of this disclosure may be implemented in one or more non-transitory computer-readable media (e.g., one or more RAMs, ROMs, optical discs, or other tangible storage media) that store code for performing (e.g., an executable encoder to perform) any embodiment of the disclosed methods or steps thereof.

[0182] While specific embodiments and applications of this disclosure have been described herein, it will be apparent to those skilled in the art that many variations of the embodiments and applications described herein are possible without departing from the scope of this disclosure.

[0183] Various aspects of the present invention can be understood from the following exemplary embodiments (EEE):

[0184] EEE 1. A method for estimating the location of a user in an environment, the method comprising:

[0185] An output signal is received from each of a plurality of microphones in the environment, each of the plurality of microphones residing at a microphone location in the environment, the output signal corresponding to the user’s current speech;

[0186] Multiple current acoustic features are determined from the output signal of each microphone;

[0187] Applying a classifier to the plurality of current acoustic features, wherein applying the classifier involves applying a model trained on previously determined acoustic features, the previously determined acoustic features being derived from multiple previous utterances spoken by the user in multiple user zones within the environment; and

[0188] The estimate of the user region where the user is currently located is determined at least in part based on the output from the classifier.

[0189] EEE 2. The method according to EEE 1, wherein at least one of the microphones is included in or configured to communicate with the smart audio device.

[0190] EEE 3. The method according to EEE 1, wherein the plurality of user areas includes a plurality of predetermined user areas.

[0191] EEE 4. The method according to EEE 1, wherein the estimation is determined without reference to the geometric positions of the plurality of microphones.

[0192] EEE 5. The method according to any one of EEEs 1 to 4, wherein the plurality of current acoustic features are determined asynchronously.

[0193] EEE 6. The method according to any one of EEE 1 to 5, wherein the current utterance and the previous utterance include a wake word utterance.

[0194] EEE 7. The method according to any of EEEs 1 to 6, wherein the user region is estimated as the class with the maximum posterior probability.

[0195] EEE 8. The method according to any of EEE 1 to 7, wherein the model is trained using training data labeled with user regions.

[0196] EEE 9. The method according to any of EEEs 1 to 7, wherein applying the classifier involves applying a model trained using unlabeled training data with unlabeled user regions.

[0197] EEE 10. The method according to any of EEEs 1 to 9, wherein applying the classifier involves applying a Gaussian mixture model trained on one or more of normalized wake word confidence, normalized average received level, or maximum received level.

[0198] EEE 11. The method according to any one of EEE 6, 8, 9 or 10, wherein the training of the model continues during the process of applying the classifier.

[0199] EEE 12. The method according to EEE 11, wherein the training is based on explicit feedback from the user.

[0200] EEE 13. The method according to EEE 11, wherein the training is based on implicit feedback of successful beamforming or microphone selection based on estimated user area.

[0201] EEE 14. The method according to EEE 13, wherein the implicit feedback includes a determination that the user has abnormally terminated the response of the voice assistant.

[0202] EEE 15. The method according to EEE 13, wherein the implicit feedback includes a command identifier returning a low-confidence result.

[0203] EEE 16. The method according to EEE 13, wherein the implicit feedback includes a second pass of the wake word detector returning a low confidence level of saying the wake word.

[0204] EEE 17. The method according to any of EEEs 1 to 16, wherein the user area comprises one or more of a sink area, a food preparation area, a refrigerator area, a dining area, a sofa area, a television area, or a porch area.

[0205] EEE 18. The method according to any of EEEs 1 to 17, further comprising selecting at least one speaker based on the estimated user area, and controlling the at least one speaker to provide sound to the estimated user area.

[0206] EEE 19. The method according to any one of EEEs 1 to 18, further comprising selecting at least one microphone based on the estimated user area and providing a signal output by the at least one microphone to a smart audio device.

[0207] EEE 20. The method according to any one of EEEs 1 to 19, wherein a first microphone of the plurality of microphones samples audio data according to a first sampling clock, and a second microphone of the plurality of microphones samples audio data according to a second sampling clock.

[0208] EEE 21. An apparatus configured to perform the method described in any of EEEs 1 to 20.

[0209] EEE 22. A system configured to perform the method described in any of EEEs 1 to 20.

[0210] EEE 23. One or more non-transitory media having software stored thereon, the software containing instructions for controlling one or more devices to perform the methods described in any of EEEs 1 to 20.

[0211] EE 24. An apparatus comprising:

[0212] An interface system configured to receive output signals from each of a plurality of microphones in an environment, each of the plurality of microphones residing at a microphone location in the environment, the output signals corresponding to a user's current speech; and

[0213] Control system, configured for:

[0214] Multiple current acoustic features are determined from the output signal of each microphone;

[0215] Applying a classifier to the plurality of current acoustic features, wherein applying the classifier involves applying a model trained on previously determined acoustic features derived from multiple previous utterances spoken by the user in multiple user zones within the environment; and

[0216] The estimate of the user region where the user is currently located is determined at least in part based on the output from the classifier.

[0217] EE 25. An apparatus comprising:

[0218] An interface system configured to receive output signals from each of a plurality of microphones in an environment, each of the plurality of microphones residing at a microphone location in the environment, the output signals corresponding to a user's current speech; and

[0219] Control components, which are used for:

[0220] Multiple current acoustic features are determined from the output signal of each microphone;

[0221] Applying a classifier to the plurality of current acoustic features, wherein applying the classifier involves applying a model trained on previously determined acoustic features derived from multiple previous utterances spoken by the user in multiple user zones within the environment; and

[0222] The estimate of the user region where the user is currently located is determined at least in part based on the output from the classifier.

[0223] EEE 26. A training method comprising:

[0224] The user is prompted to say the training utterance at least once in each of several locations within the first user area of ​​the environment.

[0225] A first output signal is received from each of a plurality of microphones in the environment, each of the plurality of microphones residing in a microphone position in the environment, the first output signal corresponding to an instance of a detected training utterance received from the first user area;

[0226] Determine a first acoustic feature from each of the first output signals; and

[0227] A classifier model is trained to correlate the first user area with the first acoustic feature, wherein the classifier model is trained without reference to the geometric positions of the plurality of microphones.

[0228] EEE 27. The training method according to EEE 26, wherein the training utterances include wake word utterances.

[0229] EEE 28. The training method according to EEE 27, wherein the first acoustic feature includes one or more of normalized wake word confidence, normalized average received level, or maximum received level.

[0230] EEE 29. The method according to any one of EEE 26 to 28, further comprising:

[0231] The user is prompted to speak the training utterance in each of multiple locations within the second to Kth user zones of the environment.

[0232] The second to the Hth output signals are received from each of a plurality of microphones in the environment, the second to the Hth output signals respectively corresponding to instances of the detected training utterance received from the second to the Kth user areas;

[0233] Determine the second to the Gth acoustic features from each of the second to the Hth output signals; and

[0234] The classifier model is trained to correlate the second to the Kth user regions with the second to the Gth acoustic features, respectively.

[0235] EEE 30. The method according to any one of EEEs 26 to 29, wherein a first microphone of the plurality of microphones samples audio data according to a first sampling clock, and a second microphone of the plurality of microphones samples audio data according to a second sampling clock.

[0236] EEE 31. A system configured to perform the method described in any of EEEs 26 to 30.

[0237] EEE 32. An apparatus configured to perform the method described in any of EEEs 26 to 30.

[0238] EEE 33. An apparatus comprising:

[0239] A component used to prompt the user to say the training utterance at least once in each of multiple locations within the first user area of ​​the environment;

[0240] A component for receiving a first output signal from each of a plurality of microphones in the environment, each of the plurality of microphones residing in a microphone position in the environment, the first output signal corresponding to an instance of a detected training utterance received from the first user area.

[0241] Components for determining a first acoustic feature from each of the first output signals; and

[0242] A component for training a classifier model to correlate the first user area with the first acoustic feature, wherein the classifier model is trained without reference to the geometric positions of the plurality of microphones.

[0243] EEE 34. An apparatus comprising:

[0244] A user interface system, comprising at least one of a display or a speaker; and

[0245] Control system, configured for:

[0246] The user interface system is controlled to prompt the user to say at least one training utterance in each of multiple locations within a first user area of ​​the environment.

[0247] A first output signal is received from each of a plurality of microphones in the environment, each of the plurality of microphones residing in a microphone position in the environment, the first output signal corresponding to an instance of a detected training utterance received from the first user area;

[0248] Determine a first acoustic feature from each of the first output signals; and

[0249] A classifier model is trained to correlate the first user area with the first acoustic feature, wherein the classifier model is trained without reference to the geometric positions of the plurality of microphones.

Claims

1. A computer-implemented method for estimating a user's location in an environment, the method comprising: An output signal is received from each of a plurality of microphones in the environment, wherein at least two of the plurality of microphones are contained in separate devices at separate locations in the environment, and the output signal corresponds to the user’s current speech. Multiple current acoustic features are determined from the output signal of each microphone; and Applying a classifier to the plurality of current acoustic features involves applying a model trained on previously determined acoustic features to associate multiple user areas in the environment with the plurality of acoustic features, the previously determined acoustic features being determined from the output signal of each microphone corresponding to multiple previous utterances spoken by the user in the multiple user areas of the environment. The output of the classifier provides an estimate of the user region where the user is currently located. Wherein the current utterance and the previous utterance each include utterances of the same wake word, and the acoustic features include a measure of wake word duration, and The first microphone of the plurality of microphones samples the audio data according to a first sampling clock, and the second microphone of the plurality of microphones samples the audio data according to a second sampling clock.

2. The computer-implemented method of claim 1, wherein at least one of the plurality of microphones is included in or configured to communicate with the intelligent audio device.

3. The computer-implemented method according to claim 1 or claim 2, wherein the plurality of user areas comprises a plurality of predetermined user areas.

4. The computer-implemented method according to claim 1 or claim 2, wherein the estimation is determined without reference to the geometric positions of the plurality of microphones.

5. The computer-implemented method according to claim 1 or claim 2, wherein the plurality of current acoustic features are determined asynchronously.

6. The computer-implemented method according to claim 1 or claim 2, wherein the classifier estimates the posterior probability of each user region, wherein the user region in which the user is currently located is estimated to be the user region with the maximum posterior probability.

7. The computer-implemented method according to claim 1 or claim 2, wherein the model is trained using training data labeled with user areas.

8. The computer-implemented method according to claim 1 or claim 2, wherein the model is trained using unlabeled training data with unlabeled user areas.

9. The computer-implemented method according to claim 1 or claim 2, wherein the acoustic feature includes at least one received level metric, wherein the received level metric indicates the sound level detected by a microphone.

10. The computer-implemented method of claim 9, wherein the model is a Gaussian mixture model trained on one or more of normalized wake word confidence, normalized average received level, or maximum received level.

11. The computer-implemented method of claim 10, wherein the normalized average received level comprises the average received level over the duration of the most trusted wake word.

12. The computer-implemented method of claim 10, wherein the maximum received level includes the maximum received level during the duration of the most trusted wake word.

13. The computer-implemented method according to claim 1 or claim 2, wherein training of the model continues during the process of applying the classifier.

14. The computer-implemented method of claim 13, wherein the training is based on explicit feedback from the user.

15. The computer-implemented method of claim 13, wherein the training is based on implicit feedback of successful beamforming or microphone selection based on estimated user area.

16. The computer-implemented method of claim 15, wherein the implicit feedback comprises at least one of the following: The user abnormally terminated the confirmation of the voice assistant's response; or The command identifier returns a low-confidence result; or The second pass traces back to the wake word detector, which returns a low confidence level for saying the wake word.

17. A computer-implemented method according to claim 1 or claim 2, comprising selecting at least one speaker based on the estimated user area and controlling the at least one speaker to provide sound to the estimated user area.

18. The computer-implemented method according to claim 1 or claim 2, comprising selecting at least one microphone based on the estimated user area and providing a signal output by the at least one microphone to a smart audio device.

19. An apparatus comprising: An interface system configured to receive an output signal from each of a plurality of microphones in an environment, at least two of which are contained in separate devices at separate locations in the environment, the output signal corresponding to the user’s current speech; and Control system, configured for: Multiple current acoustic features are determined from the output signal of each microphone; Applying a classifier to the plurality of current acoustic features involves applying a model trained on previously determined acoustic features to associate multiple user areas in the environment with the plurality of acoustic features, the previously determined acoustic features being determined from the output signal of each microphone corresponding to multiple previous utterances spoken by the user in the multiple user areas of the environment. The output of the classifier provides an estimate of the user region where the user is currently located. Wherein the current utterance and the previous utterance each include utterances of the same wake word, and the acoustic features include a measure of wake word duration, and The first microphone of the plurality of microphones samples the audio data according to a first sampling clock, and the second microphone of the plurality of microphones samples the audio data according to a second sampling clock.

20. A system configured to perform the method according to any one of claims 1 to 18.

21. One or more non-transitory media having software stored thereon, the software comprising instructions for controlling one or more devices to perform the method according to any one of claims 1 to 18.

22. A training method comprising: The user is prompted to say the training utterance at least once in each of several locations within the first user area of ​​the environment. A first output signal is received from each of a plurality of microphones in the environment, at least two of which are contained in separate devices at separate locations in the environment, the first output signal corresponding to an instance of a detected training utterance received from the first user area; Determine a first acoustic feature from each of the first output signals; and A classifier model is trained to correlate the first user area with the first acoustic feature, wherein the classifier model is trained without reference to the geometric positions of the plurality of microphones; Each training utterance includes the same wake word utterance, and the first acoustic feature includes a wake word duration measure; and The first microphone of the plurality of microphones samples the audio data according to a first sampling clock, and the second microphone of the plurality of microphones samples the audio data according to a second sampling clock.

23. The training method of claim 22, wherein the first acoustic feature includes one or more of normalized wake word confidence, normalized average received level, or maximum received level, wherein the received level indicates the sound level detected by the microphone.

24. The training method according to any one of claims 22 to 23, comprising: The user is prompted to speak the training utterance in each of multiple locations within the second to Kth user zones of the environment. The second to the Hth output signals are received from each of a plurality of microphones in the environment, the second to the Hth output signals respectively corresponding to instances of the detected training utterance received from the second to the Kth user areas; Determine the second to the Gth acoustic features from each of the second to the Hth output signals; and The classifier model is trained to correlate the second to the Kth user regions with the second to the Gth acoustic features, respectively.

25. A system configured to perform the method according to any one of claims 22 to 24.

26. An apparatus comprising: User interface system, which includes at least one of a display or a speaker; and Control system, configured for: The user interface system is controlled to prompt the user to say at least one training utterance in each of multiple locations within a first user area of ​​the environment. A first output signal is received from each of a plurality of microphones in the environment, at least two of which are contained in separate devices at separate locations in the environment, the first output signal corresponding to an instance of a detected training utterance received from the first user area; Determine a first acoustic feature from each of the first output signals; and A classifier model is trained to correlate the first user area with the first acoustic feature, wherein the classifier model is trained without reference to the geometric positions of the plurality of microphones. Each training utterance includes the same wake word utterance, and the first acoustic feature includes a wake word duration measure, and The first microphone of the plurality of microphones samples the audio data according to a first sampling clock, and the second microphone of the plurality of microphones samples the audio data according to a second sampling clock.

Citation Information

Patent Citations

  • Speech recognition models based on location indicia

    CN104509079A

  • Establishing microphone zones in a vehicle

    US20160039356A1