Multimodal intelligent audio device system attention expression
By coordinating wake word detection and attention signal generation across multiple intelligent audio devices, the problem of unnatural attention expression in multi-device environments in existing technologies is solved, enabling natural and continuous user interaction and improving device coordination and interaction efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-30
- Publication Date
- 2026-03-17
AI Technical Summary
Existing audio equipment control systems struggle to effectively coordinate user attention in multi-device environments, resulting in unnatural interactions. They are particularly prone to errors in high-noise and mobile environments, and traditional methods often force the selection of a single device while ignoring the potential attention of other devices.
By coordinating wake word detection and attention signal generation across multiple smart audio devices, and dynamically modulating device behavior to provide a continuous expression of attention based on user location and device relevance metrics, the system utilizes a mixture of light and sound signals to express the user's attention state, thus avoiding the need for a single device selection.
It enables natural and continuous attention expression in multi-device environments, improves interaction efficiency and accuracy, reduces dissonance caused by location selection, and enhances the continuity of user interaction with devices.
Smart Images

Figure CN114175145B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to the following applications: U.S. Provisional Patent Application No. 62 / 880,110, filed July 30, 2019; U.S. Provisional Patent Application No. 62 / 880,112, filed July 30, 2019; U.S. Provisional Patent Application No. 62 / 964,018, filed January 21, 2020; and U.S. Provisional Patent Application No. 63 / 003,788, filed April 1, 2020, all of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to systems and methods for use with multiple intelligent audio devices in an automated control environment. Background Technology
[0004] Audio devices, including but not limited to smart audio devices, have been widely deployed and are becoming a common feature in many homes. While existing systems and methods for controlling audio devices offer benefits, improved systems and methods will continue to be desired.
[0005] Symbols and terms
[0006] In this document, the term "smart audio device" is used to refer to a smart device, which is a single-purpose audio device or a virtual assistant (e.g., a connected virtual assistant). A single-purpose audio device is a device that includes or is coupled to at least one microphone (and in some examples may also include or be coupled to at least one speaker) and is largely or primarily designed to perform a single purpose (e.g., a smart speaker, a television (TV), or a mobile phone). Although a TV can generally play (and is considered capable of playing) audio from program material, in most cases, modern TVs run some kind of operating system on which applications (including TV-watching applications) run natively. Similarly, audio input and output in a mobile phone can do many things, but these are served by applications running on the phone. In this sense, a single-purpose audio device with (multiple) speakers and (multiple) microphones is typically configured to run local applications and / or services to directly use (multiple) speakers and (multiple) microphones. Some single-purpose audio devices can be configured to be combined to achieve audio playback in a zone or user-configured zone.
[0007] In this document, a “virtual assistant” (e.g., a connected virtual assistant) is a device (e.g., a smart speaker, smart display, or voice assistant integrated device) that includes or is coupled to at least one microphone (and optionally includes or is coupled to at least one speaker), and that provides the ability to use multiple devices (different from the virtual assistant) to enable, in some sense, cloud-based or applications not implemented in or on the virtual assistant itself. Virtual assistants can sometimes work together, for example, in a very discrete and conditionally defined manner. For example, two or more virtual assistants may work together in response to the meaning of one of them (i.e., the virtual assistant most certain that it has heard the wake word). Connected devices can form a contellation that can be managed by a master application, which may be (or include or implement) the virtual assistant.
[0008] In this document, "wake word" is used broadly to refer to any sound (e.g., a human-spoken word or other sound) in which a smart audio device is configured to wake up in response to the detection ("hearing") of a sound (using at least one microphone included in or coupled to the smart audio device, or at least one other microphone). In this context, "wake up" means that the device enters a state of waiting (i.e., listening) for a sound command. In some cases, what may be referred to as a "wake word" in this document may include more than one word, such as a phrase.
[0009] In this paper, the term "wake word detector" refers to a device (or software including instructions for configuring the device to continuously search for alignments between real-time sound (e.g., speech) features and a training model. Typically, a wake word event is triggered whenever the wake word detector determines that the probability of detecting a wake word exceeds a predefined threshold. For example, this threshold could be a predetermined threshold adjusted to provide a reasonable trade-off between a false acceptance rate and a false rejection rate. Following a wake word event, the device may enter a state (which may be referred to as a "wake-up" state or an "attention" state) in which the device listens for commands and passes the received commands to a larger, more computationally intensive recognizer.
[0010] Throughout this disclosure, including in the claims, the terms "speaker" and "loudspeaker" are used synonymously to refer to any sound-emitting transducer (or a group of transducers) driven by a single loudspeaker feed. A typical set of headphones includes two loudspeakers. A loudspeaker can be implemented to include multiple transducers (e.g., a woofer and a tweeter), all of which are driven by a single common loudspeaker feed. In some cases, the loudspeaker feed may undergo different processing in different circuit branches coupled to different transducers.
[0011] Throughout this disclosure, including in the claims, the expression “operating on” a signal or data (e.g., filtering, scaling, transforming, or applying gain to a signal or data) is used broadly to refer to directly operating on a signal or data or operating on a processed version of a signal or data (e.g., a signal version that has been pre-filtered or pre-processed before being operated on).
[0012] Throughout this disclosure, including in the claims, the term "system" is used broadly to refer to a device, system, or subsystem. For example, a subsystem implementing a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, wherein the subsystem generates M inputs, while the other XM inputs are received from an external source) may also be referred to as a decoder system.
[0013] Throughout this disclosure, including in the claims, the term "processor" is used broadly to refer to a system or device that is programmable or otherwise configurable (e.g., with software or firmware) for performing operations on data (e.g., audio, video, or other image data). Examples of processors include field-programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipelined processing of audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets. Summary of the Invention
[0014] At least some aspects of this disclosure can be implemented via methods such as those for controlling a device system in a control environment. In some cases, the methods can be implemented at least in part by control systems such as those disclosed herein. Some such methods may involve receiving an output signal from each of a plurality of microphones in an environment. Each of the plurality of microphones may be located at a microphone position in the environment. In some examples, the output signal may correspond to human utterance. According to some examples, at least one of the microphones may be included in or configured to communicate with a smart audio device. In some cases, a first microphone of the plurality of microphones may sample audio data according to a first sampling clock, and a second microphone of the plurality of microphones may sample audio data according to a second sampling clock.
[0015] Some such methods may involve determining a zone within an environment, at least in part, based on an output signal, that zone has at least a threshold probability of including the location of a person. Some such methods may involve generating multiple spatially varied attention signals within the zone. In some cases, each of the multiple attention signals may be generated by a device located within the zone. For example, each attention signal may indicate that the corresponding device is in an operating mode where the corresponding device is waiting for a command. In some examples, each attention signal may indicate a relevance metric for the corresponding device.
[0016] In some implementations, the attention signal generated by the first device can indicate a relevance measure of the second device. In some examples, the second device can be a device corresponding to the first device. In some cases, the utterance can be or may include a wake word. According to some such examples, the attention signal varies at least in part based on an estimate of the wake word confidence.
[0017] According to some examples, at least one of the attention signals may be modulation of at least one prior signal generated by a device within the area prior to the time of utterance. In some cases, the at least one prior signal may be or may include an optical signal. According to some such examples, the modulation may be color modulation, color saturation modulation, and / or light intensity modulation.
[0018] In some cases, at least one prior signal may be or may include an audio signal. According to some such examples, modulation may be level modulation. Alternatively or additionally, modulation may be a variation of one or more of the following: fan speed, flame size, motor speed, or airflow rate.
[0019] According to some examples, modulation can be what is referred to herein as a "swell". A swell can be or can include a predetermined signal modulation sequence. In some examples, a swell can include a first time interval corresponding to an increase in the signal level from a baseline level. According to some such examples, a swell can include a second time interval corresponding to a decrease in the signal level to the baseline level. In some cases, a swell can include a hold time interval after the first time interval and before the second time interval. In some cases, the hold time interval can correspond to a constant signal level. In some examples, a swell can include a first time interval corresponding to a decrease in the signal level from the baseline level.
[0020] According to some examples, the correlation measure may be based at least in part on the estimated distance to a location. In some cases, the location may be the estimated location of a person. In some examples, the estimated distance may be the estimated distance from that location to the acoustic centroid of multiple microphones within the area. According to some implementations, the correlation measure may be based at least in part on the estimated visibility of the corresponding device.
[0021] Some such methods may involve automated processes for determining whether a device is in a group of devices. According to some examples, the automated process may be based at least in part on sensor data corresponding to light and / or sound emitted by the device. In some cases, the automated process may be based at least in part on communication between a source and a receiver. For example, the source may be a light source and / or a sound source. According to some examples, the automated process may be based at least in part on communication between the source and an orchestration hub and / or between the receiver and the orchestration hub. In some cases, the automated process may be based at least in part on the switching of a light source and / or a sound source on and off over a period of time.
[0022] Some such methods may involve automatically updating the automation process based on explicit feedback from the human. Alternatively or additionally, some methods may involve automatically updating the automation process based on implicit feedback. For example, implicit feedback may be based on: successful beamforming based on the estimated region, successful microphone selection based on the estimated region, determination that the human has abnormally terminated the voice assistant's response, a command recognizer returning a low-confidence result, and / or a second-pass retrospective wakeword detector returning a low-confidence result indicating that a wakeword has been spoken.
[0023] Some methods may involve selecting at least one speaker of a device located within the area and controlling the at least one speaker to provide sound to a person. Alternatively or additionally, some methods may involve selecting at least one microphone of a device located within the area. Some such methods may involve providing a signal output by at least one microphone to a smart audio device.
[0024] Some or all of the operations, functions, and / or methods described herein can be performed by one or more means according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. Therefore, some innovative aspects of the subject matter described in this disclosure can be implemented in non-transitory media on which software is stored.
[0025] For example, the software may include instructions for controlling one or more devices to perform methods relating to a device system in a control environment. Some such methods may involve receiving an output signal from each of a plurality of microphones in the environment. Each of the plurality of microphones may be located at a microphone position in the environment. In some examples, the output signal may correspond to human speech. According to some examples, at least one of the microphones may be included in or configured to communicate with a smart audio device. In some cases, a first microphone of the plurality of microphones may sample audio data according to a first sampling clock, and a second microphone of the plurality of microphones may sample audio data according to a second sampling clock.
[0026] Some such methods may involve determining an area within the environment, at least in part, based on an output signal, that has at least a threshold probability of including the location of a person. Some such methods may involve generating multiple spatially varied attention signals within the area. In some cases, each of the multiple attention signals may be generated by a device located within the area. For example, each attention signal may indicate that the corresponding device is in an operating mode where the corresponding device is waiting for a command. In some examples, each attention signal may indicate a relevance metric for the corresponding device.
[0027] In some implementations, the attention signal generated by the first device can indicate a relevance measure of the second device. In some examples, the second device can be a device corresponding to the first device. In some cases, the utterance can be or may include a wake word. According to some such examples, the attention signal varies at least in part based on an estimate of the wake word confidence.
[0028] According to some examples, at least one of the attention signals may be modulation of at least one prior signal generated by a device within the area prior to the time of utterance. In some cases, the at least one prior signal may be or may include an optical signal. According to some such examples, the modulation may be color modulation, color saturation modulation, and / or light intensity modulation.
[0029] In some cases, at least one prior signal may be or may include an audio signal. According to some such examples, modulation may be horizontal modulation. Alternatively or additionally, modulation may be a variation of one or more of fan speed, flame size, motor speed, or airflow rate.
[0030] According to some examples, modulation can be what is referred to herein as a "bulge". A bulge can be or can include a predetermined signal modulation sequence. In some examples, a bulge can include a first time interval corresponding to an increase in the signal level from a baseline level. According to some such examples, a bulge can include a second time interval corresponding to a decrease in the signal level to the baseline level. In some cases, a bulge can include a hold time interval after the first time interval and before the second time interval. In some cases, the hold time interval can correspond to a constant signal level. In some examples, a bulge can include a first time interval corresponding to a decrease in the signal level from the baseline level.
[0031] According to some examples, the correlation measure may be based at least in part on the estimated distance to a location. In some cases, the location may be the estimated location of a person. In some examples, the estimated distance may be the estimated distance from that location to the acoustic centroids of multiple microphones within the area. According to some implementations, the correlation measure may be based at least in part on the estimated visibility of the corresponding device.
[0032] Some such methods may involve automated processes for determining whether a device is in a group of devices. According to some examples, the automated process may be based at least in part on sensor data corresponding to light and / or sound emitted by the device. In some cases, the automated process may be based at least in part on communication between a source and a receiver. For example, the source may be a light source and / or a sound source. According to some examples, the automated process may be based at least in part on communication between the source and an orchestration hub and / or between the receiver and the orchestration hub. In some cases, the automated process may be based at least in part on the switching of a light source and / or a sound source on and off over a period of time.
[0033] Some such methods may involve automatically updating the automation process based on explicit feedback from the human. Alternatively or additionally, some methods may involve automatically updating the automation process based on implicit feedback. For example, implicit feedback may be based on: successful beamforming based on the estimated region, successful microphone selection based on the estimated region, determination that the human has abnormally terminated the voice assistant's response, a command recognizer returning a low-confidence result, and / or a second retrospective wake word detector returning a low-confidence result indicating that a wake word was spoken.
[0034] Some methods may involve selecting at least one speaker of a device located within the area and controlling the at least one speaker to provide sound to a person. Alternatively or additionally, some methods may involve selecting at least one microphone of a device located within the area. Some such methods may involve providing a signal output by at least one microphone to a smart audio device.
[0035] At least some aspects of this disclosure can be implemented via apparatus. For example, one or more apparatuses are capable of performing at least partially the methods disclosed herein. In some embodiments, the apparatus may include an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof.
[0036] Details of one or more embodiments of the subject matter described herein are set forth in the following figures and description. Other features, aspects, and advantages will become apparent from the description, figures, and claims. Note that the relative dimensions in the following figures may not be drawn to scale. Attached Figure Description
[0037] Figure 1A This indicates the environment based on an example.
[0038] Figure 1B This indicates the environment based on another example.
[0039] Figure 2 An example of a wake word confidence value curve determined by three devices is shown.
[0040] Figure 3 This is a block diagram illustrating examples of components of an apparatus capable of implementing various aspects of this disclosure.
[0041] Figure 4 It outlines what can be achieved by, for example Figure 3 A flowchart of an example of a method performed by at least one device shown in the diagram.
[0042] Figure 5This is a block diagram illustrating examples of features according to some implementation methods.
[0043] Figure 6 The image shows an example of a raised area.
[0044] Figure 7 An example embodiment of a system for implementing automatic optical orchestration is shown.
[0045] Figure 8 It's a diagram. Figure 7 A set of diagrams illustrating examples of the system's operational aspects.
[0046] In the various figures, the same reference numerals and names indicate similar elements. Detailed Implementation
[0047] Some embodiments relate to a system of orchestrated smart audio devices, wherein each device can be configured to indicate (to the user) when the device has heard a “wake word” and is listening for voice commands from the user (i.e., voice-indicated commands).
[0048] One type of implementation involves using voice-based interfaces in various environments (e.g., relatively large living environments) where user interaction or the user interface does not have a single focus. As technology advances towards widespread Internet of Things (IoT) automation and connected devices, many things around and on us exhibit the ability to accept sensory input and convey information by altering or translating signals into the environment. In the case of automation in our living or working spaces, intelligence (e.g., provided at least in part by (multiple) automated assistants) can manifest itself in a very pervasive or ubiquitous sense in our living or working environment. There may be a sense that the assistant is somewhat omnipresent and non-intrusive, which in itself can create a certain paradoxical aspect to the user interface.
[0049] Home automation and assistants in our personal and living spaces no longer need to reside in, control, or embody a single device. There may be many devices attempting to collectively deliver a universal service or existential design goal. However, for this to feel natural, we need to engage with and trigger a natural sense of interaction and recognition through interaction with this personal assistant.
[0050] We naturally engage with this interface primarily through voice. According to some embodiments, it is envisioned to use voice to initiate interactions (e.g., with at least one smart audio device) and to engage with at least one smart audio device (e.g., an assistant). In some applications, speech can be the most accessible and high-bandwidth method for specifying more details in a request and / or providing ongoing interaction and confirmation.
[0051] However, while human communication is based on language, it is actually built upon the first stage of signal transmission and confirmation of attention. We typically don't issue commands or voice messages until we first sense that the receiver is available, ready, and interested. There are many ways we can control attention, although currently in system design and user interfaces, the way systems demonstrate attentional responses is reflected more in the computation of the individual interface's text space than in the efficiency and naturalness of the interaction. When devices are primarily the nearest microphone or user console, most systems mainly involve simple visual indicators (lights), which are ill-suited to the foreseeable future living environments with more pervasive system integration and environmental computing.
[0052] Signal transmission and attention expression are crucial components of communication, where the user indicates their expectation to interact with at least one intelligent audio device (e.g., a virtual assistant), and each device demonstrates the user's awareness and initial and ongoing attention to understanding and support. In traditional designs, the interaction suffers from several discordant aspects, where the assistant is often treated more as a discrete device interface. These aspects include:
[0053] - In situations where there are multiple points or devices that may potentially be ready to accept input and give attention, it is best to express attention not just from the device closest to the user;
[0054] -Given the extensive ergonomics of living and flexible workspaces, a user's visual attention may not align with any lighting response that indicates confirmation;
[0055] -While the voice can come from discrete places, it is often actually the house or place of residence we are visiting and seeking support from, and it gives a more general sense of attention than a single device that must change discretely or abruptly between interactions.
[0056] - In high noise and echo conditions, errors may occur when locating users to express attention to a specific area, location, or device;
[0057] - In many cases, users may be moving to or from a specific area, and therefore, decisions about boundaries would be discordant if forced to choose a location or device.
[0058] - Typically, the form of expression of concern has very discrete time boundaries in terms of whether something clearly occurs.
[0059] Therefore, we envision that interactions between a user and one or more smart audio devices typically begin with a call to attention (e.g., a wake word spoken by the user) initiated by the user, and continue with at least one indication (or signal or expression) of “attention” from (multiple) smart audio devices or from devices associated with the smart audio devices. We also envision that, in some embodiments, at least one smart audio device (e.g., a prompting assistant) may continuously listen to sound signals (e.g., indicating the type of user activity) or may continuously be sensitive to other activities (not necessarily sound signals), and the smart audio device will enter a state or operating mode that awaits a command (e.g., a voice command) from the user upon detecting a predetermined type of sound (or activity). Upon entering the latter state or operating mode, each such device exhibits attention (e.g., in any manner described herein).
[0060] It is well known that smart audio devices are configured in discrete physical areas to detect a user (who has uttered a wake word already detected by the device) and respond to the wake word by transmitting visual and / or auditory signals that the user can see or hear within the area. Some disclosed embodiments deviate from this known approach by configuring one or more smart audio devices (of the system) to treat the user's location as indeterminate (within some indeterminate volume or region) and by using all available smart audio devices within the indeterminate volume (or region) to provide an expression of the spatial variation of the system's "attention" through one or more (e.g., all) states or operating modes of the devices. In some embodiments, the goal is not to select the single device closest to the user and override its current settings, but to modulate the behavior of all devices according to a correlation metric, which in some examples may be at least partially based on the estimated proximity of the devices to the user. This gives the impression that the system is focusing its attention on a local area, eliminating the discordant experience of distant devices and suggesting that the system is listening while the user is trying to attract the attention of a closer device in the system.
[0061] Some embodiments provide (or are configured to provide) coordinated utilization of an environment or zone of an environment by defining and implementing the ability of each device to generate attention signals (e.g., in response to a wake word). In some embodiments, some or all devices may be configured to “blend” attention signals into the current configuration (and / or generate attention signals at least partially determined by the current configuration of all devices). In some embodiments, each device may be configured to determine a probabilistic estimate of distance to a location, such as the distance of a device from a user's location. Some such embodiments can provide a combined, orchestrated expression of system behavior in a user-perceived manner.
[0062] For a smart audio device that includes (or is coupled to) at least one speaker, the attention signal can be sound emitted from at least one such speaker. Alternatively or additionally, the attention signal can be some other type (e.g., light). In some examples, the attention signal can be or includes two or more components (e.g., emitted sound and light).
[0063] In this article, we sometimes use the phrase “attention indication” or “attention expression” interchangeably with the phrase “attention signal”.
[0064] In one type of embodiment, multiple smart audio devices can be coordinated (or orchestrated), and each device can be configured to generate an attention signal in response to a wake word. In some implementations, a first device may provide an attention signal corresponding to a second device. In some examples, attention signals corresponding to all devices are coordinated. Aspects of some embodiments relate to implementing smart audio devices and / or coordinating smart audio devices.
[0065] According to some embodiments, in a system, multiple smart audio devices can respond in a coordinated manner (e.g., by emitting light signals) to the determination of a common operating point (or operating state) of the system. For example, an operating point could be an attentional state entered in response to a wake word from a user, where all devices have an estimate of the user's location (e.g., with at least one degree of uncertainty), and where devices emit light of different colors based on their estimated distance from the user.
[0066] Following user research and interactive experiments, the inventors have identified certain rules or guidelines that can be applied to a wide-area life assistant that expresses concern, and support some publicly available implementations.
[0067] These include the following:
[0068] - Attention can indicate sustained and responsive improvements or personal signal transmission. This provides a better indication and closure of the signal transmission work required for training and creates a more natural interaction. It can be useful to note the range of signal transmission intensity (e.g., from a soft, gentle request to a loud curse) and to determine the associated impedance matching response (e.g., from a response to a gentle raised glance to a response to standing up for attention);
[0069] - Signaling of attention can similarly perpetuate uncertainty and ambiguity regarding the user's location and focus. Incorrect item or object responses produce a highly discontinuous and detached sense of interaction and attention. Therefore, forced selection should be avoided;
[0070] - More (rather than less) generalized signal transmission and transducers are generally preferred to complement any single voice response point, where continuous control is often an important component; and
[0071] - For expressions of concern, the ability to rise and fall naturally back to the baseline setting or environment can be advantageous, providing a sense of companionship and presence rather than a purely transactional and information-based interface.
[0072] It is well known that some things can be quickly personified, and subtle aspects of time and continuity have a significant impact. Some disclosed embodiments implement continuous control of output devices in the environment to record some sensory effects on the user, and control the device in a natural up-and-down manner to express attention and release, while avoiding incongruous hard decisions around the position of interaction thresholds and binary decision-making.
[0073] Figure 1A This is a diagram of the environment (living space) of the system, which includes a set of intelligent audio devices (device 1.1) for audio interaction, a speaker (1.3) for audio output, a microphone (1.5), and a controllable light (1.2). As with the other figures in this application, Figure 1A The specific elements and their arrangement shown are merely examples. All of these features may not be required to perform the various disclosed embodiments. For example, for at least some disclosed embodiments, a controllable light 1.2, a speaker 1.3, etc., are optional. In some cases, one or more microphones 1.5 may be part of or associated with one of the devices 1.1, 1.2, or 1.3. Alternatively or additionally, one or more microphones 1.5 may be attached to another part of the environment, such as a wall, ceiling, furniture, appliance, or another device in the environment. In the example, each smart audio device 1.1 includes at least one microphone 1.5 (and / or is configured to communicate with at least one microphone). Figure 1A The system can be configured to implement embodiments of this disclosure. Various methods can be used to... Figure 1A The microphone 1.5 gathers information and provides that information to a device configured to provide an estimate of the location of the user who utters the wake word.
[0074] In living spaces (e.g., Figure 1A In a living space, there exists a set of natural activity zones where people will perform tasks or activities, or cross thresholds. In some examples, these zones (which may be referred to as user zones in this paper) can be defined by the user without specifying the coordinates or other markings of their geometric locations. Figure 1A In the example shown, the user area may include:
[0075] 1. Kitchen sink and food preparation area (in the upper left area of the living space);
[0076] 2. Refrigerator door (to the right of the sink and food preparation area);
[0077] 3. Dining area (located in the lower left area of the living space);
[0078] 4. Open areas of the living space (to the right of the sink and food preparation area and the dining area);
[0079] 5. TV sofa (on the right side of the open area);
[0080] 6. TV itself;
[0081] 7. Table; and
[0082] 8. Door area or entrance passage (in the upper right area of the living space).
[0083] According to some embodiments, a system for estimating where a sound (e.g., a wake word or other signal of interest) originates or is located may have a certain level of confidence (or multiple assumptions) in that estimation. For example, if a user happens to be near the boundary between zones of the system environment, an uncertain estimate of the user's location may include a certain level of confidence for the user in each zone. In some conventional implementations of voice interfaces, the voice assistant's voice is required to be emitted from only one location at a time, which forces a single selection of a single location (e.g., Figure 1A (One of the eight speaker positions (1.1 and 1.3) in the system). However, based on simple hypothetical role-playing, it is clear that (in such a conventional implementation) the selected position of the assistant's voice source (e.g., the position of the speaker included in the assistant or configured to communicate with the assistant) is likely to be the focus or a natural return response to express attention.
[0084] Next, refer to Figure 1B We describe another environment 100 (acoustic space) including a user (101) speaking direct speech 102, and an example of a system including a set of intelligent audio devices (103, 105, and 107), speakers for audio output, and a microphone. The system can be configured according to embodiments of this disclosure. The speech spoken by user 101 (sometimes referred to herein as the speaker) can be recognized as a wake word by the system's components(s).
[0085] More specifically, Figure 1B The system components include:
[0086] 102: Direct local voice (generated by user 101);
[0087] 103: Voice assistant device (coupled to one or more loudspeakers). Device 103 is positioned closer to user 101 than device 105 or device 107, and therefore device 103 is sometimes referred to as a “near” device, device 105 may be referred to as a “mid-range” device and device 107 may be referred to as a “far” device;
[0088] 104: Multiple microphones in (or coupled to) the proximity device 103;
[0089] 105: Mid-range voice assistant device (coupled to one or more loudspeakers);
[0090] 106: Multiple microphones in (or coupled to) the mid-range device 105;
[0091] 107: Remote voice assistant device (coupled to one or more loudspeakers);
[0092] 108: Multiple microphones in (or coupled to) remote device 107;
[0093] 109: Household appliances (such as lamps); and
[0094] 110: Multiple microphones in (or coupled to) household appliance 109. In some examples, each microphone of microphone 110 may be configured to communicate with a device configured to implement one or more of the disclosed methods, and in some cases, the device may be at least one of devices 103, 105, or 107.
[0095] Figure 1B The system may include at least one device configured to implement one or more methods disclosed herein. For example, device 103, device 105, and / or device 107 may be configured to implement one or more of these methods. Alternatively or additionally, another device configured to communicate with device 103, device 105, and / or device 107 may be configured to implement one or more of these methods. In some examples, one or more of the disclosed methods may be implemented by another local device (e.g., a device within environment 100), while in other examples, one or more of the disclosed methods may be implemented by a remote device located outside environment 100 (e.g., a server).
[0096] When speaker 101 utters a sound 102 indicating a wake word in the acoustic space, the sound is received by a nearby device 103, a mid-range device 105, and a distant device 107. In this example, each of devices 103, 105, and 107 is (or includes) a wake word detector, and each of devices 103, 105, and 107 is configured to determine when the wake word probability (the probability that a wake word has been detected by the device) exceeds a predefined threshold. Over time, the wake word probability determined by each device can be plotted as a function of time.
[0097] Figure 2 An example of a wake word confidence value curve determined by three devices is shown. Figure 2 The dashed curve 205a shown indicates the likelihood of a wake word as a function of time, as determined by the near device 103. The short dashed curve 205b indicates the likelihood of a wake word as a function of time, as determined by the medium-distance device 105. The solid curve 205c indicates the likelihood of a wake word as a function of time, as determined by the distant device 107.
[0098] from Figure 2 The examination clearly shows that, over time, the probability of a wake word determined by each of devices 103, 105, and 107 increases and then decreases (e.g., as the wake word probability enters and exits the history buffer of a related device). In some cases, the wake word confidence of remote devices ( Figure 2 The real curve in the middle can indicate the wake word confidence of a mid-range device. Figure 2 If the threshold is exceeded before the dashed curve in the diagram, it can also be within the wake word confidence level of the near device. Figure 2 The short dashed curve in the image indicates that the threshold has been exceeded. When the wake-word confidence of a near-device reaches its local maximum (e.g., the short dashed curve in the image), the threshold is exceeded. Figure 2 When the correlation curve reaches its maximum value, this event is typically ignored (using traditional methods) to support the selection of devices where the wake word confidence (wake word probability) first exceeds a threshold. Figure 2 (The device in the example is located at a distance).
[0099] Based on some examples, a local maximum value can be determined after determining that the wake-word confidence value exceeds a wake-word detection start threshold, which can be a predetermined threshold. For example, refer to... Figure 2 In some such examples, a local maximum can be determined after the wake-word confidence value exceeds the wake-word detection start threshold 215a. In some such examples, a local maximum can be determined by detecting a decrease in the wake-word confidence value after the previous wake-word confidence value has already exceeded the wake-word detection start threshold.
[0100] In some such implementations, a local maximum can be determined by detecting the decrease in the wake-word confidence value of an audio frame after a previous wake-word confidence value has exceeded a wake-word detection start threshold, compared to the wake-word confidence value of a previous audio frame. In some cases, the previous audio frame can be the most recent audio frame or one of the most recent audio frames. For example, a local maximum can be determined by detecting the decrease in the wake-word confidence value of audio frame n after a previous wake-word confidence value has exceeded a wake-word detection start threshold, compared to the wake-word confidence value of audio frame nk, where k is an integer.
[0101] According to some such implementations, some methods may involve initiating a local maximum determination time interval after the rising edge of the wake word confidence value of a first device, a second device, or another device exceeds a wake word detection start threshold. Some such methods may involve terminating the local maximum determination time interval after the wake word confidence value of the first device, a second device, or another device drops below a wake word detection end threshold.
[0102] For example, refer to again Figure 2 In some such examples, the local maximum determination time interval can be initiated at start time A when the wake-word confidence value of any device in a set of devices exceeds the wake-word detection start threshold 215a. In this example, the distant device is the first device with a wake-word confidence value exceeding the wake-word detection start threshold, and its time A is the time when curve 205c exceeds the wake-word detection start threshold 215a. According to this example, threshold 215b is the wake-word detection end threshold. In this example, the wake-word detection end threshold 215b is less than (lower than) the wake-word detection start threshold 215a. In some alternative examples, the wake-word detection end threshold 215b can be equal to the wake-word detection start threshold 215a. In still other examples, the wake-word detection end threshold 215b can be greater than the wake-word detection start threshold 215a.
[0103] Based on some examples, the local maximum determination time interval can terminate after the wake word confidence value of all devices in the group drops below the wake word detection termination threshold of 215b. For example, see reference... Figure 2 When the wake-word confidence value of a nearby device drops below the wake-word detection end threshold 215b, the local maximum determination time interval can be equal to K time units and can terminate at the end time A+K. By the end time A+K, the wake-word confidence values of distant and mid-distance devices have dropped below the wake-word detection end threshold 215b. According to some examples, the local maximum determination time interval can end when the wake-word confidence values of all devices in the group drop below the wake-word detection end threshold 215b or after the maximum time interval has elapsed, whichever comes first.
[0104] Figure 3 This is a block diagram illustrating examples of components of a device capable of implementing various aspects of the present disclosure. According to some examples, device 300 may be or may include a smart audio device (such as...). Figure 1A One of the smart audio devices 1.1 shown or Figure 1B (One of the smart audio devices 103, 105, and 107 shown herein), the smart audio device is configured to perform at least some of the methods disclosed herein. In other embodiments, device 300 may be or may include another device configured to perform at least some of the methods disclosed herein, as referenced below. Figure 7 The described smart home hub 740 includes laptop computers, cellular phones, tablet devices, motor controllers (e.g., controllers for fans or other devices capable of moving air in an environment, controllers for garage doors, etc.), controllers for gas fireplaces (e.g., controllers configured to change the flame level of a gas fireplace), etc. In some such embodiments, device 300 may be or may include a server.
[0105] In this example, device 300 includes an interface system 305 and a control system 310. In some embodiments, the interface system 305 may be configured to receive input from each of a plurality of microphones in the environment. Interface system 305 may include one or more network interfaces and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). According to some embodiments, interface system 305 may include one or more wireless interfaces. Interface system 305 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, interface system 305 may include a control system 310 and a memory system (such as... Figure 3 One or more interfaces between the optional memory systems 315 shown. However, the control system 310 may include a memory system.
[0106] The control system 310 may include, for example, a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic and / or discrete hardware components.
[0107] In some implementations, the control system 310 may reside in more than one device. For example, a portion of the control system 310 may be located in... Figure 1A and Figure 1BThe control system 310 may be located in a device within one of the environments depicted, and another part of the control system 310 may be located in a device outside the environment, such as a server, mobile device (e.g., a smartphone or tablet computer), etc. In other examples, a part of the control system 310 may be located in... Figure 1A and Figure 1B The control system 310 may be located in one of the devices within the environment depicted, and another part of the control system 310 may be located in another device within the environment. For example, as noted below, in some cases, one device within the environment (e.g., a light) may provide an attention signal corresponding to another device (e.g., an IoT device). In some such examples, the interface system 305 may also be located in more than one device.
[0108] In some implementations, the control system 310 may be configured to at least partially perform the methods disclosed herein. According to some examples, the control system 310 may be configured to implement methods for generating attention signals with multiple spatial variations, such as those disclosed herein. In some such examples, the control system 310 may be configured to determine a correlation metric for at least one device.
[0109] Some or all of the methods described herein can be executed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media can include memory devices as described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. One or more non-transitory media can, for example, be located on... Figure 3 The optional memory system 315 and / or control system 310 shown herein. Therefore, various innovative aspects of the subject matter described herein can be implemented in one or more non-transitory media on which software is stored. For example, the software may include instructions for controlling at least one device to process audio data. For example, the software may be provided by, for example, Figure 3 The control system 310 and other control system components are used to perform the operation.
[0110] In some examples, device 300 may include Figure 3 The optional microphone system 320 shown may include one or more microphones.
[0111] In some embodiments, device 300 may include Figure 3 The optional lighting system 325 shown may include one or more lamps, such as light-emitting diodes. According to some embodiments, the device 300 may include… Figure 3The optional speaker system 330 is shown. The optional speaker system 330 may include one or more speakers. In some examples, the control system may control the optional lighting system 325 and / or the optional speaker system 330 to generate an attention signal. In some such examples, the attention signal may indicate a correlation metric of device 300, or a correlation metric of another device.
[0112] According to some such examples, device 300 may be or may include a smart audio device. In some such embodiments, device 300 may be or may include a wake word detector. For example, device 300 may be or may include a virtual assistant.
[0113] Figure 4 It outlines what can be achieved by, for example Figure 3 The flowchart illustrates an example of a method performed by at least one device. As with other methods described herein, the blocks of method 400 need not be performed in the indicated order. In some examples, one or more blocks of method 400 may be performed simultaneously. According to some such examples, one or more blocks of method 400 may be performed simultaneously by multiple devices, for example, by... Figure 3 The apparatus or other devices shown are used for execution. Furthermore, this method may include more or fewer boxes than those shown and / or described.
[0114] In this example, box 405 relates to receiving an output signal from each of a plurality of microphones in the environment. In this example, each of the plurality of microphones is located at a microphone position in the environment, and the output signal corresponds to a person's speech. In some examples, the speech may be (or include) a wake word. At least one of the microphones may be included in or configured to communicate with a smart audio device.
[0115] In some implementations, in block 405, a single device can receive output signals from each of a plurality of microphones in the environment. According to some such examples, the single device may be located within the environment. However, in other examples, the single device may be located outside the environment. For example, in some cases, at least a portion of method 400 may be performed by a remote device such as a server.
[0116] In other embodiments, in block 405, multiple devices can receive output signals. In some such embodiments, in block 405, the control system of each of the multiple smart devices can receive output signals from multiple microphones of each smart device.
[0117] Depending on the specific implementation, microphones from different devices in the environment may or may not be synchronized microphones. As used herein, a microphone may be referred to as "synchronized" if the sound detected by the microphone is digitally sampled using the same sampling clock or a synchronized sampling clock. For example, a first microphone (or a first group of microphones, such as all microphones of a first smart device) in the environment may sample audio data according to a first sampling clock, and a second microphone (or a second group of microphones, such as all microphones of a second smart device) may sample audio data according to the first sampling clock.
[0118] According to some alternative implementations, at least some of the microphones or microphone systems in the environment can be "asynchronous." As used herein, a microphone can be referred to as "asynchronous" if the sound detected by the microphone is digitally sampled using different sampling clocks. For example, a first microphone (or a first group of microphones, such as all microphones of a first smart device) in the environment can sample audio data according to a first sampling clock, and a second microphone (or a second group of microphones, such as all microphones of a second smart device) can sample audio data according to a second sampling clock. In some cases, the microphones in the environment can be randomly located, or at least distributed in an irregular and / or asymmetrical manner within the environment.
[0119] exist Figure 4 In the example shown, box 410 involves determining, at least in part, a region within the environment that has a threshold probability of including a person, based on the output signal. For example, refer to... Figure 1A In some examples, only device 1.1 contains a microphone and can therefore receive audio data used to estimate the location of the user (1.4) who issued the wake-word command. Various methods can be used to obtain information from these device clusters to provide location estimates (e.g., fine-grained location estimates) of the user who issued (e.g., spoke) the wake-word. Related methods include direction of arrival (DOA) estimation methods such as time difference of arrival (TDOA) methods, beamforming methods (e.g., maximum variance beamformer (MVB) with delay-sum beamforming (DSB)), and multi-source localization methods such as multi-signal classification (MUSIC, an algorithm for frequency estimation and radio direction finding), steering response power phase transformation (SRP-PHAT, a beamforming-based method for searching candidate locations that maximize the output of the steering delay-sum beamformer), and signal parameter estimation via rotation-invariant techniques (ESPRIT, a technique for determining the parameters of sinusoidal mixing in background noise).
[0120] In this living space, there is a set of natural activity zones where people will perform tasks or activities, or cross thresholds. These action zones (areas) are places where the user's position (e.g., determining an uncertain position) or the role of the situation can be estimated to assist other aspects of the interface. Figure 1A In the example, the key action area is:
[0121] • Kitchen sink and food preparation area (in the upper left area of the living space);
[0122] • Refrigerator door (to the right of the sink and food preparation area);
[0123] • Dining area (located in the lower left area of the living space);
[0124] • Open areas of the living space (to the right of the sink and food preparation area and the dining area);
[0125] • TV sofa (on the right side of the open area);
[0126] TV itself;
[0127] • Table; and
[0128] • Doorway or entrance passage (in the upper right area of the living space).
[0129] Clearly, there are typically a similar number of lights with similar positioning to accommodate the area of operation. Some or all of the lights may be individually controllable networked agents.
[0130] In some examples, the goal is not to estimate the precise geometric location of a user, but rather to form a robust estimate of a discrete region (e.g., in the presence of strong noise and residual echoes). As used herein, the “geometric location” of an object or user in the environment refers to a location based on a coordinate system, whether the coordinate system is referenced to GPS coordinates, the entire environment (e.g., a Cartesian or polar coordinate system with its origin at a point in the environment), or a specific device within the environment (e.g., a Cartesian or polar coordinate system with the device as its origin), such as a smart audio device. According to some examples, an estimate of a user’s location in the environment can be determined without referencing the geometric locations of multiple microphones.
[0131] In some examples, a user's zone can be estimated using a data-driven approach, involving multiple high-order acoustic features obtained at least partially from at least one of the wake word detectors. In some implementations, these acoustic features (which may include wake word confidence and / or reception levels) can consume very little bandwidth and can be asynchronously transmitted to a device implementing a classifier with very low network load. Examples such as Figure 1D are disclosed in U.S. Provisional Patent Application No. 62 / 950,004, filed December 18, 2019, entitled "Acoustic Zoning with Distributed Microphones". Figure 2 And the corresponding discussion on page 15, line 8 through page 21, line 29, is incorporated herein by reference. Depending on the particular implementation, data on the geometric location of the microphones may or may not be provided to the classifier. As noted elsewhere herein, in some examples, an estimate of the user's position in the environment may be determined without referring to the geometric locations of multiple microphones.
[0132] Some such methods may involve receiving output signals from each of multiple microphones in an environment. Each of the multiple microphones may be located at a microphone position in the environment. In some examples, the output signal may correspond to the user's current utterance.
[0133] Some such methods may involve determining multiple current acoustic features from the output signal of each microphone and applying a classifier to these current acoustic features. Applying the classifier may involve applying a model trained on previously determined acoustic features derived from multiple previous utterances made by the user in multiple user zones within the environment. Some such methods may involve determining an estimate of the user zone in which the user is currently located, at least in part, based on the output from the classifier. User zones may, for example, include a sink area, a food preparation area, a refrigerator area, a dining area, a sofa area, a television area, and / or a doorway area.
[0134] In some examples, a first microphone among a plurality of microphones may sample audio data according to a first sampling clock, and a second microphone among a plurality of microphones may sample audio data according to a second sampling clock. In some examples, at least one of the microphones may be included in or configured to communicate with a smart audio device. According to some examples, the plurality of user areas may involve a plurality of predetermined user areas.
[0135] Based on some examples, estimates can be determined without referencing the geometric positions of multiple microphones. In some examples, multiple current acoustic features can be determined asynchronously.
[0136] In some cases, the current utterance and / or previous utterances may include wake-up words. In some examples, the user region may be estimated as the category with the highest posterior probability.
[0137] According to some implementations, the model can be trained using training data labeled by the user area. In some cases, the classifier may involve applying a model trained using unlabeled training data that is not labeled by the user area. In some examples, applying the classifier may involve applying a Gaussian Mixture Model trained on one or more of normalized wake-word confidence, normalized average reception level, or maximum reception level.
[0138] In some examples, model training can continue during the application of the classifier. For example, training can be based on explicit feedback from the user. Alternatively or additionally, training can be based on implicit feedback, such as implicit feedback regarding the success (or failure) of beamforming or microphone selection based on the estimated user area. In some examples, implicit feedback may include determining that the user has abnormally terminated the voice assistant's response. According to some implementations, implicit feedback may include a command recognizer returning a low-confidence result. In some cases, implicit feedback may include a second, retrospective wake word detector returning a low-confidence wake word that has been spoken.
[0139] Back Figure 4 In this example, box 415 relates to generating multiple spatially varying attention signals within the area. According to this example, each attention signal is generated by a device located within the area, and each attention signal indicates that the corresponding device is in an operating mode where the corresponding device is waiting for a command. Furthermore, in this example, each attention signal indicates a "relevance metric" for the corresponding device.
[0140] Depending on the specific implementation, the "corresponding device" may or may not be a device that provides attention signals. For example, a virtual assistant may include a speaker system and / or a light system, and may be configured to generate attention signals that indicate a relevance metric of the virtual assistant via the speaker system and / or the light system.
[0141] In some alternative examples, the attention signal generated by the first device can indicate a relevance measure for the second device. In this example, the second device is the "corresponding device" mentioned in box 415. (See reference) Figure 1AAs noted above, one or more microphones 1.5 may be part of or associated with one of the lamps 1.2 and / or the speaker 1.3. Furthermore, one or more microphones 1.5 may be attached to another device in the appliance or environment, some of which may be “smart devices” capable of being controlled at least partially according to voice commands. In some such examples, one or more of the lamps 1.2 and / or the speaker 1.3 (determined to be within the area based on output signals from the associated microphones, as mentioned in boxes 405 and 410) may be configured to generate attention signals for corresponding appliances or other devices (e.g., IoT devices) in the environment within the area.
[0142] In some examples, the relevance measure may be based at least in part on the estimated distance to a location. In some examples, the location may be the estimated location of the person uttering the utterance, as mentioned in box 405. According to some such examples, the relevance measure may be based at least in part on the estimated distance from the person to the device corresponding to the attention signal.
[0143] In some implementations, the estimated distance can be the estimated distance from a location (e.g., the location of a light, the location of a smart device, etc.) to the acoustic centroids of multiple microphones within the area. For example, the estimated distance can be the estimated Euclidean distance from the acoustic centroids of the microphones within the area. In other cases, the estimated distance can be the estimated Mahalanobis distance from the acoustic centroids of the microphones within the area. In still other cases, the correlation measure can be the posterior probability that a given light would be classified as associated in a given area if it is a microphone.
[0144] In some implementations, the control system can be configured, for example, to estimate the posterior probability p(C) of the feature set W(j) corresponding to the utterance, by using a classifier. k |W(j)). In some such implementations, the classifier can be a Bayesian classifier. Probability p(C k |W(j)) can instruct the user to perform each zone C k The probability of (for the j-th utterance and the k-th region, for each region C) k (and each utterance). These probabilities are examples of the output of this classifier.
[0145] In some examples, the amount of attention expressed can be related to p(C). k |W(j)) correlation (e.g., monotonic correlation). For example, in some cases, if the lighting device of interest may not include any microphones, then the classifier can determine or estimate the agent based on the relative positions of the lighting device and nearby microphones.
[0146] Based on some examples, the process of building and / or updating a region location model may include the following:
[0147] 1. Collect a set of classification posterior p(C) corresponding to the most recent set of utterances j = 1...J (e.g., the set of the 200 most recent wake words spoken in the home). k |W(j)), and the estimated position x of the speaker during each utterance in the group (e.g., in 3D Cartesian space). j ;
[0148] 2. Calculate (e.g., in 3D Cartesian space) the "acoustic centroid" μ for each region k. k As a weighted average as well as
[0149] 3. Optionally, for example, in the case of assuming a multivariate Gaussian distribution in Cartesian space, the “acoustic size and shape” of each region can be calculated. In some such examples, the process may involve calculating a weighted covariance matrix, for example, as follows:
[0150]
[0151] Then, given a new position y, the control system can be configured to utilize the zone position model to perform one or more of the following:
[0152] 1. Calculate Euclidean distance And use d k (For example, in meters) as a measure of relevance. Some such examples could involve using d k Mapped to a monotonic function f(d) in the range [0, 1] k ) to transmit d k .
[0153] 2. Calculate the Mahalanobis distance And use m k (Measured in standard deviation from the centroid) as a correlation measure. Some examples of this could involve using m k Mapped to the monotonic function g(m) of the range [0, 1] k ) to pass m k .
[0154] 3. Evaluate the probability density of the multivariate Gaussian k-model for location y: Some such examples could involve normalizing the probability density of each region y to a posterior probability. Some such implementations may involve directly using posterior p. k As a measure of regional correlation within the range [0, 1].
[0155] According to some examples, the relevance measure can be based at least in part on the estimated visibility of the corresponding device. In some such examples, the relevance measure can be based at least in part on the elevation of the corresponding device, for example, the height of the corresponding device above the ground in the environment. According to some such examples, if the estimated distances from a person to two devices are the same or substantially the same (e.g., within threshold percentages such as 10%, 8%, 5%, etc.) and the elevation of one device is greater than that of the other, the device with the higher elevation will be assigned a higher relevance measure. In some such examples, the weighting factor of the relevance measure can be based on the estimated visibility of the corresponding device. For example, the weighting factor can correspond to the relative distance from the floor to the aforementioned device. In other examples, the estimated visibility of the corresponding device and the corresponding weighting factor can be determined based on the person's relative position and one or more features of the environment such as interior walls, furniture, etc. For example, the weighting factor can correspond to the probability of seeing the corresponding device from the person's estimated position, for example, based on the known environmental layout, wall positions, furniture positions, counter positions, etc.
[0156] According to some implementations, the relevance measure may be based at least in part on an estimate of the wake word confidence. In some such examples, the relevance measure may correspond to an estimate of the wake word confidence. According to some such examples, the wake word confidence unit may be a percentage, a number in the range [0, 1], etc. In some cases, the wake word detector may use a logarithmic implementation. In some such logarithmic implementations, a wake word confidence of zero means that the probability of saying the wake word is the same as the probability of not saying the wake word (e.g., depending on a particular training set). In some such implementations, an increasing positive number may indicate an increase in the confidence of saying the wake word. For example, a wake word confidence score of +30 may correspond to a very high probability of saying the wake word. In some such examples, a negative number may indicate that it is unlikely to say the wake word. For example, a wake word confidence score of -100 may correspond to a high probability of not saying the wake word.
[0157] In other examples, a device-specific relevance metric can be based on an estimate of the wake word confidence level for that device and an estimated distance from the person to the device. For instance, the estimate of the wake word confidence level can be used as a weighting factor, which is then multiplied by the estimated distance to determine the relevance metric.
[0158] Attention signals can include, for example, light signals. In some such examples, the attention signal can spatially vary within the area based on color, color saturation, light intensity, etc. In some such examples, the attention signal can spatially vary within the area based on the flashing rate of the light. For example, a flashing light that flashes faster can indicate a relatively higher relevance metric for the corresponding device compared to a slower-flashing light.
[0159] Alternatively or additionally, the attention signal may include, for example, sound waves. In some such examples, the attention signal may spatially vary within the area based on frequency, volume, etc. In some such examples, the attention signal may spatially vary within the area based on the rate at which a series of sounds are produced, for example, the number of beeps or chirps within a time interval. For example, a sound produced at a higher rate may indicate a relatively higher relevance metric for the corresponding device compared to a sound produced at a lower rate.
[0160] Refer again Figure 4 In some implementations, option 420 may involve selecting a device for subsequent audio processing based at least in part on a comparison of a relevance metric. In some such implementations, method 400 may involve selecting at least one speaker of a device located within the area and controlling the at least one speaker to provide sound to a person. In some such implementations, it may involve selecting at least one microphone of a device located within the area and providing a signal output by the at least one microphone to a smart audio device. In some implementations, the selection process may be automatic, while in other examples, the selection may be based on user input, for example, from the person uttering the speech.
[0161] According to some examples, attention signals may include modulation of at least one prior signal generated by a device within the area prior to the time of utterance. For example, if a luminaire or light source system has previously emitted a light signal, the modulation may be color modulation, color saturation modulation, and / or light intensity modulation. If the prior signal is already an audio signal, the modulation may include level or volume modulation, frequency modulation, etc. In some examples, modulation may be a change in fan speed, a change in flame size, a change in motor speed, and / or a change in airflow.
[0162] According to some implementations, modulation can be a “bulge.” A bulge can be or can include a predetermined signal modulation sequence. Some detailed examples are described below. Some such implementations may involve the use of variable output devices (in some cases, continuously variable output devices) in a system environment (e.g., lights, speakers, fans, fireplaces, etc. in a living space), which may be used for other purposes but are capable of modulation around its current point of operation. Some examples may provide variable attention indications (e.g., variable attention signals with bulges), for example, to indicate a changing expression of attention across a set of devices (e.g., the amount of change). Some implementations may be configured to control the variable attention signals (e.g., bulges) based on an estimated strength of user signal transmission and / or a function of confidence in the user’s location (multiple user locations).
[0163] Figure 5This is a block diagram illustrating an example of features according to some implementation methods. In this example, Figure 5 Indicates the variable probability of variable signal transmission strength 505 (e.g., the signal transmission strength of a wake word spoken by the user) and the location 510 of a variable signal source. Figure 5 It also indicates responses to variable signal transmissions for different smart audio devices (e.g., virtual assistants). The devices are in device groups 520 and 525, and these devices include or are associated with activatable lights (e.g., configured to communicate with activatable lights). Figure 5 As indicated in the document, each device can be included in a different group. Figure 5 The “equipment groups” are based on corresponding zones such as lounges and kitchens. A zone can contain multiple audio devices and / or lights. Zones can overlap, so any audio device, light, etc., can be located in multiple zones. Therefore, instead of being associated with devices, lights, audio devices, etc., can be associated with zones. Some lights, audio devices, etc., can be more strongly (or less strongly) associated with each zone, and therefore can be associated with different percentages of elevation. In some examples, the elevation percentage can correspond to a correlation metric. In some implementations, these correlation metrics can be manually set and captured in a table, for example, such as... Figure 5 As shown in the example. In other examples, the correlation metric can be automatically determined from the heuristic or probability approach, as described above.
[0164] For example, in response to a wake word (with a defined intensity and an origin position determined with uncertainty), two different lights on the device, or two different lights associated with the device, can be activated to generate a time-varying attention signal. Because in this example, the attention signal is also spatially variable, it is based in part on an estimated distance between the device and the origin position of the wake word, which varies depending on the location of each device.
[0165] exist Figure 5 In the example shown, the signal transmission strength (505) can correspond to, for example, the “wake-up word confidence” discussed above. In this example, the location probability of all zones (kitchen, lounge, etc.) 510 corresponds to the zone probabilities discussed above (e.g., within the range [0,1]). Figure 5 Examples of different behaviors for each light corresponding to each zone (which may correspond to a "correlation metric") are shown. If lights, audio devices, etc., are associated with multiple zones, in some implementations, the control system can be configured to determine the maximum output for each associated zone.
[0166] Variable output devices
[0167] Without loss of generality, Table 1 (hereinafter) indicates examples of devices that can be used as variable output devices and, in some cases, as continuously variable output devices (e.g., smart audio devices, where each smart audio device includes or is associated with elements that can controllably emit light, sound, heat, move, or vibrate (e.g., are configured to communicate with them)). In these examples, the output of each variable output device is a time-varying attention signal. Table 1 indicates some modulation ranges of sound, light, heat, air motion, or vibration emitted from or generated by each device (each as an attention signal). Although single numbers are used to indicate some ranges, single numbers indicate the maximum change during the “boom” and thus indicate the range from the baseline condition to the indicated maximum or minimum value. These ranges are by way of example only and not limiting. However, each range provides an example of the minimum detectable change and the maximum (commanded) attention indication in the indication.
[0168] For example, after an “attention signal” (e.g., within the range [0,1]) is determined for each modality, there may be an “attention-to-elevation” mapping from that attention signal. In some examples, the attention-to-elevation mapping may be a monotonic mapping.
[0169] In some cases, the mapping of attention to a bulge can be set tentatively or experimentally (e.g., on a demographically representative group of test subjects) to make the mapping appear “natural,” at least to the group of individuals providing feedback during the testing procedure. For example, for a color-changing modality, an attention of 0.1 might correspond to a hue of +20 nm, while an attention of 1 might correspond to a hue of +100 nm. Variable color lamps typically do not change the transducer’s frequency but can instead have separate R, G, B LEDs that can be controlled with different intensities; therefore, the above are only rough examples. Table 1 provides some examples of natural mappings of attention to the resulting physical phenomena, which typically vary from modality to modality.
[0170] Table 1
[0171]
[0172] Figure 6This is a diagram illustrating an example of a bulge. As with other diagrams provided herein, the time intervals, amplitudes, etc., shown in Figure 600 are merely examples. In this document, we define a “bulge” (referring to a bulge in an attention signal) as a defined (e.g., predetermined) sequence of signal modulation, such as attention signal modulation. In some cases, a bulge may comprise different envelopes of attention signal modulation. A bulge can be designed to provide the timing of attention signal modulation that reflects the natural rhythm of attention (or focus). The trajectory of a bulge is sometimes designed to avoid any sense of abrupt change at edge points (e.g., at the beginning and end of the bulge).
[0173] exist Figure 6 In the example shown, Figure 600 provides an example of a bulge-varying envelope of the attention signal, also referred to herein as the bulge envelope. The bulge envelope 601 includes an attack 605, which is an increase in the attention signal level from a baseline level 603 to a local maximum level 607 during a first time interval. In this example, the first time interval is from time = 0 to approximately time = 500 ms. As noted in Table 1, the local maximum level 607 can vary depending on the type of attention signal (e.g., light, sound, or others), how the signal will be modulated (e.g., changes in light intensity, color, or color saturation), and whether the attention signal is intended to correspond to a "detectable" condition or a "commanded" condition. In other examples, such as the sound example shown in Table 1, the first time interval of the bulge could correspond to a decrease in the attention signal level from a baseline level 603 to a local minimum level.
[0174] exist Figure 6 In the example shown, the raised envelope 601 includes a release 620, which is the reduction of the attention signal level to a baseline level 603. According to this example, the release 620 begins at approximately time N seconds and lasts for approximately 2 seconds. Both N and the duration of the release 620 can vary depending on the specific implementation. In some examples, N can be 4 seconds, 5 seconds, 6 seconds, 7 seconds, 8 seconds, 9 seconds, 10 seconds, etc. In some cases, N can respond to environmental conditions. For example, the release 620 can begin if the person who said the wake word has left the area where the corresponding device is located. In other examples, the duration of the release 620 can be greater than or less than 2 seconds.
[0175] according to Figure 6 In the example shown, the bulge envelope 601 includes attenuation 610, which is a reduction in the attention signal level from a local maximum level 607 to an intermediate or moderate level 615 between the local maximum level 607 and the baseline level 603. According to this example, attenuation 610 occurs between approximately 500 milliseconds and approximately 1 second.
[0176] In this case, the raised envelope 601 also includes a retention 617 during which the attention signal level remains unchanged. In some embodiments, the attention signal level may remain substantially the same during the retention 617, for example, it may be maintained within a certain percentage of the attention signal level at the start of the retention 617 (e.g., within 1%, within 2%, within 3%, within 4%, within 5%, etc.). Figure 6 In the example shown, 617 is held from approximately 1 second to approximately N seconds.
[0177] Estimated intensity
[0178] In some example implementations, the normalized intensity of the attention signal can vary from 0 (for threshold detection of wake words) to 1 (for wake words with estimated vocalization effects that result in a speech level 15-20 dB higher than normal).
[0179] Function for modulation device bulge
[0180] An example of a function used to modulate a bulge in an attention signal with an initial intensity "output" is:
[0181] Output = Output + Bump * Confidence * Intensity
[0182] The parameters bulge, confidence level, and intensity can change over time.
[0183] Before the introduction of raised steps for expressing attention, the control of a large number of Internet of Things (IoT) devices, such as lights, was inherently complex. With this in mind, several embodiments have been designed where, for example, a raised step is typically a short-term additional increment to any setup that occurs as a result of controlling a broader scene or spatial situation.
[0184] In some implementations, scene control may involve occupancy and may be additionally shaped during voice commands associated with the control of the system selected to express attention. For example, if there is more than one person in the area, the audio attention signal may be kept within a relatively low amplitude range.
[0185] Some embodiments provide a way to implement this scene control from a raised implementation. In some implementations, the raised attention signals of multiple devices can be controlled according to a separate protocol (in other words, separately from other protocols used to control device functions), thereby enabling devices to participate in human attention cycles and for the environment of the living space to be controlled.
[0186] Some aspects of the embodiments may include the following:
[0187] - Continuous output actuator;
[0188] - Assign smart audio devices to active groups. In some cases, devices may be assigned to more than one group.
[0189] - A ridge with one or more designed time envelopes;
[0190] - The extent of the bulge is controlled by a simple function of activation intensity and confidence in the region (or location).
[0191] Examples of how to control a virtual assistant (or other smart audio device) to demonstrate environments that were not well represented in previous systems and how to create testable standards may include the following:
[0192] - A confidence score (such as a wake word confidence score) calculated based on an estimate of the user’s intent or a call from a virtual assistant with specific contextual information (such as the location and / or zone where the wake word is spoken) may be published (e.g., shared among smart devices in the environment), and in at least some examples, this confidence score is not directly used to control the device;
[0193] - It can control appropriately equipped devices with continuous electrical control to use the information to "boost" their existing state, thereby responding naturally and mutually;
[0194] - Self-delegation of devices used to perform "bulges" (e.g., device auto-discovery and / or dynamic zone updates) can generate emergency responses that do not require manual tables of location and "zones," as well as additional robustness provided by low user setup requirements; and
[0195] - The confidence gained through continuous estimation, publication, and growth of statistical samples (e.g., via explicit or implicit user feedback) enables the system to create representations of existence. In some examples, these representations of existence can move naturally across space, and in others, they can be modulated based on the actions the user takes to help the assistant solve the problem.
[0196] Figure 7 An example embodiment of a system for implementing automatic optical orchestration is shown.
[0197] Figure 7 The components include:
[0198] • 700: This illustration shows an example home with automated optical orchestration, specifically a two-bedroom apartment.
[0199] 701: Living room;
[0200] 702: Bedroom;
[0201] • 703: The wall between the living room and the bedroom. According to this example, light cannot pass between the two rooms;
[0202] • 704: Living room window. During the daytime hours, sunlight illuminates the living room through this window;
[0203] • 705A-C: Multiple smart ceiling lights (e.g., LED) that illuminate the living room;
[0204] • 705D-F: Each ceiling light is programmed and communicates with the smart home hub 740 via Wi-Fi (or another protocol);
[0205] 706: Living room table;
[0206] ·707: A smart speaker device for the living room that incorporates a light sensor;
[0207] • 707A: Device 707 is orchestrated by Wi-Fi (or another protocol) and communicates with the smart home hub 740 via Wi-Fi (or another protocol);
[0208] • 708A-C: Controlled light propagation from lamp 705A-C to device 707;
[0209] • 709: Uncontrolled light propagation from window 704 to device 707;
[0210] ·710: A smart ceiling LED light that illuminates the bedroom;
[0211] • 710A: The bedroom light is programmed via Wi-Fi (or another protocol) and communicates with the smart home hub 740 via Wi-Fi (or another protocol);
[0212] 711: Potted plants;
[0213] ·712: An IoT (Internet of Things) automatic watering system incorporating a light sensor;
[0214] • 712A: The IoT watering device is orchestrated by Wi-Fi (or other protocols) and communicates with the smart home hub 740 via Wi-Fi (or other protocols);
[0215] 713: Bedroom table;
[0216] ·714: A smart bedroom speaker device incorporating a light sensor;
[0217] • 714A: The bedroom smart speaker is programmed via Wi-Fi or another protocol and communicates with the smart home hub 740 via Wi-Fi or another protocol;
[0218] • 715: Controlled light propagation from bedroom light 710 to IoT watering device 712; and
[0219] ·716: Controlled light propagation from the bedroom light 710 to the bedroom smart speaker 714.
[0220] Based on this example, the smart home hub 740 is the reference mentioned above. Figure 3 An example of the described device 300.
[0221] Figure 8 It's a diagram. Figure 7 A set of diagrams illustrating examples of the system's operational aspects. Figure 8 The components include:
[0222] 800: Display Figure 7 Figure 800 depicts a graph of the continuous values of the light intensity settings (810, 805A, 805B, and 805C) for a set of example smart lighting devices (710, 705A, 705B, and 705C, respectively). Figure 800 also shows, on the same time axis... Figure 7 The example optical sensors (712, 714 and 707) depicted in the figure have continuous optical sensor readings (812, 814 and 807 respectively);
[0223] 810: Continuously controlled light intensity output of the intelligent lighting device 710. The value at 6:00 PM corresponds to the lamp being completely off;
[0224] 805A: Continuously controlled light intensity output of the 705A intelligent lighting device. The value at 6:00 PM corresponds to the lamp being completely off;
[0225] 805B: Continuously controlled light intensity output of the 705B intelligent lighting device. The value at 6:00 PM corresponds to the lamp being completely off;
[0226] 805C: Continuously controlled light intensity output of the 705C intelligent lighting device. The value at 6:00 PM corresponds to the lamp being completely off;
[0227] 812: Continuous light sensor readings for example light sensor 712. The reading is low at 6:00 PM;
[0228] 814: Continuous light sensor readings for example light sensor 714. The reading is low at 6:00 PM;
[0229] 807: Continuous light sensor readings for example light sensor 707. The reading at 6:00 PM is high;
[0230] 830: As sunlight (709) enters through the window (704), the continuous light sensor reading is initially high. As dusk falls, the ambient light intensity decreases until 7:30 PM;
[0231] 820: An event occurring at 7:30 PM when two smart lighting devices (705A, 705B) are turned on by a user in response to dim lighting conditions in room (706). As shown by traces 820A and 820B, the light intensity of smart lighting devices 705A and 705B increases. Simultaneously, the continuous light sensor reading at 820C increases with a significantly similar response;
[0232] 821: Event 820 ends when smart lighting devices 705A and 706B are turned off. Accordingly, traces 820A and 820B return to fully off, and the light sensor reading 807 returns low;
[0233] 820A: The increase and decrease in light output when the intelligent lighting device 705A is turned on and then off;
[0234] 820B: The increase and decrease of light output when the intelligent lighting device 705B is turned on and then off;
[0235] 820C: The increase and decrease of the light output in response to the light sensor readings of the light sensor 707, which are activated and deactivated in response to the activation and deactivation of lamps 705A and 705B;
[0236] 822: An event that occurs at 8:00 PM when the intelligent lighting device 710 is turned on and then off (823). The light intensity of the device is modulated with response 822A. Then, the light sensor readings 812 and 822 are modulated with remarkably similar responses 822B and 822C;
[0237] 824: An event that occurred at 8:30 PM when the new intelligent lighting device 705C was connected to the system. The light output is modulated either automatically or manually by the user in the on / off mode shown in 824A.
[0238] 824A: Modulated output mode for lamp 705C;
[0239] 824B: In response to the modulation of the smart light 705C, the continuous light sensor 707 reads a significantly similar response 824B;
[0240] 825: Event 824 has ended;
[0241] 826: In response to a user request, the lights in room 701 are activated to a dim setting of approximately 50% intensity. These lights are 705A, 705B, and 705C, whose 50% output intensity is shown in traces 826A, 826B, and 826C, respectively. Accordingly, the continuous light sensor reading of sensor 707 is modulated with a remarkably similar response; and
[0242] 827: Event 826 has ended.
[0243] With the rapid proliferation of such devices, managing and registering connected devices in homes and workplaces presents increasing challenges. Lighting, furniture, appliances, mobile phones, and wearables are becoming increasingly connected, and current manual methods for installing and configuring such devices are unsustainable. Providing network authentication details and pairing devices with user accounts and other services is just one example of the types of devices that need to be registered during initial installation. Another common step in registration and installation is assigning specific “zones” or “groups” to a set of devices, organizing them into logical categories that are typically associated with specific physical spaces such as rooms. Lighting and appliances, which are usually statically installed, most often fall into this category. The manual and additional installation steps associated with assigning these “zones” or “groups” to devices create usability challenges for users and reduce their appeal as a commercial product.
[0244] This disclosure recognizes that while such logical groupings and zones are meaningful in a home automation environment, they may be too rigid to provide the level of expression and fluidity expected of human / machine interaction as a user navigates the space. In some examples, the ability to modulate and raise continuously variable output parameters of a collection of devices to best express interest may require the system to possess some knowledge about the distribution or relevance of such devices that is more finely granular or relevant than typical rigid and manually assigned “zones.” In this paper, we describe an inventive method for automatically mapping such distributions by aggregating both readings generated by multiple sensors and opportunistic sampling of the continuous output configurations of multiple smart devices. In this paper, we use the example of light to facilitate the discussion, thus employing one or more photosensitive components with digitizable output readings attached to one or more smart devices, and self-reported light intensity and hue output parameters from multiple smart lighting devices. However, it should be understood that other modalities such as sound, temperature (with temperature measuring components, and smartly connected heating and cooling appliances) are also possible embodiments of this method and approach.
[0245] refer to Figure 7 and Figure 8 We illustrate an example scenario using light as a modality to create a mapping that associates luminescent smart devices with smart assistant devices employing integrated or otherwise physically attached light sensors. For clarity, in the explanation below, Figure 7 An example environment divided into two discrete regions is described. Figure 8 The system measures signals for analytical purposes in order to determine the mapping that associates a controllable light-emitting device with an intelligent auxiliary device employing a light sensor.
[0246] In our example, all smart lighting devices (710, 705A, 705B, and 705C) were initially not emitting light at 6:00 PM and were visible in traces 810, 805A-C, respectively. Devices 710, 705A, and 705B are currently installed and mapped, while 705C is a new device that has not yet been mapped to the system. Light sensor readings for three smart devices (712, 714, and 707) are also depicted. It should be understood that the vertical axis and horizontal axis ( Figure 8 (The image is not drawn to scale, and in this case, the light sensor readings may not be scaled to the same level as the smart light output parameters.) It should also be understood that only light intensity is shown here as an example, and some disclosed embodiments also cover light tone output parameters and multiple light sensors sampling different portions of the spectrum.
[0247] In our example, room 702 is a bedroom, and room 701 is a living room. Room 702 contains a smart light-emitting device 710 and two light-sensing smart devices 712 (IoT watering device) and 714 (smart speaker). Room 701 contains two initially installed and mapped smart lights 705A and 705B, and a new unmapped smart light 705C. Room 702 also contains a light-sensing smart speaker device 707. A window 704 is also present in room 702, generating an uncontrolled amount of ambient light.
[0248] In our example, all smart devices are equipped to communicate via WiFi or some other communication protocol through a home or local network, and information collected or stored at one device can be transmitted to the orchestration hub device 740. At 6:00 PM, none of the smart lighting devices 710, 705A-C produce light, but light is emitted through window 704 in room 701. Therefore, the light sensor reading in room 702 is low, while the reading in room 701 is high.
[0249] A series of events corresponding to changes in lighting conditions will occur, and it will be demonstrated that the corresponding changes in the light sensor readings will be sufficient to establish a basic mapping between the intelligent sensing device and the intelligent light-emitting device. Trace 820 depicts the decrease in the sensor readings of device 707 as the sun sets and the decrease in the amount of light (709) generated through window 704. Event 820 occurs at 7:30 PM when the user turns on the light in living room 701. As a result, light outputs 805A and 805B increase, as shown by curves 820A and 820B. Accordingly, the light sensor reading 807 increases with curve 820C. It is noteworthy that the light sensor readings 812 and 814 corresponding to devices 712 and 714 in adjacent rooms do not change due to this event. The event ends at the horizontal time marked by 821 when the light is turned off again.
[0250] Similar to event 820, event 822 begins at 8:00 PM when the bedroom light is turned on. During this event, the continuously variable output parameter (810) of the bedroom light (710) increases with curve 822A. The light sensor readings (812 and 814) of smart devices 712 and 714 are also modulated with curves 822B and 822C respectively, in a corresponding manner. Notably, light sensor reading 807 is unaffected because this reading is in an adjacent room. At 823, the event ends when the bedroom light 710 is turned off.
[0251] At 8:30 PM, the unmapped living room light 805C periodically toggles on and off for short periods. This toggling can be initiated automatically by the lighting device itself, upon request from the smart hub 740, manually by the user using a physical switch, or by alternatively supplying and removing power to the device. Regardless of how this output modulation (identifiable by curve 824A) is achieved, the reported output intensity (805C) of device 705C is transmitted via the network for aggregation with light sensor readings 812, 814, and 807. As in event 820, the only sensor in the living room (attached to device 707) reflects output modulation 824A, which has a remarkably similar pattern 824B in the sensor readings. This event ends shortly after its inception, as indicated by label 825.
[0252] Using the data aggregated by the system so far, it can be inferred that the unmapped smart lamp 705C is closely related to lamps 705A and 705B. This is because the extent to which the transmission of light 705A and 705B through light 708A and 708B affects the light sensor readings (807) is very similar to the extent to which the light emitted by 705C (708C) affects the same sensor. This similarity (determined by the convolution process, to be discussed in more detail) determines the extent to which the light is co-localized and context-sensitive. This soft decision-making and approximate relationship mapping provides an example of how to provide more granular “zoning” and spatial awareness for intelligent assistant systems.
[0253] With the smart light 705C now effectively mapped, an example of a user requesting to turn on all “living room” lights to 50% intensity is depicted in event 826. All three living room lights 705A-C are enabled at 50% output, depicted in the output traces 805A-C and following curve 826A-C. Accordingly, the light sensor reading 807 is also modulated along curve 826D. As the accumulation of the relevant modulations observed in the device output and the readings of the sensors in question, the degree to which the device is “mapped” will increase with confidence over time. Therefore, even though the new device 705C is already understood to coexist with at least 705A and 705B, further analysis of events such as 826 that occur after the initial setup period should be understood as data for constructing an increasingly detailed and reliable spatial map of the space, which can be used to facilitate expressive personal assistant interactions as previously discussed in this disclosure.
[0254] It should be understood that light sensors can incorporate specific filters to more selectively sense light generated by consumer and commercial LED lighting devices, removing the spectrum generated by uncontrollable light sources such as the sun.
[0255] It should be understood that, from a system perspective, the 824 event in the example is optional. However, in this example, the rate at which the device maps to the system is proportional to the frequency of its modulation output parameters. With this in mind, it is expected that the device can be integrated into the system's mapping more quickly via highly discriminative modulation events such as 824, which, from an information theory perspective, encode highly relevant information.
[0256] Some embodiments can be configured to implement continuous (or at least continuous and / or periodic) remapping and refinement. Through Figure 7 and Figure 8 The example described captures both routine user use of "already mapped" devices and the installation of new lighting equipment. For automated setup and mapping methods to be implemented, the system should preferably require no user intervention or manual operation of the lighting equipment. This is why event 824 (in...) Figure 8The modulation event (in the middle) can be initiated by the user, or equivalently by the smart light itself through its own judgment or by instructions from the central control or other external orchestration devices. This clearly detectable modulation event carries a high degree of information and helps to quickly introduce new devices into the system's mapping.
[0257] We will now discuss a more subtle and complementary form of modulation, which is clearly not driven by user intervention and is referred to in this paper as "general refinement." The system can continuously adjust the output parameters of individual smart devices in a slow, incremental manner—minimal detectable to the user but discernible to the smart sensors—to establish increasingly higher fidelity mappings. Unlike operating systems that rely on users to correlate information to generate explicit data, the system can control and execute its own modulation of individual smart output devices, again in a manner minimally detectable to the user and still discernible to the sensors.
[0258] Many examples of this method are possible (with optical modal focusing). Examples are shown in the table below:
[0259]
[0260]
[0261]
[0262] Under the premise and operation of the embodiments described above, we will now further describe in detail the evolution (over time) of the mapping between continuous output devices and sensor-equipped intelligent devices. We define the “mapping” H as a normalized similarity metric between sensor-equipped intelligent devices and all continuous output devices in the system. For sensor-equipped intelligent devices D{i} and intelligent output devices L{j}, we can define the continuous similarity metric G as:
[0263] 0<=G(D{i}, L{j})<=1,
[0264] Where H is a set of all G for all D{i} and L{j} in the system: H = {G(D{i}, L{j})} for all i and j.
[0265] Given this, it can be seen that selecting the discrete region near D{i} can be achieved using a binary threshold D between 0 and 1:
[0266] Z = all j such that G(D{i},L{j})>d.
[0267] A continuous similarity metric G has been established, allowing the concept of regions to become fluid, and we do not need to confine ourselves to discrete regions to express attention. Therefore, different values of d can be chosen based on the degree of attention or expression expected by the virtual assistant during the interaction.
[0268] Refer again Figure 8 The known light activations 810, 805A, 805B, and 805C of four smart lighting devices L{j}(710, 705A, 705B, 705C) with j = 1...4 can be represented as I{j}[t]. In this example, the light reading traces 812, 814, and 807 from the light sensors on other smart devices D{i}(712, 714, 707) with i = l...3 can be represented as S{i}[t].
[0269] G(D{i}, L{j}) can be calculated from discretely sampled time series I[t] and S[t], output device parameters transmitted over the network, and sensor readings, respectively. I and S can be sampled at sufficiently close periodic intervals to allow for meaningful comparisons. Many similarity measures typically assume zero-mean signals. However, a constant environmental offset is often present in environmental sensors (e.g., ambient lighting conditions).
[0270] Therefore, signals I[t]' and S[t]' can also be obtained from I[t] and S[t], and G can be calculated from these obtained signals.
[0271] For example, a smoothed sample-to-sample increment can be represented as follows:
[0272] I[t]'=(1-a)*I[t-1]'+a*(I[t]-I[t-1]); for 0<a<1
[0273] The similarity between two time series with a recent time interval T can be established using many methods familiar to those skilled in signal processing and statistics, for example, through the following methods:
[0274] 1. The Pearson correlation coefficient (PCC or "r") between I[t] and S[t], set G = (1 + PCC) / 2, for example, as As described at http: / / mathworld.wolfram.com / CorrelationCoefficient.html, its By reference, it is incorporated into this article;
[0275] 2. The version obtained by describing the method as in 1, but using the time increments of I and S;
[0276] 3. As described in 1, but using the average deleted version of I and S; and / or
[0277] 4. Dynamic time warping on both I and S, for example, such as https: / / en.wikipedia.org / wiki / The description of Dynamic_time_warping (which is incorporated herein by reference) The resulting distance metric is used as G.
[0278] Some implementations may involve automatically updating automated processes for determining whether a device is in a device group, whether a device is in an area, and / or whether a person is in an area. Some such implementations may involve updating the automated processes based on implicit feedback, which is based on one or more of the following: successful beamforming based on the estimated area, successful microphone selection based on the estimated area, determination that a person has abnormally terminated the voice assistant's response, a command recognizer returning a low-confidence result, or a second retrospective wake word detector returning a low-confidence that a wake word has been spoken.
[0279] The goal of predicting the user's location can be to inform the microphone selection or adaptation of beamforming schemes that attempt to pick up sound more effectively from the user's acoustic zone, for example, to better identify commands following a wake word.
[0280] In this scenario, implicit techniques for obtaining feedback on the quality of the region prediction can include:
[0281] • Punishment for misrecognition of commands following the wake word. Alternatives that could indicate misrecognition could include the user shortening the voice assistant's response to commands, for example, by using a counter-command phrase like "Amanda, stop!"
[0282] • Penalty results in a low-confidence prediction that the speech recognizer has successfully identified the command. Many automatic speech recognition systems have the ability to return the confidence level, and the results can be used for this purpose;
[0283] • The penalty caused the second wake word detector to fail to retrospectively detect the wake word prediction with high confidence; and / or
[0284] • Enhanced predictions enable high-confidence identification of wake words and / or correct identification of user commands.
[0285] Below is an example of a second-pass wakeword detector failing to retrospectively detect a wakeword with high confidence. Assume that after obtaining an output signal corresponding to the current utterance from a microphone in the environment, and after determining acoustic features based on that output signal (e.g., via multiple first-pass wakeword detectors configured to communicate with the microphone), the acoustic features are provided to a classifier. In other words, the acoustic features are assumed to correspond to the detected wakeword utterance. Further assume that the classifier determines the person uttering the current utterance is most likely in zone 3, which in this example corresponds to a reading chair. For example, there might be specific microphones or known combinations of microphones known to be best suited for hearing human speech when a person is in zone 3, for example, to be sent to a cloud-based virtual assistant service for voice command recognition.
[0286] Further assuming that after determining which microphone(s) will be used for speech recognition, but before the person's speech is actually sent to the virtual assistant service, a second wake-word detector operates on the microphone signals corresponding to the speech you will submit for command recognition detected by the selected microphone(s) in zone 3(s). If this second wake-word detector is inconsistent with the multiple first wake-word detectors that actually utter the wake-word, this is likely because the classifier mispredicted the zone. Therefore, the classifier should be penalized.
[0287] Techniques for post-hoc updates to the region mapping model after one or more wake words have been spoken can include:
[0288] • Maximum a posteriori (MAP) fitting of Gaussian mixture model (GMM) or nearest neighbor model; and / or
[0289] • Reinforcement learning, such as reinforcement learning of neural networks, is performed, for example, by associating appropriate “one-hot” (in correct predictions) or “one-cold” (in incorrect predictions) real data labels with the SoftMax output and applying online backpropagation to determine new network weights.
[0290] In this context, some examples of MAP adaptation could involve adjusting the mean in the GMM each time the wake word is spoken. In this way, the mean can become more like the acoustic features observed when subsequent wake words are spoken. Alternatively or additionally, such examples could involve adjusting the variance / covariance or mixed weighting information in the GMM each time the wake word is spoken.
[0291] For example, a MAP adaptation solution could be as follows:
[0292] μ i,new =μi i,old*α+x*(1-α)
[0293] In the aforementioned equation, μ i,old Let represent the average of the i-th Gaussian in the mixture, α represent a parameter controlling the positivity of MAP adaptation (α can be in the range [0.9, 0.999]), and x represent the feature vector of the new wake-word utterance. The index "i" will correspond to the mixture element, which returns the highest prior probability containing the speaker's position at the wake-word time.
[0294] Alternatively, each mixing element can be adjusted based on its prior probability of containing a wake-up word, for example, as follows:
[0295] M i,new =μ i,old *β i *x(1-β i )
[0296] In the aforementioned equation, β i =α*(1-P(i)), where P(i) represents the prior probability of the observation x being due to the mixed element i.
[0297] In a reinforcement learning example, there can be three user regions. Suppose that for a given wake word, the model predicts the probabilities of the three user regions as [0.2, 0.1, 0.7]. If a second information source (e.g., a second wake word detector) confirms that the third region is correct, the true data label could be [0, 0, 1] (“one-hot”). The posterior update of the region mapping model can involve backpropagating the error through the neural network, which effectively means that if the same input is shown again, the neural network will predict region 3 more strongly. Conversely, if the second information source shows that region 3 is an incorrect prediction, in one example, the true data label could be [0.5, 0.5, 0.0]. If the same input is shown in the future, backpropagating the error through the neural network will make the model less likely to predict region 3.
[0298] Alternatively or additionally, some implementations may involve automatically updating the automated process based on explicit feedback from humans. Explicit techniques for obtaining feedback may include:
[0299] • Use a voice user interface (UI) to ask the user if the prediction is correct. (For example, you could provide the user with a voice indicating the following: "I think you are on the sofa, please say 'yes' or 'no'").
[0300] • Incorrect predictions can be corrected at any time using the voice UI. (For example, you could provide the user with a voice indicating something like, "I can now predict where you are when you talk to me. If I'm wrong, I'll say something like, 'Amanda, I'm not on the sofa. I'm sitting in the reading chair.'")
[0301] • Notifying users that correct predictions can be rewarded at any time using the voice UI. (For example, you could provide a voice prompt indicating something like, "I am now able to predict where you are when you talk to me. If I predict correctly, you can say something like, 'Amanda, yes. I am on the sofa,' to help further improve my predictions.")
[0302] • This includes physical buttons or other UI elements that users can interact with to provide feedback (e.g., thumb up and / or thumb down buttons on a physical device or in a smartphone app).
[0303] While specific embodiments and applications of this disclosure have been described herein, it will be apparent to those skilled in the art that many changes can be made to the embodiments and applications described herein without departing from the scope of this disclosure.
Claims
1. A method of controlling a system of devices in an environment, the method comprising: receiving an output signal from each of a plurality of microphones in the environment, each of the plurality of microphones being located at a microphone location of the environment, the output signal corresponding to a human utterance; determining a zone within the environment based at least in part on the output signal, the zone having at least a threshold probability of including a location of the human; generating a plurality of spatially-varying attention signals within the zone, each of the plurality of attention signals being generated by a device located within the zone, each attention signal indicating that a corresponding device is in an operational mode in which the corresponding device is awaiting a command, each attention signal indicating a relevance metric of the corresponding device, wherein the relevance metric is based at least in part on an estimated distance from a location to an acoustic centroid of a plurality of microphones within the zone, and wherein the acoustic centroid of the plurality of microphones within the zone is a weighted average of estimated locations of speakers from the plurality of microphones within the zone relative to a probability of the speakers being within the zone.
2. The method of claim 1, wherein, an attention signal generated by a first device indicates a relevance metric of a second device, the second device being the corresponding device.
3. The method of claim 1, wherein, the location is an estimated location of the human.
4. The method of claim 1 or claim 2, wherein, the relevance metric is based at least in part on an estimated visibility of the corresponding device.
5. The method of claim 1 or claim 2, wherein, the utterance includes a wake word.
6. The method of claim 3, wherein, the attention signal varies at least in part as a function of an estimate of wake word confidence.
7. The method of claim 1 or claim 2, wherein, at least one of the attention signals includes a modulation of at least one prior signal generated by a device within the zone prior to a time of the utterance.
8. The method of claim 7, wherein, the at least one prior signal includes a light signal, and wherein the modulation includes at least one of color modulation, color saturation modulation, or light intensity modulation.
9. The method of claim 7, wherein, the at least one prior signal includes a sound signal, and wherein the modulation includes level modulation.
10. The method of claim 7, wherein, the modulation includes a change in one or more of fan speed, flame size, motor speed, or air flow rate.
11. The method of claim 7, wherein, the modulation includes a bump, the bump including a predetermined sequence of signal modulations.
12. The method of claim 11, wherein, the bump includes a first time interval corresponding to an increase in signal level from a baseline level.
13. The method of claim 12, wherein, the bump includes a second time interval corresponding to a decrease in signal level to the baseline level.
14. The method of claim 13, wherein, the bump includes a hold time interval after the first time interval and before the second time interval, the hold time interval corresponding to a constant signal level.
15. The method of claim 11, wherein, the bump includes a first time interval corresponding to a decrease in signal level from a baseline level.
16. The method of claim 1 or claim 2, wherein, at least one of the microphones is included in or configured for communication with a smart audio device.
17. The method of claim 1 or claim 2, further comprising determining an automated process of a device in a group of devices.
18. The method of claim 17, wherein, the automated process is based at least in part on sensor data corresponding to at least one of light or sound emitted by the device.
19. The method of claim 17, wherein, the automated process is based at least in part on communication between a source and a receiver.
20. The method of claim 19, wherein, the source is a light source or a sound source.
21. The method of claim 17, wherein, The automated process is based at least in part on at least one of a communication between a source and the orchestration hub device or a communication between a receiver and the orchestration hub device.
22. The method of claim 17, wherein, The automated process is based at least in part on a light source or a sound source being turned on and off over a period of time.
23. The method of claim 17, further comprising automatically updating the automated process based on explicit feedback from the person.
24. The method of claim 17, further comprising automatically updating the automated process based on implicit feedback based on one or more of: a successful beamforming based on an estimated zone, a successful microphone selection based on the estimated zone, determining that the person has terminated a response of a voice assistant abnormally, a command recognizer returning a low confidence result, or a second pass retrospective wake word detector returning a low confidence that a wake word has been spoken.
25. The method of claim 1 or claim 2, further comprising selecting at least one speaker of a device located within the zone and controlling the at least one speaker to provide sound to the person.
26. The method of claim 1 or claim 2, further comprising selecting at least one microphone of a device located within the zone and providing a signal output by the at least one microphone to a smart audio device.
27. The method of claim 1 or claim 2, wherein, A first microphone of the plurality of microphones samples audio data according to a first sampling clock and a second microphone of the plurality of microphones samples audio data according to a second sampling clock.
28. An apparatus configured to perform the method of any one of claims 1-27.
29. A system configured to perform the method of any one of claims 1-27.
30. One or more non-transitory media having stored thereon software comprising instructions for controlling one or more devices to perform the method of any one of claims 1-27.
31. A computer program product comprising computer executable instructions for performing the method of any one of claims 1-27 by one or more processors.
Citation Information
Patent Citations
Arbitration-based voice recognition
US10181323B2
Outputting notifications using device groups
US10425781B1
Location based voice association system
US20180047394A1
Arbitration-Based Voice Recognition
US20180108351A1
Systems, Methods, and Devices for Activity Monitoring via a Home Assistant
US20180330589A1