Context-based model selection
Context-aware model selection in devices dynamically adapts to environmental changes by selecting appropriate models based on sensor data, enhancing performance and reducing resource consumption in sound event classification and natural language processing systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-19
- Publication Date
- 2026-04-06
AI Technical Summary
Existing systems for sound event classification, noise reduction, and natural language processing are difficult and expensive to update, consume significant memory and processing resources, and struggle to adapt to diverse environments and situations.
Implement context-aware model selection that dynamically chooses from a library of models based on the device's context, using sensors to determine location, activity, and acoustic environment, allowing for efficient switching and updating of models to improve accuracy and reduce resource consumption.
Enhances system performance by improving accuracy, reducing memory and processing requirements, and enabling faster convergence in dynamic environments, while maintaining user experience and adaptability.
Smart Images

Figure 0007840941000001 
Figure 0007840941000002 
Figure 0007840941000003
Abstract
Description
Claim of Priority
[0001]
[0001] This application claims the benefit of priority of U.S. Non-Provisional Patent Application No. 17 / 102,748, owned by the same applicant and filed on November 24, 2020, the entire content of which is hereby incorporated by reference in its entirety.
Technical Field
[0002]
[0002] This disclosure generally relates to context-based model selection.
Background Art
[0003]
[0003] Technological advancements have led to smaller and more powerful computing devices. For example, there are now various portable personal computing devices, including wireless telephones such as mobile phones and smartphones, tablets, and laptop computers, which are small, lightweight, and easily carried by users. These devices can communicate voice and data packets via wireless networks. Additionally, many such devices incorporate additional features such as digital still cameras, digital video cameras, digital recorders, and audio file players. Also, such devices can process executable instructions, including software applications such as web browser applications that can be used to access the Internet. Thus, these devices, including, for example, sound event classification (SEC) systems, noise reduction systems, automatic speech recognition (ASR) systems, natural language processing (NLP) systems, etc., that attempt to recognize sound events (such as closing a door with a noise, a car horn, etc.) in an audio signal can include significant computing power.
[0004]
[0004] Systems that perform operations such as SEC, noise reduction, ASR, and NLP may use models that are trained to give a broad scope of application but are difficult or expensive to update. For example, an SEC system is generally trained using supervised machine learning techniques to recognize a particular set of sounds identified in labeled training data. Thus, each SEC system tends to be domain-specific (e.g., capable of classifying a predetermined set of sounds). Once an SEC system is trained, it is difficult to update the SEC system to recognize new sound classes that were not identified in the labeled training data. Furthermore, some sound classes that the SEC system is trained to detect may represent sound events that have more variant forms than are represented in the labeled training data. While the user experience can be improved by updating the model to improve accuracy for environments that the user's device is generally exposed to, the training involved in updating the model can be time-consuming and require a lot of data, and the number of distinct classes (e.g., new distinct sounds for the SEC system) that the model is updated to recognize can grow rapidly, potentially consuming a lot of memory on the device. [Overview of the project]
[0005]
[0005] In certain embodiments, the device includes one or more processors configured to receive sensor data from one or more sensor devices. The one or more processors are also configured to determine the context of the device based on the sensor data. The one or more processors are further configured to select a model based on the context. The one or more processors are also configured to process an input signal using the model to generate a context-specific output.
[0006]
[0006] In a particular embodiment, the method includes receiving sensor data from one or more sensor devices in one or more processors of the device. The method includes determining the context of the device based on the sensor data in one or more processors. The method includes selecting a model based on the context in one or more processors. The method also includes processing an input signal using the model to generate a context-specific output in one or more processors.
[0007]
[0007] In certain embodiments, the device includes means for receiving sensor data. The device includes means for determining a context based on the sensor data. The device includes means for selecting a model based on the context. The device also includes means for processing an input signal using the model to generate a context-specific output.
[0008]
[0008] In certain embodiments, the non-temporary computer-readable storage medium includes instructions that, when executed by a processor, cause the processor to receive sensor data from one or more sensor devices. When executed by a processor, the instructions cause the processor to determine the context on the sensor data. When executed by a processor, the instructions cause the processor to select a model based on the context. When executed by a processor, the instructions also cause the processor to process an input signal using the model to generate a context-specific output.
[0009]
[0009] Other aspects, advantages, and features of the present disclosure will become apparent after reviewing the entire application, including the following sections, namely the brief description of the drawings, the modes for carrying out the invention, and the claims. [Brief explanation of the drawing]
[0010] [Figure 1]
[0010] A block diagram of an example of a system including a device configured to perform context-based model selection, as shown in some examples of the present disclosure. [Figure 2]
[0011] A diagram illustrating the operation of the components of Figure 1, using some examples from this disclosure. [Figure 3]
[0012] Figure 1 illustrates some examples of how a model update can be performed by the device shown in this disclosure. [Figure 4]
[0013] A figure illustrating some examples of how to update a sound event classification model to account for drift in this disclosure. [Figure 5]
[0014] A figure illustrating, by some examples of this disclosure, aspects of updating a sound event classification model to take new sound classes into consideration. [Figure 6]
[0015] Figure 1 illustrates a specific example of updating a model, which can be performed by the device shown in Figure 1, by some examples of the disclosure. [Figure 7]
[0016] Figure 1 illustrates another specific example of updating a model, which can be performed by the device shown in Figure 1, by some examples of the present disclosure. [Figure 8]
[0017] A block diagram illustrating a specific example of the device in Figure 1, based on several examples of the present disclosure. [Figure 9]
[0018] Figures illustrating some examples of vehicles incorporating the device configuration of Figure 1, according to several examples of the present disclosure. [Figure 10]
[0019] Figures illustrating virtual reality, mixed reality, or augmented reality headsets incorporating the device configuration of Figure 1, as shown in some examples of the present disclosure. [Figure 11]
[0020] A figure illustrating a wearable electronic device incorporating an embodiment of the device shown in Figure 1, according to some examples of the present disclosure. [Figure 12]
[0021] A diagram showing an audio control speaker system incorporating the aspects of the device of FIG. 1 according to some examples of the present disclosure. [Figure 13]
[0022] A diagram showing a camera incorporating the aspects of the device of FIG. 1 according to some examples of the present disclosure. [Figure 14]
[0023] A diagram showing a mobile device incorporating the aspects of the device of FIG. 1 according to some examples of the present disclosure. [Figure 15]
[0024] A diagram showing a hearing aid incorporating the aspects of the device of FIG. 1 according to some examples of the present disclosure. [Figure 16]
[0025] A diagram showing an aircraft incorporating the aspects of the device of FIG. 1 according to some examples of the present disclosure. [Figure 17]
[0026] A diagram showing a headset incorporating the aspects of the device of FIG. 1 according to some examples of the present disclosure. [Figure 18]
[0027] A diagram showing a device incorporating the aspects of the device of FIG. 1 according to some examples of the present disclosure. [Figure 19]
[0028] A flowchart showing an example of a method of operating the device of FIG. 1 according to some examples of the present disclosure. [Figure 20]
[0029] A flowchart showing an example of another method of operating the device of FIG. 1 according to some examples of the present disclosure. [Figure 21]
[0030] A flowchart showing an example of another method of operating the device of FIG. 1 according to some examples of the present disclosure. [Figure 22]
[0031] A flowchart showing an example of another method of operating the device of FIG. 1 according to some examples of the present disclosure. [Figure 23]
[0032] A flowchart showing an example of another method of operating the device of FIG. 1 according to some examples of the present disclosure. [Figure 24]
[0033] A flowchart illustrating an example of another way in which the device in Figure 1 operates, based on several examples of this disclosure. [Modes for carrying out the invention]
[0011]
[0034] Systems that perform operations such as SEC, noise reduction, ASR, and NLP may use models that are trained to give a broad range of applicability, but which are difficult or expensive to update. While updating the model to improve accuracy for environments that users or their devices are commonly exposed to can improve the user experience, retraining such models can be time-consuming and require large amounts of data. Furthermore, since users encounter numerous environments and situations in their daily lives, updating the model to adapt to new environments and situations can cause the model to consume an ever-increasing amount of memory.
[0012]
[0035] The disclosed systems and methods, in some embodiments, employ context-aware model selection to choose from multiple models based on the context in which the selected model should be used (e.g., an acoustic environment). For example, an appropriate acoustic model may be selected for a particular acoustic environment to provide a better user experience in most voice user interface applications. As an example, ambient noise must be adequately estimated to provide a pleasant listening experience for hearing aid users. An appropriate model may be selected based on a specific location of the user, such as a particular street corner, building, or restaurant, or based on the type of environment of the user, such as inside a vehicle, in a park, or in a subway station, as an exemplary and non-limiting example. A specific context may be identified based on one or more of various techniques, such as location data (e.g., from a Global Positioning System (GPS) system), activity detection data, camera recognition, audio classification, user input, one or more other techniques, or any combination thereof.
[0013]
[0036] In some embodiments, the acoustic characteristics of noisy areas such as shopping malls, restaurants, and stadiums may be known, and models of these areas may be made available to the user. When a user walks into one of these locations, they may be granted access to the model for that location. After the user leaves, the model may be removed or "prune" from the user's device. In another embodiment, as a user moves around (e.g., via driving, walking, or public transport), the user's device may swap models so that the most appropriate model for each location or setting the user encounters is used. In some embodiments, a library of available models is provided to enable the discovery and downloading of appropriate models. Some of these models, such as updated models trained based on the exposure of other users' devices to various environments, may be uploaded by other users and made available (or available with specific access permissions) as part of a crowdsourced, context-aware model library.
[0014]
[0037] In certain embodiments, models can be combined or “generalized” by grouping related categories of classes into a single model to create an ensemble of various source models. For example, in an SEC application, related sound classes may be grouped based on location, sound type, or one or more other characteristics. For illustrative purposes, one model may include a group of sound classes to generally represent crowded areas such as public squares, shopping malls, and subways, while another model may include a group of sound classes related to domestic activities. These generalized models enable generally improved performance based on the broad categories of the generalized model, while using reduced memory compared to the amount of memory used for multiple specific models to adapt to each specific environment or activity. Furthermore, if a particular model is unavailable due to privacy issues or other accessibility restrictions, a more general model may be used instead. For example, if a user arrives at a busy restaurant for which there is no specific published model available for use, a generic model for crowded areas or a general model for crowded restaurants may be downloaded to the user's device and used instead.
[0015]
[0038] By modifying the model based on the device context, systems using such models can perform with higher accuracy compared to using a single model for all contexts. Furthermore, by modifying the model, such systems can perform with increased accuracy without incurring the power consumption, memory requirements, and processing resource usage associated with retraining an existing model from scratch on the device for a specific context. The use of generalized context-based models enables improved system performance compared to using a default model, and also enables reduced bandwidth, memory, and processing resource usage compared to downloading and switching between multiple high-accuracy, context-specific models. In addition, the operation of such systems using context-based models enables improved operation of the device itself, such as enabling faster convergence when performing iterative or dynamic processes (e.g., in noise cancellation techniques) to use higher-accuracy models specific to a particular context.
[0016]
[0039] Specific aspects of this disclosure are described below with reference to the drawings. In the description, common features are designated by common reference numbers. As used herein, various terms are used solely to describe specific implementations and are not intended to limit the implementations. For example, the singular forms “a,” “an,” and “the” also include the plural form unless the context makes otherwise clear. Furthermore, some features described herein are singular in some implementations and plural in others. For illustrative purposes, Figure 1 shows a device 101 containing one or more processors (the “processor 110” in Figure 1), which indicates that in some implementations, the device 101 contains a single processor 110, and in other implementations, the device 101 contains multiple processors 110. For ease of reference herein, such features are generally introduced as “one or more” features and are then referred to in the singular or optional plural form (generally indicated by terms ending in “s”) unless an aspect relating to multiple features is described thereafter.
[0017]
[0040] The terms “comprise,” “comprises,” and “comprising” are used herein interchangeably with “include,” “includes,” or “including.” Furthermore, the term “wherein” is used interchangeably with “where.” As used herein, “exemplary” indicates an example, an implementation, and / or a form, and should not be construed as limiting, or indicating a preferred or suitable implementation. As used herein, ordinal numbers used to modify elements such as structure, components, and behavior (e.g., “first,” “second,” “third,” etc.) do not by themselves indicate any priority or order of that element over another element, but rather (apart from the use of ordinal numbers) merely distinguish that element from another element having the same name. As used herein, the term “set” refers to one or more of a particular element, and the term “plurality” refers to more than one of a particular element (e.g., two or more).
[0018]
[0041] As used herein, “coupled” may include “communicatively coupled,” “electrically coupled,” or “physically coupled,” and (or alternatively) any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or a combination thereof). Two electrically coupled devices (or components) may, in exemplary, non-limiting examples, be contained within the same device or in different devices, and may be connected via electronic circuits, one or more connectors, or inductive coupling. Two communicatively coupled devices (or components), such as those communicating telecommunicates, may directly or indirectly send and receive electrical signals (digital or analog signals), such as via one or more wires, buses, networks, etc. As used herein, “directly coupled” refers to two devices that are coupled without any intermediary components (for example, communicatively coupled, electrically coupled, or physically coupled).
[0019]
[0042] In this disclosure, terms such as “determining,” “calculating,” “estimating,” “shifting,” and “adjusting” may be used to describe how one or more actions are performed. It should be noted that such terms should not be interpreted as restrictive, and other techniques may be used to perform similar actions. Additionally, as used herein, “generating,” “calculating,” “estimating,” “using,” “selecting,” “accessing,” and “determining” may be used interchangeably. For example, “generating,” “calculating,” “estimating,” or “determining” a parameter (or signal) may refer to actively generating, estimating, calculating, or determining the parameter (or signal), or it may refer to using, selecting, or accessing a parameter (or signal) that has already been generated, such as by another component or device.
[0020]
[0043] Figure 1 is a block diagram of an example system including a device 100 configured to perform context-based model selection. The device 100 includes one or more processors 110 coupled to memory 108. Memory 108 includes L available models 114 (where L is an integer greater than 1) that can be selected by one or more processors 110, indicated as one or more additional models 118 including a first model 116 and the Lth model.
[0021]
[0044] One or more processors 110 are configured to receive sensor data 138 from one or more sensor devices 134 and to determine the context 142 of device 100 based on the sensor data 138. Although the sensor devices 134 are shown as coupled to device 100, in other implementations, one or more (or all) of the sensor devices 134 may be integrated with or included within device 100.
[0022]
[0045] One or more sensor devices 134 include one or more microphones 104 coupled to one or more processors 110, and sensor data 138 includes audio data 105 from one or more microphones 104. In one example, the audio data 105 corresponds to an audio scene, and the context 142 is at least partially based on the audio scene. For illustrative purposes, based on the amount and type of noise detected in the audio data, as well as acoustic characteristics such as reverberation and absorption, the audio scene may indicate that the device 100 is in a confined, noisy space, a large enclosed space, a large outdoor space, a moving vehicle, etc.
[0023]
[0046] One or more sensor devices 134 include a location sensor 152 coupled to one or more processors 110, and the sensor data 138 includes location data 153 from the location sensor 152, such as a global positioning sensor, which provides global location data to device 100. In one example, the location data 153 indicates the location of device 100, and the context 142 is at least partially based on the location.
[0024]
[0047] One or more sensor devices 134 include a camera 150 coupled to one or more processors 110, and the sensor data 138 includes image data 151 from the camera 150 (e.g., still image data, video data, or both). In one example, the image data 151 corresponds to a visual scene, and the context 142 is at least partially based on the visual scene.
[0025]
[0048] One or more sensor devices 134 include an activity detector 154 coupled to one or more processors 110, and the sensor data 138 includes motion data such as activity data 155 from the activity detector 154. In one example, the motion data corresponds to the movement of device 100, and the context 142 is at least partially based on the movement of device 100.
[0026]
[0049] One or more sensor devices 134 may also include one or more other sensors 156 that provide one or more processors 110 with additional sensor data 157 for use in determining the context 142. The other sensors 156 may be coupled to or contained within device 100 and used to generate sensor data 157 that is useful in determining the context 142 related to device 100 at a particular time, and may include, for example, orientation sensors, magnetometers, light sensors, contact sensors, temperature sensors, or any other sensors. As another example, the other sensors 156 may include wireless network detectors that can be used to determine the context 142 by detecting when device 100 is near a recognized wireless network location (for example, by detecting a home or business WiFi® network, or a Bluetooth® network related to a family friend of the user of device 100).
[0027]
[0050] One or more processors 110 include a context detector 140, a model selector 190, and a model-based application 192. In a particular implementation, the context detector 140 is a neural network trained to determine the context 142 based on sensor data 138. In other implementations, the context detector 140 is a classifier trained using a different machine learning technique. For example, the context detector 140 may include, or correspond to, a decision tree, random forest, support vector machine, or another classifier trained to produce an output indicating the context 142 based on the sensor data 138. In yet another implementation, the context detector 140 uses heuristics to determine the context 142 based on the sensor data 138. In yet another implementation, the context detector 140 uses a combination of artificial intelligence and heuristics to determine the context 142 based on the sensor data 138. For example, the sensor data 138 may include image data, video data, or both, and the context detector 140 may include an image recognition model trained using machine learning techniques to detect specific objects, motion, background, or other image or video information. In this example, the output of the image recognition model may be evaluated via one or more heuristics to determine the context 142.
[0028]
[0051] The model selector 190 is configured to select a model 112 based on the context 142. In some implementations, the model 112 is selected from among several available models 114 stored in memory 108. In some implementations, the model 112 is selected from a model library 162 accessible via the network 160, such as a cloud-based library of models available for exploration and download. An example of the model library 162 is described in more detail with reference to Figure 2.
[0029]
[0052] As used herein, “downloading” and “uploading” models involve the transfer of data (e.g., compressed data) corresponding to the model via a wired link, a wireless link, or a combination thereof. For example, a wireless local area network ("WLAN") may be used instead of or in addition to a wired network. Wireless technologies such as Bluetooth ("Bluetooth") and Wireless Fidelity “Wi-Fi®” or variations of Wi-Fi (e.g., Wi-Fi Direct®) enable high-speed communication between mobile electronic devices (e.g., cellular phones, watches, headphones, remote controls, etc.) within relatively short distances from each other (e.g., 100 to 200 meters or less depending on the specific wireless technology). Wi-Fi is often used to connect devices with access points (e.g., routers) and devices that are Wi-Fi enabled to exchange information. Examples of such devices include smart televisions, laptops, thermostats, personal digital assistants, home automation devices, wireless speakers, and other similar devices. Similarly, Bluetooth is also used to connect devices together. Examples of such devices include mobile phones, computers, digital cameras, wireless headsets, keyboards, mice or other input peripherals, and similar devices.
[0030]
[0053] Devices (for example, those described above) may have both Bluetooth and Wi-Fi capabilities, or other wireless means for communicating with each other. Internet-connected devices may have wireless means for communicating with each other and may also be connected based on a variety of cellular communication systems, such as Long-Term Evolution (LTE®) systems, Code Division Multiple Access (CDMA) systems, Global System for Mobile Communications (GSM®) systems, Wireless Local Area Network (WLAN) systems, or any other wireless systems. CDMA systems may implement Broadband CDMA (WCDMA®), CDMA 1X, Evolution Data Optimized (EVDO), Time Division Synchronized CDMA (TD-SCDMA), or any other version of CDMA. As used herein, “wireless” means one or more of the technologies listed above, one or more other technologies that enable the transfer of information other than over wires, or a combination thereof.
[0031]
[0054] The model-based application 192 is configured to process the input signal 106 using a model 112 selected to generate a context-specific output 122. Although a single model-based application 192 is shown, in other implementations, one or more processors 110 may run multiple model-based applications 192 using different models 112 selected based on the context 142 to perform various operations.
[0032]
[0055] For example, in some implementations, model 112 includes a sound event detection model, and the input signal 106 includes audio signals such as audio data 105 from a microphone 104, audio data retrieved from an audio file in memory 108, audio signals received via wireless transmission such as a telephone call or streaming audio session, or any combination thereof. The model-based application 192 is configured to process the input signal 106 using the sound event detection model to generate a context-specific output 122 that includes a classification of sound events in the audio signal.
[0033]
[0056] In some implementations, Model 112 includes a noise reduction model, and the input signal 106 includes an audio signal such as audio data 105 from a microphone 104, audio data retrieved from an audio file in memory 108, an audio signal received via a wireless transmission such as a telephone call or streaming audio session, or any combination thereof. The model-based application 192 is configured to process the input signal 106 using the noise reduction model to generate a context-specific output 122 which includes an audio signal with reduced noise based on the audio signal.
[0034]
[0057] In some implementations, Model 112 includes an Automatic Speech Recognition (ASR) model, and the input signal 106 includes audio signals such as audio data 105 from a microphone 104, audio data extracted from an audio file in memory 108, audio signals received via wireless transmission such as a telephone call or streaming audio session, or any combination thereof. The model-based application 192 is configured to process the input signal 106 using the ASR model to generate a context-specific output 122 containing text data representing the speech in the audio signal.
[0035]
[0058] In some implementations, model 112 includes a natural language processing (NLP) model, and the input signal 106 includes text data such as text data generated by the NLP model, user keyboard input, and text messages. The model-based application 192 is configured to process the input signal 106 using the NLP model to generate context-specific output 122 which includes NLP output data based on the text data.
[0036]
[0059] In some implementations, Model 112 is associated with automatic adjustment of the device's operating mode. For example, Model 112 can map or associate user input (e.g., voice commands, gestures, touchscreen selections, etc.) to adjustment of the operating mode based on context 142. For illustrative purposes, when Model 112 is selected based on context 142 corresponding to a public area, Model 112 can cause a model-based application 192 to map the user command "play music" to playback operation in the user's earphones, and context-specific output 122 can include a signal to adjust the device's operating mode to initiate audio playback and route the output audio signal to the user's earphones. However, when Model 112 is selected based on context 142 corresponding to the user's home, Model 112 can cause a model-based application 192 to map the "play music" command to playback operation in the user's home entertainment system, and context-specific output 122 can include a signal to adjust the device's operating mode to initiate audio playback and route the output audio signal to the loudspeaker of the home entertainment system.
[0037]
[0060] In some implementations, the model selector 190 is configured to prune the model 112 in response to detection of a change in context 142, such as when a change in context 142 results in the model 112 no longer being suitable for the changed context 142, or becoming less suitable for the changed context 142 than another available model. As used herein, “prune” a model includes replacing a model with another model in a model-based application 192, permanently removing a model from memory 108, removing a model from the model-based application 192, marking a model as unused, or making a model inaccessible (for example, by removing permissions for the model as further described with reference to Figure 2), or removing a model from use by a combination thereof. In some implementations, pruning a model can reduce memory usage, processor cycles, one or more other processes or system resources, or a combination thereof, and thus improve the overall functionality of one or more processors 110 or devices 100 (e.g., increased speed, reduced power consumption). In an implementation where the model is updated in device 100, pruning the model may include maintaining updates to the model. For example, if a sound model is created or modified during use in device 100 to identify new sound classes, as illustrated with reference to Figures 3-5, the new sound classes may be maintained in memory 108 or elsewhere, or uploaded to the model library 162, in an exemplary, non-limiting example.
[0038]
[0061] During operation, the context detector 140 may monitor sensor data 138 and update the context 142 based on detecting changes in the sensor data 138. In response to the change in context 142, the model selector 190 accesses available models 114 in memory 108, model library 162, or both, to select one or more models that may be more suitable for the updated context 142 than the current model 112. For example, the model selector 190 may query the model library 162 to identify one or more aspects of context 142, such as geographical location, venue name, description of an acoustic or visual scene, one or more other aspects, or a combination thereof. The model selector 190 may receive the results of the query, select a specific model based on the query results, have the device 100 download the specific model from the model library 162, and store the downloaded model in memory 108. In some implementations, in response to detecting a change in context 142, the currently selected model 112 is pruned (for example, the model selector 190 removes model 112 from use by the model-based application 192) and replaced with a model downloaded from memory 108.
[0039]
[0062] By modifying the model based on context 142, device 100 can run model-based applications 192 with greater accuracy compared to using a single default model. Therefore, the user experience with device 100 can be improved.
[0040]
[0063] In some implementations, device 100 is also configured to modify or customize a selected model 112 to further improve accuracy, such as in response to the detection of new sound events or variations of existing classifications that may not be accurately identified by model 112. An example of modifying model 112 based on data acquired by device 100 (e.g., sensor data 138) is further illustrated with reference to Figures 3-5. After modifying model 112, device 100 may upload the modified model 112 to a model library 162 for use by other devices. In some implementations, model 112 includes trained models uploaded to library 162 from another user device. Thus, device 100 may utilize and contribute to a crowdsourced library of models in a distributed context-aware system.
[0041]
[0064] Device 100 includes, corresponds to, or may contain, voice-activated devices, audio devices, wireless speakers and voice-activated devices, portable electronic devices, cars, vehicles, computing devices, communication devices, Internet of Things (IoT) devices, virtual reality (VR) devices, augmented reality (AR) devices, mixed reality (MR) devices, hearing aids, smart speakers, mobile computing devices, mobile communication devices, smartphones, cellular phones, laptop computers, computers, tablets, personal digital assistants, display devices, televisions, game consoles, equipment, music players, wireless devices, digital video players, digital video disc (DVD) players, tuners, cameras, navigation devices, or any combination thereof. In certain embodiments, one or more processors 110, memory 108, or combinations thereof are included in the integrated circuit. Various implementation forms, including embodiments of Device 100, will be further described with reference to Figures 7 to 16.
[0042]
[0065] While device 100 is described as storing available models 114 in memory 108 and accessing the model library 162, in other implementations, device 100 may not store models in memory 108 and instead retrieve models from the model library 162 in response to detecting a change in the context 142. In other implementations, device 100 operates without accessing the model library 162 and instead selects from the available models 114 in memory 108. In some implementations, the models 114 available in memory 108 represent a locally stored portion of a distributed context-aware system and may be accessible to other devices as part of the model library 162, such as in peer-to-peer models sharing the configuration.
[0043]
[0066] Although sensor 134 is shown to include a microphone 104, a camera 150, a location sensor 152, an activity detector 154, and other sensors 156, other implementations may omit one or more of the microphone 104, camera 150, location sensor 152, activity detector 154, or other sensors 156. In the illustrative example, the context detector 140 operates using only audio data 105, image data 151, location data 153, activity data 155, or other sensor data 157. In another illustrative example, the context detector 140 operates using two of the following: audio data 105, image data 151, location data 153, activity data 155, and other sensor data 157; using three of the following: audio data 105, image data 151, location data 153, activity data 155, and other sensor data 157; or using four of the following: audio data 105, image data 151, location data 153, activity data 155, and other sensor data 157.
[0044]
[0067] Although the context detector 140, the model selector 190, and the model-based application 192 are described as separate components, in other implementations, the context detector 140 and the model selector 190 may be combined into a single component, or the model selector 190 and the model-based application 192 may be combined into a single component, or the context detector 140, the model selector 190, and the model-based application 192 may be combined into a single component. In some implementations, each of the context detector 140, the model selector 190, and the model-based application 192 may be implemented via processor execution instructions, separate hardware or circuitry, or a combination of both.
[0045]
[0068] Figure 2 shows a specific example of the operation of device 100 and model library 162. Device 100 is shown within building 202, which includes rooms 204 and 206, and an elevator 208, such as an office building. Device 100 communicates wirelessly with model library 162.
[0046]
[0069] Model Library 162 includes various types of models, such as a representative sound event detection model 220, a representative noise reduction model 222, a representative ASR model 224, a representative NLP model 226, a representative operating mode adjustment model 228, and various acoustic models 250. While only one of each of the sound event detection model 220, noise reduction model 222, ASR model 224, NLP model 226, and operating mode adjustment model 228 is shown for clarity in the figure, it should be understood that Model Library 162 can include multiple versions of each of different types of models, such as models trained for different contexts, different personalizations, and different levels of generality, in a manner similar to that described below for the acoustic model 250.
[0047]
[0070] The acoustic model 250 is presented in a configuration where a model for a general category (e.g., a generalized model) is shown as the root of a tree structure, and models for more specific contexts are shown as branches or leaves of the tree structure. For example, the “crowded area” model 252 is a general category model with branches including the “highway” model, the “metro center” model, the “suburbs” model, the “theme park” model, the “shopping mall” model, and the “public square” model. Although not illustrated, each branch model can function as a general category model for various more specific models. For example, model library 162 may contain acoustic models for several specific theme parks. If a user's device (e.g., device 100) is in a specific theme park, the device may request an acoustic model for that specific theme park. If an acoustic model is not available for that specific theme park, the user's device may request a “theme park” model, which is generally applicable to theme parks but not specifically applicable to any particular theme park. If the "theme park" model is not available, the user's device may request the "crowded area" model, which is generally applicable to crowded areas but not specifically applicable to theme parks. Therefore, the user's device can search for and select the most specific model available in the model library 162 for the specific context of its device.
[0048]
[0071] Acoustic model 250 also includes a "limited space" model 254, which is a general category model with branches corresponding to the "office building" model 262, the "residential" model 264, and the "vehicle" model 266. The "office building" model 262 is a general category model for various locations within an office building. More specific models for various locations within an office building include the "lobby" model 270, the "elevator" model 272, and the "office" model 274. The "residential" model 264 is a general category model for various locations within a residential building. More specific models for locations within a residential building include the "kitchen" model, the "room" model, and the "garage" model. The "vehicle" model 266 is a general category model for various locations within a vehicle. More specific models for locations within a vehicle include the "driver's seat" model, the "passenger seat" model, and the "rear seat" model.
[0049]
[0072] Please understand that the illustrated models are shown for illustrative and explanatory clarity. In other implementations, the model library 162 may have any number of models (e.g., hundreds, thousands, millions, etc.) configured for any number of different contexts and any number of different applications at any number of levels of generality. Also, while the acoustic model 250 is organized according to a tree structure, please understand that in other implementations, the model library 162 may utilize one or more other data structures or categorical classification techniques instead of or in addition to a tree structure.
[0050]
[0073] As illustrated, device 100 determines, based on sensor 134, that the context 142 of device 100 is inside room 204. Device 100 may determine whether the available acoustic models 250 in model library 162 include an acoustic model that is specific to the particular acoustic environment associated with context 142 and is available to device 100 (e.g., model selector 190). For example, device 100 transmits data indicating the acoustic environment 210 of device 100, such as location data, the name of building 202, a general description of building 202 (e.g., "office building"), a general description of room 204 (e.g., "office"), or any combination thereof. In a particular implementation, in response to the model library 162 not having an acoustic model specific to the particular acoustic environment 210 and available to one or more processors 110, device 100 (e.g., model selector 190) determines whether an acoustic model is available for a general category of the particular acoustic environment 210.
[0051]
[0074] As an exemplary example, device 100 may send data indicating the acoustic environment 210 by transmitting the location coordinates of device 100. If model library 162 has an acoustic model specific to (e.g., matching) the location of device 100 (e.g., an acoustic model corresponding to the geofence of the area containing the location of device 100), device 100 downloads the location-specific acoustic model 212. Otherwise, in response that model library 216 does not have a model specific to the location coordinates (e.g., not specific to building 202), device 100 may transmit additional data indicating the acoustic environment 210, such as an “office” descriptor for room 204. In response to the decision that model library 162 includes an “office” model 274 as available to device 100, device 100 downloads the “office” model 274 as the acoustic model 212 for use within room 204. If the “office” model 274 is not available, device 100 may request the more general “office building” model 262 or the even more general “confined space” model 254. In some implementations, the model library 162 is configured to automatically locate the most specific model corresponding to the acoustic environment 210 and send it to device 100, instead of device 100 sending a series of requests for increasingly generalized models until a suitable model is located.
[0052]
[0075] In some implementations, device 100 may also receive one or more permissions 214 that allow device 100 to access acoustic models 212. Permissions 214 may allow device 100 to prohibit models that it expects to use. For example, prohibited models may be downloaded from model library 162 to device 100's memory 108 prior to expected use. Downloading such models in advance may be scheduled based on available bandwidth (e.g., during periods of reduced network traffic) or to reduce latency in accessing the models when a change in device 100's context 142 is detected. Each prohibited model remains inaccessible (e.g., encrypted) in memory 108 until a corresponding permission 214 (e.g., an encryption key) for the model is received from model library 162 or another permission management system. For example, device 100 may receive permissions 214 for model 112 that are at least partially based on a location of device 100 that matches a specific location associated with model 112 (e.g., a specific location).
[0053]
[0076] In a specific example, device 100 transmits data indicating the acoustic environment 210 of room 204, receives the “office” model 274 (and any associated permissions 214) from the model library 162, stores a copy of the “office” model 274 in memory 108, and uses the “office” model 274 in a model-based application 197 for purposes such as noise reduction. When device 100 moves from room 204 to elevator 208 within building 202, device 100 detects the new context 142 and requests an acoustic model for the acoustic environment 210 corresponding to elevator 208. In response, the “elevator” model 272 is sent to device 100 as the acoustic model 212, and permissions 214 for the “elevator” model 272 may also be sent. Device 100 replaces the “office” model 274 with the “elevator” model 272 in the model-based application 192. In some implementations, when the available storage capacity in memory 108 is limited, device 100 removes the "office" model 274 from memory 108.
[0054]
[0077] Upon exiting elevator 208 and entering room 206, device 100 may search for available models 114 in memory 108 for the “office” model 274. If the “office” model is not available in memory 108, device 100 transmits data indicating the acoustic environment 210 of room 206, retrieves the “office” model 274 (and any associated permissions 214) from the model library 162, stores a copy of the “office” model 274 in memory 108, and uses the “office” model 274 in the model-based application 197. Thus, as device 100 moves from one location to the next, the model is replaced with a more appropriate model for the changing context 142 of device 100. In some implementations, when device 100 leaves building 202, any stored models in memory 108 that are specific to building 202 may be deleted or archived to conserve storage space in memory 108, or made inaccessible in response to access permissions 214 imposing geographical or other restrictions on the use of the models.
[0055]
[0078] In conjunction with the various embodiments described with reference to Figures 1 and 2, device 100 includes one or more processors 110 configured to select an acoustic model corresponding to a specific room in the building where device 100 is located, such as an acoustic model 212 corresponding to a room 204 in building 202, and to process an input audio signal using the acoustic model 212. For example, the input signal 106 may include audio data 105 generated by a microphone 104 and processed in a model-based application 192 to perform noise reduction.
[0056]
[0079] In some implementations, one or more processors 110 are configured to download acoustic model 212 from the acoustic model library 162 in response to a decision that device 100 has entered a particular room 204. In some implementations, one or more processors 110 are further configured to remove acoustic model 212 in response to device 100 leaving the particular room 204. As a non-limiting example, removing acoustic model 212 may include replacing acoustic model 212 with another model in a model-based application 192, removing acoustic model 212 from memory 108, removing acoustic model 212 from the model-based application 192, marking acoustic model 212 as unused, or making acoustic model 212 inaccessible (for example, by removing the permission 214 for acoustic model 212).
[0057]
[0080] One or more sensor devices 134 (also called “sensors 134”) coupled to one or more processors 110 are configured to generate sensor data 138 indicating the location of device 100, and one or more processors 110 are configured to select an acoustic model 212 based on the sensor data 138. For example, device 100 may include a modem, as described with reference to Figure 6, coupled to one or more processors 110 and configured to receive location data 153 indicating the location of device 100, and one or more processors 110 are configured to select an acoustic model 212 based on the location data 153.
[0058]
[0081] In some implementations, model selection may be performed predictively based on context 142. For example, based on sensor data 138 (e.g., activity detection, GPS analysis, camera recognition, audio classification, or a combination thereof), one or more processors 110 may determine that the user of device 100 is moving to a new location (e.g., New York City) where user-related assistive Internet of Things (IoT) devices can exhibit improved performance with updated settings for the new location. Thus, an appropriate source model (e.g., an acoustic model for traffic) may be retrieved from memory 108 or from one or more model libraries (e.g., model library 162) and used. Upon leaving the new location, the source model for the new location may be removed and the previous source model restored.
[0059]
[0082] In conjunction with the various embodiments described with reference to Figures 1 and 2, and as further described with reference to Figure 7, in response to device 100 entering the vehicle, one or more processors 110 are configured to select a personal acoustic model for the user of device 100 from among a plurality of personal acoustic models corresponding to the vehicle, and to process the input audio signal using the personal acoustic model. For example, device 100 may train or generate a model specific to a particular user of device 100, as further described with reference to Figures 3 to 5, and may access the personal acoustic model from the model library 162, memory 108, or both. For illustrative purposes, one or more processors 110 may be configured to download a personal acoustic model (e.g., acoustic model 250 in model library 162) from the acoustic model library in response to the decision that device 100 has entered the vehicle. In some implementations, one or more processors 110 are further configured to remove the personal acoustic model in response to device 100 leaving the vehicle.
[0060]
[0083] For example, one or more processors 110 may be configured to determine that device 100 has entered the vehicle based on sensor data 138. For illustrative purposes, one or more processors 110 may determine that device 100 has entered (or left) the vehicle based on location data 153.
[0061]
[0084] In conjunction with the various embodiments described with reference to Figures 1 and 2, one or more processors 110 of device 100 are configured to download an acoustic model corresponding to a specific location where device 100 is located, process an input audio signal using the acoustic model, and remove the acoustic model in response to device 100 leaving the location. In an exemplary example, the location corresponds to a specific restaurant, and in response to the decision that device 100 has entered the specific restaurant, an acoustic model (for example, acoustic model 250 in model library 162) is downloaded from the acoustic model library.
[0062]
[0085] In conjunction with the various embodiments described with reference to Figures 1 and 2, one or more processors 110 of device 100 are configured to select an acoustic model corresponding to a specific location, to receive permission for the acoustic model based at least partially on the location of device 100 that matches the specific location, and to process an input audio signal using the acoustic model.
[0063]
[0086] In some implementations, the device 100 in Figures 1 and 2 may be further configured to update one or more models, such as to personalize a model for a particular user or to improve the accuracy of a model for environments that the device 100 frequently encounters, as an exemplary and non-limiting example. Figures 3–5 show exemplary examples of how the device 100 is configured to update models. While Figures 3–5 illustrate updating a sound event classification model as a specific example, the techniques described are generally applicable to updating any type of model that may be used by the device 100.
[0064]
[0087] Figure 3 is a block diagram of an example of the components of a device 100 configured to generate sound identification information data in response to an audio data sample 310 and to update a sound event classification model. The device 100 in Figure 3 includes one or more microphones 304 (e.g., microphone 104) configured to generate an audio signal 306 (e.g., audio data 105) based on a sound 302 detected in an acoustic environment. Microphones 304 are coupled to a feature extractor 308 that generates an audio data sample 310 based on the audio signal 306. For example, an audio data sample 310 may include an array or matrix of data elements, each data element corresponding to a feature detected in the audio signal 306. In a particular example, an audio data sample 310 may correspond to a mel-spectral feature extracted from one second of audio signal 306. In this example, an audio data sample 310 may include a 128 × 128 element matrix of feature values. In other examples, other audio data sample configurations or sizes may be used.
[0065]
[0088] An audio data sample 310 is provided to a sound event classification (SEC) engine 320 (for example, a model-based application 192). The SEC engine 320 is configured to perform inference operations based on one or more SEC models, such as SEC model 312. If the sound class of the audio data sample 310 is recognized by SEC model 312, “inference operations” refer to assigning the audio data sample 310 to a sound class. For example, the SEC engine 320 may include or correspond to software that implements a machine learning runtime environment, such as the Qualcomm Neural Processing SDK, which is available from Qualcomm Technologies, Inc. in San Diego, California, USA. In certain embodiments, SEC model 312 is one of several SEC models (for example, available SEC model 314) that are available to the SEC engine 320.
[0066]
[0089] In a particular example, each of the available SEC models 314 (stored, for example, in memory 108 or model library 162) contains or corresponds to a neural network trained as a sound event classifier. For illustrative purposes, SEC model 312 (and each of the other available SEC models 314) may include an input layer, one or more hidden layers, and an output layer. In this example, the input layer is configured to correspond to an array or matrix of values of audio data samples 310 generated by the feature extractor 308. For illustrative purposes, if the audio data samples 310 contain 15 data elements, the input layer may contain 15 nodes (for example, one for each data element). The output layer is configured to correspond to a sound class trained for recognition by SEC model 312. The particular configuration of the output layer may vary depending on the information given as output. For example, the SEC model 312 may be trained to output an array containing one bit for each sound class, where the output layer performs "one-hot coding" such that all but one bit of the output array has a value of 0, and the bit corresponding to the detected sound class has a value of 1. Another output method may be used, for example, to indicate a confidence metric value for each sound class, where the confidence metric value represents the probability estimate that the audio data sample 310 corresponds to each sound class. For illustrative purposes, if the SEC model 312 is trained to recognize four sound classes, the SEC model 312 may generate output data containing four values (one for each sound class), where each value may represent the probability estimate that the audio data sample 310 corresponds to each sound class.
[0067]
[0090] Each hidden layer contains multiple nodes, each node interconnected (via links) with other nodes in the same or different layers. Each input link of a node is associated with a link weight. During operation, a node receives input values from the other nodes it is linked to, weights the input values based on the corresponding link weights to determine the combined value, and then applies the combined value to an activation function to produce the node's output value. The output value is provided to one or more other nodes via the node's output link. A node may also contain bias values used to generate the combined value. Nodes can be linked in various configurations and may include various other features (e.g., memory of previous values) to facilitate the processing of specific data. For audio data samples, a convolutional neural network (CNN) may be used. For illustrative purposes, one or more of the SEC Model 312 may contain three linked CNNs, each CNN may contain a two-dimensional (2D) convolutional layer, a Maximum Pulin layer, and a Batch Normalization layer. In other implementations, hidden layers may contain a different number of CNNs or other layers. Training a neural network involves modifying link weights to reduce the output error of the neural network.
[0068]
[0091] During operation, the SEC engine 320 may provide an audio data sample 310 as input to a single SEC model (e.g., SEC model 312), to multiple selected SEC models (e.g., SEC model 312 and the Kth SEC model 318 of the available SEC models 314), or to each of the SEC models (e.g., SEC model 312, the first SEC model 316, the Kth SEC model 318, and any other SEC model of the available SEC models 314). For example, the SEC engine 320 (or another component of device 100) may select SEC model 312 from the available SEC models 314 based on, for example, user input, device settings associated with device 100, sensor data, the time at which the audio data sample 310 is received, or other factors. In this example, the SEC engine 320 may choose to use only SEC model 312, or it may choose to use two or more of the available SEC models 314. For illustrative purposes, the device settings may indicate that SEC model 312 and the first SEC model 316 will be used during a particular time frame. In another example, the SEC engine 320 may provide audio data samples 310 (for example, sequentially or in parallel) to each of the SEC models 314 available for generating outputs from each. In certain embodiments, the SEC models are trained to recognize different sound classes, recognize the same sound class in different acoustic environments, or both. For example, SEC model 312 may be configured to recognize a first set of sound classes, and a first SEC model 316 may be configured to recognize a second set of sound classes, where the first set of sound classes is different from the second set of sound classes.
[0069]
[0092] In certain embodiments, the SEC engine 320 determines, based on the output of SEC model 312, whether SEC model 312 has recognized the sound class of the audio data sample 310. If the SEC engine 320 provides the audio data sample 310 to multiple SEC models, the SEC engine 320 may determine, based on the output of each SEC model, whether any of the SEC models has recognized the sound class of the audio data sample 310. If SEC model 312 (or another of the available SEC models 314) has recognized the sound class of the audio data sample 310, the SEC engine 320 generates an output 324 indicating the sound class 322 of the audio data sample 310. For example, output 324 may be sent to a display to notify the user of the detection of sound class 322 associated with sound 302, or it may be sent to another device or another component of device 100 to trigger an action (for example, sending a command to activate a light in response to recognizing the sound of a door closing).
[0070]
[0093] If the SEC engine 320 determines that SEC model 312 (and other available SEC models 314 given audio data sample 310) did not recognize the sound class of audio data sample 310, the SEC engine 320 gives a trigger signal 326 to the drift detector 328. For example, the SEC engine 320 may set a trigger flag in the memory of device 100. In some implementations, the SEC engine 320 may also provide other data to the drift detector 328. For example, if SEC model 312 generates confidence metric values for each sound class it has been trained to recognize, one or more of these confidence metric values may be given to the drift detector 328. For example, if SEC model 312 is trained to recognize three sound classes, the SEC engine 320 may give the drift detector 328 the highest confidence value out of the three confidence values output by SEC model 312 (one for each of the three sound classes).
[0071]
[0094] In certain embodiments, the SEC engine 320 determines whether the SEC model 312 has recognized a sound class of the audio data sample 310 based on a confidence metric value. In this particular embodiment, the confidence metric value for a particular sound class indicates the probability that the audio data sample 310 is associated with that particular sound class. For illustrative purposes, if the SEC model 312 is trained to recognize four sound classes, the SEC model 312 may produce an array as output containing one of the four confidence metric values for each sound class. In some implementations, if the confidence metric value for sound class 322 is greater than a detection threshold, the SEC engine 320 determines that the SEC model 312 has recognized sound class 322 of the audio data sample 310. For example, if the confidence metric value for sound class 322 is greater than the detection threshold of 0.90 (e.g., 90% confidence), 0.95 (e.g., 95% confidence), or any other value, the SEC engine 320 determines that the SEC model 312 recognized sound class 322 for the audio data sample 310. In some implementations, if the confidence metric value for each sound class that the SEC model 312 was trained to recognize is less than the detection threshold, the SEC engine 320 determines that the SEC model 312 did not recognize the sound class for the audio data sample 310. For example, if each value of the confidence metric is less than the detection threshold of 0.90 (e.g., 90% confidence), 0.95 (e.g., 95% confidence), or any other value, the SEC engine 320 determines that the SEC model 312 did not recognize sound class 322 for the audio data sample 310.
[0072]
[0095] The drift detector 328 is configured to determine whether an SEC model 312, which was unable to recognize the sound class of the audio data sample 310, corresponds to an audio scene 342 associated with the audio data sample 310. In the example shown in Figure 1, the scene detector 340 (e.g., context detector 140) is configured to receive scene data 338 (e.g., a portion of sensor data 138) and to use the scene data 338 to determine an audio scene 342 (e.g., context 142) associated with the audio data sample 310. In certain embodiments, the scene data 338 is generated based on configuration data 330 indicating one or more device settings associated with device 100, the output of clock 332, sensor data from one or more sensors 334 (e.g., sensor 134), inputs received via input device 336, or a combination thereof. In some embodiments, the scene detector 340 uses different information to determine the audio scene 342 than the information used by the SEC engine 320 to select an SEC model 312. For illustrative purposes, if the SEC engine 320 selects an SEC model 312 based on time, the scene detector 340 may use position sensor data from the position sensor of sensor 334 to determine the audio scene 342. In some embodiments, the scene detector 340 uses at least some of the same information that the SEC engine 320 uses to select an SEC model 312, plus additional information. For illustrative purposes, if the SEC engine 320 selects an SEC model 312 based on time and setting data 330, the scene detector 340 may use position sensor data and setting data 330 to determine the audio scene 342. Thus, the scene detector 340 uses a different audio scene detection mode than that used by the SEC engine 320 to select an SEC model 312.
[0073]
[0096] In certain implementations, the scene detector 340 is a neural network trained to determine an audio scene 342 based on scene data 338. In other implementations, the scene detector 340 is a classifier trained using different machine learning techniques. For example, the scene detector 340 may include, or correspond to, a decision tree, random forest, support vector machine, or another classifier trained to produce an output indicating an audio scene 342 based on scene data 338. In yet another implementation, the scene detector 340 uses heuristics to determine an audio scene 342 based on scene data 338. In yet another implementation, the scene detector 340 uses a combination of artificial intelligence and heuristics to determine an audio scene 342 based on scene data 338. For example, the scene data 338 may include image data, video data, or both, and the scene detector 340 may include an image recognition model trained using machine learning techniques to detect specific objects, motion, background, or other image or video information. In this example, the output of the image recognition model may be evaluated via one or more heuristics to determine an audio scene 342.
[0074]
[0097] The drift detector 328 compares information describing the SEC model 312 with the audio scene 342 indicated by the scene detector 340 to determine whether the SEC model 312 is associated with the audio scene 342 of the audio data sample 310. If the drift detector 328 determines that the SEC model 312 is associated with the audio scene 342 of the audio data sample 310, the drift detector 328 stores the drift data 344 as model update data 348. In a particular implementation, the drift data 344 includes the audio data sample 310 and a label, where the label identifies the SEC model 312, indicates a sound class associated with the audio data sample 310, or both. If the drift data 344 indicates a sound class associated with the audio data sample 310, the sound class may be selected based on the highest value of the confidence metric generated by the SEC model 312. As an exemplary example, if the SEC engine 320 uses a detection threshold of 0.90 and the highest reliability metric output by the SEC model 312 is 0.85 for a particular sound class, the SEC engine 320 determines that the sound class of the audio data sample 310 was not recognized and sends a trigger signal 326 to the drift detector 328. In this example, if the drift detector 328 determines that the SEC model 312 corresponds to an audio scene 342 of the audio data sample 310, the drift detector 328 stores the audio data sample 310 as drift data 344 associated with a particular sound class. In certain embodiments, metadata associated with the SEC model 314 includes information specifying one or more audio scenes associated with each SEC model 314. For example, the SEC model 312 may be configured to detect sound events in a user's home, in which case the metadata associated with the SEC model 312 may indicate that the SEC model 312 is associated with the “home” audio scene.In this example, if the audio scene 342 indicates that device 100 is at the home location (for example, based on location information, user input, detection of the home wireless network signal, or image or video data representing the home location), the drift detector 328 determines that SEC model 312 corresponds to the audio scene 342.
[0075]
[0098] In some implementations, the drift detector 328 also stores some audio data samples 310 as model update data 348 and designates them as unknown data 346. As a first example, if the drift detector 328 determines that the SEC model 312 does not correspond to the audio scene 342 of the audio data sample 310, the drift detector 328 may store unknown data 346. As a second example, if the reliability metric value output by the SEC model 312 fails to meet the drift threshold, the drift detector 328 may store unknown data 346. In this example, the drift threshold is smaller than the detection threshold used by the SEC engine 320. For example, if the SEC engine 320 uses a detection threshold of 0.95, the drift threshold could be 0.80, 0.75, or some other value lower than 0.95. In this example, if the highest value of the reliability metric for audio data sample 310 is less than the drift threshold, the drift detector 328 determines that audio data sample 310 belongs to a sound class that the SEC model 312 has not been trained to recognize, and designates audio data sample 310 as unknown data 346. In a particular embodiment, if the drift detector 328 determines that the SEC model 312 corresponds to audio scene 342 of audio data sample 310, the drift detector 328 simply stores unknown data 346. In another example, the drift detector 328 stores unknown data 346 regardless of whether the drift detector 328 determines that the SEC model 312 corresponds to audio scene 342 of audio data sample 310.
[0076]
[0099] After the model update data 348 is stored, the model update data 352 can access the model update data 348 and use it to update one of the available SEC models 314 (for example, SEC model 312). For example, each entry in the model update data 348 indicates the SEC model to which the entry is associated, and the model update data 352 uses the entry as training data to update the corresponding SEC model. In certain embodiments, the model update data 352 updates the SEC model when the update criteria are met, or when a model update is initiated by a user or another party (for example, the vendor of device 100, SEC engine 320, SEC model 314, etc.). The update criteria may be met when a certain number of entries are available in model update data 348, when a certain number of entries for a particular SEC model are available in model update data 348, when a certain number of entries for a particular sound class are available in model update data 348, when a certain amount of time has elapsed since the previous update, when another update is performed (for example, when a software update related to device 100 is performed), or based on the occurrence of another event.
[0077]
[0100] Model update data 352 uses drift data 344 as labeled training data to update the training of SEC model 312 using backpropagation or a similar machine learning optimization process. For example, model update data 352 provides audio data samples from drift data 344 of model update data 348 as input to SEC model 312, determines the value of the error function (also called the loss function) based on the output of SEC model 312 and the labels associated with the audio data samples (as shown in drift data 344 stored by drift detector 328), and determines updated link weights for SEC model 312 using gradient descent (or some variation thereof) or another machine learning optimization process.
[0078]
[0101] Model update data 352 may also provide the SEC model 312 with other audio data samples (in addition to the audio data samples from drift data 344) during update training. For example, model update data 348 may include one or more known audio data samples (such as a subset of the audio data samples initially used to train the SEC model 312), which may reduce the likelihood of update training causing the SEC model 312 to forget previous training (where “forgetting” here means losing confidence in detecting sound classes previously trained for the SEC model 312 to recognize). Since the sound classes associated with the audio data samples in drift data 344 are indicated by the drift detector 328, update training that takes drift into account can be achieved automatically (e.g., without user input). Thus, the functionality of device 100 (e.g., accuracy in recognizing sound classes) can be improved over time with fewer computing resources that would otherwise be used to generate a new SEC model from scratch, without user intervention. A specific example of a transfer learning process in which model-up data 352 can be used to update the SEC model 312 based on drift data 344 is described with reference to Figure 4.
[0079]
[0102] In some embodiments, the model update data 352 can also use the unknown data 346 of the model update data 348 to update the training of the SEC model 312. For example, periodically or from time to time, such as when an update criterion is met, the model update data 352 may prompt the user to label the sound class of the entries in the unknown data 346 in the model update data 348. If the user chooses to label the sound class of the entries in the unknown data 346, device 100 (or another device) may play out the sound corresponding to the audio data sample of the unknown data 346. The user can provide one or more labels 350 (for example, via input device 336) that identify the sound class of the audio data sample. If the sound class indicated by the user is a sound class that the SEC model 312 has been trained to recognize, the unknown data 346 is reclassified as drift data 344 associated with the user-specified sound class and the SEC model 312. Depending on the configuration of the model update data 352, if the sound class indicated by the user is a sound class that has not been trained for the SEC model 312 to recognize (for example, a new sound class), the model update data 352 may discard the unknown data 346, send the unknown data 346 and the user-specified sound class to another device for use in generating a new or updated SEC model, or use the unknown data 346 and the user-specified sound class to update the SEC model 312. A specific example of a transfer learning process in which the model update data 352 can be used to update the SEC model 312 based on the unknown data 346 and the user-specified sound class is described with reference to Figure 5.
[0080]
[0103] The updated SEC model 354, generated by the model update data 352, is added to the available SEC models 314 to make available to the updated SEC model 354 for evaluating audio data samples 310 received after the updated SEC model 354 was generated. Thus, the set of available SEC models 314 that can be used to evaluate sounds is dynamic. For example, one or more of the available SEC models 314 may be automatically updated to take drift data 344 into consideration. Furthermore, one or more of the available SEC models 314 may be updated to take unknown sound classes using transfer learning operations, which use fewer computing resources (e.g., memory, processing time, and power) than training a new SEC model from scratch.
[0081]
[0104] Figure 4 shows an embodiment of updating SEC Model 408 to account for drift according to a specific example. SEC Model 408 in Figure 4 includes or corresponds to one of the available SEC Models 314 in Figure 3, associated with drift data 344. For example, if SEC Engine 320 generates a trigger signal 326 in response to the output of SEC Model 312, then drift data 344 is associated with SEC Model 312, and SEC Model 408 corresponds to or includes SEC Model 312. As another example, if SEC Engine 320 generates a trigger signal 326 in response to the output of the Kth SEC Model 318, then drift data 344 is associated with the Kth SEC Model 318, and SEC Model 408 corresponds to or includes the Kth SEC Model 318.
[0082]
[0105] In the example shown in Figure 4, training data 402 is used to update the SEC model 408. Training data 402 includes drift data 344 and one or more labels 404. Each entry in the drift data 344 contains an audio data sample (e.g., audio data sample 406) and is associated with a corresponding label among the (one or more) labels 404. The audio data sample of an entry in the drift data 344 contains a set of values representing features extracted from or determined based on sounds that were not recognized by the SEC model 408. The label 404 corresponding to an entry in the drift data 344 identifies the sound class to which the sound is expected to belong. For example, in response to the SEC model 408 determining that the audio data sample corresponds to the audio scene to which it was generated, the label 404 corresponding to an entry in the drift data 344 may be assigned by the drift detector 328 in Figure 3. In this example, the drift detector 328 may assign the audio data sample to the sound class associated in the output of the SEC model 408 with the highest confidence metric value.
[0083]
[0106] In Figure 4, an audio data sample 406 corresponding to a sound is given to the SEC model 408, which generates an output 410 indicating the sound class to which the audio data sample 406 is assigned, one or more values of a reliability metric, or both. Model update data 352 uses the output 410 and the label 404 corresponding to the audio data sample 406 to determine updated link weights 412 for the SEC model 408. The SEC model 408 is updated based on the updated link weights 412, and the training process is repeated iteratively until the training termination condition is met. During training, each entry of the drift data 344 (e.g., one entry per iteration) may be given to the SEC model 408. Furthermore, in some implementations, other audio data samples (e.g., audio data samples previously used to train the SEC model 408) may also be given to the SEC model 408 to reduce the likelihood that the SEC model 408 will forget previous training.
[0084]
[0107] The training termination condition may be met when all of the drift data 344 has been given to the SEC model 408 at least once, when a certain number of training iterations have been performed, when the convergence metric meets the convergence threshold, or when any other condition indicating the end of training is met. When the training termination condition is met, the model update data 352 stores the updated SEC model 414, where the updated SEC model 414 corresponds to the SEC model 408 with link weights based on the updated link weights 412 applied during training.
[0085]
[0108] Figure 5 shows an embodiment of updating the SEC model 510 based on training data 502 to account for unknown data according to a specific example. The SEC model 510 in Figure 5 includes or corresponds to a specific one of the available SEC models 314 in Figure 3 that is associated with the unknown data 346. For example, if the SEC engine 320 generates a trigger signal 326 in response to the output of SEC model 312, then the unknown data 346 is associated with SEC model 312, and SEC model 510 corresponds to or includes SEC model 312. In another example, if the SEC engine 320 generates a trigger signal 326 in response to the output of the kth SEC model 318, then the unknown data 346 is associated with the kth SEC model 318, and SEC model 510 corresponds to or includes the kth SEC model 318.
[0086]
[0109] In the example in Figure 5, the model update data 352 generates an updated model 506. The updated model 506 includes the SEC model 510 to be updated, an incremental model 508, and one or more adapter networks 512. The incremental model 508 is a copy of the SEC model 510 with a different output layer. In particular, the output layer of the incremental model 508 contains more output nodes than the output layer of the SEC model 510. For example, the output layer of the SEC model 510 contains a first number of nodes (e.g., N nodes, where N is a positive integer corresponding to the number of sound classes that the SEC model 510 is trained to recognize), and the output layer of the incremental model 508 contains a second number of nodes (e.g., N+M nodes, where M is a positive integer corresponding to the number of new sound classes that the updated SEC model 524 will be trained to recognize, which the SEC model 510 is not trained to recognize). The first number of nodes corresponds to the number of sound classes in the first set of sound classes trained for SEC model 510 to recognize (for example, the first set of sound classes contains N distinct sound classes that SEC model 510 can recognize), and the second number of nodes corresponds to the number of sound classes in the second set of sound classes that will be trained for the updated SEC model 524 to recognize (for example, the second set of sound classes contains N+M distinct sound classes that will be trained for the updated SEC model 524 to recognize). The second set of sound classes includes the first set of sound classes (for example, N classes) plus one or more additional sound classes (for example, M classes). The model parameters of incremental model 508 (for example, link weights) are initialized to be equal to the model parameters of SEC model 510.
[0087]
[0110] The adapter network 512 includes a neural adapter and a merger adapter. The neural adapter includes one or more adapter layers configured to receive input from the SEC model 510 and to generate an output that can be merged with the output of the incremental model 508. For example, the SEC model 510 generates a first output corresponding to a first number of classes of a first set of sound classes. In a particular embodiment, the first output contains one data element per node of the output layer of the SEC model 510 (e.g., N data elements). In contrast, the incremental model 508 generates a second output corresponding to a second number of classes of a second set of sound classes. For example, the second output contains one data element per node of the output layer of the incremental model 508 (e.g., N+M data elements). In this example, the adapter layers of the adapter network 512 receive the output of the SEC model 510 as input and generate an output having a second number of data elements (e.g., N+M). In a specific example, the adapter layer of adapter network 512 includes two fully connected layers (for example, an input layer containing N nodes where each node in the input layer is connected to every node in the output layer, and an output layer containing N+M nodes).
[0088]
[0111] The merger adapter of the adapter network 512 is configured to generate the output 514 of the updated model 506 by merging the output of the adapter layer with the output of the incremental model 508. For example, the merger adapter combines the output of the adapter layer and the output of the incremental model 508 element by element to generate a combined output, and then applies an activation function (such as a sigmoid function) to the combined output to generate output 514. Output 514 indicates the sound class to which the audio data sample 504 is assigned by the updated model 506, one or more confidence metric values determined by the updated model 506, or both.
[0089]
[0112] The model update data 352 uses the output 514 and the labels 350 corresponding to the audio data samples 504 to determine the updated link weights 516, adapter network 512, or both for the incremental model 508. The link weights of the SEC model 510 remain constant during training. The training process is repeated iteratively until the training termination condition is met. During training, each of the entries of the unknown data 346 (e.g., one entry per iteration) may be provided to the update model 506. Furthermore, in some implementations, other audio data samples (e.g., audio data samples previously used to train the SEC model 510) may also be provided to the update model 506 to reduce the likelihood that the incremental model 508 will forget previous training of the SEC model 510.
[0090]
[0113] The training termination condition may be met when all of the unknown data 346 have been given to the updated model 506 at least once, when a certain number of training iterations have been performed, when the convergence metric meets the convergence threshold, or when any other condition indicating the end of training is met. When the training termination condition is met, the model checker 520 selects the updated SEC model 524 from among the incremental model 508 and the updated model 506 (for example, a combination of SEC model 510, incremental model 508, and adapter network 512).
[0091]
[0114] In certain embodiments, the model checker 520 selects an updated SEC model 524 based on the precision of the sound class 522 assigned by the incremental model 508 and the precision of the sound class 522 assigned by the SEC model 510. For example, the model checker 520 may determine the F1 score for the incremental model 508 (based on the sound class 522 assigned by the incremental model 508) and the F1 score for the SEC model 510 (based on the sound class 522 assigned by the SEC model 510). In this example, if the value of the F1 score for the incremental model 508 is greater than or equal to the value of the F1 score for the SEC model 510, the model checker 520 selects the incremental model 508 as the updated SEC model 524. In some implementations, if the F1 score of the incremental model 508 is greater than or equal to the F1 score of the SEC model 510 (or less than the F1 score of the SEC model 510 by a threshold amount), the model checker 520 selects the incremental model 508 as the updated SEC model 524. If the F1 score of the incremental model 508 is less than the F1 score for the SEC model 510 (or less than the F1 score for the SEC model 510 by a threshold amount), the model checker 520 selects the updated model 506 as the updated SEC model 524. If the incremental model 508 is selected as the updated SEC model 524, the SEC model 510, the adapter network 512, or both may be discarded.
[0092]
[0115] In some implementations, the model checker 520 is omitted or integrated with the model update data 352. For example, after training the update model 506, the update model 506 may be stored as the updated SEC model 524 (without, for example, a choice between the update model 506 and the incremental model 508). As an example, while training the update model 506, the model update data 352 may determine the accuracy metric for the incremental model 508. In this example, the training termination condition may be based on the accuracy metric for the incremental model 508, so that after training, the incremental model 508 is stored as the updated SEC model 524 (without, for example, a choice between the update model 506 and the incremental model 508).
[0093]
[0116] Using the transfer learning technique described with reference to Figure 5, the model checker 520 enables device 100 in Figure 3 to update the SEC model to recognize previously unknown sound classes. Furthermore, the transfer learning technique described uses significantly fewer computing resources (e.g., memory, processing time, and power) than would be used to train the SEC model from scratch to recognize previously unknown sound classes.
[0094]
[0117] In some implementations, the operation described with reference to Figure 4 (e.g., generating an updated SEC model 414 based on drift data 344) is performed in device 100 in Figure 3 (e.g., on one or more processors 110), while the operation described with reference to Figure 5 (e.g., generating an updated SEC model 524 based on unknown data 346) is performed in a different device (e.g., a remote computer device 818 in Figure 8). For illustrative purposes, the unknown data 346 and label 350 may be captured in device 100 and sent to a second device with more available computing resources. In this example, the second device generates the updated SEC model 524, and device 100 downloads or receives a transmission or data representing the updated SEC model 524 from the second device. Generating an updated SEC model 524 based on unknown data 346 is a more resource-intensive process than generating an updated SEC model 414 based on drift data 344 (e.g., using more memory, power, and processor time). Therefore, separating the operation described with reference to Figure 4 and the operation described with reference to Figure 5 between different devices can save resources on device 100.
[0095]
[0118] Figure 6 shows a specific example of the operation of device 100 in Figure 1, where the determination of whether the active SEC model (e.g., SEC model 312) corresponds to the audio scene in which the audio data sample 310 is captured is based on comparing the current audio scene with the previous audio scene.
[0096]
[0119] In Figure 6, audio data captured by microphone 104 is used to generate audio data samples 310. Audio data samples 310 are used to perform audio classification 602. For example, one or more of the available SEC models 314 are used as the active SEC model by the SEC engine 320 in Figure 3. In a particular embodiment, the active SEC model is selected from the available SEC models 314 based on the audio scene indicated by the scene detector 340 during a previous sampling period, also called the previous audio scene 608.
[0097]
[0120] The audio classification 602 generates result 604 based on an analysis of the audio data sample 310 using an active SEC model. Result 604 may indicate the sound class associated with the audio data sample 310, the probability that the audio data sample 310 corresponds to a specific sound class, or that the sound class of the audio data sample 310 is unknown. If result 604 indicates that the audio data sample 310 corresponds to a known sound class, then in block 606, a decision is made to generate output 324 indicating the sound class 322 associated with the audio data sample 310. For example, the SEC engine 320 in Figure 1 may generate output 324.
[0098]
[0121] If result 604 indicates that the audio data sample 310 does not correspond to a known sound class, a decision is made in block 606 to generate a trigger 326. The trigger 326 activates a drift detection scheme, which in Figure 6 involves causing the scene detector 340 to identify the current audio scene 607 based on data from sensor 134.
[0099]
[0122] In block 610, the current audio scene 607 is compared to the previous audio scene 608 to determine whether the audio scene has been changed since the active SEC model was selected. In block 612, a determination is made as to whether the sound class of the audio data sample 310 was not recognized due to drift. For example, if the current audio scene 607 does not correspond to the previous audio scene 608, the determination in block 612 is that drift was not the reason the sound class of the audio data sample 310 was not recognized. In this situation, the audio data sample 310 may be discarded or stored as unknown data in block 614.
[0100]
[0123] If the current audio scene 607 corresponds to the previous audio scene 608, the decision in block 612 is that the active SEC model corresponds to the current audio scene 607, and therefore the sound class of audio data sample 310 was not recognized due to drift. In this situation, the drifted sound class is identified in 616, and in block 618, the audio data sample 310 and the sound class identifier are stored as drift data.
[0101]
[0124] Once sufficient drift data has been stored, the SEC model is updated in block 620 to generate an updated SEC model 354. The updated SEC model 354 is added to the available SEC model 314. In some implementations, the updated SEC model 354 replaces the active SEC model that generated result 604.
[0102]
[0125] Figure 6 shows another specific example of the operation of device 100 in Figure 1, where the determination of whether the active SEC model (e.g., SEC model 312) corresponds to the audio scene in which the audio data sample 310 is captured is based on comparing the current audio scene with information describing the active SEC model.
[0103]
[0126] In Figure 6, the audio data captured by the microphone 104 is used to generate an audio data sample 310. The audio data sample 310 is used to perform audio classification 602. For example, one or more of the available SEC models 314 are used as the active SEC model by the SEC engine 320 in Figure 3. In certain embodiments, the active SEC model is selected from among the available SEC models 314. In some implementations, an ensemble of available SEC models 314 is used rather than selecting one or more SEC models 314 as the active SEC model.
[0104]
[0127] The audio classification 602 generates a result 604 based on an analysis of the audio data sample 310 using one or more of the available SEC models 314. The result 604 may indicate the sound class associated with the audio data sample 310, the probability that the audio data sample 310 corresponds to a particular sound class, or that the sound class of the audio data sample 310 is unknown. If the result 604 indicates that the audio data sample 310 corresponds to a known sound class, then in block 606, a decision is made to generate an output 324 indicating the sound class 322 associated with the audio data sample 310. For example, the SEC engine 320 in Figure 3 may generate an output 324.
[0105]
[0128] If result 604 indicates that the audio data sample 310 does not correspond to a known sound class, a decision is made in block 606 to generate a trigger 326. Trigger 326 activates a drift detection scheme in Figure 7, which involves causing the scene detector 340 to identify the current audio scene based on data from sensor 134 and to determine whether the current audio scene corresponds to the SEC model that generated result 604, prompting trigger 326 to be sent.
[0106]
[0129] In block 612, a decision is made as to whether the sound class of audio data sample 310 was not recognized due to drift. For example, if the current audio scene does not correspond to the SEC model that produced result 604, the decision in block 612 will be that drift was not the reason the sound class of audio data sample 310 was not recognized. In this situation, audio data sample 310 may be discarded or stored as unknown data in block 614.
[0107]
[0130] If the current audio scene corresponds to the SEC model that generated result 604, the decision in block 612 is that the sound class of audio data sample 310 was not recognized due to drift. In this situation, in 616, the drifted sound class is block-identified, and in block 618, the audio data sample 310 and the sound class identifier are stored as drift data.
[0108]
[0131] Once sufficient drift data has been stored, the SEC model is updated in block 620 to generate an updated SEC model 354. The updated SEC model 354 is added to the available SEC model 314. In some implementations, the updated SEC model 354 replaces the active SEC model that generated result 604.
[0109]
[0132] The operation for updating models, as described with reference to Figures 3-7, can be used in conjunction with context-based model selection, as described with reference to Figures 1-2. For example, a first person and a second person living together in a first residence (e.g., a rural house) may have devices that use a common model relevant to the first residence, and such a model can be updated by training the model to adapt to new sound events and drifts relevant to the first residence. After the first person moves to a second residence (e.g., a college dormitory), the first person's device may update one or more models for improved accuracy in the second residence, and thus the models used by the first person's and second person's devices may diverge significantly. When the first person returns to the first residence, the first person's device may select and download models to be used by the second person's device to achieve higher accuracy during use in the first residence. For example, the second person may grant permission to share one or more models with the first person's device, such as via peer-to-peer transfer between devices or via a local wireless home network. Upon leaving the first residence, the first person's device may remove the shared model associated with the first residence and revert to the model associated with the second residence.
[0110]
[0133] Figure 8 is a block diagram showing a specific example of device 100 in Figure 1. In various implementations, device 100 may have more or fewer components than those shown in Figure 8.
[0111]
[0134] In certain implementations, device 100 includes a processor 804 (e.g., a central processing unit (CPU)). Device 100 may include one or more additional processors 806 (e.g., one or more digital signal processors (DSPs)). Processor 804, processor 806, or both may correspond to one or more processors 110. For example, in Figure 8, processor 806 includes a context detector 140, a model selector 190, and a model-based application 192.
[0112]
[0135] In Figure 8, device 100 also includes memory 108 and codec 824. Memory 108 stores instructions 860 that can be executed by processor 804 or processor 806 to implement one or more operations described with reference to Figures 1 to 7. In one example, memory 108 corresponds to a non-temporary computer-readable medium that stores instructions 860 that can be executed by one or more processors 110, and the instructions 860 include or correspond to a context detector 140, a model selector 190, a model-based application 192, or a combination thereof (for example, that can be executed by a processor to perform an operation resulting therefrom). Memory 108 may also store available models 114.
[0113]
[0136] In Figure 8, the speaker 822 and microphone 104 can be coupled to the codec 824. In the example shown in Figure 8, the codec 824 includes a digital-to-analog converter (DAC 826) and an analog-to-digital converter (ADC 828). In a particular implementation, the codec 824 receives an analog signal from the microphone 104, uses the ADC 828 to convert the analog signal to a digital signal, and feeds the digital signal to the processor 806. In a particular implementation, the processor 806 feeds the digital signal to the codec 824, which uses the DAC 826 to convert the digital signal to an analog signal and feeds the analog signal to the speaker 822.
[0114]
[0137] In Figure 8, device 100 also includes an input device 336. Device 100 may also include a display 820 coupled to a display controller 810. In certain embodiments, the input device 336 includes sensors, keyboards, pointing devices, etc. In some implementations, the input device 336 and the display 820 are combined in a touchscreen or similar touch or motion-sensitive display.
[0115]
[0138] In some implementations, device 100 also includes a modem 812 coupled to a transceiver 814. In Figure 8, the transceiver 814 is coupled to an antenna 816 to enable wireless communication with other devices, such as a remote computer device 818 (for example, a server or network memory storing at least a portion of the model library 162). For example, the modem 812 may be configured to receive, at least partially based on the location of device 100 matching a specific location via wireless transmission, a model 112, permission for a model 112, or both. In other examples, the transceiver 814 is coupled to a communication port (for example, an Ethernet® port) to enable wired communication with other devices, such as the remote computer device 818.
[0116]
[0139] In Figure 8, device 100 includes a clock 332 and a sensor 134. As a specific example, the sensor 134 includes one or more cameras 150, one or more location sensors 152, a microphone 104, an activity detector 154, other sensors 156, or a combination thereof.
[0117]
[0140] In certain embodiments, the clock 332 generates a clock signal that can be used to assign a timestamp to a particular sensor data sample to indicate when that particular sensor data sample was received. In this embodiment, the model selector 190 can use the timestamp to select the model to be used to process the input data. Additionally or alternatively, the timestamp can be used by the context detector 140 to determine the context 142 associated with a particular sensor data sample.
[0118]
[0141] In certain embodiments, one or more cameras 150 generate image data, video data, or both. A model selector 190 can use the image data, video data, or both to select a specific model to be used to analyze the input data. Additionally or alternatively, the image data, video data, or both may be used by a context detector 140 to determine a context 142 associated with a particular sensor data sample. For example, a particular model 112 may be designated for outdoor use, and the image data, video data, or both may be used to verify that the device 100 is located in an outdoor environment.
[0119]
[0142] In certain embodiments, one or more location sensors 152 generate location data, such as global location data indicating the location of device 100. A model selector 190 can use the location data to select a specific model to be used to analyze the input data. Additionally or alternatively, location data may be used by a context detector 140 to determine a context 142 associated with a particular sensor data sample. For example, a particular model 112 may be designated for home use, and location data may be used to verify that device 100 is located at the home location. The location sensors 852 may include receivers for satellite-based positioning systems, receivers for local positioning system receivers, inertial navigation systems, landmark-based positioning systems, or a combination thereof.
[0120]
[0143] (One or more) other sensors 156 may be coupled to or included within device 100 and used to generate sensor data useful for determining the context 142 related to device 100 at a particular time, for example, an orientation sensor, a magnetometer, a light sensor, a contact sensor, a temperature sensor, or any other sensor.
[0121]
[0144] In certain implementations, device 100 is included in a system-in-package or system-on-chip device 802. In a certain implementation, memory 108, processor 804, (one or more) processors 806, display controller 810, codec 824, modem 812, and transceiver 814 are included in a system-in-package or system-on-chip device 802. In certain implementations, input device 336 and power supply 830 are coupled to the system-on-chip device 802. Furthermore, in certain implementations, as shown in Figure 8, the display 820, input device 336, (one or more) speakers 822, sensor 134, clock 332, antenna 816, and power supply 830 are outside the system-on-chip device 802. In certain implementations, each of the display 820, input device 336, speaker 822, sensor 134, clock 332, antenna 816, and power supply 830 may be coupled to components of the system-on-chip device 802, such as an interface or controller.
[0122]
[0145] Device 100 includes, corresponds to, or may contain, a voice-activated device, an audio device, a wireless speaker and voice-activated device, a portable electronic device, a car, a vehicle, a computing device, a communication device, an Internet of Things (IoT) device, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, a smart speaker, a mobile computing device, a mobile communication device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a game console, equipment, a music player, a wireless device, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, or any combination thereof. In certain embodiments, a processor 804, a processor 806, or a combination thereof is included in the integrated circuit.
[0123]
[0146] Figure 9 is an exemplary example of a vehicle 900 incorporating an embodiment of the device 100 of Figure 1. According to one implementation, the vehicle 900 is an autonomous vehicle. According to other implementations, the vehicle 900 is a car, truck, motorcycle, aircraft, water vehicle, etc. In Figure 9, the vehicle 900 includes a display 820, one or more of the sensors 134, the device 100 including a context detector 140, a model selector 190, a model-based application 192, or a combination thereof. The sensors 134, the context detector 140, the model selector 190, and the model-based application 192 are indicated using dotted lines to show that these components may not be visible to passengers of the vehicle 900. The device 100 may be integrated into or coupled to the vehicle 900.
[0124]
[0147] In certain embodiments, device 100 is coupled to display 820 and provides outputs to display 820 in response to model-based applications 192, such as detecting or recognizing various events described herein (e.g., sound events). For example, device 100 provides display 820 with output 324 in Figure 3, indicating the sound class of sound 302 (such as a car horn) in audio data 105 received from microphone 104. In some implementations, device 100 can perform an action in response to recognizing a sound event, such as warning a vehicle operator or activating one of the sensors 134. In certain examples, device 100 provides an output indicating whether an action is being performed in response to a recognized sound event. In certain embodiments, a user can select options displayed on display 820 to enable or disable the performance of an action in response to a recognized sound event.
[0125]
[0148] In certain implementations, the sensor 134 includes the microphone 104 in Figure 1, a vehicle occupancy sensor, an eye-tracking sensor, a location sensor 152, or an external environment sensor (e.g., a LiDAR sensor or a camera). In certain embodiments, the sensor input of sensor 134 indicates the user's location. For example, sensor 134 is associated with various locations within the vehicle 900.
[0126]
[0149] Therefore, the techniques described with respect to Figures 1 to 8 enable the user of vehicle 900 to select a model that will be used based on the specific context in which device 100 operates.
[0127]
[0150] Figure 10 shows an example of a device 100 coupled to or integrated into a headset 1002, such as a virtual reality headset, augmented reality headset, mixed reality headset, extended reality headset, head-mounted display, or a combination thereof. A visual interface device, such as a display 820, is positioned in front of the user's eyes to enable the display of augmented reality, mixed reality, or virtual reality images or scenes to the user while the headset 1002 is worn. In a particular example, the display 820 is configured to display the output of device 100. The headset 1002 includes sensors 134, such as (one or more) microphones 104, (one or more) cameras 150, (one or more) location sensors 152, other sensors 156, or a combination thereof. Although shown in a single location, in other implementations, sensors 134 may be located in other locations on the headset 1002, such as an array of one or more microphones and one or more cameras distributed around the headset 1002 to detect multimodal inputs.
[0128]
[0151] Sensor 134 enables the detection of sensor data that device 100 uses to detect the context of the headset 1002 and update a model based on the detected context. For example, a model-based application 192 (e.g., SEC Engine 320) may use one or more models to generate sound event classification data that can be given to display 820 to indicate that a recognized sound event, such as a car horn, has been detected in an audio data sample received from sensor 134. In some implementations, device 100 may perform actions in response to recognizing a sound event, such as activating another of the cameras or sensors 134 or providing haptic feedback to the user.
[0129]
[0152] Figure 11 shows an example of device 100 integrated into a wearable electronic device 1102, indicated as a “smartwatch,” which includes a display 820 and a sensor 134. Sensor 134 enables context detection that device 100 can use to update one or more models used by a model-based application 192, based on modalities such as location, video, audio, and gestures. Sensor 134 also enables detection of sounds and other events in the environment around the wearable electronic device 1102 that device 100 can detect or interpret using the model-based application 192. For example, device 100 gives the display 820 output 324 in Figure 3, indicating that a recognized sound event is detected in an audio data sample received from sensor 134. In some implementations, device 100 can perform actions in response to recognizing a sound event, such as activating a camera or another of the sensors 134 or giving haptic feedback to the user.
[0130]
[0153] Figure 12 shows an exemplary example of a voice-controlled speaker system 1200. The voice-controlled speaker system 1200 may have wireless network connectivity and may be configured to perform auxiliary operations. In Figure 12, device 100 is included in the voice-controlled speaker system 1200. The voice-controlled speaker system 1200 also includes a speaker 1202 and a sensor 134. The sensor 134 includes one or more microphones 104 in Figure 1 to receive voice input or other audio input.
[0131]
[0154] During operation, in response to receiving verbal commands or recognized sound events, the voice-controlled speaker system 1200 can perform auxiliary actions. These auxiliary actions may include adjusting the temperature, playing music, turning on lights, etc. Sensor 134 enables the detection of data samples that device 100 can use to update the context of the voice-controlled speaker system 1200 and to update one or more models based on the context. Furthermore, the voice-controlled speaker system 1200 can perform several actions based on events recognized by device 100. For example, if device 100 recognizes the sound of a door closing, the voice-controlled speaker system 1200 may turn on one or more lights.
[0132]
[0155] Figure 13 shows a camera 1300 incorporating an embodiment of device 100 of Figure 1. In Figure 13, device 100 is incorporated into or coupled to camera 1300. Camera 1300 includes an image sensor 1302 and one or more other sensors (e.g., sensor 134), such as the (one or more) microphones 104 of Figure 1. Furthermore, camera 1300 includes device 100 configured to determine the context of camera 1300 and update one or more models based on the context. In certain embodiments, camera 1300 is configured to perform one or more actions in response to a recognized sound event. For example, camera 1300 may cause image sensor 1302 to capture an image in response to device 100 detecting a particular sound event in an audio data sample from sensor 134.
[0133]
[0156] Figure 14 shows a mobile device 1400 incorporating an embodiment of device 100 of Figure 1. In Figure 14, the mobile device 1400 includes or is coupled to device 100 of Figure 1. The mobile device 1400 includes, as an exemplary and non-limiting example, a telephone or a tablet. The mobile device 1400 includes a display 820 and sensors 134 such as (one or more) microphones 104, (one or more) cameras 150, (one or more) location sensors 152, or (one or more) other sensors 156. While operating, the mobile device 1400 may perform certain actions in response to device 100 recognizing certain sound events. For example, the action may include sending commands to other devices such as a thermostat, a home automation system, or another mobile device.
[0134]
[0157] Figure 15 shows a hearing aid 1500 incorporating an embodiment of device 100 of Figure 1. In Figure 15, the hearing aid 1500 includes or is coupled to device 100 of Figure 1. The hearing aid 1500 includes sensors 134 such as a microphone 104, a camera 150, a location sensor 152, or other sensors 156. During operation, the hearing aid 1500 may update one or more models in response to device 100 recognizing the context of the hearing aid 1500, such as the acoustic environment of the hearing aid 1500 for use by a model-based application 192 to process audio data, such as for noise reduction specific to the location.
[0135]
[0158] Figure 16 shows a flight device 1600 incorporating an embodiment of device 100 of Figure 1. In Figure 16, the flight device 1600 includes or is coupled to device 100 of Figure 1. The flight device 1600 is a manned, unmanned, or remotely controlled flight device (e.g., a package delivery drone). The flight device 1600 includes a control system 1602 and sensors 134 such as (one or more) microphones 104, (one or more) cameras 150, (one or more) location sensors 152, or (one or more) other sensors 156. The control system 1602 controls various operations of the flight device 1600, such as cargo release, sensor activation, takeoff, navigation, landing, or a combination thereof. For example, the control system 1602 may control the flight of the flight device 1600 between a designated point and cargo deployment at a specific location. During operation, the flight device 1600 may update one or more models in response to device 100 becoming aware of the context of the flight device 1600, such as the location or acoustic environment of the flight device 1600, for use by a model-based application 192 to detect events. For illustrative purposes, the control system 1602 may initiate a safe landing protocol in response to device 100 detecting an aircraft engine.
[0136]
[0159] Figure 17 shows a headset 1700 incorporating an embodiment of device 100 of Figure 1. In Figure 17, the headset 1700 includes or is coupled to device 100 of Figure 1. The headset 1700 includes one or more of the (one or more) microphones 104 of Figure 1, which are positioned primarily to capture the user's voice. The headset 1700 may also include one or more additional microphones positioned primarily to capture ambient sound (for example, for noise cancellation operation), and one or more of the sensors 134, such as (one or more) cameras 150, (one or more) location sensors 152, or (one or more) other sensors 156. In certain embodiments, the headset 1700 may update one or more models in response to device 100 recognizing a change in the context of the headset 1700, such as the location or acoustic environment of the headset 1700, for use by a model-based application 192 to perform operations such as noise cancellation features.
[0137]
[0160] Figure 18 shows a device 1800 incorporating an embodiment of device 100 in Figure 1. In Figure 18, device 1800 is a lamp, but in other implementations, device 1800 includes other Internet of Things devices such as a refrigerator, coffee maker, oven, or other household appliance. Device 1800 includes or is coupled to device 100 in Figure 1. Device 1800 includes sensors 134 such as (one or more) microphones 104, (one or more) cameras 150, (one or more) location sensors 152, activity detectors 154, or (one or more) other sensors 156. In certain embodiments, device 1800 may update one or more models in response to device 100 recognizing a change in the context of device 1800 used by a model-based application 192 to perform actions such as device 100 activating a light in response to device 100 detecting that a door is closed.
[0138]
[0161] Figure 19 is a flowchart illustrating an example of a method 1900 of operation for device 100 in Figure 1. Method 1900 can be initiated, controlled, or implemented by device 100. For example, processor 110 may execute instructions such as instruction 860 in Figure 8 from memory 108 to implement context-based model selection.
[0139]
[0162] Method 1900 includes, in block 1902, receiving sensor data from one or more sensor devices in one or more processors of the device. For example, a context detector 140 in one or more processors 110 receives sensor data 138 from one or more sensor devices 134. In some implementations, the sensor data includes location data of the device's location, such as location data 153, and the context is at least partially based on the location. In some implementations, the sensor data includes image data corresponding to a visual scene, such as image data 151, and the context is at least partially based on the visual scene. In some implementations, the sensor data includes audio corresponding to an audio scene, such as audio data 105, and the context is at least partially based on the audio scene. In some implementations, the sensor data includes motion data corresponding to the device's movement, such as activity data 155, and the context is at least partially based on the device's movement.
[0140]
[0163] In block 1904, method 1900 includes determining the context of a device based on sensor data in one or more processors. For example, a context detector 140 in one or more processors 110 receives sensor data 138 from one or more sensor devices 134 and determines the context 142 based on the sensor data 138.
[0141]
[0164] In 1906, method 1900 includes selecting a model based on context in one or more processors. For example, model selector 190 selects model 112 based on context 142. In certain implementations, the model is selected from among several models stored in the device's memory, such as available model 114. According to some implementations, the model is downloaded from a library, such as model library 162, which corresponds to a library of shared models. The model may include trained models uploaded to the library from another user device. In one example, the library corresponds to a crowdsourced library of models. The library may be contained within a distributed context-aware system.
[0142]
[0165] In block 1908, method 1900 includes processing an input signal using a model to generate a context-specific output in one or more processors. For example, model-based application 192 processes an input signal 106 using a model 112 selected to generate a context-specific output 122.
[0143]
[0166] In some implementations, method 1900 includes pruning the model in response to a determination that the context has changed. For example, the model selector 190 permanently removes the model in response to detecting that the context 142 has changed (e.g., device 100 has been moved to a different location) and that the current model is no longer suitable for the new context, or that another model is more suitable for the new context.
[0144]
[0167] In some implementations, the model includes a sound event detection model, where the input signal includes an audio signal, and the context-specific output includes the classification of sound events within the audio signal. In some implementations, the model includes an automatic speech recognition model, where the input signal includes an audio signal, and the context-specific output includes text data representing the speech within the audio signal. In some implementations, the model includes a natural language processing (NLP) model, where the input signal includes text data, and the context-specific output includes NLP output data based on the text data. In some implementations, the model includes a noise reduction model, where the input signal includes an audio signal, and the context-specific output includes a noise-reduced audio signal based on the audio signal. In some implementations, the model is associated with automatic tuning of the device operating mode, where the context-specific output includes a signal for tuning the device operating mode.
[0145]
[0168] In some implementations, method 1900 includes receiving a model from a second device via wireless transmission. For example, the context may correspond to the device's location, and the model may include an acoustic model corresponding to a specific location. In some implementations, method 1900 includes receiving permission for a model that is at least partially based on the device's location matching a specific location.
[0146]
[0169] In some implementations, the context includes a specific acoustic environment, and method 1900 includes determining whether the library of available acoustic models is specific to the specific acoustic environment and includes acoustic models available to the device, and, in response to the fact that acoustic models specific to the specific acoustic environment are not available to the device, determining whether acoustic models for a general category of the specific acoustic environment are available to the device. For example, when device 100 is located in room 204, in response to the model library 162 not having an “office” model 274 available for device 100, device 100 may request an “office building” model 262 for a general model of the acoustic environment for room 204.
[0147]
[0170] By selecting a model based on the device context, Method 1900 enables the device to perform with higher accuracy compared to using a single model for all contexts. Furthermore, by modifying the model, the device can perform with increased accuracy without incurring the power consumption, memory requirements, and processing resource usage associated with retraining an existing model from scratch in the device for a specific context. In addition, the device's operation using context-based models enables improved operation of the device itself, such as enabling faster convergence when performing iterative or dynamic processes (e.g., in noise cancellation techniques) by using a more accurate model specific to a particular context.
[0148]
[0171] Figure 18 is a flowchart illustrating an example of a method 1800 of operation for device 100 in Figure 1. Method 1800 can be initiated, controlled, or performed by device 100. For example, processor 110 may execute instructions such as instruction 660 in Figure 6 from memory 108 to perform context-based model selection.
[0149]
[0172] Method 1800 includes, in block 1802, selecting an acoustic model in one or more processors of the device that corresponds to a specific room in the building where the device is located. For example, the model selector 190 selects the acoustic model 212 in Figure 2 that corresponds to room 204 where device 100 is located.
[0150]
[0173] Method 2000, in block 2004, includes processing an input audio signal using an acoustic model in one or more processors. As an exemplary example, model-based application 192 uses an acoustic model 212 to perform noise reduction on an input signal 106 (e.g., audio data 105 from a microphone 104) to generate a noise-reduced audio signal as a context-specific output 122.
[0151]
[0174] In some implementations, Method 2000 includes downloading an acoustic model from a library of acoustic models in response to a decision that a device has entered a particular room. In some implementations, Method 2000 includes pruning (e.g., removing) the acoustic model in response to a device leaving a particular room. In some implementations, Method 2000 includes selecting an acoustic model based on sensor data indicating the location of the device, such as detecting the location of the device 100 by analyzing image data 151. In some implementations, Method 2000 includes selecting an acoustic model based on location data indicating the location of the device, such as location data 153.
[0152]
[0175] By selecting an acoustic model based on the device context, Method 2000 enables the device to perform with higher accuracy compared to using a single acoustic model for all contexts. Furthermore, by modifying the acoustic model, the device can perform with increased accuracy without incurring the power consumption, memory requirements, and processing resource usage associated with retraining an existing acoustic model from scratch within the device for a specific context. In addition, the device's operation using a context-based acoustic model enables improved operation of the device itself, such as enabling faster convergence when performing iterative or dynamic processes (e.g., in noise cancellation techniques) by using a more accurate acoustic model specific to a particular context.
[0153]
[0176] Figure 21 is a flowchart illustrating an example of a method 2100 of operation of device 100 in Figure 1 within a vehicle, such as being integrated into vehicle 900 in Figure 9. Method 2100 can be initiated, controlled, or performed by device 100. For example, processor 110 may execute instructions such as instruction 860 in Figure 8 from memory 108 to perform context-based model selection.
[0154]
[0177] Method 2100, in block 2102, includes, in one or more processors of the device, selecting a personal acoustic model for a user from a plurality of personal acoustic models corresponding to the vehicle in response to detection that a user has entered the vehicle. For example, device 100 in vehicle 900 in Figure 9 may store a plurality of personalized acoustic models for each user in vehicle 900 and may select a personal model or set of models for a particular user entering vehicle 900. In some implementations, Method 2100 includes determining that a user has entered the vehicle based on sensor data indicating the user's location. For illustrative purposes, sensor 134 in Figure 9 may determine a user in vehicle 900 through facial recognition, voice recognition, gestures, voice, input of identification information via interaction with an input device (e.g., a touchscreen in vehicle 900), or one or more other techniques for identifying a user.
[0155]
[0178] Method 2100 includes, in block 2104, processing an input audio signal using a personal acoustic model in one or more processors. For example, in some implementations, the personal acoustic model corresponds to an ASR model trained for a specific user, and the personal acoustic model is used by a model-based application 192 in the vehicle 900 to perform speech recognition on the user's voice captured through one or more microphones in the vehicle 900. For illustrative purposes, the ASR model is used to improve the accuracy of a voice interface to control one or more operations of the vehicle 900 (e.g., a navigation system, an entertainment system, a thermostat, driver assistance or autonomous driving settings). In another example, the personal acoustic model may include an SEC model personalized for a specific user. For illustrative purposes, if a particular user frequently takes their dog on road trips, the user's personal SEC model for use in the vehicle 900 may be trained to recognize an additional sound class, such as "dog barking in the vehicle."
[0156]
[0179] In some implementations, Method 2100 includes downloading a personal acoustic model from a library of acoustic models in response to a decision that the user has entered the vehicle. In some implementations, Method 2100 includes pruning (e.g., removing) the personal acoustic model in response to the device leaving the vehicle.
[0157]
[0180] By selecting a personal acoustic model in response to detecting a user entering a vehicle, Method 2100 enables the device to perform with higher accuracy compared to using a single acoustic model for all contexts. Furthermore, by changing the acoustic model, the device can perform with increased accuracy without incurring the power consumption, memory requirements, and processing resource usage associated with retraining an existing acoustic model from scratch in the device for a specific context. In addition, the device's operation using context-based acoustic models enables improved operation of the device itself, such as enabling faster convergence when performing iterative or dynamic processes (e.g., in noise cancellation techniques) by using a more accurate acoustic model specific to a particular context.
[0158]
[0181] Figure 22 is a flowchart illustrating an example of a method 2200 of operation for device 100 in Figure 1. Method 2200 can be initiated, controlled, or implemented by device 100. For example, processor 110 may execute instructions such as instruction 860 in Figure 8 from memory 108 to implement context-based model selection.
[0159]
[0182] Method 2200 includes, in block 2202, downloading an acoustic model corresponding to a specific location where the device is located in one or more processors of the device. In some implementations, Method 2200 includes determining that a user has entered a specific location based on sensor data indicating the location of the device. In one example, device 100 downloads an acoustic model 212 (e.g., an "office" model 274 corresponding to room 204) in response to device 100 entering room 204 or determining (e.g., based on location data 153) that the device's location is within room 204.
[0160]
[0183] Method 2200 includes processing an input audio signal using an acoustic model in one or more processors in block 2204. In one example, device 100 corresponds to the hearing aid 1500 in Figure 15 and uses an acoustic model 212 (for example, the “office” model 274 corresponding to room 204) to perform noise reduction in a model-based application 192.
[0161]
[0184] Method 2200, in block 2206, includes removing an acoustic model in one or more processors in response to a device leaving a location. In some implementations, Method 2200 includes determining that a user has entered a particular location based on location data indicating the location of a device. In one example, device 100 prunes an acoustic model 212 (e.g., an "office" model 274 corresponding to room 204) in response to device 100 leaving room 204 or determining (e.g., based on location data 153) that the device's location is no longer within room 204.
[0162]
[0185] While the example given above illustrates Method 2200 implemented in Building 202 of Figure 2, it should be understood that Method 2200 is not limited to any specific location or location type. For example, in some implementations, the location corresponds to a specific restaurant, and the acoustic model is downloaded from a library of acoustic models in response to the decision that the device has entered a specific restaurant. In other implementations, the location could correspond to a public park, a subway station, a car, a train, an airplane, a specific room in a user's home, a specific room in a museum, an auditorium, or a concert hall, etc.
[0163]
[0186] By selecting an acoustic model that corresponds to a specific location of the device, Method 2200 enables the device to perform with higher accuracy compared to using a single acoustic model for all locations. Furthermore, by changing the acoustic model, the device can perform with increased accuracy without incurring the power consumption, memory requirements, and processing resource usage associated with retraining an existing acoustic model from scratch in the device for a specific context. Removing the acoustic model in response to the device leaving a location improves the device's operation by freeing up memory and resources associated with continuing to store the acoustic model when it is no longer in use, thereby reducing the power consumption associated with storing the model and the latency associated with subsequent retrieval of the acoustic model stored in the device. In addition, the device's operation using a location-based acoustic model enables improved operation of the device itself, such as enabling faster convergence when performing iterative or dynamic processes (e.g., in noise cancellation techniques) by using a more accurate acoustic model specific to a particular location.
[0164]
[0187] Figure 23 is a flowchart illustrating an example of a method 2300 of operation for device 100 in Figure 1. Method 2300 can be initiated, controlled, or implemented by device 100. For example, processor 110 may execute instructions such as instruction 860 in Figure 8 from memory 108 to implement context-based model selection.
[0165]
[0188] Method 2300 includes, in block 2302, selecting an acoustic model corresponding to a specific location in one or more processors of the device.
[0166]
[0189] Method 2300, in block 2304, includes receiving access permissions for an acoustic model based at least partially on the location of a device matching a specific location in one or more processors. In some implementations, Method 2300 includes receiving access permissions in response to the discovery of a device in a specific location. For example, device 100 may receive access permissions 214 to use the “office” model 274 in Figure 2 in response to the discovery of device 100 in room 204. Device 100 may transmit data indicating the acoustic environment 210, such as location data, and in response, the permission management system may send access permissions to device 100 as described with reference to Figure 2.
[0167]
[0190] Method 2300 includes, in block 2306, processing an input audio signal using an acoustic model in one or more processors.
[0168]
[0191] By selecting an acoustic model corresponding to a specific location of the device, Method 2300 enables the device to perform with greater precision compared to using a single acoustic model for all locations. Furthermore, receiving access permissions for the acoustic model based on location matching allows for maintaining the security of the acoustic model by preventing its use when the device is not in a specific location. Since the acoustic model is already stored in the device and does not need to be downloaded when entering a specific location, such security makes it possible to download the acoustic model to the device as prohibited data to reduce peak network bandwidth usage and latency associated with using the acoustic model. In addition, the operation of the device using a location-based acoustic model enables improved operation of the device itself by allowing for faster convergence when performing iterative or dynamic processes (e.g., in noise cancellation techniques) by using a more accurate acoustic model specific to a particular location.
[0169]
[0192] Figure 24 is a flowchart illustrating an example of a method 2400 of operation for device 100 in Figure 1. Method 2400 can be initiated, controlled, or implemented by device 100. For example, processor 110 may execute instructions such as instruction 860 in Figure 8 from memory 108 to implement context-based model selection.
[0170]
[0193] Method 2400, in block 2402, includes detecting the device context in one or more processors of the device. For illustrative purposes, in some implementations, one or more processors 110 are configured to detect a context such as a context 142 detected by a context detector 142 based on sensor data 138. In illustrative, non-limiting examples, the context may correspond to a location or an activity such as driving a car.
[0171]
[0194] Method 2400 includes sending a context-indicating request to a remote device in block 2404. For illustrative purposes, in some implementations, one or more processors 110 are configured to initiate sending a context-indicating request (e.g., acoustic environment 210) to a remote device such as a remote computer device 818 (e.g., a server or network memory that stores at least a portion of the model library 162, which may be part of a service).
[0172]
[0195] Method 2400 includes receiving a context-dependent model in block 2406. For illustrative purposes, in some implementations, one or more processors 110 are configured to receive a model 112 from a remote computer device 818 or another server or network memory storing at least a portion of the model library 162 in response to sending a request indicating the context. In some implementations, the model 112 is received as a compressed source model and decompressed by one or more processors 110 for use in device 100. In some implementations, the model is received based on private access. In an exemplary example, an acoustic model 212 is received along with an access permission 214 that authenticates device 100 accessing the acoustic model 212 (for example, along with authorization to use the model 212). In an exemplary example, the model is received based on access permitted by a family member or friend of the user of device 100.
[0173]
[0196] Method 2400, in block 2408, includes using a model in one or more processors while a context is being discovered. For illustrative purposes, in some implementations, one or more processors 110 are configured to receive a model 112 in response to sending a request indicating a context 142, and in response to continuing to use the model 112 in a model-based application 192 while the context 142 remains immutable (for example, using the temporarily received model 112 as long as the context 142 is discovered).
[0174]
[0197] Method 2400 includes pruning a model in block 2410 in response to detecting a change in context in one or more processors. For illustrative purposes, in some implementations, one or more processors 110 are configured to prune a model 112 in response to detecting a change in context 142 (for example, pruning model 112 when the context changes). In some implementations, pruning a model includes permanently deleting the model.
[0175]
[0198] In some implementations, method 2400 includes generating at least one new sound class while the context is being detected, and pruning the model includes maintaining at least one new sound class. For illustrative purposes, in some implementations, one or more processors 110 are configured to generate new sound classes, such as by generating an updated or new model (e.g., updated model 506), as described with reference to Figures 3-5, which is maintained by storing the updated or new model in memory 108 or by uploading it to a model library 162.
[0176]
[0199] In relation to the described implementation, the apparatus includes means for receiving sensor data. For example, means for receiving sensor data include device 100, instruction 860, processor 804, processor 806, context detector 140, microphone 104, camera 150, location sensor 152, activity detector 154, other sensors 156, codec 824, one or more other circuits or components configured to receive sensor data, or any combination thereof.
[0177]
[0200] The apparatus also includes means for determining context based on sensor data. For example, means for determining context based on sensor data include device 100, instruction 860, processor 804, processor 806, context detector 140, one or more other circuits or components configured to determine context based on sensor data, or any combination thereof.
[0178]
[0201] The apparatus also includes means for selecting a model based on context. For example, the means for selecting a model based on context includes device 100, instruction 860, processor 804, processor 806, model selector 190, one or more other circuits or components configured to select a model based on context, or any combination thereof.
[0179]
[0202] The apparatus also includes means for processing an input signal using a model to generate a context-specific output. For example, means for processing an input signal using a model to generate a context-specific output include device 100, instruction 860, processor 804, processor 806, model-based application 192, one or more other circuits or components configured to process an input signal using a model to generate a context-specific output, or any combination thereof.
[0180]
[0203] Furthermore, those skilled in the art will understand that the various exemplary logic blocks, configurations, modules, circuits, and algorithmic steps described in relation to the implementations disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or a combination of both. Various exemplary components, blocks, configurations, modules, circuits, and steps have been described above in general terms of their function. Whether such functions are implemented as hardware or as processor-executable instructions depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functions in various ways for each specific application, but such implementation decisions should not be construed as causing a departure from the scope of this disclosure.
[0181]
[0204] Steps of methods or algorithms described in relation to the implementations disclosed herein may be implemented directly in hardware, in software modules executed by a processor, or in a combination of both. Software modules may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM®), registers, hard disks, removable disks, compact disk read-only memory (CD-ROM), or any other form of non-temporary storage medium known in the art. An exemplary storage medium is coupled to a processor so that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integrated with the processor. The processor and storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or user terminal. Alternatively, the processor and storage medium may reside as separate components in a computing device or user terminal.
[0182]
[0205] Specific aspects of this disclosure are described below in the first set of interrelated provisions.
[0183]
[0206] According to Clause 1, the device includes one or more processors configured to receive sensor data from one or more sensor devices, determine the context of the device based on the sensor data, select a model based on the context, and process an input signal using the model to generate a context-specific output.
[0184]
[0207] Clause 2 includes the device of Clause 1, further including a location sensor coupled to one or more processors, wherein the sensor data includes location data from the location sensor, the location data indicates the location of the device, and herein the context is at least in part based on the location.
[0185]
[0208] Clause 3 includes the device of Clause 1 or Clause 2, further including a camera coupled to one or more processors, wherein the sensor data includes image data from the camera, the image data corresponds to a visual scene, and herein the context is at least partially based on the visual scene.
[0186]
[0209] Clause 4 includes any of the devices of Clauses 1 to 3, further including a microphone coupled to one or more processors, wherein the sensor data includes audio data from the microphone, the audio data corresponds to an audio scene, and the context is at least partially based on the audio scene.
[0187]
[0210] Clause 5 includes any of the devices of Clauses 1 to 4, further including an activity detector coupled to one or more processors, wherein the sensor data includes motion data from the activity detector, the motion data corresponds to the motion of the device, and herein the context is at least partially based on the motion of the device.
[0188]
[0211] Clause 6 includes any of the devices of Clauses 1 to 5, and further includes memory coupled to one or more processors, wherein the model is selected from a plurality of models stored in the memory.
[0189]
[0212] Clause 7 includes any of the devices in Clauses 1 through 6, where the model includes a sound event detection model, the input signal includes an audio signal, and the context-specific output includes the classification of sound events within the audio signal.
[0190]
[0213] Clause 8 includes any of the devices in Clauses 1 through 7, where the model includes an automatic speech recognition model, the input signal includes an audio signal, and the context-specific output includes text data representing the speech in the audio signal.
[0191]
[0214] Clause 9 includes any device from Clauses 1 through 8, where the model includes a natural language processing (NLP) model, the input signal includes text data, and the context-specific output includes NLP output data based on the text data.
[0192]
[0215] Clause 10 includes any device from Clauses 1 through 9, where the model includes a noise reduction model, the input signal includes an audio signal, and the context-specific output includes a noise reduction audio signal based on the audio signal.
[0193]
[0216] Clause 11 includes any device from Clauses 1 through 10, where the model is associated with automatic adjustment of the device operating mode, and the context-specific output includes signals for adjusting the device operating mode.
[0194]
[0217] Clause 12 further includes a modem comprising any of the devices described in Clauses 1 through 11, coupled to one or more processors, and configured to receive a model from a second device via wireless transmission.
[0195]
[0218] Clause 13 includes the device of Clause 12, where context corresponds to the location of the device, and hereby model includes an acoustic model corresponding to a specific location.
[0196]
[0219] Clause 14 includes the devices of Clause 13, and one or more processors are further configured to receive access permissions for a model that is at least partially based on the location of a device matching a specific location via the modem.
[0197]
[0220] Clause 15 includes any of the devices in Clauses 1 through 14, wherein one or more processors are further configured to prune the model in response to a determination that the context has changed.
[0198]
[0221] Clause 16 includes any device from Clauses 1 through 15, and the model is downloaded from the shared model library.
[0199]
[0222] Clause 17 includes the devices of Clause 16, and the models include trained models uploaded to the library from another user device.
[0200]
[0223] Clause 18 includes the devices of Clause 16 or Clause 17, and the library corresponds to the crowdsourced library of models.
[0201]
[0224] Clause 19 includes any device from Clauses 16 through 18, and the library is included in the distributed context-aware system.
[0202]
[0225] Clause 20 includes any device of Clauses 1 to 19, where the context includes a particular acoustic environment, wherein one or more processors are configured to determine whether a library of available acoustic models is specific to a particular acoustic environment and includes acoustic models available to one or more processors, and, in response to the absence of an acoustic model specific to a particular acoustic environment available to one or more processors, whether an acoustic model for a general category of the particular acoustic environment is available to one or more processors.
[0203]
[0226] Clause 21 includes any device from Clauses 1 through 20, wherein one or more processors are integrated into an integrated circuit.
[0204]
[0227] Clause 22 includes any of the devices described in Clauses 1 through 20, and one or more processors are integrated into the vehicle.
[0205]
[0228] Clause 23 includes any of the devices described in Clauses 1 through 20, wherein one or more processors are integrated into at least one of the following: a mobile phone, a tablet computer device, a virtual reality headset, an augmented reality headset, a mixed reality headset, a wireless speaker device, a wearable device, a camera device, or a hearing aid.
[0206]
[0229] Specific aspects of this disclosure are described below in the second set of interrelated provisions.
[0207]
[0230] According to Clause 24, the method includes: in one or more processors of the device receiving sensor data from one or more sensor devices; in one or more processors determining the context of the device based on the sensor data; in one or more processors selecting a model based on the context; and in one or more processors processing an input signal using the model to generate a context-specific output.
[0208]
[0231] Clause 25 includes the method of Clause 24, wherein the sensor data includes location data of the device's location, and herein the context is at least in part based on the location.
[0209]
[0232] Clause 26 includes the methods of Clause 24 or Clause 25, wherein the sensor data includes image data corresponding to a visual scene, and herein the context is at least partially based on the visual scene.
[0210]
[0233] Clause 27 includes any of the methods described in Clauses 24 to 26, wherein the sensor data includes audio corresponding to the audio scene, and the context is at least partially based on the audio scene.
[0211]
[0234] Clause 28 includes any of the methods described in Clauses 24 to 27, wherein the sensor data includes motion data corresponding to the movement of the device, and herein the context is based at least in part on the movement of the device.
[0212]
[0235] Clause 29 includes any of the methods described in Clauses 24 through 28, wherein the model is selected from among several models stored in the device's memory.
[0213]
[0236] Clause 30 includes any method of Clauses 24 to 29, wherein the model includes a sound event detection model, the input signal includes an audio signal, and the context-specific output includes the classification of sound events in the audio signal.
[0214]
[0237] Clause 31 includes any method of Clauses 24 to 30, wherein the model includes an automatic speech recognition model, the input signal includes an audio signal, and the context-specific output includes text data representing the speech in the audio signal.
[0215]
[0238] Clause 32 includes any method of Clauses 24 through 31, wherein the model includes a natural language processing (NLP) model, the input signal includes text data, and the context-specific output includes NLP output data based on the text data.
[0216]
[0239] Clause 33 includes any method of Clauses 24 to 32, wherein the model includes a noise reduction model, the input signal includes an audio signal, and the context-specific output includes a noise reduction audio signal based on the audio signal.
[0217]
[0240] Clause 34 includes any method of Clauses 24 to 33, wherein the model is associated with automatic adjustment of the device operating mode, and herein the context-specific output includes a signal for adjusting the device operating mode.
[0218]
[0241] Clause 35 further includes receiving a model from a second device via wireless transmission, including any of the methods described in Clauses 24 through 34.
[0219]
[0242] Clause 36 includes any of the methods described in Clauses 24 through 35, wherein the context corresponds to the location of the device, and herein the model includes an acoustic model corresponding to a specific location.
[0220]
[0243] Clause 37 includes, in the manner of Clause 36, and further includes, receiving permission for a model that is at least partially based on the location of a device that matches a specific location.
[0221]
[0244] Clause 38 further includes pruning the model in response to a determination that the context has changed, including any of the methods described in Clauses 24 through 37.
[0222]
[0245] Clause 39 includes any of the methods described in Clauses 24 through 38, where the model is downloaded from a shared model library.
[0223]
[0246] Clause 40 includes the method of Clause 39, and the model includes a trained model uploaded to the library from another user device.
[0224]
[0247] Clause 41 includes the methods of Clause 39 or Clause 40, and the library corresponds to a crowdsourced library of models.
[0225]
[0248] Clause 42 includes any of the methods described in Clauses 39 through 41, and the library is included in the distributed context-aware system.
[0226]
[0249] Clause 43 includes any method of Clauses 24 to 42, wherein the context includes a particular acoustic environment, and further includes determining whether the library of available acoustic models is specific to that particular acoustic environment and includes acoustic models available to the device, and in response to the absence of acoustic models specific to that particular acoustic environment, determining whether acoustic models for a general category of that particular acoustic environment are available to the device.
[0227]
[0250] Specific aspects of this disclosure are described below in the third set of interrelated provisions.
[0228]
[0251] According to Clause 44, the device includes means for receiving sensor data, means for determining a context based on the sensor data, means for selecting a model based on the context, and means for processing an input signal using the model to generate a context-specific output.
[0229]
[0252] Clause 45 includes the device of Clause 44, and the sensor data includes location data of the device's location, wherein the context is at least in part based on the location.
[0230]
[0253] Clause 46 includes the device of Clause 44 or Clause 45, wherein the sensor data includes image data corresponding to a visual scene, and herein the context is at least partially based on the visual scene.
[0231]
[0254] Clause 47 includes any of the devices described in Clauses 44 to 46, wherein the sensor data includes audio corresponding to an audio scene, and the context is at least partially based on the audio scene.
[0232]
[0255] Clause 48 includes any device of Clauses 44 to 47, wherein the sensor data includes motion data corresponding to the movement of the device, and herein the context is at least in part based on the movement of the device.
[0233]
[0256] Clause 49 includes any device of Clauses 44 to 48, further including means for storing a model, wherein the model is selected from a plurality of models stored in the means for storing the model.
[0234]
[0257] Clause 50 includes any device from Clauses 44 to 49, where the model includes a sound event detection model, the input signal includes an audio signal, and the context-specific output includes the classification of sound events within the audio signal.
[0235]
[0258] Clause 51 includes any device from Clauses 44 to 50, where the model includes an automatic speech recognition model, the input signal includes an audio signal, and the context-specific output includes text data representing the speech in the audio signal.
[0236]
[0259] Clause 52 includes any device from Clauses 44 to 51, where the model includes a natural language processing (NLP) model, the input signal includes text data, and the context-specific output includes NLP output data based on the text data.
[0237]
[0260] Clause 53 includes any device from Clauses 44 to 52, where the model includes a noise reduction model, the input signal includes an audio signal, and the context-specific output includes a noise reduction audio signal based on the audio signal.
[0238]
[0261] Clause 54 includes any device from Clauses 44 to 53, where the model is associated with automatic adjustment of the device operating mode, and the context-specific output includes a signal for adjusting the device operating mode.
[0239]
[0262] Clause 55 includes any of the devices described in Clauses 44 through 54, and the model is received from a second device via wireless transmission.
[0240]
[0263] Clause 56 includes any device from Clauses 44 to 55, where context corresponds to the location of the device, and hereby model includes an acoustic model corresponding to a specific location.
[0241]
[0264] Clause 57 includes the devices of Clause 56, and access permissions for the models are received on at least partially based on the location of the device matching the specific location.
[0242]
[0265] Clause 58 includes the devices of Clause 56 and further includes means for removing the model in response to the device leaving a particular location.
[0243]
[0266] Clause 59 includes any device from Clauses 44 through 58, and the model is downloaded from the shared model library.
[0244]
[0267] Clause 60 includes the devices of Clause 59, and the models include trained models uploaded to the library from another user device.
[0245]
[0268] Clause 61 includes the devices of Clause 59 or Clause 60, and the library corresponds to the crowdsourced library of models.
[0246]
[0269] Clause 62 includes any device from Clauses 59 to 61, and the library is included in the distributed context-aware system.
[0247]
[0270] Specific aspects of this disclosure are described below in the fourth set of interrelated provisions.
[0248]
[0271] According to Article 63, a non-temporary computer-readable storage medium comprising instructions, when executed by the device's processor, causing the processor to receive sensor data from one or more sensor devices, determine the context on the sensor data, select a model based on the context, and process an input signal using the model to produce a context-specific output.
[0249]
[0272] Clause 64 includes the non-temporary computer-readable storage medium of Clause 63, and the sensor data includes location data of the device's location, wherein the context is at least in part based on the location.
[0250]
[0273] Clause 65 includes a non-temporary computer-readable storage medium as defined in Clause 63 or Clause 64, wherein the sensor data includes image data corresponding to a visual scene, and herein the context is at least partially based on the visual scene.
[0251]
[0274] Clause 66 includes a non-temporary computer-readable storage medium as defined in any of Clauses 63 to 65, wherein the sensor data includes audio corresponding to an audio scene, and herein the context is at least partially based on the audio scene.
[0252]
[0275] Clause 67 includes a non-temporary computer-readable storage medium as defined in any of Clauses 63 to 66, wherein the sensor data includes motion data corresponding to the movement of the device, and herein the context is based at least in part on the movement of the device.
[0253]
[0276] Clause 68 includes a non-temporary computer-readable storage medium as defined in any of Clauses 63 to 67, wherein the model is selected from among several models stored in the device's memory.
[0254]
[0277] Clause 69 includes a non-temporary computer-readable storage medium as defined in any of Clauses 63 to 68, the model includes a sound event detection model, the input signal includes an audio signal, and the context-specific output includes the classification of sound events in the audio signal.
[0255]
[0278] Clause 70 includes a non-temporary computer-readable storage medium as defined in any of Clauses 63 to 69, where the model includes an automatic speech recognition model, the input signal includes an audio signal, and the context-specific output includes text data representing the speech in the audio signal.
[0256]
[0279] Clause 71 includes a non-temporary computer-readable storage medium as defined in any of Clauses 63 to 70, the model includes a natural language processing (NLP) model, the input signal includes text data, and the context-specific output includes NLP output data based on the text data.
[0257]
[0280] Clause 72 includes a non-temporary computer-readable storage medium as defined in any of Clauses 63 to 71, the model includes a noise reduction model, the input signal includes an audio signal, and the context-specific output includes a noise reduction audio signal based on the audio signal.
[0258]
[0281] Clause 73 includes any non-temporary computer-readable storage medium of Clauses 63 to 72, the model of which is associated with automatic adjustment of the device operating mode, wherein context-specific output includes signals for adjusting the device operating mode.
[0259]
[0282] Clause 74 includes a non-temporary computer-readable storage medium as defined in any of Clauses 63 to 73, and the instruction further causes the processor to receive a model from a second device via wireless transmission.
[0260]
[0283] Clause 75 includes a non-temporary computer-readable storage medium as defined in any of Clauses 63 to 74, where context corresponds to the location of the device, and model includes an acoustic model corresponding to a specific location.
[0261]
[0284] Clause 76 includes the non-temporary computer-readable storage medium of Clause 75, and the instruction further causes the processor to obtain permission for a k model based at least in part on the location of a device matching a particular location.
[0262]
[0285] Clause 77 includes the non-temporary computer-readable storage medium of Clause 75, and the instruction causes the processor to further remove the model in response to the device leaving a particular location.
[0263]
[0286] Clause 78 includes any non-temporary computer-readable storage medium specified in Clauses 63 through 77, and the model is downloaded from a library of shared models.
[0264]
[0287] Clause 79 includes non-temporary computer-readable storage media as defined in Clause 78, and the model includes a trained model uploaded to the library from another user device.
[0265]
[0288] Clause 80 includes non-temporary computer-readable storage media as defined in Clause 78 or Clause 79, and the library corresponds to a crowdsourced library of models.
[0266]
[0289] Clause 81 includes any non-temporary computer-readable storage medium of Clauses 78 to 80, and the library is included in a distributed context-aware system.
[0267]
[0290] Clause 82 includes a non-temporary computer-readable storage medium as defined in any of Clauses 63 to 81, wherein the context includes a particular acoustic environment, wherein the instruction causes the processor to determine whether a library of available acoustic models is specific to a particular acoustic environment and includes acoustic models available to the device, and, in response to the absence of an acoustic model specific to a particular acoustic environment, to determine whether an acoustic model for a general category of the particular acoustic environment is available to the device.
[0268]
[0291] Specific aspects of this disclosure are described below in the fifth set of interrelated provisions.
[0269]
[0292] According to Article 83, the device includes one or more processors configured to select an acoustic model corresponding to a particular room in the building in which the device is located, and to process an input audio signal using the acoustic model.
[0270]
[0293] Clause 84 includes the device of Clause 83, and one or more processors are configured to download an acoustic model from a library of acoustic models in response to a determination that the device has entered a particular room.
[0271]
[0294] Clause 85 includes the device of Clause 83 or 84, and one or more processors are further configured to remove the acoustic model in response to a determination that the device has left a particular room.
[0272]
[0295] Clause 86 includes any of the devices of Clauses 83 to 85 and further includes one or more microphones configured to generate an input audio signal.
[0273]
[0296] Clause 87 includes any of the devices of Clauses 83 to 86 and further includes one or more sensor devices coupled to one or more processors and configured to generate sensor data indicating the location of the device, where the one or more processors are configured to select an acoustic model based on the sensor data.
[0274]
[0297] Clause 88 includes any of the devices of Clauses 83 to 87 and further includes a modem coupled to one or more processors and configured to receive location data indicating the location of the device, where the one or more processors are configured to select an acoustic model based on the location data.
[0275]
[0298] For certain aspects of the present disclosure, they are described below in a sixth set of related clauses.
[0276]
[0299] According to Article 89, the method includes selecting an acoustic model in one or more processors of the device that corresponds to a particular room in the building in which the device is located, and processing an input audio signal using the acoustic model in one or more processors.
[0277]
[0300] Clause 90 includes the methods of Clause 89, and further includes downloading an acoustic model from a library of acoustic models in response to a decision that a device has entered a particular room.
[0278]
[0301] Clause 91 further includes removing an acoustic model in response to the device leaving a particular room, including the methods of Clause 89 or Clause 90.
[0279]
[0302] Clause 92 includes, further, selecting an acoustic model based on sensor data indicating the location of the device, in any of the methods described in Clauses 89 to 91.
[0280]
[0303] Clause 93 further includes selecting an acoustic model based on location data indicating the location of a device, including any method described in Clauses 89 through 91.
[0281]
[0304] According to Article 94, the device includes means for selecting an acoustic model corresponding to a particular room in a building in which the device is located, and means for processing an input audio signal using the acoustic model.
[0282]
[0305] According to Article 95, a non-temporary computer-readable storage medium includes instructions, when executed by the device's processor, that cause the processor to select an acoustic model corresponding to a particular room in the building in which the device is located, and to process an input audio signal using the acoustic model.
[0283]
[0306] Specific aspects of this disclosure are described below in the seventh set of interrelated provisions.
[0284]
[0307] According to Article 96, the device includes one or more processors configured to, in response to the device entering the vehicle, select a personal acoustic model for the user of the device from among a plurality of personal acoustic models corresponding to the vehicle, and process an input audio signal using the personal acoustic model.
[0285]
[0308] Clause 97 includes the device of Clause 96, wherein one or more processors are configured to download a personal acoustic model from a library of acoustic models in response to a decision that the device has entered the vehicle.
[0286]
[0309] Clause 98 includes the device of Clause 96 or Clause 97, wherein one or more processors are further configured to remove a personal acoustic model in response to the device leaving the vehicle.
[0287]
[0310] Clause 99 further includes one or more microphones configured to generate an input audio signal, which include any of the devices in Clauses 96 to 98.
[0288]
[0311] Clause 100 further includes one or more sensor devices, which include any of the devices in Clauses 96 to 99, and are coupled to one or more processors and configured to generate sensor data indicating the location of the device, wherein the one or more processors are configured to determine, based on the sensor data, that the device has entered the vehicle.
[0289]
[0312] Clause 101 further includes a modem comprising any device of Clauses 96 to 100, coupled to one or more processors and configured to receive location data indicating the location of the device, wherein one or more processors are configured to determine, based on the location data, that the device has entered a vehicle.
[0290]
[0313] For certain aspects of the present disclosure, the following is described in the eighth set of related clauses.
[0291]
[0314] According to clause 102, the method includes selecting, in one or more processors of a device, a personal acoustic model for a user from among a plurality of personal acoustic models corresponding to a vehicle, in response to detecting that the user has entered the vehicle, and processing an input audio signal using the personal acoustic model in one or more processors.
[0292]
[0315] Clause 103 includes the method of clause 102 and further includes downloading a personal acoustic model from a library of acoustic models in response to a determination that the user has entered the vehicle.
[0293]
[0316] Clause 104 includes the method of clause 102 or clause 103 and further includes removing the personal acoustic model in response to the user leaving the vehicle.
[0294]
[0317] Clause 105 includes any of the methods from clause 102 to 104 and further includes determining that the user has entered the vehicle based on sensor data indicating the user's location.
[0295]
[0318] According to clause 106, the device includes means for selecting, in response to detecting that the user has entered the vehicle, a personal acoustic model for the user from among a plurality of personal acoustic models corresponding to the vehicle, and means for processing an input audio signal using the personal acoustic model.
[0296]
[0319] According to clause 107, a non - transitory computer - readable storage medium includes instructions that, when executed by a processor of a device, cause the processor to select, in response to detecting that the user has entered the vehicle, a personal acoustic model for the user from among a plurality of personal acoustic models corresponding to the vehicle, and process an input audio signal using the personal acoustic model.
[0297]
[0320] Specific aspects of this disclosure are described below in the ninth set of interrelated provisions.
[0298]
[0321] According to Clause 108, the device includes one or more processors configured to download an acoustic model corresponding to a specific location where the device is located, process an input audio signal using the acoustic model, and remove the acoustic model in response to the device leaving the location.
[0299]
[0322] Clause 109 includes the device of Clause 108, where the location corresponds to a specific restaurant, and herein, the acoustic model is downloaded from the acoustic model library in response to the decision that the device has entered the specific restaurant.
[0300]
[0323] Clause 110 further includes one or more microphones configured to generate an input audio signal, which include the devices of Clause 108 or 109.
[0301]
[0324] Clause 111 includes any of the devices of Clauses 108 to 110, and further includes one or more sensor devices coupled to one or more processors and configured to generate sensor data indicating the location of the device, wherein one or more processors are configured to determine, based on the sensor data, that the device has entered a particular location.
[0302]
[0325] Clause 112 further includes a modem comprising any of the devices of Clauses 108 to 110, coupled to one or more processors and configured to receive location data indicating the location of the device, wherein one or more processors are configured to determine, based on the location data, that the device has entered a particular location.
[0303]
[0326] Specific aspects of this disclosure are described below in the tenth set of interrelated provisions.
[0304]
[0327] According to Article 113, the method includes downloading an acoustic model in one or more processors of the device that corresponds to a specific location where the device is located; processing an input audio signal using the acoustic model in one or more processors; and removing the acoustic model in one or more processors in response to the device leaving the location.
[0305]
[0328] Clause 114 includes the method of Clause 113, wherein the location corresponds to a specific restaurant, and herein the acoustic model is downloaded from a library of acoustic models in response to the decision that the device has entered a specific restaurant.
[0306]
[0329] Clause 115 further includes determining that a device has entered a particular location based on sensor data indicating the location of the device, including the methods of Clause 113 or Clause 114.
[0307]
[0330] Clause 116 further includes determining that a device has entered a particular location based on location data indicating the device's location, including the methods of Clause 113 or 114.
[0308]
[0331] According to Article 117, the device includes means for downloading an acoustic model corresponding to a specific location where the device is located, means for processing an input audio signal using the acoustic model, and means for removing the acoustic model in response to the device leaving the location.
[0309]
[0332] According to Article 118, a non-temporary computer-readable storage medium includes instructions, when executed by the device's processor, that cause the processor to download an acoustic model corresponding to a specific location where the device is located, to process an input audio signal using the acoustic model, and to remove the acoustic model in response to the device leaving the location.
[0310]
[0333] Specific aspects of this disclosure are described below in the eleventh set of interrelated provisions.
[0311]
[0334] According to Clause 119, the device includes one or more processors configured to select an acoustic model corresponding to a particular location, to receive permission for an acoustic model that is at least partially based on the device's location matching the particular location, and to process an input audio signal using the acoustic model.
[0312]
[0335] Clause 120 includes the devices of Clause 119, further including a modem, wherein one or more processors are further configured to receive access permissions via the modem in response to the discovery of a device in a particular location.
[0313]
[0336] Specific aspects of this disclosure are described below in 12 sets of interrelated provisions.
[0314]
[0337] According to Clause 121, the method includes selecting an acoustic model corresponding to a specific location in one or more processors of the device, receiving permission in one or more processors for an acoustic model based at least in part on a device location matching the specific location, and processing an input audio signal using the acoustic model in one or more processors.
[0315]
[0338] Clause 122 includes the methods of Clause 121, and further includes receiving access permissions in response to the discovery of a device in a specific location.
[0316]
[0339] According to Clause 123, the device includes means for selecting an acoustic model corresponding to a particular location, means for receiving permission for an acoustic model that is at least partially based on the device's location matching the particular location, and means for processing an input audio signal using the acoustic model.
[0317]
[0340] According to Article 124, a non-temporary computer-readable storage medium includes instructions, when executed by the device's processor, that cause the processor to select an acoustic model corresponding to a particular location, to obtain permission for an acoustic model based at least in part on the device's location matching the particular location, and to process an input audio signal using the acoustic model.
[0318]
[0341] Specific aspects of this disclosure are described below in 13 sets of interrelated provisions.
[0319]
[0342] According to Clause 125, the device includes one or more processors configured to detect the context of the device, send a request to a remote device indicating the context for removing the device, receive a model corresponding to the context, use the model while the context is detected, and prune the model in response to detecting a change in the context.
[0320]
[0343] Clause 126 further comprises the device of Clause 125, wherein one or more processors are configured to generate at least one new sound class while the context is detected, wherein pruning the model includes maintaining at least one new sound class.
[0321]
[0344] Clause 127 includes the devices of Clause 125 or Clause 126, and pruning a model includes permanently deleting a model.
[0322]
[0345] Clause 128 includes any device from Clauses 125 to 127, wherein one or more processors are further configured to receive a model based on private access.
[0323]
[0346] According to Clause 129, the method includes, in one or more processors of the device, detecting the context of the device; sending a request to a remote device indicating the context; receiving a model corresponding to the context; in one or more processors, using the model while the context is detected; and in one or more processors, pruning the model in response to detecting a change in the context.
[0324]
[0347] Clause 130 includes the method of Clause 129, further comprising generating at least one new sound class while the context is being detected, wherein pruning the model includes maintaining at least one new sound class.
[0325]
[0348] Clause 131 includes the methods of Clause 129 or Clause 130, and pruning a model includes permanently deleting a model.
[0326]
[0349] Clause 132 includes any of the methods described in Clauses 129 through 131, and the model is received on the basis of private access.
[0327]
[0350] According to Article 133, the apparatus includes means for detecting the context of a device, means for sending a request indicating the context to a remote device, means for receiving a model corresponding to the context, means for using the model while the context is detected, and means for pruning the model in response to detecting a change in the context.
[0328]
[0351] According to Article 134, a non-temporary computer-readable storage medium includes instructions, when executed by the device's processor, that cause the processor to detect the device's context, send a request to a remote device indicating the context, receive a model corresponding to the context, use the model while the context is detected, and prune the model in response to detecting a change in the context.
[0329]
[0352] The above description of the disclosed embodiments is provided to enable those skilled in the art to manufacture or use the disclosed embodiments. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other embodiments without departing from the scope of this disclosure. Accordingly, this disclosure is not intended to be limited to the embodiments shown herein and should be given the broadest possible scope that coincides with principles and novel features as defined by the following claims. The invention described in the original claims of this application is listed below. [C1] It is a device, The system comprises one or more processors, and the one or more processors are Receiving sensor data from one or more sensor devices, Determining the context of the device based on the aforementioned sensor data, Selecting a model based on the aforementioned context, The model is used to process the input signal in order to generate context-specific output, A device configured to perform the following actions. [C2] The system further comprises a location sensor coupled to one or more of the aforementioned processors, The sensor data includes location data from the location sensor, and the location data indicates the location of the device. The aforementioned context is the device described in C1, which is at least partially based on the aforementioned location. [C3] The system further comprises a camera coupled to one or more of the aforementioned processors, The sensor data includes image data from the camera, and the image data corresponds to a visual scene. The context is the device described in C1, which is at least partially based on the visual scene. [C4] The system further comprises a microphone coupled to one or more of the aforementioned processors, The sensor data includes audio data from the microphone, and the audio data corresponds to an audio scene. The aforementioned context is the device described in C1, which is at least partially based on an audio scene. [C5] The device further comprises an activity detector coupled to one or more processors, The sensor data includes motion data from the activity detector, and the recorded motion data corresponds to the movement of the device. The context is the device described in C1, which is at least partially based on the movement of the device. [C6] The memory further comprises one or more processors coupled to the aforementioned processors, The aforementioned model is selected from among a plurality of models stored in the memory, and is the device described in C1. [C7] The device according to C1, wherein the model includes a sound event detection model, the input signal includes an audio signal, and the context-specific output includes a classification of sound events in the audio signal. [C8] The device according to C1, wherein the model includes an automatic speech recognition model, the input signal includes an audio signal, and the context-specific output includes text data representing the speech in the audio signal. [C9] The device according to C1, wherein the model includes a natural language processing (NLP) model, the input signal includes text data, and the context-specific output includes NLP output data based on the text data. [C10] The device according to C1, wherein the model includes a noise reduction model, the input signal includes an audio signal, and the context-specific output includes a noise-reduced audio signal based on the audio signal. [C11] The aforementioned model is associated with automatic adjustment of the device operating mode, The device according to C1, wherein the context-specific output includes a signal for adjusting the device operating mode. [C12] The device according to C1, further comprising a modem coupled to one or more processors and configured to receive the model from a second device via wireless transmission. [C13] The device described in C12, wherein the context corresponds to the location of the device, and the model includes an acoustic model corresponding to a specific location. [C14] The device according to C13, wherein one or more processors are further configured to receive access permissions for the model of the device that match the specific location, at least in part, based on the location of the device, via the modem. [C15] The device according to C1, wherein one or more processors are further configured to prune the model in response to a determination that the context has changed. [C16] The device according to C1, wherein one or more processors are integrated into an integrated circuit. [C17] The device described in C1, wherein one or more processors are integrated into a vehicle. [C18] The device according to C1, wherein the one or more processors are integrated into at least one of the following: a mobile phone, a tablet computer device, a virtual reality headset, an augmented reality headset, a mixed reality headset, a wireless speaker device, a wearable device, a camera device, or a hearing aid. [C19] A context-based model selection method, One or more processors in the device receive sensor data from one or more sensor devices, In the one or more processors, the context of the device is determined based on the sensor data, In the one or more processors, a model is selected based on the context, In the one or more processors, the model is used to process an input signal in order to generate a context-specific output, A method that includes [a certain feature]. [C20] The method according to C19 further comprises receiving the model from a second device via wireless transmission. [C21] The method according to C20, wherein the context corresponds to the location of the device, and the model includes an acoustic model corresponding to a specific location. [C22] The method according to C21, further comprising receiving access permissions for the model based at least in part on the location of the device that matches the specific location. [C23] The method of C19, further comprising pruning the model in response to a determination that the context has changed. [C24] The aforementioned model is downloaded from a shared model library, using the method described in C19. [C25] The method described in C24, wherein the model includes a trained model uploaded to the library from another user device. [C26] The aforementioned library corresponds to the method of C25, which is a model crowdsourcing library. [C27] The aforementioned library is included in a distributed context-aware system, as described in C25. [C28] The aforementioned context includes a specific acoustic environment, and the aforementioned method is Determining whether the library of available acoustic models is specific to the particular acoustic environment and whether it includes acoustic models available for the device, In response to the fact that an acoustic model specific to the aforementioned particular acoustic environment is not available to the device, determine whether an acoustic model for a general category of the aforementioned particular acoustic environment is available to the device. A method using C19, which further includes the following features. [C29] It is a device, A means for receiving sensor data, Means for determining the context based on the aforementioned sensor data, Means for selecting a model based on the aforementioned context, Means for processing an input signal using the model to generate context-specific output, A device equipped with the following features. [C30] A non-temporary computer-readable storage medium, which, when executed by a processor, provides to the processor, Receiving sensor data from one or more sensor devices, Determining the context based on the aforementioned sensor data, Selecting a model based on the aforementioned context, The model is used to process the input signal in order to generate context-specific output, A non-temporary computer-readable storage medium equipped with instructions to perform the following actions.
Claims
1. It is a device, Memory configured to store one or more models, The system comprises one or more processors coupled to the memory, and the one or more processors are Receiving sensor data from one or more sensor devices, wherein the sensor data includes image data from a camera, and the image data corresponds to a visual scene. The context of the device is determined based on the sensor data, and the context is determined at least partially based on the visual scene. Selecting a specific model based on the aforementioned context, wherein the specific model comprises a sound event classification (SEC) model, and the selection of the specific model causes the device to download the specific model from a cloud-based library of models and store the downloaded specific model in the memory. Processing the input signal using the aforementioned specific model to generate context-specific output, A device configured to perform the following, wherein one or more processors are further configured to update the context based on detecting changes in the sensor data, and to disable the particular model in response to determining that the context has changed.
2. The system further comprises a location sensor coupled to one or more processors, The sensor data includes location data from the location sensor, and the location data indicates the location of the device. The device according to claim 1, wherein the context is at least partially based on the location.
3. The system further comprises a microphone coupled to one or more of the aforementioned processors, The sensor data includes audio data from the microphone, and the audio data corresponds to an audio scene. The device according to claim 1, wherein the context is at least partially based on an audio scene.
4. The device further comprises an activity detector coupled to one or more processors, The sensor data includes motion data from the activity detector, and the recorded motion data corresponds to the movement of the device. The device according to claim 1, wherein the context is at least partially based on the movement of the device.
5. The device according to claim 1, wherein the input signal includes an audio signal, and the context-specific output includes a classification of sound events in the audio signal.
6. The one or more processors are further configured to select a second specific model based on the context, the second specific model is Automatic speech recognition model, wherein the input signal includes an audio signal, and the context-specific output includes text data representing the speech in the audio signal. A natural language processing (NLP) model, wherein the input signal includes text data, and the context-specific output includes NLP output data based on the text data, or Noise reduction model, wherein the input signal includes an audio signal, and the context-specific output includes a noise-reduced audio signal based on the audio signal. The device according to claim 1, including the device described in claim 1.
7. The one or more processors are further configured to select a second specific model based on the context, The second specific model is associated with automatic adjustment of the device operating mode, The device according to claim 1, wherein the context-specific output includes a signal for adjusting the device operating mode.
8. The device according to claim 1, further comprising a modem coupled to one or more processors and configured to receive the specific model from the library via wireless transmission.
9. The device according to claim 8, wherein the context corresponds to the location of the device, and the specific model includes an acoustic model corresponding to the specific location.
10. The device according to claim 9, wherein the one or more processors are further configured to receive access permissions for the particular model of the device that matches the particular location, at least in part, based on the location of the device, via the modem.
11. The device according to claim 1, wherein the one or more processors are integrated into an integrated circuit or a vehicle.
12. The device according to claim 1, wherein the one or more processors are integrated into at least one of the following: a mobile phone, a tablet computer device, a virtual reality headset, an augmented reality headset, a mixed reality headset, a wireless speaker device, a wearable device, a camera device, or a hearing aid.
13. A context-based model selection method, One or more processors of the device receive sensor data from one or more sensor devices, and the sensor data includes image data from a camera, and the image data corresponds to a visual scene. In the one or more processors, the context of the device is determined based on the sensor data, and the context is determined at least partially based on the visual scene. In the one or more processors, the selection of a specific model is performed based on the context, wherein the specific model comprises a sound event classification (SEC) model, and the selection of the specific model causes the device to download the specific model from a cloud-based library of models and to store the downloaded specific model in the device's memory. In the one or more processors, the input signal is processed using the specific model to generate context-specific output, Updating the context based on detecting a change in the sensor data, In response to the determination that the aforementioned context has changed, the model is made unusable, A method that includes [a certain feature].
14. A non-temporary computer-readable storage medium comprising, when executed by a processor, an instruction causing the processor to perform the method according to claim 13.
Citation Information
Patent Citations
Method and Apparatus for Identifying Acoustic Background Environments Based on Time and Speed to Enhance Automatic Speech Recognition
US20140303972A1
User-specific acoustic models
US20180330737A1
Contextual sound filter
US20180336000A1
Device and a method for classifying an acoustic environment
US20190206418A1
Signal processing device, signal processing method, and computer-readable recording medium
WO2017217412A1