Adaptive sound event classification
By employing transfer learning techniques and adapter networks, the sound event classification model is automatically updated to adapt to environmental drift and identify new sound categories. This solves the problems of difficult updates and resource waste in existing sound event classification systems, and improves recognition accuracy and resource utilization efficiency.
Patent Information
- Application Number
- CN202180077242.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-24
- Filing Date
- 2021-11-19
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2041-11-19
AI Technical Summary
Existing sound event classification systems struggle to be updated to identify new sound categories not identified in labeled training data. Furthermore, as the number of audio data samples increases over time, processing becomes more difficult, leading to wasted computing resources and decreased recognition accuracy.
By employing transfer learning techniques, the sound event classification model is automatically updated to adapt to environmental drift and identify new sound categories by combining scene data and audio data samples. The model is dynamically updated using transfer learning techniques and adapter networks, reducing the consumption of computing resources.
It enables dynamic updating and adaptation of sound event classification models under limited computing resources, improves the ability to identify new sound categories and adapt to environmental changes, and reduces the demand for computing and storage resources.
Smart Images

Figure CN116457879B_ABST
Abstract
Description
[0001] I. Priority Claim
[0002] This application claims priority to commonly owned U.S. Nonprovisional Patent Application No. 17 / 102,724, filed November 24, 2020, the contents of which are expressly incorporated by reference in their entirety.
[0003] II. Field
[0004] The present disclosure relates generally to adaptive sound event classification.
[0005] III. Related Art
[0006] Advances in technology have resulted in smaller and more powerful computing devices. For example, there currently exist a variety of portable personal computing devices, including wireless telephones such as mobile and smart phones, tablets, and laptop computers, that are small, lightweight, and easily carried by users. These devices can communicate voice and data packets over wireless networks. Further, many such devices incorporate additional functionality such as digital still cameras, digital video cameras, digital recorders, and audio file players. Still further, such devices can process executable instructions including software applications, such as web browser applications that can be used to access the Internet. As such, these devices can include significant computing capabilities including, for example, sound event classification (SEC) systems that attempt to recognize sound events (e.g., a slamming door, a car horn, etc.) in audio signals.
[0007] SEC systems are generally trained using supervised machine learning techniques to recognize a particular set of sounds identified in labeled training data. As a result, each SEC system tends to be specific to a particular domain (e.g., capable of classifying a predetermined set of sounds). After the SEC system is trained, it is difficult to update the SEC system to recognize new sound classes that were not identified in the labeled training data. Additionally, some sound classes that the SEC system is trained to detect can represent sound events with more variants than represented in the labeled training data. To illustrate, the labeled training data can include audio data samples of many different doorbells, but is unlikely to include all existing variants of the doorbell sound. Re-training the SEC system to recognize a new sound not represented in the training data used to train the SEC system can involve completely re-training the SEC system using a new set of labeled training data that includes examples of the new sound in addition to the original training data. As a result, training the SEC system to recognize new sounds (whether for a new sound class or for a variant of an existing sound class) requires approximately the same computational resources (e.g., processor cycles, memory, etc.) as generating an entirely new SEC system. Furthermore, over time, as more sounds are added to be recognized, the number of audio data samples that must be maintained and used to train the SEC system can become intractable.
[0008] IV. SUMMARY
[0009] In particular aspects, an apparatus includes one or more processors configured to provide audio data samples to a sound event classification model and receive an output of the sound event classification model responsive to the audio data samples. The one or more processors are further configured to determine, based on the output, whether a sound class of the audio data samples is recognized by the sound event classification model. The one or more processors are further configured to determine, based on a determination that the sound class is not recognized, whether the sound event classification model corresponds to an audio scene associated with the audio data samples. The one or more processors are further configured to store, based on a determination that the sound event classification model corresponds to the audio scene associated with the audio data samples, model update data based on the audio data samples. In particular aspects, a method includes providing, by one or more processors, audio data samples as input to a sound event classification model. The method further includes determining, by the one or more processors, based on an output of the sound event classification model responsive to the audio data samples, whether a sound class of the audio data samples is recognized by the sound event classification model. The method further includes determining, by the one or more processors, based on a determination that the sound class is not recognized, whether the sound event classification model corresponds to an audio scene associated with the audio data samples. The method further includes storing, by the one or more processors, based on a determination that the sound event classification model corresponds to the audio scene associated with the audio data samples, model update data based on the audio data samples.
[0010] In particular aspects, an apparatus includes means for providing audio data samples to a sound event classification model. The apparatus further includes means for determining, based on an output of the sound classification model, whether a sound class of the audio data samples is recognized by the sound event classification model. The apparatus further includes means for determining, in response to a determination that the sound class is not recognized, whether the sound event classification model corresponds to an audio scene associated with the audio data samples. The apparatus further includes means for storing, in response to a determination that the sound event classification model corresponds to the audio scene associated with the audio data samples, model update data based on the audio data samples.
[0011] In a particular aspect, a non-transitory computer-readable storage medium includes instructions that, when executed by a processor, cause the processor to provide audio data samples as input to a sound event classification model. The instructions, when executed by the processor, also cause the processor to determine whether a sound class of the audio data samples is recognized by the sound event classification model based on an output of the sound event classification model responsive to the audio data samples. The instructions, when executed by the processor, further cause the processor to determine whether the sound event classification model corresponds to an audio scene associated with the audio data samples based on a determination that the sound class is not recognized. The instructions, when executed by the processor, also cause the processor to store model update data based on the audio data samples based on a determination that the sound event classification model corresponds to the audio scene associated with the audio data samples.
[0012] Other aspects, advantages, and features of the present disclosure will become apparent after review of the entire application, including the following sections: BRIEF
[0013] V. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 is a block diagram of an example of a device configured to generate sound identification data in response to audio data samples and configured to update a sound event classification model.
[0015] Figure 2 is a diagram illustrating aspects of updating a sound event classification model to account for a drift, according to a particular example.
[0016] Figure 3 is a diagram illustrating aspects of updating a sound event classification model to account for a new sound class, according to a particular example.
[0017] Figure 4 is a diagram illustrating a particular example of operation of the device of Figure 1
[0018] Figure 5 is a diagram illustrating another particular example of operation of the device of Figure 1
[0019] Figure 6 is a block diagram of a particular example of a device of Figure 1
[0020] Figure 7 is an illustrative example of a vehicle incorporating aspects of the device of Figure 1
[0021] Figure 8 incorporating aspects of the device of Figure 1 a virtual reality, mixed reality, or augmented reality headset incorporating aspects of the devices of
[0022] Figure 9 a wearable electronic device incorporating aspects of the devices of Figure 1
[0023] Figure 10 a voice-controlled speaker system incorporating aspects of the devices of Figure 1
[0024] Figure 11 a camera incorporating aspects of the devices of Figure 1
[0025] Figure 12 a mobile device incorporating aspects of the devices of Figure 1
[0026] Figure 13 an aerial device incorporating aspects of the devices of Figure 1
[0027] Figure 14 a headset incorporating aspects of the devices of Figure 1
[0028] Figure 15 an appliance incorporating aspects of the devices of Figure 1
[0029] is a flowchart of an example of a method of operation of the devices of Figure 16 Figure 1
[0030] VI. DETAILED DESCRIPTION
[0031] A sound event classification model can be trained using machine learning techniques. For example, a neural network can be trained as a sound event classifier using backpropagation or other machine learning training techniques. A neural network trained in this manner is referred to herein as a "sound event classification model." A sound event classification model trained in this manner can be small enough (in terms of storage space occupied) and simple enough (in terms of computing resources used during operation) for a portable computing device to store and use the sound event classification model. The process of training a sound event classification model uses many more processing resources than are used to perform sound event classification using a sound event classification model. Additionally, the training process uses a large labeled training dataset that includes many audio data samples for each sound class that the sound event classification model is being trained to detect. Training a sound event classification model from scratch on a portable computing device or another resource-constrained computing device can be prohibitive in terms of memory utilization or other computing resources. As a result, a user desiring to use a sound event classification model on a portable computing device can be limited to downloading a pre-trained sound event classification model onto the portable computing device from a less resource-constrained computing device or a library of pre-trained sound event classification models. Thus, the user has limited customization options.
[0032] The disclosed systems and methods use transfer learning techniques to update a sound event classification model in a manner that uses much less computing resources than training a sound event classification model from scratch. According to particular aspects, transfer learning techniques can be used to update a sound event classification model to account for drift within a sound class or to recognize a new sound class. In this context, "drift" refers to a change within a sound class. For example, a sound event classification model can be able to recognize some examples of the sound class, but can not be able to recognize other examples of the sound class. To illustrate, a sound event classification model trained to recognize a car horn sound class can be able to recognize many different types of car horns, but can not be able to recognize certain examples of car horns. Drift can also occur due to changes in the acoustic environment. To illustrate, a sound event classification model can be trained to recognize a bass drum played in a concert hall, but can not be able to recognize the bass drum in the case where the bass drum is played by a military band in an outdoor parade. The transfer learning techniques disclosed herein facilitate updating a sound event classification model to account for such drift, which enables the sound event classification model to detect a wider range of sounds within a sound class. Since drift can correspond to sounds that a user device encounters but that the user device does not recognize, updating a sound event classification model to accommodate these encountered sound class changes enables the user device to more accurately identify particular sound class changes that the particular user device commonly encounters.
[0033] According to particular aspects, when a sound event classification model is determined to not have recognized a sound class of a sound (based on an audio data sample of the sound), a determination is made as to whether the sound was not recognized due to drift or because the sound event classification model did not recognize a type of sound class associated with the sound. For example, information different from the audio data sample (such as a timestamp, location data, image data, video data, user input data, settings data, other sensor data, etc.) is used to determine scene data indicative of a sound environment (or audio scene) associated with the audio data sample. The scene data is used to determine whether the sound event classification model corresponds to the audio scene (e.g., was trained to recognize sound events in the audio scene). If the sound event classification model corresponds to the audio scene, the audio data sample is saved as model update data and indicated as drift data. In some aspects, if the sound event classification model does not correspond to the audio scene, the audio data sample is discarded as unknown, or saved as model update data and indicated as being associated with an unknown sound class (e.g., unknown data).
[0034] Periodically or occasionally (e.g., when initiated by a user or when an update condition is met), the sound event classification models are updated using the model update data. For example, to account for the drift data, the sound event classifier can be trained using backpropagation or other similar machine learning techniques (e.g., by further training from a sound event classifier that has already been trained). In this example, the drift data is associated with labels of sound classes that have been recognized by the sound event classification model, and the drift data and corresponding labels are used as labeled training data. Updating the sound event classification model using the drift data can be augmented by adding other examples of sound classes to the labeled training data, such as examples taken from training data originally used to train the sound event classification model. In some aspects, the device automatically (e.g., without user input) updates one or more sound event classification models when drift data is available. Thus, the sound event classification system can automatically adapt to account for drift within sound classes, using much less computational resources than would be used to train the sound event classification model from scratch.
[0035] To account for unknown data, the sound event classification model can be trained using more complex transfer learning techniques. For example, when unknown data is available, a user can be asked to indicate whether the user desires to update the sound event classification model. Audio representing the unknown data can be played to the user, and the user can indicate that the unknown data is to be discarded without updating the sound event classification model, can indicate that the unknown data corresponds to a known sound class (e.g., reclassifying the unknown data as drift data), or can assign a new sound class label to the unknown data. If the user reclassifies the unknown data as drift data, the machine learning technique(s) for updating the sound event classification model to account for the drift data is initiated, as described above.
[0036] If the user assigns a new sound class label to the unknown data, the label and the unknown data are used as labeled training data to generate an updated sound event classification model. According to particular aspects, the transfer learning technique for updating the sound event classification model includes generating a copy of the sound event classifier model that includes an output node associated with the new sound class. The copy of the sound event classifier model is referred to as an incremental model. The transfer learning technique also includes connecting the sound event classification model and the incremental model to one or more adapter networks. The adapter network(s) facilitate generation of a merged output based on both the output of the sound event classification model and the output of the incremental model. Audio data samples including the unknown data and one or more audio data samples corresponding to known sound classes (e.g., sound classes that the sound event classifier was previously trained to recognize) are provided to the sound event classification model and the incremental model to generate the merged output. The merged output indicates a sound class assigned to the audio data samples based on analysis by the sound event classification model, the incremental model, and the one or more adapter networks. During training, the merged output is used to update link weights of the incremental model and the adapter network(s). When training is complete, if the incremental model is sufficiently accurate, the sound event classifier can be discarded. If the incremental model alone is not sufficiently accurate, the sound event classification model, the incremental model, and the adapter network(s) are retained together and used as a single updated sound event classification model. Thus, the techniques disclosed herein enable customization and updating of sound event classification models in a less resource-intensive (in terms of memory resources, processor time, and power) manner than training a neural network from scratch. Additionally, in some aspects, the disclosed techniques enable automatic updating of sound event classification models to account for drift.
[0037] The disclosed systems and methods provide a context-aware system that can detect dataset shift, associate the shifted data with a corresponding class (e.g., by leveraging available multi-modal inputs), and refine / fine-tune a SEC model with the shifted data with little to no supervision and without the need to train a new SEC model from scratch. In some aspects, prior to refining / fine-tuning the SEC model, the SEC model is trained to recognize a plurality of variants of a particular sound class, and refining / fine-tuning the SEC model modifies the SEC model to enable the SEC model to recognize additional variants of the particular sound class.
[0038] In some aspects, the disclosed systems and methods can be used in applications that suffer from dataset shift during testing. For example, these systems and methods can detect dataset shift and refine a SEC model without the need to retrain a previously learned sound class from scratch. In some aspects, the disclosed systems and methods can be used to add a new sound class to an existing SEC model (e.g., a SEC model that has already been trained for certain sound classes) without the need to retrain the SEC model from scratch, without the need to access all of the training data that was originally used to train the SEC model, and without introducing any performance degradation with respect to the sound classes for which the SEC model was originally trained.
[0039] In some aspects, the disclosed systems and methods can be used in applications that desire continuous learning capabilities under low footprint constraints. In some implementations, the disclosed systems and methods can access a database of various detection models (e.g., SEC models) for a wide variety of applications (e.g., various sound environments). In such implementations, a SEC model can be selected based on a sound environment during operation, and the SEC model can be loaded and used as a source model.
[0040] Particular aspects of the present disclosure are described below with reference to the accompanying drawings. Like reference numerals are used to denote like features in the various Figure 1 One or more sensors (e.g., a microphone) are depicted. Figure 1The device 100 in which the“sensor(s)” 134 are included indicates that in some implementations the device 100 includes a single sensor 134, while in other implementations the device 100 includes multiple sensors 134. For ease of reference herein, such features are generally introduced as“one or more” features, and are subsequently referred to in the singular or optionally in the plural (typically indicated by a term ending in“(s)”), unless an aspect related to multiple ones of these features is being described.
[0041] The terms“include,”“have,” and“comprise” are used interchangeably in this disclosure. In addition, the term“wherein” is used interchangeably with“whereby” herein. As used herein, “exemplary” indicates an example, an implementation, and / or an aspect, and should not be construed as a limitation or as indicating a preference or a preferred implementation. As used herein, ordinal terms (e.g.,“first,”“second,”“third,” etc.) are used interchangeably with“one,”“one,” and“one,” and are not meant to indicate any priority or order. As used herein, the term“set” refers to one or more particular elements, while the term“plurality” refers to multiple particular elements (e.g., two or more particular elements).
[0042] As used herein,“coupled” can include“communicatively coupled,”“electrically coupled,” or“physically coupled,” and can additionally (or alternatively) include any combination thereof. Two devices (or components) can be coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) directly or indirectly via one or more other devices, components, wires, buses, networks (e.g., a wired network, a wireless network, or a combination thereof), etc. Two devices (or components) that are electrically coupled can be included in the same device or different devices, and can be connected via electronics, one or more connectors, or inductive coupling, as illustrative, non-limiting examples. In some implementations, two devices (or components) that are communicatively coupled (such as in electrical communication) can send and receive electrical signals (digital signals or analog signals) directly or indirectly via one or more wires, buses, networks, etc. As used herein,“directly coupled” indicates that two devices are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without intervening components.
[0043] In this disclosure, terms such as “determine,” “calculate,” “estimate,” “shift,” “adjust,” and the like can be used to describe how one or more operations are performed. It should be noted that such terms are not to be construed as limiting and that other techniques can be utilized to perform similar operations. Additionally, as referenced herein, “generate,” “calculate,” “estimate,” “use,” “select,” “access,” and “determine” can be used interchangeably. For example, “generating,” “calculating,” “estimating,” or “determining” a parameter (or signal) can refer to actively generating, estimating, calculating, or determining the parameter (or signal), or can refer to using, selecting, or accessing a parameter (or signal) that has been generated (such as by another component or device).
[0044] Figure 1 is a block diagram of an example of a device 100 configured to generate sound identification data in response to audio data samples 110 and configured to update a sound event classification model. Figure 1 The device 100 of FIG. 1 includes one or more microphones 104 configured to generate audio signals 106 based on sounds 102 detected within an acoustic environment. The microphone(s) 104 are coupled to a feature extractor 108 that generates audio data samples 110 based on the audio signals 106. For example, the audio data samples 110 can include an array or matrix of data elements, where each data element corresponds to a feature detected in the audio signals 106. As a specific example, the audio data samples 110 can correspond to Mel-spectral features extracted from one second of audio signals 106. In this example, the audio data samples 110 can include a 128x128 matrix of feature value elements. In other examples, other audio data sample configurations or sizes can be used.
[0045] The audio data samples 110 are provided to a sound event classification (SEC) engine 120. The SEC engine 120 is configured to perform inference operations based on one or more SEC models, such as SEC model 112. An “inference operation” refers to assigning the audio data samples 110 to a sound class in which the audio data samples 110 are identified by the SEC model 112. For example, the SEC engine 120 can include or correspond to software that implements a machine learning runtime environment, such as the Qualcomm Neural Processing SDK, which is available from Qualcomm Technologies, Inc. of San Diego, California. In a particular aspect, the SEC model 112 is one of a plurality of SEC models (e.g., available SEC models 114) available to the SEC engine 120.
[0046] In a particular example, each of the available SEC models 114 includes or corresponds to a neural network trained as a sound event classifier. To illustrate, a SEC model 112 (as well as each of the other available SEC models 114) can include an input layer, one or more hidden layers, and an output layer. In this example, the input layer is configured to correspond to an array or matrix of values of the audio data sample 110 generated by the feature extractor 108. To illustrate, if the audio data sample 110 includes 15 data elements, the input layer can include 15 nodes (e.g., one node per data element). The output layer is configured to correspond to the sound classes that the SEC model 112 is trained to recognize. The particular arrangement of the output layer can vary depending on the information to be provided as output. As one example, the SEC model 112 can be trained to output an array including one bit per sound class, where the output layer performs “one-hot encoding” such that all but one bit of the output array has a value of 0, and the bit corresponding to the detected sound class has a value of 1. Other output schemes can be used to indicate, for example, a confidence metric value for each sound class, where the confidence metric value indicates a probability estimate that the audio data sample 110 corresponds to the respective sound class. To illustrate, if the SEC model 112 is trained to recognize four sound classes, the SEC model 112 can generate output data including four values (one value per sound class), and each value can indicate a probability estimate that the audio data sample 110 corresponds to the respective sound class.
[0047] Each hidden layer includes a plurality of nodes, and each node is interconnected (via links) with other nodes in the same layer or different layers. Each input link of a node is associated with a link weight. During operation, a node receives input values from other nodes to which it is linked, weights these input values based on the corresponding link weights to determine a combined value, and feeds the combined value into an activation function to generate an output value for the node. The output value is provided to one or more other nodes via an output link of the node. A node can also include a bias value that is used in generating the combined value. The nodes can be linked in various arrangements, and can include various other features (e.g., prior value memory) to facilitate processing of particular data. In the case of audio data samples, convolutional neural networks (CNNs) can be used. To illustrate, one or more of the SEC models 112 can include three linked CNNs, and each CNN can include two-dimensional (2D) convolutional layers, max-pooling layers, and batch normalization layers. In other implementations, the hidden layers include different numbers of CNNs or other layers. Training a neural network includes modifying the link weights to reduce an output error of the neural network.
[0048] During operation, the SEC engine 120 can provide the audio data sample 110 as input to a single SEC model (e.g., the SEC model 112), to multiple selected SEC models (e.g., the SEC model 112 and the Kth SEC model 118 of the available SEC models 114), or to each SEC model (e.g., to the SEC model 112, the first SEC model 116 of the available SEC models 114, the Kth SEC model 118, and any other SEC models). For example, the SEC engine 120 (or another component of the device 100) can select the SEC model 112 from among the available SEC models 114 based on, for example, user input, a device setting associated with the device 100, sensor data, a time at which the audio data sample 110 is received, or other factors. In this example, the SEC engine 120 can select to use only the SEC model 112, or can select to use two or more of the available SEC models 114. To illustrate, a device setting can indicate that the SEC model 112 and the first SEC model 116 are to be used during a particular timeframe. In another example, the SEC engine 120 can provide the audio data sample 110 to each of the available SEC models 114 (e.g., sequentially or in parallel) to generate an output from each SEC model. In particular aspects, the SEC models are trained to identify different sound classes, to identify the same sound classes in different acoustic environments, or both. For example, the SEC model 112 can be configured to identify a first set of sound classes, and the first SEC model 116 can be configured to identify a second set of sound classes, where the first set of sound classes is different from the second set of sound classes.
[0049] In particular aspects, the SEC engine 120 determines whether the SEC model 112 identified a sound class of the audio data sample 110 based on the output of the SEC model 112. If the SEC engine 120 provided the audio data sample 110 to multiple SEC models, the SEC engine 120 can determine whether any of the SEC models identified a sound class of the audio data sample 110 based on the output of each SEC model. If the SEC model 112 (or another of the available SEC models 114) identified a sound class of the audio data sample 110, the SEC engine 120 generates an output 124 that indicates the sound class 122 of the audio data sample 110. For example, the output 124 can be sent to a display to notify a user that the sound class 122 associated with the sound 102 was detected, or can be sent to another device or another component of the device 100 and used to trigger an action (e.g., send a command to activate a light in response to identifying a sound of a closing door).
[0050] If the SEC engine 120 determines that the SEC model 112 (and other SEC models 114 for which the audio data sample 110 was provided, if available) does not identify a sound class of the audio data sample 110, the SEC engine 120 provides a trigger signal 126 to the drift detector 128. For example, the SEC engine 120 can set a trigger flag in memory of the device 110. In some implementations, the SEC engine 120 can also provide other data to the drift detector 128. To illustrate, if the SEC model 112 generates a confidence metric value for each sound class that the SEC model 112 is trained to identify, one or more of these confidence metric values can be provided to the drift detector 128. For example, if the SEC model 112 is trained to identify three sound classes, the SEC engine 120 can provide the highest confidence value among the three confidence values (one for each of the three sound classes) output by the SEC model 112 to the drift detector 128.
[0051] In a particular aspect, the SEC engine 120 determines whether the SEC model 112 identifies a sound class of the audio data sample 110 based on the confidence metric value. In this particular aspect, the confidence metric value for a particular sound class indicates a probability that the audio data sample 110 is associated with the particular sound class. To illustrate, if the SEC model 112 is trained to identify four sound classes, the SEC model 112 can generate an array including four confidence metric values (one for each sound class) as output. In some implementations, the SEC engine 120 determines that the SEC model 112 identifies the sound class 122 of the audio data sample 110 if the confidence metric value for the sound class 122 is greater than a detection threshold. For example, the SEC engine 120 determines that the SEC model 112 identifies the sound class 122 of the audio data sample 110 if the confidence metric value for the sound class 122 is greater than 0.90 (e.g., 90% confidence), 0.95 (e.g., 95% confidence), or some other value for the detection threshold. In some implementations, the SEC engine 120 determines that the SEC model 112 does not identify a sound class of the audio data sample 110 if the confidence metric value for each sound class that the SEC model 112 is trained to identify is less than a detection threshold. For example, the SEC engine 120 determines that the SEC model 112 does not identify the sound class 122 of the audio data sample 110 if each confidence metric value is less than 0.90 (e.g., 90% confidence), 0.95 (e.g., 95% confidence), or some other value for the detection threshold.
[0052] The drift detector 128 is configured to determine whether the SEC model 112 that fails to identify a sound class of the audio data sample 110 corresponds to the audio scene 142 associated with the audio data sample 110. In a particular aspect, the drift detector 128 determines whether the SEC model 112 that fails to identify a sound class of the audio data sample 110 corresponds to the audio scene 142 associated with the audio data sample 110 based on the trigger signal 126 and the other data provided by the SEC engine 120. For example, the drift detector 128 can determine whether the SEC model 112 that fails to identify a sound class of the audio data sample 110 corresponds to the audio scene 142 associated with the audio data sample 110 based on the trigger signal 126 and the highest confidence value provided by the SEC engine 120. Figure 1In the example illustrated, the scene detector 140 is configured to receive the scene data 138 and use the scene data 138 to determine an audio scene 142 associated with the audio data samples 110. In particular aspects, the scene data 138 is generated based on setting data 130 indicative of one or more device settings associated with the device 100, an output of a clock 132, sensor data from one or more sensors 134, inputs received via an input device 136, or a combination thereof. In some aspects, the scene detector 140 uses different information to determine the audio scene 142 than the information used by the SEC engine 120 to select the SEC model 112. To illustrate, if the SEC engine 120 selects the SEC model 112 based on a time of day, the scene detector 140 can use location sensor data from a location sensor of the sensor(s) 134 to determine the audio scene 142. In some aspects, the scene detector 140 uses at least some of the same information used by the SEC engine 120 to select the SEC model 112, and uses additional information. To illustrate, if the SEC engine 120 selects the SEC model 112 based on a time of day and the setting data 130, the scene detector 140 can use the location sensor data and the setting data 130 to determine the audio scene 142. Thus, the scene detector 140 uses a different audio scene detection mode than the audio scene detection mode used by the SEC engine 120 to select the SEC model 112.
[0053] In particular implementations, the scene detector 140 is a neural network trained to determine the audio scene 142 based on the scene data 138. In other implementations, the scene detector 140 is a classifier trained using a different machine learning technique. For example, the scene detector 140 can include or correspond to a decision tree, a random forest, a support vector machine, or another classifier trained to generate an output indicative of the audio scene 142 based on the scene data 138. In yet other implementations, the scene detector 140 determines the audio scene 142 based on the scene data 138 using heuristics. In still other implementations, the scene detector 140 determines the audio scene 142 based on the scene data 138 using a combination of artificial intelligence and heuristics. For example, the scene data 138 can include image data, video data, or both, and the scene detector 140 can include an image recognition model trained using a machine learning technique to detect particular objects, motions, backgrounds, or other image or video information. In this example, the output of the image recognition model can be evaluated via one or more heuristics to determine the audio scene 142.
[0054] Drift detector 128 compares audio scene 142 indicated by scene detector 140 to information describing SEC model 112 to determine whether SEC model 112 is associated with audio scene 142 of audio data sample 110. If drift detector 128 determines that SEC model 112 is associated with audio scene 142 of audio data sample 110, drift detector 128 causes drift data 144 to be stored as model update data 148. In particular implementations, drift data 144 includes audio data sample 110 and a label that identifies SEC model 112, indicates a sound class associated with audio data sample 110, or both. If drift data 144 indicates a sound class associated with audio data sample 110, the sound class can be selected based on a highest confidence measure value generated by SEC model 112. As an illustrative example, if SEC engine 120 uses a detection threshold of 0.90, and for a particular sound class, the highest confidence measure value output by SEC model 112 is 0.85, SEC engine 120 determines that the sound class of audio data sample 110 is not recognized, and sends a trigger signal 126 to drift detector 128. In this example, if drift detector 128 determines that SEC model 112 corresponds to audio scene 142 of audio data sample 110, drift detector 128 stores audio data sample 110 as drift data 144 associated with the particular sound class. In particular aspects, metadata associated with SEC models 114 includes information specifying one or more audio scenes associated with each SEC model 114. For example, SEC model 112 can be configured to detect sound events in a user's home, in which case metadata associated with SEC model 112 can indicate that SEC model 112 is associated with a "home" audio scene. In this example, if audio scene 142 indicates that device 100 is in a home location (e.g., based on positioning information, user input, detection of a home wireless network signal, image or video data representative of a home location, etc.), drift detector 128 determines that SEC model 112 corresponds to audio scene 142.
[0055] In some implementations, the drift detector 128 also causes some of the audio data samples 110 to be stored as model update data 418 and designated as unknown data 146. As a first example, the drift detector 128 can store unknown data 146 if the drift detector 128 determines that the SEC model 112 does not correspond to the audio scene 142 of the audio data sample 110. As a second example, the drift detector 128 can store unknown data 146 if the confidence metric value output by the SEC model 112 fails to satisfy a drift threshold. In this example, the drift threshold is less than the detection threshold used by the SEC engine 120. For example, if the SEC engine 120 uses a detection threshold of 0.95, the drift threshold can have a value of 0.80, 0.75, or some other value less than 0.95. In this example, if the highest confidence metric value for the audio data sample 110 is less than the drift threshold, the drift detector 128 determines that the audio data sample 110 belongs to a sound class for which the SEC model 112 is not trained to recognize and designates the audio data sample 110 as unknown data 146. In a particular aspect, the drift detector 128 stores unknown data 146 only if the drift detector 128 determines that the SEC model 112 corresponds to the audio scene 142 of the audio data sample 110. In another particular aspect, the drift detector 128 stores unknown data 146 regardless of whether the drift detector 128 determines that the SEC model 112 corresponds to the audio scene 142 of the audio data sample 110.
[0056] After the model update data 148 is stored, the model updater 152 can access the model update data 148 and use the model update data 148 to update one of the available SEC models 114 (e.g., the SEC model 112). For example, each entry of the model update data 148 indicates a SEC model associated with the entry, and the model updater 152 uses the entry as training data to update the corresponding SEC model. In a particular aspect, the model updater 152 updates the SEC models when an update criterion is satisfied, or when a user or another party (e.g., a vendor of the device 100, the SEC engine 120, the SEC models 114, etc.) initiates a model update. The update criterion can be satisfied when a particular number of entries in the model update data 148 are available, when a particular number of entries in the model update data 148 are available for a particular SEC model, when a particular number of entries in the model update data 148 are available for a particular sound class, when a particular amount of time has passed since a previous update, when another update occurs (e.g., when a software update associated with the device 100 occurs), or based on the occurrence of another event.
[0057] The model updater 152 updates the training of the SEC model 112 using the drift data 144 as labeled training data using backpropagation or a similar machine learning optimization process. For example, the model updater 152 provides audio data samples from the drift data 144 of the model update data 148 as input to the SEC model 112, determines a value of an error function (also referred to as a loss function) based on an output of the SEC model 112 and a label associated with the audio data samples (as indicated in the drift data 144 stored by the drift detector 128), and uses a gradient descent operation (or some variation thereof) or another machine learning optimization process to determine updated link weights of the SEC model 112.
[0058] The model updater 152 can also provide other audio data samples to the SEC model 112 during the update training in addition to the audio data samples of the drift data 144. For example, the model update data 148 can include one or more unknown audio data samples (such as a subset of the audio data samples originally used to train the SEC model 112), which can reduce the chance that the update training causes the SEC model 112 to forget previously trained sounds classes (where “forget” here refers to a loss of reliability in detecting sound classes that the SEC model 112 was previously trained to recognize). Since the sound classes associated with the audio data samples of the drift data 114 are indicated by the drift detector 128, the update training to account for the drift can be done automatically (e.g., without user input). As a result, the functionality of the device 100 (e.g., the accuracy of recognizing sound classes) can improve over time without user intervention and using less computational resources than would be used to generate a new SEC model from scratch. See Figure 2 A particular example of a transfer learning process that the model updater 152 can use to update the SEC model 112 based on the drift data 144 is described.
[0059] In some aspects, the model updater 152 can also use unknown data 146 of the model update data 148 to update training of the SEC model 112. For example, periodically or occasionally (such as when an update criterion is met), the model updater 152 can prompt a user to request that the user label a sound class of an item of unknown data 146 in the model update data 148. If the user chooses to label the sound class of the item of unknown data 146, the device 100 (or another device) can play a sound corresponding to an audio data sample of the unknown data 146. The user can provide one or more labels 150 identifying a sound class of the audio data sample (e.g., via the input device 136). If the sound class indicated by the user is a sound class that the SEC model 112 is trained to recognize, the unknown data 146 is reclassified as drift data 144 associated with the user-specified sound class and the SEC model 112. Depending on the configuration of the model updater 152, if the sound class indicated by the user is a sound class that the SEC model 112 is not trained to recognize (e.g., is a new sound class), the model updater 152 can discard the unknown data 146, send the unknown data 146 and the user-specified sound class to another device for use in generating a new or updated SEC model, or can use the unknown data 146 and the user-specified sound class to update the SEC model 112. See Figure 3 Particular examples of a transfer learning process by which the model updater 152 can update the SEC model 112 based on unknown data 146 and user-specified sound classes are described.
[0060] The updated SEC model 154 generated by the model updater 152 is added to the available SEC models 114 so that the updated SEC model 154 is available for use in evaluating audio data samples 110 received after the updated SEC model 154 is generated. Thus, the set of available SEC models 114 that can be used to evaluate sounds is dynamic. For example, one or more of the available SEC models 114 can be automatically updated to account for drift data 144. Additionally, one or more of the available SEC models 114 can be updated to account for unknown sound classes using fewer computational resources (e.g., memory, processing time, and power) than would be used to train a new SEC model from scratch.
[0061] Figure 2 is a diagram illustrating aspects of updating a SEC model 208 to account for drift, according to a particular example. Figure 2 The SEC model 208 of Figure 1The specific SEC model associated with drift data 144 in the available SEC models 114. For example, if SEC engine 120 generates trigger signal 126 in response to the output of SEC model 112, then drift data 144 is associated with SEC model 112, and SEC model 208 corresponds to or includes SEC model 112. As another example, if SEC engine 120 generates trigger signal 126 in response to the output of the Kth SEC model 118, then drift data 144 is associated with the Kth SEC model 118, and SEC model 208 corresponds to or includes the Kth SEC model 118.
[0062] exist Figure 2 In the example explained, training data 202 is used to update the SEC model 208. Training data 202 includes drift data 144 and one or more labels 204. Each entry in drift data 144 includes an audio data sample (e.g., audio data sample 206) and is associated with a corresponding label in the labels 204. The audio data sample of each entry in drift data 144 includes a set of values representing features extracted from or determined based on sounds never recognized by the SEC model 208. The label 204 corresponding to each entry in drift data 144 identifies the sound category to which the expected sound belongs. As an example, the label 204 corresponding to each entry in drift data 144 could be derived from... Figure 1 The drift detector 128 is assigned in response to determining that the SEC model 208 corresponds to the audio scene from which the audio data samples were generated. In this example, the drift detector 128 may assign the audio data samples to the sound category associated with the highest confidence metric in the output of the SEC model 208.
[0063] exist Figure 2 In this process, audio data samples 206 corresponding to sounds are provided to the SEC model 208, and the SEC model 208 generates an output 210 indicating the sound category, one or more confidence metrics, or both, to which the audio data sample 206 is assigned. The model updater 152 uses the output 210 and the label 204 corresponding to the audio data sample 206 to determine the updated link weights 212 of the SEC model 208. The SEC model 208 is updated based on the updated link weights 212, and the training process is iteratively repeated until the training termination condition is met. During training, each entry of drift data 144 may be provided to the SEC model 208 (e.g., one entry per iteration). Additionally, in some implementations, other audio data samples (e.g., audio data samples previously used to train the SEC model 208) may also be provided to the SEC model 208 to reduce the chance of the SEC model 208 forgetting previous training data.
[0064] The training termination condition can be satisfied when all of the drift data 144 has been provided to the SEC model 208 at least once, after a particular number of training iterations have been performed, when a measure of convergence satisfies a convergence threshold, or when some other condition indicative of the end of training is satisfied. When the training termination condition is satisfied, the model updater 152 stores an updated SEC model 214, where the updated SEC model 214 corresponds to the SEC model 208 with link weights based on the updated link weights 212 applied during training.
[0065] Figure 3 is a diagram illustrating aspects of updating the SEC model 310 to account for unknown data, according to a particular example. Figure 3 The SEC model 310 includes or corresponds to Figure 1 A particular SEC model of the available SEC models 114 with which the unknown data 146 is associated. For example, if the SEC engine 120 generated the trigger signal 126 in response to an output of the SEC model 112, the unknown data 146 is associated with the SEC model 112, and the SEC model 310 corresponds to or includes the SEC model 112. As another example, if the SEC engine 120 generated the trigger signal 126 in response to an output of the Kth SEC model 118, the unknown data 146 is associated with the Kth SEC model 118, and the SEC model 310 corresponds to or includes the Kth SEC model 118.
[0066] In some examples, the SEC model 310 is updated to account for the unknown data 146 by updating the link weights of the SEC model 310 based on the drift data 144 and the known data 142. Figure 3In the example of FIG. 3, the model updater 152 generates an updated model 306. The updated model 306 includes the SEC model 310 to be updated, a delta model 308, and one or more adapter networks 312. The delta model 308 is a copy of the SEC model 310 that has a different output layer than the SEC model 310. In particular, the output layer of the delta model 308 includes more output nodes than the output layer of the SEC model 310. For example, the output layer of the SEC model 310 includes a first count of nodes (e.g., N nodes, where N is a positive integer corresponding to a number of sound classes that the SEC model 310 is trained to recognize), and the output layer of the delta model 308 includes a second count of nodes (e.g., N+M nodes, where M is a positive integer corresponding to a number of new sound classes that the updated SEC model 324 is trained to recognize that the SEC model 310 is not trained to recognize). The first count of nodes corresponds to a sound class count of a first set of sound classes that the SEC model 310 is trained to recognize (e.g., the first set of sound classes includes N different sound classes that the SEC model 310 is able to recognize), and the second count of nodes corresponds to a sound class count of a second set of sound classes that the updated SEC model 324 is to be trained to recognize (e.g., the second set of sound classes includes N+M different sound classes that the updated SEC model 324 is to be trained to recognize). The second set of sound classes includes the first set of sound classes (e.g., N classes) plus one or more additional sound classes (e.g., M classes). The model parameters (e.g., link weights) of the delta model 308 are initialized to be equal to the model parameters of the SEC model 310.
[0067] The adapter network(s) 312 include a neural adapter and a merge adapter. The neural adapter includes one or more adapter layers configured to receive input from the SEC model 310 and generate an output that can be merged with the output of the delta model 308. For example, the SEC model 310 generates a first output corresponding to a first class count of the first set of sound classes. In particular aspects, the first output includes one data element for each node of the output layer of the SEC model 310 (e.g., N data elements). In contrast, the delta model 308 generates a second output corresponding to a second class count of the second set of sound classes. For example, the second output includes one data element for each node of the output layer of the delta model 308 (e.g., N+M data elements). In this example, the adapter layer(s) of the adapter network(s) 312 receive the output of the SEC model 310 as input and generate an output having the second count (e.g., N+M) of data elements. In particular examples, the adapter layer(s) of the adapter network(s) 312 include two fully connected layers (e.g., including an input layer of N nodes and an output layer of N+M nodes, where each node of the input layer is connected to each node of the output layer).
[0068] The merge adapter of the adapter network(s) 312 is configured to generate the output 314 of the updated model 306 by merging the output of the adapter layer(s) with the output of the delta model 308. For example, the merge adapter combines the output of the adapter layer(s) with the output of the delta model 308 in an element-wise manner to generate a combined output, and applies an activation function (such as a sigmoid function) to the combined output to generate the output 314. The output 314 indicates the sound class to which the audio data sample 304 is assigned by the updated model 306, one or more confidence metric values determined by the updated model 306, or both.
[0069] The model updater 152 uses the output 314 and the labels 150 corresponding to the audio data samples 304 to determine updated link weights 316 of the delta model 308, the adapter network(s) 312, or both. The link weights of the SEC model 310 are not changed during training. The training process is iteratively repeated until a training termination condition is satisfied. During training, each entry of the unknown data 146 can be provided to the updated model 306 (e.g., one entry per iteration). Additionally, in some implementations, other audio data samples (e.g., audio data samples previously used to train the SEC model 310) can also be provided to the updated model 306 to reduce the chance that the delta model 308 forgets the previous training of the SEC model 310.
[0070] The training termination condition can be satisfied when all of the unknown data 146 has been provided to the updated model 306 at least once, after a certain number of training iterations have been performed, when a convergence metric satisfies a convergence threshold, or when some other condition indicating the end of training is satisfied. When the training termination condition is satisfied, the model checker 320 selects the updated SEC model 324 (e.g., a combination of the SEC model 310, the delta model 308, and the adapter network(s) 312) from between the delta model 308 and the updated model 306.
[0071] In a specific aspect, model checker 320 selects the updated SEC model 324 based on the accuracy of the sound category 322 assigned by incremental model 308 and the accuracy of the sound category 322 assigned by SEC model 310. For example, model checker 320 may determine the F1 score of incremental model 308 (based on the sound category 322 assigned by incremental model 308) and the F1 score of SEC model 310 (based on the sound category 322 assigned by SEC model 310). In this example, if the F1 score of incremental model 308 is greater than or equal to the F1 score of SEC model 310, then model checker 320 selects incremental model 308 as the updated SEC model 324. In some implementations, if the F1 score of incremental model 308 is greater than or equal to the F1 score of SEC model 310 (or less than a threshold amount of the F1 score of SEC model 310), then model checker 320 selects incremental model 308 as the updated SEC model 324. If the F1 score of incremental model 308 is less than the F1 score of SEC model 310 (or less than the F1 score of SEC model 310 by more than a threshold amount), then model checker 320 selects updated model 306 as updated SEC model 324. If incremental model 308 is selected as updated SEC model 324, then SEC model 310, adapter networks 312, or both can be discarded.
[0072] In some implementations, the model checker 320 is omitted or integrated with the model updater 152. For example, after training the updated model 306, the updated model 306 may be stored as the updated SEC model 324 (e.g., there is no choice between the updated model 306 and the incremental model 308). As an example, when training the updated model 306, the model updater 152 may determine the accuracy metric of the incremental model 308. In this example, the training termination condition may be based on the accuracy metric of the incremental model 308, such that after training, the incremental model 308 is stored as the updated SEC model 324 (e.g., there is no choice between the updated model 306 and the incremental model 308).
[0073] Using reference Figure 3 The described transfer learning technique, model checker 320 enables Figure 1 The device 100 is able to update the SEC model to identify previously unknown sound categories. Additionally, the described transfer learning technique uses far fewer computer resources (e.g., memory, processing time, and power) than would be used to train the SEC model from scratch to identify previously unknown sound categories.
[0074] In some implementations, refer to Figure 2 The described operation (e.g., generating an updated SEC model 214 based on drift data 144) inFigure 1 The equipment was executed at 100 locations, and referenced Figure 3 The described operations (e.g., generating an updated SEC model 324 based on unknown data 146) are performed on different devices (such as...) Figure 6 The process is performed at a remote computing device 618. For illustration, unknown data 146 and tags 150 can be captured at device 100 and transmitted to a second device with more available computing resources. In this example, the second device generates an updated SEC model 324, and device 100 downloads or receives transmissions or data representing the updated SEC model 324 from the second device. Generating an updated SEC model 324 based on unknown data 146 is a more resource-intensive process (e.g., using more memory, power, and processor time) compared to generating an updated SEC model 214 based on drift data 144. Therefore, the reference is partitioned between different devices. Figure 2 The described operations and references Figure 3 The described operation can save 100% of the device's resources.
[0075] Figure 4 It is an explanation Figure 1 A diagram illustrating a specific example of how the device operates. Figure 4 The explanation explains that the determination of whether an active SEC model (e.g., SEC model 112) corresponds to the audio scene captured by audio data sample 110 is based on an implementation that compares the current audio scene with previous audio scenes.
[0076] exist Figure 4 In this process, audio data captured by microphones 104 is used to generate audio data samples 110. Audio data samples 110 are then used to perform audio classification 402. For example, one or more of the available SEC models 114 are... Figure 1 The SEC engine 120 is used as the active SEC model. In a specific aspect, the active SEC model is selected from the available SEC models 114 based on the audio scene indicated by the scene detector 140 (also referred to as the previous audio scene 408) during the previous sampling period.
[0077] Audio classification 402 uses an active SEC model to generate a result 404 based on the analysis of audio data sample 110. Result 404 may indicate the sound category associated with audio data sample 110, the probability that audio data sample 110 corresponds to a specific sound category, or that the sound category of audio data sample 110 is unknown. If result 404 indicates that audio data sample 110 corresponds to a known sound category, a decision is made at box 406 to generate output 124, which indicates the sound category 122 associated with audio data sample 110. For example, Figure 1 The SEC engine 120 can generate an output of 124.
[0078] If the result 404 indicates that the audio data sample 110 does not correspond to a known sound class, a decision is made at block 406 to generate a trigger 126. The trigger 126 activates a drift detection scheme; in Figure 4 which the drift detection scheme includes causing the scene detector 140 to identify a current audio scene 407 based on data from the sensor(s) 134.
[0079] The current audio scene 407 is compared to a previous audio scene 408 at block 410 to determine whether an audio scene change has occurred since the selected active SEC model was selected. At block 412, a determination is made as to whether the sound class of the audio data sample 110 was not recognized due to drift. For example, if the current audio scene 407 does not correspond to the previous audio scene 408, the determination at block 412 is that drift was not the reason the sound class of the audio data sample 110 was not recognized. In this case, the audio data sample 110 can be discarded, or stored as unknown data at block 414.
[0080] If the current audio scene 407 corresponds to the previous audio scene 408, the determination at block 412 is that the sound class of the audio data sample 110 was not recognized due to drift, because the active SEC model corresponds to the current audio scene 407. In this case, the sound class that has drifted is identified at block 416, and the audio data sample 110 and the identifier of the sound class are stored as drift data at block 418.
[0081] When sufficient drift data has been stored, the SEC model is updated at block 420 to generate an updated SEC model 154. The updated SEC model 154 is added to the available SEC models 114. In some implementations, the updated SEC model 154 replaces the active SEC model that generated the result 404.
[0082] Figure 5 is a diagram illustrating another particular example of operation of a device. Figure 1 is a diagram illustrating another particular example of operation of a device. Figure 5 Implementations are illustrated in which the determination as to whether the active SEC model (e.g., SEC model 112) corresponds to the audio scene in which the audio data sample 110 was captured is based on comparing a current audio scene to information describing the active SEC model.
[0083] In Figure 5 implementations, the audio data captured by the microphone(s) 104 is used to generate the audio data sample 110. The audio data sample 110 is used to perform the audio classification 402. For example, one or more of the available SEC models 114 are Figure 1The SEC engine 120 uses as an active SEC model. In particular aspects, the active SEC model is selected from the available SEC models 114. In some implementations, a set of the available SEC models 114 is used, rather than selecting one or more of the available SEC models 114 as the active SEC model.
[0084] The audio classification 402 generates results 404 based on analysis of the audio data sample 110 using one or more of the available SEC models 114. The results 404 can indicate a sound class associated with the audio data sample 110, a probability that the audio data sample 110 corresponds to a particular sound class, or that the sound class of the audio data sample 110 is unknown. If the results 404 indicate that the audio data sample 110 corresponds to a known sound class, a decision is made at block 406 to generate an output 124 that indicates the sound class 122 associated with the audio data sample 110. For example, Figure 1 The SEC engine 120 can generate the output 124.
[0085] If the results 404 indicate that the audio data sample 110 does not correspond to a known sound class, a decision is made at block 406 to generate a trigger 126. The trigger 126 activates a drift detection scheme; in Figure 5 The drift detection scheme includes causing the scene detector 140 to identify a current audio scene based on data from the sensor(s) 134 and determine whether the current audio scene corresponds to the SEC model that generated the results 404 that caused the trigger 126 to be sent.
[0086] At block 412, a determination is made as to whether the sound class of the audio data sample 110 was not identified due to drift. For example, if the current audio scene does not correspond to the SEC model that generated the results 404, the determination at block 412 is that drift was not the reason that the sound class of the audio data sample 110 was not identified. In this case, the audio data sample 110 can be discarded, or stored as unknown data at block 414.
[0087] If the current audio scene corresponds to the SEC model that generated the results 404, the determination at block 412 is that the sound class of the audio data sample 110 was not identified due to drift. In this case, the sound class that has drifted is identified at block 416, and the audio data sample 110 and an identifier for the sound class are stored as drift data at block 418.
[0088] When sufficient drift data has been stored, the SEC model is updated at block 420 to generate an updated SEC model 154. The updated SEC model 154 is added to the available SEC models 114. In some implementations, the updated SEC model 154 replaces the active SEC model that generated the results 404.
[0089] Figure 6 It is an explanation Figure 1 A block diagram of a specific example of device 100. Figure 6 In this configuration, device 100 is set to respond to audio data samples (e.g., Figure 1 The audio data sample 110 is used as input to generate sound identification data (e.g., Figure 1 The output is 124). Additionally, Figure 6 Device 100 is configured to update one or more SEC models 114 based on model update data 148. For example, device 100 is configured as described in reference... Figure 2 The described method uses drift data 144 to update the SEC model 114, configured as referenced. Figure 3 The description uses unknown data 146 to update the SEC model 114, or both. In some implementations, the remote computing device 618 updates the SEC model 114 in some cases. For illustration, device 100 may use drift data 144 to update the SEC model 114, and the remote computing device 618 may use unknown data 146 to update the SEC model 114. In various implementations, device 100 may have more than Figure 6 More or fewer components as explained in the text.
[0090] In a particular implementation, device 100 includes a processor 604 (e.g., a central processing unit (CPU)). Device 100 may include one or more additional processors 606 (e.g., one or more digital signal processors (DSPs)). Processor 604, processor(s), or both may be configured to generate voice identification data, update SEC model 114, or both. For example, in Figure 6 In this processor 606, there is a SEC engine 120. The SEC engine 120 is configured to analyze audio data samples using one or more SEC models 114.
[0091] exist Figure 6 In addition, device 100 also includes memory 608 and CODEC 624. Memory 608 stores instructions 660 executable by processor 604 or processor(s) 606 to implement reference Figures 1-5 One or more operations are described. In one example, instruction 660 includes or corresponds to feature extractor 108, SEC engine 120, scene detector 140, drift detector 128, model updater 152, model checker 320, or a combination thereof. Memory 608 may also store setup data 130, SEC model 114, and model update data 148.
[0092] exist Figure 6In particular aspects, the speakers 622 and the microphone(s) 104 can be coupled to the CODEC 624. In particular aspects, the microphone(s) 104 are configured to receive audio representative of an acoustic environment associated with the device 100 and generate audio signals that a feature extractor uses to generate audio data samples. In Figure 6 In the example illustrated in FIG. 6, the CODEC 624 includes a digital-to-analog converter (DAC 626) and an analog-to-digital converter (ADC 628). In particular implementations, the CODEC 624 receives analog signals from the microphone(s) 104, converts the analog signals to digital signals using the ADC 628, and provides the digital signals to the processor(s) 606. In particular implementations, the processor(s) 606 provide digital signals to the CODEC 624, and the CODEC 624 converts the digital signals to analog signals using the DAC 626 and provides the analog signals to the speaker(s) 622.
[0093] In Figure 6 The device 100 also includes an input device 136 in particular aspects. The device 100 can also include a display 620 coupled to the display controller 610. In particular aspects, the input device 136 includes a sensor, a keyboard, a pointing device, and the like. In some implementations, the input device 136 and the display 620 are combined in a touch screen or similar touch-sensitive or motion-sensitive display. The input device 136 can be used to provide labels associated with unknown data 146 to generate training data 302. The input device 136 can also be used to initiate a model update operation, such as starting the model update process described with reference to Figure 2 the model update process described with reference to Figure 3 In some implementations, the input device 136 can additionally or alternatively be used to select a particular SEC model of the available SEC models 114 to be used by the SEC engine 120. In particular aspects, the input device 136 can be used to configure the settings data 130 (which can be used to select a SEC model to be used by the SEC engine 120), determine the audio scene 142, or both. The display 620 can be used to display analysis results (e.g., the output 124) of one of the SEC models, display a prompt to a user to provide labels associated with unknown data 146, or both. Figure 1
[0094] In some implementations, the device 100 also includes a modem 612 coupled to the transceiver 614. In Figure 6 In particular aspects, the transceiver 614 is coupled to an antenna 616 to enable wireless communication with other devices, such as the remote computing device 618. In other examples, the transceiver 614 is additionally or alternatively coupled to a communication port (e.g., an Ethernet port) to enable wired communication with other devices, such as the remote computing device 618.
[0095] In Figure 6 In particular aspects, the device 100 includes a clock 132 and a sensor 134. As specific examples, the sensor 134 includes one or more cameras 650, one or more positioning sensors 652, the microphone(s) 104, other sensor(s) 654, or a combination thereof.
[0096] In particular aspects, the clock 132 generates a clock signal that can be used to assign a timestamp to a particular audio data sample to indicate when the particular audio data sample was received. In this aspect, the SEC engine 120 can use the timestamp to select a particular SEC model 114 to be used to analyze the particular audio data sample. Additionally or alternatively, the timestamp can be used by the scene detector 140 to determine an audio scene 142 associated with the particular audio data sample.
[0097] In particular aspects, the camera(s) 650 generate image data, video data, or both. The SEC engine 120 can use the image data, video data, or both to select a particular SEC model 114 to be used to analyze an audio data sample. Additionally or alternatively, the image data, video data, or both can be used by the scene detector 140 to determine an audio scene 142 associated with the particular audio data sample. For example, a particular SEC model 114 can be designated for outdoor use, and the image data, video data, or both can be used to confirm whether the device 100 is located in an outdoor environment.
[0098] In particular aspects, the positioning sensor(s) 652 generate positioning data, such as global positioning data that indicates a location of the device 100. The SEC engine 120 can use the positioning data to select a particular SEC model 114 to be used to analyze an audio data sample. Additionally or alternatively, the positioning data can be used by the scene detector 140 to determine an audio scene 142 associated with the particular audio data sample. For example, a particular SEC model 114 can be designated for use at home, and the positioning data can be used to confirm whether the device 100 is located at a home location. The positioning sensor(s) 652 can include a receiver for a satellite-based positioning system, a receiver for a local positioning system, an inertial navigation system, a receiver for a landmark-based positioning system, or a combination thereof.
[0099] The other sensor(s) 654 can include, for example, an orientation sensor, a magnetometer, a light sensor, a contact sensor, a temperature sensor, or any other sensor coupled to or included within the device 100 and that can be used to generate scene data 138 useful in determining an audio scene 142 associated with the device 100 at a particular time.
[0100] In a particular implementation, device 100 is included in a system-in-package or system-on-a-chip device 602. In a particular implementation, memory 608, processor 604, processor(s) 606, display controller 610, CODEC 624, modem 612, and transceiver 614 are included in a system-in-package or system-on-a-chip device 602. In a particular implementation, input device 136 and power supply 630 are coupled to system-on-a-chip device 602. Furthermore, in a particular implementation, such as Figure 6 As explained, the display 620, input device 136, speakers(s) 622, sensor 134, clock 132, antenna 616, and power supply 630 are external to the system-on-chip device 602. In a particular implementation, each of the display 620, input device 136, speakers(s) 622, sensor 134, clock 132, antenna 616, and power supply 630 may be coupled to components of the system-on-chip device 602 (such as interfaces or controllers).
[0101] Device 100 may include, correspond to, or be included in the following: voice-activated devices, audio devices, wireless speakers and voice-activated devices, portable electronic devices, automobiles, vehicles, computing devices, communication devices, Internet of Things (IoT) devices, virtual reality (VR) devices, augmented reality (AR) devices, mixed reality (MR) devices, smart speakers, mobile computing devices, mobile communication devices, smartphones, cellular phones, laptops, computers, tablets, personal digital assistants, display devices, televisions, game consoles, electrical appliances, music players, radios, digital video players, digital video disc (DVD) players, tuners, cameras, navigation devices, or any combination thereof. In a particular aspect, processor 604, (various) memories 606, or combinations thereof are included in an integrated circuit.
[0102] Figure 7 Is it included Figure 1 The illustrative example of various aspects of the vehicle 700 in the device 100. According to one implementation, the vehicle 700 is an autonomous vehicle. According to other implementations, the vehicle 700 is a car, truck, motorcycle, aircraft, watercraft, etc. Figure 7 In this vehicle 700, there is a display 620, one or more of sensors 134, device 100, or a combination thereof. Sensors 134 and device 100 are shown using dashed lines to indicate that these components may not be visible to passengers of the vehicle 700. Device 100 may be integrated into or coupled to the vehicle 700.
[0103] In particular aspects, the device 100 is coupled to the display 620 and provides output to the display 620 in response to one of the SEC models 114 detecting or recognizing various events (e.g., sound events) described herein. For example, the device 100 provides the output 124 of the Figure 1 sound class of the sound 102 (such as a car horn) recognized by one of the SEC models 114 in the audio data samples 110 received from the microphone(s) 104 to the display 620. In some implementations, the device 100 can perform an action in response to recognizing a sound event, such as alerting an operator of the vehicle or activating one of the sensors 134. In a particular example, the device 100 provides output indicating whether an action is being performed in response to a recognized sound event. In particular aspects, a user can select an option displayed on the display 620 to enable or disable performance of an action in response to a recognized sound event.
[0104] In particular implementations, the sensors 134 include Figure 1 the microphone(s) 104, a vehicle occupancy sensor, an eye tracking sensor, the positioning sensor(s) 652, or an external environment sensor (e.g., a lidar sensor or a camera). In particular aspects, the sensor input of the sensors 134 indicates a location of a user. For example, the sensors 134 are associated with various locations within the vehicle 700.
[0105] Figure 7 The device 100 in the vehicle 700 includes the SEC models 114, the SEC engine 120, the drift detector 128, the scene detector 140, and the model updater 152. In other implementations, the device 100 omits the model updater 152 when installed in or used in the vehicle 700. To illustrate, Figure 6 The model update data 148 of the device 100 can be sent to the remote computing device 618, and the remote computing device 618 can update one of the SEC models 114 based on the model update data 148. In such implementations, the updated SEC model 114 can be downloaded to the vehicle 700 for use by the SEC engine 120. In some implementations, the device 100 further includes the model checker 320 of the Figure 3 when installed in or used in the vehicle 700.
[0106] Thus, the techniques described with reference to Figures 1-5 enable a user of the vehicle 700 to generate an updated SEC model to account for drifts that are specific to (or perhaps unique to) the acoustic environment in which the device 100 operates. In some implementations, the device 100 can generate the updated SEC model without user intervention. Moreover, the techniques described with reference to Figures 1-5The described technology enables users of vehicle 700 to generate updated SEC models to detect one or more new sound categories. Furthermore, the SEC model can be updated without excessively utilizing the computing resources loaded on vehicle 700. For example, vehicle 700 does not need to store all training data used to train the SEC model from scratch in local memory.
[0107] Figure 8 Examples of devices 100 coupled to or integrated within head-mounted devices 802 (such as virtual reality head-mounted devices, augmented reality head-mounted devices, mixed reality head-mounted devices, extended reality head-mounted devices, head-mounted displays, or combinations thereof) are depicted. A visual interface device (such as display 620) is placed in front of a user's eyes to enable the display of augmented reality, mixed reality, or virtual reality images or scenes to the user when the head-mounted device 802 is worn. In a particular example, display 620 is configured to display the output of device 100, such as indications of recognized sound events (e.g., Figure 1 The output 124). The head-mounted device 802 includes sensors 134, such as Figure 1 104 microphones Figure 6 (various) cameras 650, Figure 6 Positioning sensors 652, Figure 6 Other sensors 654, or combinations thereof. Although interpreted in a single location, in other implementations, sensor 134 may be positioned at other locations on head-mounted device 802 (such as an array of one or more microphones and one or more cameras distributed around head-mounted device 802) to detect multimodal input.
[0108] Sensor 134 detects audio data, which device 100 uses to detect sound events or update the SEC model 114. For example, SEC engine 120 uses one or more SEC models 114 to generate sound event classification data, which can be provided to display 620 to indicate that an identified sound event, such as a car horn, has been detected in a sample of audio data received from sensor 134. In some implementations, device 100 may perform actions in response to the recognition of a sound event, such as activating a camera or another sensor 134 or providing haptic feedback to the user.
[0109] exist Figure 8 In the example described, device 100 includes a SEC model 114, a SEC engine 120, a drift detector 128, a scene detector 140, and a model updater 152. In other implementations, the model updater 152 is omitted when device 100 is installed in or used in head-mounted device 802. For illustrative purposes, Figure 6The model update data 148 can be sent to a remote computing device 618, and the remote computing device 618 can update one of the SEC models 114 based on the model update data 148. In such implementations, the updated SEC model 114 can be downloaded to the head-mounted device 802 for use by the SEC engine 120. In some implementations, the device 100, when installed in or used in the head-mounted device 802, further includes Figure 3 the model checker 320.
[0110] Figure 9 An example of the device 100 integrated into a wearable electronic device 902, illustrated as a "smart watch," which includes a display 620 and sensors 134, is depicted. The sensors 134 enable detection of, for example, user input and audio scenes based on modalities such as location, video, speech, and gesture. The sensors 134 also enable detection of sounds in the acoustic environment surrounding the wearable electronic device 902, which the device 100 uses to detect sound events or update the SEC models 114. For example, the device 100 provides the output 124 to the display 620 that indicates that a recognized sound event was detected in audio data samples received from the sensors 134. In some implementations, the device 100 can perform an action in response to recognizing a sound event, such as activating a camera or another sensor 134 or providing haptic feedback to the user. Figure 1
[0111] In the example illustrated in Figure 9 , the device 100 includes the SEC models 114, the SEC engine 120, the drift detector 128, the scene detector 140, and the model updater 152. In other implementations, the device 100, when installed in or used in the wearable electronic device 902, omits the model updater 152. To illustrate, Figure 6 the model update data 148 can be sent to a remote computing device 618, and the remote computing device 618 can update one of the SEC models 114 based on the model update data 148. In such implementations, the updated SEC model 114 can be downloaded to the wearable electronic device 902 for use by the SEC engine 120. In some implementations, the device 100, when installed in or used in the wearable electronic device 902, further includes Figure 3 the model checker 320.
[0112] Figure 10 is an illustrative example of a voice-controlled speaker system 1000. The voice-controlled speaker system 1000 can have wireless network connectivity and be configured to perform an auxiliary operation. In Figure 10 In this system, device 100 is included in a voice-controlled speaker system 1000. The voice-controlled speaker system 1000 also includes a speaker 1002 and a sensor 134. Sensor 134 includes... Figure 1 The microphones 104 are used to receive voice input or other audio input.
[0113] During operation, in response to a received verbal command or a recognized sound event, the voice-controlled speaker system 1000 can perform auxiliary operations. Auxiliary operations may include adjusting the temperature, playing music, turning on lights, etc. Sensor 134 enables the detection of audio data samples, which device 100 uses to detect sound events or update one or more SEC models 114. Additionally, the voice-controlled speaker system 1000 can perform operations based on sound events recognized by device 100. For example, if device 100 recognizes the sound of a door closing, the voice-controlled speaker system 1000 can turn on one or more lights.
[0114] exist Figure 10 In the example described, device 100 includes SEC model 114, SEC engine 120, drift detector 128, scene detector 140, and model updater 152. In other implementations, the model updater 152 is omitted when device 100 is installed in or used in voice-controlled speaker system 1000. For explanation purposes, Figure 6 The model update data 148 can be sent to a remote computing device 618, and the remote computing device 618 can update one of the SEC models 114 based on the model update data 148. In such an implementation, the updated SEC model 114 can be downloaded to the voice-controlled speaker system 1000 for use by the SEC engine 120. In some implementations, the device 100, when installed in or used in the voice-controlled speaker system 1000, further includes Figure 3 Model checker 320.
[0115] Figure 11 The explanation included Figure 1 The equipment includes 100 cameras of various aspects, 1100. Figure 11 In this embodiment, device 100 is incorporated into or coupled to camera 1100. Camera 1100 includes image sensor 1102 and one or more other sensors (e.g., sensor 134), such as Figure 1Microphone 104. Additionally, camera 1100 includes device 100 configured to: identify sound events based on audio data samples, and update one or more of the SEC models 114. In a particular aspect, camera 1100 is configured to perform one or more actions in response to an identified sound event. For example, camera 1100 may cause image sensor 1102 to capture an image in response to device 100 detecting a specific sound event in audio data samples from sensor 134.
[0116] exist Figure 11 In the example described, device 100 includes a SEC model 114, a SEC engine 120, a drift detector 128, a scene detector 140, and a model updater 152. In other implementations, the model updater 152 is omitted when device 100 is mounted in or used in camera 1100. For illustrative purposes, Figure 6 The model update data 148 can be sent to a remote computing device 618, and the remote computing device 618 can update one of the SEC models 114 based on the model update data 148. In such an implementation, the updated SEC model 114 can be downloaded to the camera 1100 for use by the SEC engine 120. In some implementations, the device 100, when installed in or used in the camera 1100, further includes... Figure 3 Model checker 320.
[0117] Figure 12 The explanation included Figure 1 The device comprises 100 mobile devices of various aspects, totaling 1200. Figure 12 In the middle, the mobile device 1200 includes or is coupled to Figure 1 Device 100. By way of illustrative and not limiting example, mobile device 1200 includes a telephone or tablet device. Mobile device 1200 includes a display 620 and sensors 134, such as… Figure 1 104 microphones Figure 6 (various) cameras 650, Figure 6 Positioning sensors 652, or Figure 6 Other sensors 654. During operation, mobile device 1200 may perform specific actions in response to device 100 recognizing a specific sound event. For example, these actions may include sending commands to other devices, such as thermostats, home automation systems, another mobile device, etc.
[0118] exist Figure 12In the example described, device 100 includes a SEC model 114, a SEC engine 120, a drift detector 128, a scene detector 140, and a model updater 152. In other implementations, the model updater 152 is omitted when device 100 is installed in or used in mobile device 1200. For illustrative purposes, Figure 6 The model update data 148 can be sent to a remote computing device 618, and the remote computing device 618 can update one of the SEC models 114 based on the model update data 148. In such an implementation, the updated SEC model 114 can be downloaded to a mobile device 1200 for use by the SEC engine 120. In some implementations, the device 100, when installed or used in the mobile device 1200, further includes... Figure 3 Model checker 320.
[0119] Figure 13 The explanation included Figure 1 The equipment includes 100 types of aerial equipment and 1300 types of various types. Figure 13 In the middle, aerial equipment 1300 includes or is coupled to Figure 1 Device 100. Aerial device 1300 is a manned, unmanned, or remotely controlled aerial device (e.g., a package delivery drone). Aerial device 1300 includes a control system 1302 and sensors 134, such as Figure 1 104 microphones Figure 6 (various) cameras 650, Figure 6 Positioning sensors 652, or Figure 6 Other sensors 654. Control system 1302 controls various operations of the airborne equipment 1300, such as cargo release, sensor activation, takeoff, navigation, landing, or combinations thereof. For example, control system 1302 can control the flight of airborne equipment 1300 between designated points and the deployment of cargo at a specific location. In a particular aspect, control system 1302 performs one or more actions in response to the detection of a specific acoustic event by equipment 100. For illustration, control system 1302 can initiate a safe landing protocol in response to equipment 100 detecting an aircraft engine.
[0120] exist Figure 13 In the example described, device 100 includes SEC model 114, SEC engine 120, drift detector 128, scene detector 140, and model updater 152. In other implementations, the model updater 152 is omitted when device 100 is installed in or used in air device 1300. For explanation purposes, Figure 6Model update data 148 can be sent to remote computing device 618, and remote computing device 618 can update one of the SEC models 114 based on model update data 148. In such implementations, the updated SEC model 114 can be downloaded to air device 1300 for use by SEC engine 120. In some implementations, device 100, when installed or used in air device 1300, further includes Figure 3 Model checker 320.
[0121] Figure 14 The explanation included Figure 1 The device comprises 100 head-mounted devices in various aspects, totaling 1400. Figure 14 In the middle, the head-mounted device 1400 includes or is coupled to Figure 1 Device 100. Head-mounted device 1400 includes... Figure 1 The headset 1400 may also include: one or more additional microphones positioned to primarily capture ambient sound (e.g., for noise cancellation); and one or more of the sensors 134 (such as...). Figure 6 (The cameras 650, positioning sensors 652, or other sensors 654). In a particular aspect, the head-mounted device 1400 performs one or more actions in response to a specific sound event detected by the device 100. For example, the head-mounted device 1400 may activate a noise cancellation feature in response to the device 100 detecting a gunshot. The head-mounted device 1400 may also update one or more of the SEC models 114.
[0122] exist Figure 14 In the example described, device 100 includes a SEC model 114, a SEC engine 120, a drift detector 128, a scene detector 140, and a model updater 152. In other implementations, the model updater 152 is omitted when device 100 is installed in or used in head-mounted device 1400. For illustrative purposes, Figure 6 The model update data 148 can be sent to a remote computing device 618, and the remote computing device 618 can update one of the SEC models 114 based on the model update data 148. In such an implementation, the updated SEC model 114 can be downloaded to the head-mounted device 1400 for use by the SEC engine 120. In some implementations, the device 100, when installed or used in the head-mounted device 1400, further includes Figure 3 Model checker 320.
[0123] Figure 15 The explanation included Figure 1 The equipment includes 100 electrical appliances of various types. Figure 15In this implementation, appliance 1500 is a table lamp; however, in other implementations, appliance 1500 includes another IoT appliance, such as a refrigerator, coffee maker, oven, or another household appliance. Appliance 1500 includes or is coupled to... Figure 1 Device 100. Electrical appliance 1500 includes sensor 134, such as Figure 1 104 microphones Figure 6 (various) cameras 650, Figure 6 Positioning sensors 652, or Figure 6 Other sensors 654. In a particular aspect, appliance 1500 performs one or more actions in response to a specific sound event detected by device 100. For example, appliance 1500 may activate a light in response to device 100 detecting a door being closed. Appliance 1500 may also update one or more of the SEC models 114.
[0124] exist Figure 15 In the example described, device 100 includes a SEC model 114, a SEC engine 120, a drift detector 128, a scene detector 140, and a model updater 152. In other implementations, the model updater 152 is omitted when device 100 is installed in or used in appliance 1500. For illustrative purposes, Figure 6 The model update data 148 can be sent to a remote computing device 618, and the remote computing device 618 can update one of the SEC models 114 based on the model update data 148. In such an implementation, the updated SEC model 114 can be downloaded to the appliance 1500 for use by the SEC engine 120. In some implementations, the device 100, when installed or used in the appliance 1500, further includes... Figure 3 Model checker 320.
[0125] Figure 16 It is an explanation Figure 1 A flowchart illustrating an example of an operation method 1600 for device 100 is provided. Method 1600 can be initiated, controlled, or executed by device 100. For example, Figure 6 The processors 604 or 606 can execute instructions 606 from memory 608 to cause the drift detector 128 to generate model update data 148.
[0126] In box 1602, method 1600 includes providing audio data samples as input to a sound event classification model. For example, Figure 1 SEC Engine 120 (or Figure 6 The processor 606(s) executing instructions 660 corresponding to the SEC engine 120 can provide audio data samples 110 as input to the SEC model 112. In some implementations, method 1600 also includes capturing audio data corresponding to the audio data samples. For example,Figure 1 Microphone(s) 104 can generate audio signals 106 based on sound 102 detected by microphone(s) 104. Further, in some implementations, method 1600 includes selecting a sound event classification model from among a plurality of sound event classification models stored at memory. For example, Figure 1 SEC engine 120 (or Figure 6 Processor(s) 606 executing instructions 660 corresponding to SEC engine 120 can select SEC model 112 from among available SEC models 114 based on sensor data associated with the audio data sample, input identifying the audio scene or SEC model 112, when the audio data sample was received, setting data, or a combination thereof.
[0127] In block 1604, method 1600 includes determining whether a sound class of the audio data sample is recognized by the sound event classification model based on an output of the sound event classification model in response to the audio data sample. For example, Figure 1 SEC engine 120 (or Figure 6 Processor(s) 606 executing instructions 660 corresponding to SEC engine 120 can determine whether a sound class 122 of audio data sample 110 is recognized by SEC model 112. To illustrate, SEC model 112 can generate a confidence metric associated with each sound class that SEC model 112 is trained to recognize, and the determination as to whether a sound class is recognized by the SEC model can be based on the value of the confidence metric(s). In particular aspects, based on a determination that sound class 122 is recognized by SEC model 112, SEC engine 120 generates output 124 indicating sound class 122 associated with audio data sample 110.
[0128] In block 1606, method 1600 includes determining whether the sound event classification model corresponds to an audio scene associated with the audio data sample based on a determination that the sound class is not recognized. For example, Figure 1 Drift detector 128 (or Figure 6 Processor(s) 606 executing instructions 660 corresponding to SEC engine 128 can determine whether SEC model 112 corresponds to an audio scene 142 associated with audio data sample 110. To illustrate, scene detector 140 can determine audio scene 142 based on input received via input device 136, sensor data from sensor(s) 134, a timestamp associated with audio data sample 110 indicated by clock 132, setting data 130, or a combination thereof. In implementations in which SEC model 112 is selected from among available SEC models, the determination of audio scene 142 can be based on different information than that used to select SEC model 112.
[0129] In box 1608, method 1600 includes: storing model update data based on audio data samples, based on determining that the sound event classification model corresponds to an audio scene associated with the audio data samples. For example, Figure 1 Drift detector 128 (or Figure 6 The processor 606(s) executing instructions 660 corresponding to drift detector 128 may store model update data 148 based on audio data sample 110. In a particular aspect, if drift detector 128 determines that SEC model 112 corresponds to audio scene 142 associated with audio data sample 110, drift detector 128 stores drift data 144 as model update data 148, while if drift detector 128 determines that SEC model 112 does not correspond to audio scene 142 associated with audio data sample 110, drift detector 128 stores unknown data 146 as model update data 148.
[0130] Method 1600 may also include updating the SEC model based on model update data. For example, Figure 1 Model updater 152 (or Figure 6 The processor 606(s) that executes the instructions 660 corresponding to the model updater 152 can be referred to as follows Figure 2 The place described, as referenced Figure 3 The SEC model 112 is updated based on model update data 148, as described or as described in both.
[0131] In a particular aspect, method 1600 includes: after storing model update data, determining whether a threshold number of model update data has been accumulated. For example, model updater 152 may determine... Figure 1 The method 1600 may include whether the model update data 148 includes sufficient data (e.g., a threshold number of model update data 148 associated with a particular SEC model or a particular sound category) to initiate an update training of the SEC model 112. The method 1600 may also include: based on determining that a threshold number of model update data has been accumulated, using the accumulated model update data to initiate an automatic update of the sound event classification model. For example, the model updater 152 may initiate a model update without input from a user. In a particular implementation, the automatic update fine-tunes the SEC model 112 to generate an updated SEC model 154. For example, prior to the automatic update, the SEC model 112 is trained to recognize multiple variants of a particular sound category, and the automatic update modifies the SEC model 112 so that the SEC model 112 can recognize additional variants of that particular sound category as corresponding to that particular sound category.
[0132] Due to the manner in which training occurs, SEC models are often closed sets. That is, the number and type of sound classes that a SEC model can recognize is fixed and limited during training. After training, a SEC model typically has a static relationship between inputs and outputs. This static relationship between inputs and outputs means that the mapping learned during training is valid in the future (e.g., when evaluating new data), and the relationship between inputs and outputs does not change. However, it is difficult to collect a complete set of training samples for each sound class, and it is difficult to correctly annotate all available training data to train a comprehensive and complex SEC model.
[0133] In contrast, during use, a SEC model can face open set problems. For example, during use, a SEC model can be provided with data samples associated with both known and unknown sound events. Additionally, the distribution of sounds or sound features in each sound class that a SEC model is trained to recognize can change over time, or can not be fully represented in the available training data. For example, for traffic sounds, sound differences based on location, time, busy or non-busy intersections, etc. can not be explicitly captured in the training data for the traffic sound class. Due to these and other reasons, there can be a difference between the training data used to train a SEC model and the set of data that the SEC model is provided during use. Such differences (e.g., dataset shift or drift) depend on various factors, such as location, time, device that is capturing the sound signal, etc. Dataset shift can result in poor prediction results from a SEC model. By adapting a SEC model to detect such shifted data with little or no supervision, the disclosed systems and methods overcome these and other problems. Additionally, in some aspects, a SEC model can be updated to recognize new sound classes without forgetting previously trained sound classes.
[0134] In particular aspects, training of the SEC model is not performed when the system is operating in inference mode. Rather, during operation in inference mode, existing knowledge in the form of one or more previously trained SEC models is used to analyze detected sounds. More than one SEC model can be used to analyze sounds. For example, a set of SEC models can be used during operation in inference mode. A particular SEC can be selected from the set of available SEC models based on detection of a trigger condition. To illustrate, whenever a certain trigger (or certain triggers) is activated, a particular SEC model will be used as the active SEC model, which can also be referred to as the "source SEC model." The trigger(s) can be based on location, sound, camera information, other sensor data, user input, and the like. For example, a particular SEC model can be trained to recognize sound events related to crowded areas, such as theme parks, outdoor shopping centers, public squares, and the like. In this example, the particular SEC model can be used as the active SEC model when global positioning data indicates that the device capturing the sound is at any of these locations. In this example, the trigger is based on the location of the device capturing the sound, and the active SEC model is selected and loaded when the device is detected to be at that location (e.g., in addition to or in place of a previous active SEC model).
[0135] In conjunction with the described implementations, an apparatus includes means for providing an audio data sample to a sound event classification model. For example, the means for providing an audio data sample to a sound event classification model includes the device 100, the instructions 660, the processor 604, the processor(s) 606, the SEC engine 120, the feature extractor 108, the microphone(s) 104, the CODEC 624, one or more other circuits or components configured to provide an audio data sample to a sound event classification model, or any combination thereof.
[0136] The apparatus also includes means for determining whether a sound class of the audio data sample is recognized by the sound event classification model based on an output of the sound classification model. For example, the means for determining whether a sound class of the audio data sample is recognized by the sound event classification model includes the device 100, the instructions 660, the processor 604, the processor(s) 606, the SEC engine 120, one or more other circuits or components configured to determine whether a sound class of the audio data sample is recognized by the sound event classification model, or any combination thereof.
[0137] The device also includes means for determining, in response to determining that the sound class is not identified, whether the sound event classification model corresponds to an audio scene associated with the audio data sample. For example, the means for determining whether the sound event classification model corresponds to an audio scene associated with the audio data sample includes the device 100, the instructions 660, the processor(s) 604, the processor(s) 606, the drift detector 128, the scene detector 140, one or more other circuits or components configured to determine whether the sound event classification model corresponds to an audio scene associated with the audio data sample, or any combination thereof.
[0138] The device also includes means for storing model update data based on the audio data sample, in response to determining that the sound event classification model corresponds to an audio scene associated with the audio data sample. For example, the means for storing model update data includes the remote computing device 618, the device 100, the instructions 660, the processor(s) 604, the processor(s) 606, the drift detector 128, the memory 608, one or more other circuits or components configured to store model update data, or any combination thereof.
[0139] In some implementations, the device includes means for selecting the sound event classification model from among a plurality of sound event classification models based on selection criteria. For example, the means for selecting the sound event classification model includes the device 100, the instructions 660, the processor(s) 604, the processor(s) 606, the SEC engine 120, one or more circuits or components configured to select the sound event classification model, or any combination thereof.
[0140] In some implementations, the device includes means for updating the sound event classification model based on the model update data. For example, the means for updating the sound event classification model based on the model update data includes the remote computing device 618, the device 100, the instructions 660, the processor(s) 604, the processor(s) 606, the model updater 152, the model checker 320, one or more other circuits or components configured to update the sound event classification model, or any combination thereof.
[0141] Those of skill would further appreciate that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or combinations of both. Various illustrative components, blocks, configurations, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or processor executable instructions depends upon the particular application and design constraints imposed on the overall system. Skilled persons can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0142] The steps of a method or algorithm described in connection with the implementations disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in Random Access Memory (RAM), flash memory, Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and
[0143] Certain aspects of the present disclosure are described in the following first set of interrelated clauses:
[0144] According to Clause 1, an apparatus includes one or more processors configured to provide audio data samples to a sound event classification model. The one or more processors are further configured to determine whether a sound class of the audio data samples is recognized by the sound event classification model based on an output of the sound event classification model responsive to the audio data samples. The one or more processors are also configured to determine whether the sound event classification model corresponds to an audio scene associated with the audio data samples based on a determination that the sound class is not recognized. The one or more processors are further configured to store model update data based on the audio data samples based on a determination that the sound event classification model corresponds to the audio scene associated with the audio data samples.
[0145] Clause 2 includes the apparatus of Clause 1, and further includes a microphone coupled to the one or more processors and configured to capture audio data corresponding to the audio data samples.
[0146] Clause 3 includes the apparatus of Clause 1 or Clause 2, and further includes a memory coupled to the one or more processors and configured to store a plurality of sound event classification models, wherein the one or more processors are configured to select the sound event classification model from among the plurality of sound event classification models.
[0147] Clause 4 includes the device of clause 3, and further includes one or more sensors configured to generate sensor data associated with the audio data samples, wherein the one or more processors are configured to select the sound event classification model based on the sensor data.
[0148] Clause 5 includes the device of clause 4, wherein the one or more sensors include a camera and a positioning sensor.
[0149] Clause 6 includes the device of any of clauses 3-5, and further includes one or more input devices configured to receive input identifying the audio scene, wherein the one or more processors are configured to select the sound event classification model based on the audio scene.
[0150] Clause 7 includes the device of any of clauses 3-6, wherein the one or more processors are configured to select the sound event classification model based on when the audio data samples are received.
[0151] Clause 8 includes the device of any of clauses 3-8, wherein the memory further stores setting data indicative of one or more device settings, and wherein the one or more processors are configured to select the sound event classification model based on the setting data.
[0152] Clause 9 includes the device of any of clauses 1-8, wherein the one or more processors are further configured to generate output indicative of the sound class associated with the audio data samples based on determining that the sound class is identified.
[0153] Clause 10 includes the device of any of clauses 1-9, wherein the one or more processors are further configured to, based on determining that the sound event classification model does not correspond to the audio scene associated with the audio data samples, store audio data corresponding to the audio data samples as training data for a new sound event classification model.
[0154] Clause 11 includes the device of any of clauses 1-10, wherein the sound event classification model is further configured to generate a confidence measure associated with the output, and wherein the one or more processors are configured to determine whether the sound class is identified by the sound event classification model based on the confidence measure.
[0155] Clause 12 includes the device of any of clauses 1-11, wherein the one or more processors are further configured to update the sound event classification model based on the model update data.
[0156] Clause 13 includes the device of any of clauses 1-12, and further includes one or more input devices configured to receive input identifying the audio scene, wherein the one or more processors are configured to determine whether the sound event classification model corresponds to the audio scene based on the input.
[0157] Clause 14 includes the device of any of clauses 1-13, and further includes one or more sensors configured to generate sensor data associated with the audio data samples, wherein the one or more processors are configured to determine whether the sound event classification model corresponds to the audio scene based on the sensor data.
[0158] Clause 15 includes the device of clause 14, wherein the one or more sensors include a camera and a positioning sensor.
[0159] Clause 16 includes the device of clause 14 or clause 15, wherein the one or more processors are further configured to determine whether the sound event classification model corresponds to the audio scene based on timestamps associated with the audio data samples.
[0160] Clause 17 includes the device of any of clauses 1-16, wherein the one or more processors are integrated within a mobile computing device.
[0161] Clause 18 includes the device of any of clauses 1-16, wherein the one or more processors are integrated within a vehicle.
[0162] Clause 19 includes the device of any of clauses 1-16, wherein the one or more processors are integrated within a wearable device.
[0163] Clause 20 includes the device of any of clauses 1-16, wherein the one or more processors are integrated within an augmented reality headset, a mixed reality headset, or a virtual reality headset.
[0164] Clause 21 includes the device of any of clauses 1-20, wherein the one or more processors are included in an integrated circuit.
[0165] Clause 22 includes the device of any of clauses 1-21, wherein the sound event classification model is trained to identify a particular sound class, and the model update data includes drift data representing a change in sound characteristics within the particular sound class for which the sound event classification model is not trained to identify as corresponding to the particular sound class.
[0166] Particular aspects of the present disclosure are described in the following second set of interrelated clauses:
[0167] According to Clause 23, a method includes providing, by one or more processors, audio data samples as input to a sound event classification model. The method also includes determining, by the one or more processors, whether a sound class of the audio data samples is recognized by the sound event classification model based on an output of the sound event classification model responsive to the audio data samples. The method further includes determining, by the one or more processors, whether the sound event classification model corresponds to an audio scene associated with the audio data samples based on a determination that the sound class is not recognized. The method also includes storing, by the one or more processors, model update data based on the audio data samples based on a determination that the sound event classification model corresponds to the audio scene associated with the audio data samples.
[0168] Clause 24 includes the method of Clause 23, and further includes selecting the sound event classification model from among a plurality of sound event classification models stored at a memory coupled to the one or more processors.
[0169] Clause 25 includes the method of Clause 24, wherein the sound event classification model is selected based on user input, setting data, location data, image data, video data, a timestamp associated with the audio data samples, or a combination thereof.
[0170] Clause 26 includes the method of any of Clauses 23 to 25, wherein the determination as to whether the sound event classification model corresponds to the audio scene is based on a confidence metric generated by the sound event classification model, user input, setting data, location data, image data, video data, a timestamp associated with the audio data samples, or a combination thereof.
[0171] Clause 27 includes the method of any of Clauses 23 to 26, and further includes capturing audio data corresponding to the audio data samples.
[0172] Clause 28 includes the method of any of Clauses 23 to 27, and further includes selecting the sound event classification model from among a plurality of available sound event classification models.
[0173] Clause 29 includes the method of any of Clauses 23 to 27, and further includes receiving sensor data associated with the audio data samples, and selecting the sound event classification model from among a plurality of available sound event classification models based on the sensor data.
[0174] Clause 30 includes the method of any of Clauses 23 to 27, and further includes receiving input identifying the audio scene, and selecting the sound event classification model from among a plurality of available sound event classification models based on the audio scene.
[0175] Clause 31 includes the method of any of clauses 23-27, and further includes selecting the sound event classification model from among a plurality of available sound event classification models based on when the audio data samples were received.
[0176] Clause 32 includes the method of any of clauses 23-27, and further includes selecting the sound event classification model from among a plurality of available sound event classification models based on the setting data.
[0177] Clause 33 includes the method of any of clauses 23-32, and further includes generating an output indicating the sound class associated with the audio data samples based on determining that the sound class is identified.
[0178] Clause 34 includes the method of any of clauses 23-33, and further includes storing audio data corresponding to the audio data samples as training data for a new sound event classification model based on determining that the sound event classification model does not correspond to the audio scene associated with the audio data samples.
[0179] Clause 35 includes the method of any of clauses 23-34, wherein the output of the sound event classification model includes a confidence measure, and the method further includes determining whether the sound class is identified by the sound event classification model based on the confidence measure.
[0180] Clause 36 includes the method of any of clauses 23-35, and further includes updating the sound event classification model based on the model update data.
[0181] Clause 37 includes the method of any of clauses 23-36, and further includes receiving input identifying the audio scene, wherein the determination as to whether the sound event classification model corresponds to the audio scene is based on the input.
[0182] Clause 38 includes the method of any of clauses 23-37, and further includes receiving sensor data associated with the audio data samples, wherein the determination as to whether the sound event classification model corresponds to the audio scene is based on the sensor data.
[0183] Clause 39 includes the method of any of clauses 23-38, wherein the determination as to whether the sound event classification model corresponds to the audio scene is based on a timestamp associated with the audio data samples.
[0184] Clause 40 includes the method of any of clauses 23 to 39 and further includes, after storing the model update data, determining whether a threshold amount of model update data has accumulated, and based on determining that the threshold amount of model update data has accumulated, initiating an automatic update of the sound event classification model using the accumulated model update data.
[0185] Clause 41 includes the method of any of clauses 23 to 40, wherein, prior to the automatic update, the sound event classification model is trained to recognize a plurality of variants of a particular sound class, and wherein the automatic update modifies the sound event classification model to enable the sound event classification model to recognize an additional variant of the particular sound class as corresponding to the particular sound class.
[0186] Particular aspects of the present disclosure are described in the following third set of interrelated clauses:
[0187] According to clause 42, a device includes means for providing audio data samples to a sound event classification model. The device also includes means for determining, based on an output of the sound classification model, whether a sound class of the audio data samples is recognized by the sound event classification model. The device further includes means for determining, in response to determining that the sound class is not recognized, whether the sound event classification model corresponds to an audio scene associated with the audio data samples. The device also includes means for storing model update data based on the audio data samples in response to determining that the sound event classification model corresponds to the audio scene associated with the audio data samples.
[0188] Clause 43 includes the device of clause 42 and further includes means for selecting the sound event classification model from among a plurality of sound event classification models based on selection criteria.
[0189] Clause 44 includes the device of clause 42 or clause 43 and further includes means for updating the sound event classification model based on the model update data.
[0190] Clause 45 includes the device of any of clauses 42 to 44 and further includes means for capturing audio data corresponding to the audio data samples.
[0191] Clause 46 includes the device of any of clauses 42 to 45 and further includes means for storing a plurality of sound event classification models and means for selecting the sound event classification model from among the plurality of sound event classification models.
[0192] Clause 47 includes the device of clause 46 and further includes means for receiving an input identifying the audio scene, wherein the sound event classification model is selected based on the input identifying the audio scene.
[0193] Clause 48 includes the device of clause 46, and further includes means for determining when the audio data samples were received, wherein the sound event classification model is selected based on when the audio data samples were received.
[0194] Clause 48 includes the device of clause 46, and further includes means for storing setting data indicative of one or more device settings, wherein the sound event classification model is selected based on the setting data.
[0195] Clause 50 includes the device of any of clauses 42-49, and further includes means for generating an output indicative of the sound class associated with the audio data samples based on determining that the sound class is identified.
[0196] Clause 51 includes the device of any of clauses 42-50, and further includes means for storing audio data corresponding to the audio data samples as training data for a new sound event classification model based on determining that the sound event classification model does not correspond to the audio scene associated with the audio data samples.
[0197] Clause 52 includes the device of any of clauses 42-51, wherein the sound event classification model is further configured to generate a confidence measure associated with the output, and wherein the determination as to whether the sound class is identified by the sound event classification model is based on the confidence measure.
[0198] Clause 53 includes the device of any of clauses 42-52, and further includes means for updating the sound event classification model based on the model update data.
[0199] Clause 54 includes the device of any of clauses 42-53, and further includes means for receiving an input identifying the audio scene, wherein the determination as to whether the sound event classification model corresponds to the audio scene is based on the input.
[0200] Clause 55 includes the device of any of clauses 42-54, and further includes means for generating sensor data associated with the audio data samples, wherein the determination as to whether the sound event classification model corresponds to the audio scene is based on the sensor data.
[0201] Clause 56 includes the device of any of clauses 42-55, wherein the determination as to whether the sound event classification model corresponds to the audio scene is based on a timestamp associated with the audio data samples.
[0202] Clause 57 includes the device of any of clauses 42 to 56, wherein the means for providing audio data samples to a sound event classification model, the means for receiving an output of the audio data classification model, the means for determining whether the sound class of the audio data samples is recognized by the sound event classification model, the means for determining whether the sound event classification model corresponds to the audio scene associated with the audio data samples, and the means for storing the model update data based on the audio data samples are integrated within a mobile computing device.
[0203] Clause 58 includes the device of any of clauses 42 to 56, wherein the means for providing audio data samples to a sound event classification model, the means for receiving an output of the audio data classification model, the means for determining whether the sound class of the audio data samples is recognized by the sound event classification model, the means for determining whether the sound event classification model corresponds to the audio scene associated with the audio data samples, and the means for storing the model update data based on the audio data samples are integrated within a vehicle.
[0204] Clause 59 includes the device of any of clauses 42 to 56, wherein the means for providing audio data samples to a sound event classification model, the means for receiving an output of the audio data classification model, the means for determining whether the sound class of the audio data samples is recognized by the sound event classification model, the means for determining whether the sound event classification model corresponds to the audio scene associated with the audio data samples, and the means for storing the model update data based on the audio data samples are integrated within a wearable device.
[0205] Clause 60 includes the device of any of clauses 42 to 56, wherein the means for providing audio data samples to a sound event classification model, the means for receiving an output of the audio data classification model, the means for determining whether the sound class of the audio data samples is recognized by the sound event classification model, the means for determining whether the sound event classification model corresponds to the audio scene associated with the audio data samples, and the means for storing the model update data based on the audio data samples are integrated within an augmented reality headset, a mixed reality headset, or a virtual reality headset.
[0206] Clause 61 includes the device of any of clauses 42 to 60, wherein the means for providing audio data samples to a sound event classification model, the means for receiving an output of the audio data classification model, the means for determining whether the sound class of the audio data samples is recognized by the sound event classification model, the means for determining whether the sound event classification model corresponds to the audio scene associated with the audio data samples, and the means for storing the model update data based on the audio data samples are included in an integrated circuit.
[0207] Certain aspects of the present disclosure are described in the following fourth set of interrelated clauses:
[0208] According to clause 62, a non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to provide audio data samples as input to a sound event classification model. The instructions are further executable by the processor to determine whether a sound class of the audio data samples is recognized by the sound event classification model based on an output of the sound event classification model in response to the audio data samples. The instructions are further executable by the processor to determine whether the sound event classification model corresponds to an audio scene associated with the audio data samples based on a determination that the sound class is not recognized. The instructions are further executable by the processor to store model update data based on the audio data samples based on a determination that the sound event classification model corresponds to the audio scene associated with the audio data samples.
[0209] Clause 63 includes the non-transitory computer-readable medium of clause 62, wherein the instructions further cause the processor to update the sound event classification model based on the model update data.
[0210] Clause 64 includes the non-transitory computer-readable medium of clause 62 or clause 63, wherein the instructions further cause the processor to select the sound event classification model from among a plurality of sound event classification models stored in memory.
[0211] Clause 65 includes the non-transitory computer-readable medium of clause 64, wherein the instructions cause the processor to select the sound event classification model based on sensor data.
[0212] Clause 66 includes the non-transitory computer-readable medium of clause 64, wherein the instructions cause the processor to select the sound event classification model based on input identifying an audio scene associated with the audio data samples.
[0213] Clause 67 includes the non-transitory computer-readable medium of clause 64, wherein the instructions cause the processor to select the sound event classification model based on when the audio data samples were received.
[0214] Clause 68 includes the non-transitory computer-readable medium of clause 64, wherein the instructions cause the processor to select the sound event classification model based on setting data.
[0215] Clause 69 includes the non-transitory computer-readable medium of any of clauses 62-68, wherein the instructions cause the processor to generate, based on a determination that the sound class is identified, an output indicating the sound class associated with the audio data samples.
[0216] Clause 70 includes the non-transitory computer-readable medium of any of clauses 62-69, wherein the instructions cause the processor to, based on a determination that the sound event classification model does not correspond to the audio scene associated with the audio data samples, store audio data corresponding to the audio data samples as training data for a new sound event classification model.
[0217] Clause 71 includes the non-transitory computer-readable medium of any of clauses 62-70, wherein the instructions cause the processor to generate a confidence measure associated with the output, and wherein the determination as to whether the sound class is identified by the sound event classification model is based on the confidence measure.
[0218] Clause 72 includes the non-transitory computer-readable medium of any of clauses 62-71, wherein the instructions cause the processor to update the sound event classification model based on the model update data.
[0219] Clause 73 includes the non-transitory computer-readable medium of any of clauses 62-72, wherein the determination as to whether the sound event classification model corresponds to the audio scene is based on user input indicating the audio scene.
[0220] Clause 74 includes the non-transitory computer-readable medium of any of clauses 62-73, wherein the determination as to whether the sound event classification model corresponds to the audio scene is based on sensor data.
[0221] Clause 75 includes the non-transitory computer-readable medium of any of clauses 62-74, wherein the determination as to whether the sound event classification model corresponds to the audio scene is based on a timestamp associated with the audio data samples.
[0222] The previous description of the disclosed aspects is provided to enable any persons skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A device comprising: one or more processors configured to: select a sound event classification model from among a plurality of sound event classification models; provide an audio data sample to the sound event classification model; determine, based on an output of the sound event classification model in response to the audio data sample, whether a sound class of the audio data sample is recognized by the sound event classification model; based on a determination that the sound class is not recognized, determine whether the sound event classification model corresponds to an audio scene associated with the audio data sample; and based on a determination that the sound event classification model corresponds to the audio scene associated with the audio data sample, store model update data based on the audio data sample for updating the sound event classification model.
2. The device of claim 1, further comprising a microphone coupled to the one or more processors and configured to capture audio data corresponding to the audio data sample.
3. The device of claim 1, further comprising a memory coupled to the one or more processors and configured to store the plurality of sound event classification models.
4. The device of claim 3, further comprising one or more sensors configured to generate sensor data associated with the audio data sample, wherein the one or more processors are configured to select the sound event classification model based on the sensor data.
5. The device of claim 4, wherein the one or more sensors comprise a camera and a positioning sensor.
6. The device of claim 3, further comprising one or more input devices configured to receive input identifying the audio scene, wherein the one or more processors are configured to select the sound event classification model based on the audio scene.
7. The device of claim 3, wherein the one or more processors are configured to select the sound event classification model based on when the audio data sample is received.
8. The device of claim 3, wherein the memory further stores setting data indicative of one or more device settings, and wherein the one or more processors are configured to select the sound event classification model based on the setting data.
9. The device of claim 1, wherein the one or more processors are further configured to generate, based on a determination that the sound class is recognized, output indicative of the sound class associated with the audio data sample.
10. The device of claim 1, wherein the one or more processors are further configured to, based on a determination that the sound event classification model does not correspond to the audio scene associated with the audio data sample, store audio data corresponding to the audio data sample as training data for a new sound event classification model. 11. The device of claim 1, wherein the sound event classification model is further configured to generate a confidence measure associated with the output, and wherein the one or more processors are configured to determine whether the sound class is recognized by the sound event classification model based on the confidence measure.
12. The device of claim 1, wherein the one or more processors are further configured to update the sound event classification model based on the model update data.
13. The device of claim 1, further comprising one or more input devices configured to receive input identifying the audio scene, wherein the one or more processors are configured to determine whether the sound event classification model corresponds to the audio scene based on the input.
14. The device of claim 1, further comprising one or more sensors configured to generate sensor data associated with the audio data sample, wherein the one or more processors are configured to determine whether the sound event classification model corresponds to the audio scene based on the sensor data.
15. The device of claim 14, wherein the one or more sensors comprise a camera and a positioning sensor.
16. The device of claim 14, wherein the one or more processors are further configured to determine whether the sound event classification model corresponds to the audio scene based on a timestamp associated with the audio data sample.
17. The device of claim 1, wherein the sound event classification model is trained to recognize a particular sound class, and the model update data comprises drift data representing a change in sound characteristics within the particular sound class for which the sound event classification model is not trained to recognize as corresponding to the particular sound class.
18. The device of claim 1, wherein the one or more processors are integrated within a mobile computing device, a vehicle, a wearable device, an augmented reality headset, a mixed reality headset, or a virtual reality headset.
19. The device of claim 1, wherein the one or more processors are included in an integrated circuit.
20. A method comprising: selecting, by one or more processors, a sound event classification model from among a plurality of sound event classification models; providing, by the one or more processors, an audio data sample as input to the sound event classification model; determining, by the one or more processors, whether a sound class of the audio data sample is recognized by the sound event classification model based on an output of the sound event classification model in response to the audio data sample; based on determining that the sound class is not recognized, determining, by the one or more processors, whether the sound event classification model corresponds to an audio scene associated with the audio data sample; and based on a determination that the sound event classification model corresponds to the audio scene associated with the audio data sample, store, by the one or more processors, model update data based on the audio data sample for updating the sound event classification model.
21. The method of claim 20, wherein the plurality of sound event classification models are stored at a memory coupled to the one or more processors.
22. The method of claim 21, wherein the sound event classification model is selected based on user input, setting data, location data, image data, video data, a timestamp associated with the audio data sample, or a combination thereof.
23. The method of claim 20, wherein the determination as to whether the sound event classification model corresponds to the audio scene is based on a confidence metric generated by the sound event classification model, user input, setting data, location data, image data, video data, a timestamp associated with the audio data sample, or a combination thereof.
24. The method of claim 20, further comprising, after storing the model update data: determining whether a threshold number of model update data has been accumulated, and based on a determination that the threshold number of model update data has been accumulated, initiating an automatic update of the sound event classification model using accumulated model update data.
25. The method of claim 24, wherein, prior to the automatic update, the sound event classification model is trained to recognize a plurality of variants of a particular sound class, and wherein the automatic update modifies the sound event classification model to enable the sound event classification model to recognize an additional variant of the particular sound class as corresponding to the particular sound class.
26. A device comprising: means for selecting a sound event classification model from among a plurality of sound event classification models; means for providing an audio data sample to the sound event classification model; means for determining, based on an output of the sound classification model, whether a sound class of the audio data sample is recognized by the sound event classification model; means for determining, in response to a determination that the sound class is not recognized, whether the sound event classification model corresponds to an audio scene associated with the audio data sample; and means for storing, in response to a determination that the sound event classification model corresponds to the audio scene associated with the audio data sample, model update data based on the audio data sample for updating the sound event classification model.
27. The device of claim 26, wherein the means for selecting the sound event classification model from among the plurality of sound event classification models comprises means for selecting the sound event classification model from among the plurality of sound event classification models based on selection criteria.
28. The device of claim 26, further comprising means for updating the sound event classification model based on the model update data.
29. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to: selecting a sound event classification model from among a plurality of sound event classification models; providing an audio data sample as input to the sound event classification model; determining, based on an output of the sound event classification model responsive to the audio data sample, whether a sound class of the audio data sample is recognized by the sound event classification model; based on a determination that the sound class is not recognized, determining whether the sound event classification model corresponds to an audio scene associated with the audio data sample; and based on a determination that the sound event classification model corresponds to the audio scene associated with the audio data sample, storing model update data based on the audio data sample for updating the sound event classification model.
30. The non-transitory computer-readable medium of claim 29, wherein the instructions further cause the processor to update the sound event classification model based on the model update data.
Citation Information
Patent Citations
Model-based media classification service using sensed media noise characteristics
US20170193097A1
Methods and apparatus for notifying a user of the operating condition of a household appliance
US20180158288A1