Adaptive Sound Event Classification
Transfer learning techniques enhance sound event classification systems by adapting to new sound classes and environmental variations, improving accuracy with minimal computational resources.
Patent Information
- Application Number
- JP2023529961
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-24
- Filing Date
- 2021-11-19
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2041-11-19
AI Technical Summary
Existing sound event classification systems are difficult to update to recognize new sound classes or variations within sound classes, requiring significant computing resources and are limited in customization options, especially on resource-constrained devices.
Utilizing transfer learning techniques to update sound event classification models by incorporating drift data and unknown data, allowing the models to adapt to new sound classes and variations with minimal computing resources.
Enables sound event classification systems to accurately identify a wider range of sounds with reduced computational overhead, adapting to new sound classes and environmental variations without the need for complete retraining.
Smart Images

Figure 0007757405000001 
Figure 0007757405000002 
Figure 0007757405000003
Abstract
Description
[Technical Field]
[0001] Priority claims
[0001] This application claims the benefit of priority to commonly owned U.S. Non-Provisional Patent Application No. 17 / 102,724, filed November 24, 2020, the entire contents of which are expressly incorporated herein by reference.
[0002] FIELD OF THE DISCLOSURE
[0002] This disclosure relates generally to adaptive sound event classification. [Background technology]
[0003]
[0003] Advances in technology have led to smaller and more powerful computing devices. For example, there are now a variety of portable personal computing devices, including wireless telephones such as mobile phones and smartphones, tablets, and laptop computers, that are small, lightweight, and easily carried by users. These devices can communicate voice and data packets over wireless networks. Furthermore, many such devices incorporate additional functionality, such as digital still cameras, digital video cameras, digital recorders, and audio file players. Such devices can also process executable instructions, including software applications such as web browser applications that can be used to access the Internet. Thus, for example, these devices, including sound event classification (SEC) systems that attempt to recognize sound events in audio signals (e.g., slamming doors, horns, etc.), can contain significant computing power.
[0004] SEC systems are generally trained using supervised machine learning techniques to recognize a specific set of sounds identified in labeled training data. Therefore, each SEC system tends to be domain-specific (e.g., capable of classifying a predetermined set of sounds). After a SEC system has been trained, it is difficult to update the SEC system to recognize new sound classes that were not identified in the labeled training data. Furthermore, some sound classes that the SEC system was trained to detect may represent sound events with more variants than are represented in the labeled training data. To illustrate, the labeled training data may include audio data samples for many different doorbells, but is unlikely to include all existing variations of doorbell sounds. Retraining a SEC system to recognize a new sound that was not represented in the training data used to train the SEC system may involve completely retraining the SEC system using a new set of labeled training data that includes examples for the new sound in addition to the original training data. Thus, training a SEC system to recognize new sounds (either for new sound classes or for variations on existing sound classes) requires roughly the same computing resources (e.g., processor cycles, memory, etc.) as creating a new SEC system. Furthermore, over time, as more sounds are added that need to be recognized, the number of audio data samples that must be maintained and used to train the SEC system can become prohibitive. Summary of the Invention
[0005] In certain aspects, a device includes one or more processors configured to provide an audio data sample to a sound event classification model and receive an output of the sound event classification model in response to the audio data sample. The one or more processors are also configured to determine, based on the output, whether a sound class of the audio data sample was recognized by the sound event classification model. The one or more processors are further configured to determine, based on a determination that the sound class was not recognized, whether the sound event classification model corresponds to an audio scene associated with the audio data sample. The one or more processors are also configured to store model update data based on the audio data sample, based on a determination that the sound event classification model corresponds to an audio scene associated with the audio data sample.
[0006] In certain aspects, a method includes providing, by one or more processors, an audio data sample as an input to a sound event classification model. The method also includes determining, by the one or more processors, whether a sound class of the audio data sample was recognized by the sound event classification model based on an output of the sound event classification model responsive to the audio data sample. The method further includes determining, by the one or more processors, whether the sound event classification model corresponds to an audio scene associated with the audio data sample based on a determination that the sound class was not recognized. The method also includes storing, by the one or more processors, model update data based on the audio data sample based on a determination that the sound event classification model corresponds to an audio scene associated with the audio data sample.
[0007]
[0007] In certain aspects, a device includes means for providing an audio data sample to a sound event classification model. The device also includes means for determining, based on an output of the sound event classification model, whether a sound class of the audio data sample has been recognized by the sound event classification model. The device further includes means for determining, in response to a determination that the sound class has not been recognized, whether the sound event classification model corresponds to an audio scene associated with the audio data sample. The device also includes means for storing model update data based on the audio data sample, in response to a determination that the sound event classification model corresponds to an audio scene associated with the audio data sample.
[0008] In certain aspects, a non-transitory computer-readable storage medium includes instructions that, when executed by a processor, cause the processor to provide an audio data sample as input to a sound event classification model. The instructions, when executed by the processor, also cause the processor to determine, based on an output of the sound event classification model in response to the audio data sample, whether a sound class of the audio data sample was recognized by the sound event classification model. The instructions, when executed by the processor, further cause the processor to determine, based on a determination that the sound class was not recognized, whether the sound event classification model corresponds to an audio scene associated with the audio data sample. The instructions, when executed by the processor, also cause the processor to store model update data based on the audio data sample based on a determination that the sound event classification model corresponds to an audio scene associated with the audio data sample.
[0009]
[0009] Other aspects, advantages, and features of the present disclosure will become apparent after review of the entire application, including the following sections: Brief Description of the Drawings, Form for Implementing the Invention, and Claims. [Brief explanation of the drawings]
[0010] [Figure 1]
[0010] FIG. 1 is a block diagram of an example of a device configured to generate sound identification data in response to audio data samples and configured to update a sound event classification model. [Figure 2]
[0011] FIG. 10 illustrates updating a sound event classification model to account for drift according to a particular example. [Figure 3]
[0012] FIG. 10 illustrates aspects of updating a sound event classification model to account for new sound classes according to a particular example. [Figure 4]
[0013] 2A and 2B illustrate specific examples of the operation of the device of FIG. 1. [Figure 5]
[0014] 2A and 2B illustrate another specific example of the operation of the device of FIG. 1. [Figure 6]
[0015] 2 is a block diagram illustrating a specific example of the device of FIG. 1; [Figure 7]
[0016] 2 is a diagram of an illustrative example of a vehicle incorporating aspects of the device of FIG. 1; [Figure 8]
[0017] 2 illustrates a virtual reality, mixed reality, or augmented reality headset incorporating aspects of the device of FIG. 1. [Figure 9]
[0018] FIG. 2 illustrates a wearable electronic device incorporating an embodiment of the device of FIG. 1. [Figure 10]
[0019] FIG. 2 illustrates a voice-controlled speaker system incorporating aspects of the device of FIG. 1. [Figure 11]
[0020] FIG. 2 illustrates a camera incorporating an embodiment of the device of FIG. 1. [Figure 12]
[0021] 2 illustrates a mobile device incorporating aspects of the device of FIG. 1; [Figure 13]
[0022] FIG. 2 illustrates a flying device incorporating an embodiment of the device of FIG. 1. [Figure 14]
[0023] FIG. 2 illustrates a headset incorporating aspects of the device of FIG. 1. [Figure 15]
[0024] FIG. 2 illustrates an apparatus incorporating an embodiment of the device of FIG. 1. [Figure 16]
[0025] 2 is a flow chart illustrating an example method of operation of the device of FIG. 1; DETAILED DESCRIPTION OF THE INVENTION
[0011]
[0026] The sound event classification model may be trained using machine learning techniques. For example, a neural network may be trained as a sound event classifier using backpropagation or other machine-learning training techniques. A neural network trained in this manner is referred to herein as a "sound event classification model." A sound event classification model trained in this manner may be small enough (in terms of storage space occupied) and simple enough (in terms of computing resources used during operation) for a portable computing device that stores and uses the sound event classification model. The process of training a sound event classification model uses significantly more processing resources than those used to perform sound event classification using the sound event classification model. Furthermore, the training process uses a large set of labeled training data containing many audio data samples for each sound class that the sound event classification model is trained to detect. Training a sound event classification model from scratch on a portable computing device or another resource-limited computing device can be expensive in terms of memory utilization or other computing resources. Thus, a user desiring to use a sound event classification model on a portable computing device may be limited to downloading a pre-trained sound event classification model onto the portable computing device from a less resource-constrained computing device or from a library of pre-trained sound event classification models, and thus the user has limited customization options.
[0012]
[0027] The disclosed systems and methods use transfer learning techniques to update sound event classification models in a manner that uses significantly fewer computing resources than training a sound event classification model from scratch. According to certain aspects, transfer learning techniques can be used to update sound event classification models to account for drift within sound classes or to recognize new sound classes. In this context, "drift" refers to variation within a sound class. For example, a sound event classification model may be able to recognize some examples of a sound class but may not be able to recognize other examples of the sound class. To illustrate, a sound event classification model trained to recognize a car horn sound class may be able to recognize many different types of car horns but may not be able to recognize some examples of car horns. Drift can also occur due to variation in the acoustic environment. To illustrate, a sound event classification model may be trained to recognize the sound of a bass drum played in a concert hall but may not recognize the bass drum when played by a marching band in an outdoor parade. The transfer learning techniques disclosed herein facilitate updating a sound event classification model to account for such drift, thereby enabling the sound event classification model to detect a wider range of sounds within a sound class. Because drift may correspond to sounds encountered by a user device but not recognized by the user device, updating the sound event classification model for a user device to adapt to these encountered variations in sound classes enables the user device to more accurately identify the particular variations in sound classes that that particular user device typically encounters.
[0013]
[0028] According to certain aspects, when it is determined that the sound event classification model did not recognize the sound class of a sound (based on an audio data sample of the sound), a determination is made as to whether the sound was not recognized due to drift or because the sound event classification model does not recognize a sound class of a type associated with the sound. Information separate from the audio data sample, such as a timestamp, location data, image data, video data, user input data, settings data, and other sensor data, is used to determine scene data indicative of a sound environment (or audio scene) associated with the audio data sample. The scene data is used to determine whether the sound event classification model corresponds to an audio scene (e.g., it has been trained to recognize sound events therein). If the sound event classification model corresponds to an audio scene, the audio data sample is saved as model update data and designated as drift data. In some aspects, if the sound event classification model does not correspond to the audio scene, the audio data sample is either discarded as unknown or saved as model update data and indicated as being associated with an unknown sound class (e.g., unknown data).
[0014]
[0029] Periodically or occasionally (e.g., when initiated by a user or when an update condition is met), the sound event classification model is updated using model update data. For example, to account for drift data, the sound event classifier may be trained using backpropagation or other similar machine learning techniques (e.g., further trained starting from an already trained sound event classifier). In this example, the drift data is associated with labels of sound classes already recognized by the sound event classification model, and the drift data and corresponding labels are used as labeled training data. Updating the sound event classification model using drift data may be augmented by adding other examples of sound classes to the labeled training data, such as examples taken from the training data originally used to train the sound event classification model. In some aspects, the device automatically (e.g., without user input) updates one or more sound event classification models when drift data is available. Thus, the sound event classification system can automatically adapt to account for drift within sound classes using significantly fewer computing resources than would be used to train a sound event classification model from scratch.
[0015]
[0030] To account for unknown data, the sound event classification model may be trained using more complex transfer learning techniques. For example, when unknown data is available, the user may be queried to indicate whether they wish to update the sound event classification model. Audio representing the unknown data may be played out to the user, and the user may indicate that the unknown data will be discarded without updating the sound event classification model, may indicate that the unknown data corresponds to a known sound class (e.g., to reclassify the unknown data as drift data), or may assign a new sound class designation to the unknown data. If the user reclassifies the unknown data as drift data, the machine learning techniques used to update the sound event classification model to account for the drift data are initiated as described above.
[0016]
[0031] When a user assigns a new sound class indication to unknown data, the indication and the unknown data are used as labeled training data to generate an updated sound event classification model. According to certain aspects, a transfer learning technique used to update the sound event classification model includes generating a copy of the sound event classification model including an output node associated with the new sound class. The copy of the sound event classification model is referred to as an incremental model. The transfer learning technique also includes connecting the sound event classification model and the incremental model to one or more adapter networks. The adapter network facilitates generating a merged output based on both the output of the sound event classification model and the output of the incremental model. Audio data samples including unknown data and one or more audio data samples corresponding to known sound classes (e.g., sound classes that the sound event classifier was previously trained to recognize) are provided to the sound event classification model and the incremental model to generate a merged output. The merged output indicates the sound class assigned to the audio data sample based on analysis by the sound event classification model, the incremental model, and the one or more adapter networks. During training, the merged output is used to update link weights between the incremental model and the adapter network. Once training is complete, the sound event classifier can be discarded if the incremental model is sufficiently accurate. If the incremental model alone is not sufficiently accurate, the sound event classification model, the incremental model, and the adapter network are kept together and used as a single updated sound event classification model. Thus, the techniques disclosed herein enable customization and updating of sound event classification models in a manner that requires less resources (in terms of memory resources, processor time, and power) than training a neural network from scratch. Furthermore, in some aspects, the disclosed techniques enable automatic updates of sound event classification models to account for drift.
[0017]
[0032] The disclosed systems and methods provide a context-aware system that can detect dataset drift with little or no supervision and without training a new SEC model from scratch, associate the drift data with corresponding classes (e.g., by utilizing available multi-modal input), and utilize the drift data to improve / fine-tune the SEC model. In some embodiments, before improving / fine-tuning the SEC model, the SEC model is trained to recognize multiple variants of a particular sound class, and improving / fine-tuning the SEC model modifies the SEC model to enable it to recognize additional variants of the particular sound class.
[0018]
[0033] In some aspects, the disclosed systems and methods can be used for applications that experience dataset drift during testing. For example, the systems and methods can detect dataset drift and improve a SEC model without retraining previously learned sound classes from scratch. In some aspects, the disclosed systems and methods can be used to add new sound classes to an existing SEC model (e.g., a SEC model already trained for several sound classes) without retraining the SEC model from scratch, without needing to access all the training data originally used to train the SEC model, and without introducing any performance degradation with respect to the sound classes the SEC model was originally trained to recognize.
[0019]
[0034] In some aspects, the disclosed systems and methods can be used in applications where continuous learning capabilities with low footprint constraints are desirable. In some implementations, the disclosed systems and methods can have access to a database of various detection models (e.g., SEC models) for a diverse range of applications (e.g., various sound environments). In such implementations, an SEC model can be selected during operation based on the sound environment, and the SEC model can be loaded and utilized as a source model.
[0020]
[0035] Certain aspects of the present disclosure are described below with reference to the drawings. In the description, common features are designated by common reference numerals. As used herein, various terms are used only to describe particular implementations and are not intended to limit the implementations. For example, the singular forms "a," "an," and "the" include the plural forms unless the context clearly indicates otherwise. Furthermore, some features described herein are singular in some implementations and plural in other implementations. For illustrative purposes, FIG. 1 shows a device 100 including one or more sensors ("sensor 134" in FIG. 1), indicating that in some implementations, the device 100 includes a single sensor 134 and in other implementations, the device 100 includes multiple sensors 134. For ease of reference herein, such features will generally be introduced as "one or more" features and will thereafter be referred to in the singular or optional plural (generally indicated by a term ending in "s"), unless an aspect relating to a plurality of the feature is described.
[0021]
[0036] The terms “comprise,” “comprises,” and “comprising” are used interchangeably herein with “include,” “includes,” or “including.” Furthermore, the term “wherein” is used interchangeably with “where.” As used herein, “exemplary” indicates an example, an implementation, and / or an aspect and should not be construed as limiting or as indicating a preferred or preferred implementation. As used herein, ordinal terms (e.g., “first,” “second,” “third,” etc.) used to modify an element such as a structure, component, operation, etc. do not in themselves indicate any priority or ordering of that element relative to another element, but rather merely distinguish that element from other elements having the same name (apart from the use of the ordinal term). As used herein, the term “set” refers to one or more of a particular element, and the term “plurality” refers to a plurality (e.g., two or more) of a particular element.
[0022]
[0037] As used herein, “coupled” may include “communicatively coupled,” “electrically coupled,” or “physically coupled,” and (or alternatively) may include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof), etc. Two devices (or components) that are electrically coupled may be included in the same device or in different devices and may be connected via electronic circuitry, one or more connectors, or inductive coupling, as illustrative, non-limiting examples. In some implementations, two devices (or components) that are communicatively coupled, such as in electrical communication, may send and receive electrical signals (digital or analog signals) directly or indirectly, such as via one or more wires, buses, networks, etc. As used herein, "directly coupled" refers to two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) with no intervening components.
[0023]
[0038] In this disclosure, terms such as “determining,” “calculating,” “estimating,” “shifting,” “adjusting,” and the like may be used to describe how one or more operations are performed. It should be noted that such terms should not be construed as limiting, and other techniques may be utilized to perform similar operations. Additionally, as referred to herein, “generating,” “calculating,” “estimating,” “using,” “selecting,” “accessing,” and “determining” may be used interchangeably. For example, “generating,” “calculating,” “estimating,” or “determining” a parameter (or signal) may refer to actively generating, estimating, calculating, or determining a parameter (or signal), or may refer to using, selecting, or accessing a parameter (or signal) that has already been generated, such as by another component or device.
[0024]
[0039] FIG. 1 is a block diagram of an example device 100 configured to generate sound identification data in response to audio data samples 110 and to update a sound event classification model. The device 100 of FIG. 1 includes one or more microphones 104 configured to generate an audio signal 106 based on a sound 102 detected in an acoustic environment. The microphones 104 are coupled to a feature extractor 108 that generates audio data samples 110 based on the audio signal 106. For example, the audio data samples 110 may include an array or matrix of data elements, each corresponding to a feature detected in the audio signal 106. As a particular example, the audio data samples 110 may correspond to Mel spectrum features extracted from one second of the audio signal 106. In this example, the audio data samples 110 may include a 128×128 element matrix of feature values. In other examples, other audio data sample configurations or sizes may be used.
[0025]
[0040] Audio data samples 110 are provided to a sound event classification (SEC) engine 120. SEC engine 120 is configured to perform an inference operation based on one or more SEC models, such as SEC model 112. If the sound class of audio data sample 110 is recognized by SEC model 112, then the "inference operation" refers to assigning audio data sample 110 to a sound class. For example, SEC engine 120 may include or correspond to software implementing a machine learning runtime environment, such as the Qualcomm Neural Processing SDK available from Qualcomm Technologies, Inc., San Diego, California, USA. In certain aspects, SEC model 112 is one of multiple SEC models available to SEC engine 120 (e.g., available SEC model 114).
[0026]
[0041] In a particular example, each of the available SEC models 114 includes or corresponds to a neural network trained as a sound event classifier. For illustrative purposes, the SEC model 112 (as well as each of the other available SEC models 114) may include an input layer, one or more hidden layers, and an output layer. In this example, the input layer is configured to correspond to an array or matrix of values of the audio data samples 110 generated by the feature extractor 108. For illustrative purposes, if the audio data samples 110 include 15 data elements, the input layer may include 15 nodes (e.g., one per data element). The output layer is configured to correspond to the sound classes that the SEC model 112 is trained to recognize. The specific configuration of the output layer can vary depending on the information provided as output. As an example, the SEC model 112 may be trained to output an array containing one bit per sound class, where the output layer implements "one-hot coding" such that all but one of the bits in the output array have a value of 0, and the bit corresponding to the detected sound class has a value of 1. Other output schemes may be used, for example, to indicate a confidence metric value for each sound class, where the confidence metric value indicates a probability estimate that the audio data sample 110 corresponds to the respective sound class. To illustrate, if the SEC model 112 is trained to recognize four sound classes, the SEC model 112 may generate output data including four values (one for each sound class), each value indicating a probability estimate that the audio data sample 110 corresponds to the respective sound class.
[0027]
[0042] Each hidden layer includes multiple nodes, each interconnected (via links) with other nodes in the same or different layers. Each input link of a node is associated with a link weight. In operation, a node receives input values from other nodes to which it is linked, weights the input values based on the corresponding link weights to determine a combined value, and applies the combined value to an activation function to generate an output value for the node. The output value is provided to one or more other nodes via the node's output links. A node may also include a bias value used to generate the combined value. Nodes may be linked in various configurations and may include various other features (e.g., memory of previous values) to facilitate processing of specific data. In the case of audio data samples, a convolutional neural network (CNN) may be used. By way of example, one or more of the SEC models 112 may include three linked CNNs, each of which may include a two-dimensional (2D) convolutional layer, a maxpooling layer, and a batch normalization layer. In other implementations, the hidden layer includes a different number of CNN or other layers. Training the neural network includes modifying link weights to reduce output errors of the neural network.
[0028]
[0043] During operation, the SEC engine 120 may provide the audio data sample 110 as input to a single SEC model (e.g., SEC model 112), to multiple selected SEC models (e.g., SEC model 112 and the Kth SEC model 118 of the available SEC models 114), or to each of the SEC models (e.g., to SEC model 112, the first SEC model 116, the Kth SEC model 118, and any other SEC model of the available SEC models 114). For example, the SEC engine 120 (or another component of the device 100) may select the SEC model 112 from among the available SEC models 114 based on, for example, user input, device settings associated with the device 100, sensor data, the time the audio data sample 110 is received, or other factors. In this example, the SEC engine 120 may select to use only the SEC model 112, or may select to use two or more of the available SEC models 114. To illustrate, the device settings may indicate that the SEC model 112 and the first SEC model 116 are to be used during a particular time frame. In another example, the SEC engine 120 may provide audio data samples 110 to each of the available SEC models 114 (e.g., serially or in parallel) to generate output from each. In particular aspects, the SEC models are trained to recognize different sound classes, the same sound class in different acoustic environments, or both. For example, the SEC model 112 may be configured to recognize a first set of sound classes, and the first SEC model 116 may be configured to recognize a second set of sound classes, where the first set of sound classes are different from the second set of sound classes.
[0029]
[0044] In certain aspects, the SEC engine 120 determines, based on the output of the SEC model 112, whether the SEC model 112 recognized the sound class of the audio data sample 110. If the SEC engine 120 provides the audio data sample 110 to multiple SEC models, the SEC engine 120 may determine, based on the output of each of the SEC models, whether any of the SEC models recognized the sound class of the audio data sample 110. If the SEC model 112 (or another of the available SEC models 114) recognized the sound class of the audio data sample 110, the SEC engine 120 generates an output 124 indicating the sound class 122 of the audio data sample 110. For example, the output 124 may be sent to a display to notify a user of the detection of the sound class 122 associated with the sound 102, or may be sent to another device or component of the device 100 and used to trigger an action (e.g., sending a command to activate a light in response to recognizing the sound of a door closing).
[0030]
[0045] If the SEC engine 120 determines that the SEC model 112 (and other of the available SEC models 114 given the audio data sample 110) did not recognize the sound class of the audio data sample 110, the SEC engine 120 provides a trigger signal 126 to the drift detector 128. For example, the SEC engine 120 may set a trigger flag in the memory of the device 100. In some implementations, the SEC engine 120 may also provide other data to the drift detector 128. To illustrate, if the SEC model 112 generates a confidence metric value for each sound class that the SEC model 112 is trained to recognize, one or more of the confidence metric values may be provided to the drift detector 128. For example, if the SEC model 112 is trained to recognize three sound classes, the SEC engine 120 may provide the drift detector 128 with the highest confidence value of the three confidence values output by the SEC model 112 (one for each of the three sound classes).
[0031]
[0046] In particular aspects, the SEC engine 120 determines whether the SEC model 112 recognized the sound class of the audio data sample 110 based on the value of the confidence metric. In this particular aspect, the value of the confidence metric for a particular sound class indicates the probability that the audio data sample 110 is associated with the particular sound class. To illustrate, if the SEC model 112 is trained to recognize four sound classes, the SEC model 112 may generate as output an array containing four values of the confidence metric, one for each sound class. In some implementations, if the value of the confidence metric for the sound class 122 is greater than a detection threshold, the SEC engine 120 determines that the SEC model 112 recognized the sound class 122 of the audio data sample 110. For example, if the value of the confidence metric for the sound class 122 is greater than a detection threshold of 0.90 (e.g., 90% confidence), 0.95 (e.g., 95% confidence), or some other value, the SEC engine 120 determines that the SEC model 112 recognized the sound class 122 of the audio data sample 110. In some implementations, if the value of the confidence metric for each sound class that the SEC model 112 was trained to recognize is less than the detection threshold, the SEC engine 120 determines that the SEC model 112 did not recognize the sound class 122 of the audio data sample 110. For example, if each value of the confidence metric is less than a detection threshold of 0.90 (e.g., 90% confidence), 0.95 (e.g., 95% confidence), or some other value, the SEC engine 120 determines that the SEC model 112 did not recognize the sound class 122 of the audio data sample 110.
[0032]
[0047] The drift detector 128 is configured to determine whether a SEC model 112 that was unable to recognize the sound class of the audio data sample 110 corresponds to an audio scene 142 associated with the audio data sample 110. In the example shown in FIG. 1 , the scene detector 140 is configured to receive scene data 138 and use the scene data 138 to determine an audio scene 142 associated with the audio data sample 110. In particular aspects, the scene data 138 is generated based on setting data 130 indicative of one or more device settings associated with the device 100, an output of a clock 132, sensor data from one or more sensors 134, input received via an input device 136, or a combination thereof. In some aspects, the scene detector 140 uses different information to determine the audio scene 142 than the SEC engine 120 uses to select the SEC model 112. To illustrate, if the SEC engine 120 selects the SEC model 112 based on the time of day, the scene detector 140 may use position sensor data from a position sensor of the sensors 134 to determine the audio scene 142. In some aspects, the scene detector 140 uses at least some of the same information and additional information that the SEC engine 120 uses to select the SEC model 112. To illustrate, if the SEC engine 120 selects the SEC model 112 based on the time of day and the configuration data 130, the scene detector 140 may use the position sensor data and the configuration data 130 to determine the audio scene 142. Thus, the scene detector 140 uses a different audio scene detection mode than that used by the SEC engine 120 to select the SEC model 112.
[0033]
[0048] In particular implementations, the scene detector 140 is a neural network that is trained to determine the audio scene 142 based on the scene data 138. In other implementations, the scene detector 140 is a classifier that is trained using different machine learning techniques. For example, the scene detector 140 may include or correspond to a decision tree, a random forest, a support vector machine, or another classifier that is trained to generate an output indicative of the audio scene 142 based on the scene data 138. In yet other implementations, the scene detector 140 uses heuristics to determine the audio scene 142 based on the scene data 138. In still other implementations, the scene detector 140 uses a combination of artificial intelligence and heuristics to determine the audio scene 142 based on the scene data 138. For example, the scene data 138 may include image data, video data, or both, and the scene detector 140 may include an image recognition model that is trained using machine learning techniques to detect particular objects, motion, backgrounds, or other image or video information. In this example, the output of the image recognition model may be evaluated via one or more heuristics to determine the audio scene 142.
[0034]
[0049] The drift detector 128 compares the information describing the SEC model 112 with the audio scene 142 indicated by the scene detector 140 to determine whether the SEC model 112 is associated with the audio scene 142 of the audio data sample 110. If the drift detector 128 determines that the SEC model 112 is associated with the audio scene 142 of the audio data sample 110, the drift detector 128 stores drift data 144 as model update data 148. In particular implementations, the drift data 144 includes the audio data sample 110 and an indication, where the indication identifies the SEC model 112, indicates a sound class associated with the audio data sample 110, or both. If the drift data 144 indicates a sound class associated with the audio data sample 110, the sound class may be selected based on the highest value of the confidence metric generated by the SEC model 112. As an illustrative example, if the SEC engine 120 uses a detection threshold of 0.90 and the highest confidence metric output by the SEC models 112 is 0.85 for a particular sound class, the SEC engine 120 determines that the sound class of the audio data sample 110 was not recognized and sends a trigger signal 126 to the drift detector 128. In this example, if the drift detector 128 determines that the SEC model 112 corresponds to an audio scene 142 of the audio data sample 110, the drift detector 128 stores the audio data sample 110 as drift data 144 associated with the particular sound class. In particular aspects, metadata associated with the SEC models 114 includes information specifying one or more audio scenes associated with each SEC model 114. For example, the SEC models 112 may be configured to detect sound events in a user's home, in which case the metadata associated with the SEC models 112 may indicate that the SEC model 112 is associated with a "home" audio scene.In this example, if the audio scene 142 indicates that the device 100 is at its home location (e.g., based on location information, user input, detection of a home wireless network signal, image or video data representing the home location, etc.), the drift detector 128 determines that the SEC model 112 corresponds to the audio scene 142.
[0035]
[0050] In some implementations, the drift detector 128 also stores some audio data samples 110 as model update data 148 and designates them as unknown data 146. As a first example, if the drift detector 128 determines that the SEC model 112 does not correspond to the audio scene 142 of the audio data sample 110, the drift detector 128 may store the unknown data 146. As a second example, if the value of the reliability metric output by the SEC model 112 fails to satisfy a drift threshold, the drift detector 128 may store the unknown data 146. In this example, the drift threshold is less than the detection threshold used by the SEC engine 120. For example, if the SEC engine 120 uses a detection threshold of 0.95, the drift threshold may have a value of 0.80, a value of 0.75, or some other value lower than 0.95. In this example, if the highest value of the reliability metric for the audio data sample 110 is less than the drift threshold, the drift detector 128 determines that the audio data sample 110 belongs to a sound class that the SEC model 112 is not trained to recognize and designates the audio data sample 110 as unknown data 146. In particular aspects, the drift detector 128 only stores the unknown data 146 if the drift detector 128 determines that the SEC model 112 does not correspond to the audio scene 142 of the audio data sample 110. As another example, the drift detector 128 stores the unknown data 146 regardless of whether the drift detector 128 determines that the SEC model 112 corresponds to the audio scene 142 of the audio data sample 110.
[0036]
[0051] After the model update data 148 is stored, a model updater 152 can access the model update data 148 and use it to update one of the available SEC models 114 (e.g., SEC model 112). For example, each entry in the model update data 148 indicates the SEC model with which the entry is associated, and the model updater 152 uses the entry as training data to update the corresponding SEC model. In certain aspects, the model updater 152 updates the SEC model when update criterion is met or when a model update is initiated by a user or another party (e.g., the vendor of the device 100, the SEC engine 120, the SEC model 114, etc.). The update criteria may be met when a certain number of entries are available in the model update data 148, when a certain number of entries for a particular SEC model are available in the model update data 148, when a certain number of entries for a particular sound class are available in the model update data 148, when a certain amount of time has passed since the previous update, when another update occurs (e.g., when a software update related to the device 100 occurs), or based on the occurrence of another event.
[0037]
[0052] The model updater 152 uses the drift data 144 as indicated training data to update the training of the SEC model 112 using backpropagation or a similar machine learning optimization process. For example, the model updater 152 provides the audio data samples from the drift data 144 in the model update data 148 as input to the SEC model 112, determines the value of an error function (also called a loss function) based on the output of the SEC model 112 and the indications associated with the audio data samples (as indicated in the drift data 144 stored by the drift detector 128), and determines updated link weights for the SEC model 112 using a gradient descent operation (or some variant thereof) or another machine learning optimization process.
[0038]
[0053] The model updater 152 may also provide other audio data samples (in addition to the audio data samples in the drift data 144) to the SEC model 112 during update training. For example, the model update data 148 may include one or more known audio data samples (such as a subset of the audio data samples initially used to train the SEC model 112), which may reduce the likelihood of update training causing the SEC model 112 to forget the previous training (where "forgetting" here refers to losing reliability for detecting the sound classes that the SEC model 112 was previously trained to recognize). Because the sound classes associated with the audio data samples in the drift data 144 are indicated by the drift detector 128, update training that takes drift into account may be achieved automatically (e.g., without user input). Thus, the functionality of the device 100 (e.g., accuracy in recognizing sound classes) can improve over time without user intervention and using fewer computing resources than would be used to generate a new SEC model from scratch. A specific example of a transfer learning process that the model updater 152 may use to update the SEC model 112 based on the drift data 144 is described with reference to FIG.
[0039]
[0054] In some aspects, the model updater 152 can also use the unknown data 146 in the model update data 148 to update the training of the SEC model 112. For example, from time to time, such as periodically or when update criteria are met, the model updater 152 may prompt the user to ask the user to indicate a sound class for the unknown data 146 entry in the model update data 148. If the user selects to indicate a sound class for the unknown data 146 entry, the device 100 (or another device) may play out a sound corresponding to the audio data sample of the unknown data 146. The user can provide one or more indications 150 (e.g., via the input device 136) that identify the sound class of the audio data sample. If the sound class indicated by the user is a sound class that the SEC model 112 was trained to recognize, the unknown data 146 is reclassified as drift data 144 associated with the user-specified sound class and the SEC model 112. Depending on the configuration of the model updater 152, if the sound class indicated by the user is a sound class that the SEC model 112 is not trained to recognize (e.g., a new sound class), the model updater 152 may discard the unknown data 146, send the unknown data 146 and the user-specified sound class to another device for use in generating a new or updated SEC model, or use the unknown data 146 and the user-specified sound class to update the SEC model 112. A specific example of a transfer learning process that the model updater 152 can use to update the SEC model 112 based on the unknown data 146 and the user-specified sound class is described with reference to FIG.
[0040]
[0055] Updated SEC models 154 generated by the model updater 152 are added to the available SEC models 114 to make them available for evaluating audio data samples 110 received after the updated SEC models 154 were generated. Thus, the set of available SEC models 114 that can be used to evaluate a sound is dynamic. For example, one or more of the available SEC models 114 may be automatically updated to account for drift data 144. Furthermore, one or more of the available SEC models 114 may be updated to account for unknown sound classes using a transfer learning operation, which uses fewer computing resources (e.g., memory, processing time, and power) than training a new SEC model from scratch.
[0041]
[0056] 2 illustrates updating a SEC model 208 to account for drift, according to a particular example. The SEC model 208 of FIG. 2 includes or corresponds to a particular one of the available SEC models 114 of FIG. 1 associated with drift data 144. For example, if the SEC engine 120 generates the trigger signal 126 in response to the output of the SEC model 112, the drift data 144 is associated with the SEC model 112, and the SEC model 208 corresponds to or includes the SEC model 112. As another example, if the SEC engine 120 generates the trigger signal 126 in response to the output of the Kth SEC model 118, the drift data 144 is associated with the Kth SEC model 118, and the SEC model 208 corresponds to or includes the Kth SEC model 118.
[0042]
[0057] In the example shown in FIG. 2, training data 202 is used to update SEC model 208. Training data 202 includes drift data 144 and one or more indicators 204. Each entry in drift data 144 includes an audio data sample (e.g., audio data sample 206) and is associated with a corresponding indicator in indicators 204. The audio data sample of an entry in drift data 144 includes a set of values representing features extracted from or determined based on a sound that was not recognized by SEC model 208. The indicator 204 corresponding to an entry in drift data 144 identifies a sound class to which the sound is expected to belong. As an example, in response to SEC model 208 determining that an audio data sample corresponds to a generated audio scene, the indicator 204 corresponding to the entry in drift data 144 may be assigned by drift detector 128 of FIG. 1. In this example, drift detector 128 may assign the audio data sample to the sound class associated in the output of SEC model 208 with the highest confidence metric value.
[0043]
[0058] 2 , an audio data sample 206 corresponding to a sound is provided to a SEC model 208, which generates an output 210 indicating the sound class to which the audio data sample 206 is assigned, one or more values of a confidence metric, or both. The model updater 152 uses the output 210 and the indication 204 corresponding to the audio data sample 206 to determine updated link weights 212 for the SEC model 208. The SEC model 208 is updated based on the updated link weights 212, and the training process is iteratively repeated until a training termination condition is met. During training, each of the entries in the drift data 144 (e.g., one entry per iteration) may be provided to the SEC model 208. Additionally, in some implementations, other audio data samples (e.g., audio data samples previously used to train the SEC model 208) may also be provided to the SEC model 208 to reduce the likelihood that the SEC model 208 will forget the previous training.
[0044]
[0059] The training termination condition may be met when all of the drift data 144 has been presented to the SEC model 208 at least once, when a certain number of training iterations have been performed, when a convergence metric meets a convergence threshold, or when some other condition indicating the end of training is met. When the training termination condition is met, the model updater 152 stores an updated SEC model 214, where the updated SEC model 214 corresponds to the SEC model 208 with link weights based on the updated link weights 212 applied during training.
[0045]
[0060] 3 illustrates updating a SEC model 310 to account for unknown data according to a particular example. The SEC model 310 of FIG. 3 includes or corresponds to a particular one of the available SEC models 114 of FIG. 1 associated with the unknown data 146. For example, if the SEC engine 120 generates the trigger signal 126 in response to the output of the SEC model 112, the unknown data 146 is associated with the SEC model 112, and the SEC model 310 corresponds to or includes the SEC model 112. As another example, if the SEC engine 120 generates the trigger signal 126 in response to the output of the Kth SEC model 118, the unknown data 146 is associated with the Kth SEC model 118, and the SEC model 310 corresponds to or includes the Kth SEC model 118.
[0046]
[0061] 3 , the model updater 152 generates an updated model 306. The updated model 306 includes the SEC model 310 to be updated, an incremental model 308, and one or more adapter networks 312. The incremental model 308 is a copy of the SEC model 310 with a different output layer than the SEC model 310. In particular, the output layer of the incremental model 308 includes more output nodes than the output layer of the SEC model 310. For example, the output layer of the SEC model 310 includes a first number of nodes (e.g., N nodes, where N is a positive integer corresponding to the number of sound classes that the SEC model 310 is trained to recognize), and the output layer of the incremental model 308 includes a second number of nodes (e.g., N+M nodes, where M is a positive integer corresponding to the number of new sound classes that the updated SEC model 324 is to be trained to recognize that the SEC model 310 is not trained to recognize). The first number of nodes corresponds to the number of sound classes in the first set of sound classes that the SEC model 310 was trained to recognize (e.g., the first set of sound classes includes N distinct sound classes that the SEC model 310 is capable of recognizing), and the second number of nodes corresponds to the number of sound classes in the second set of sound classes that the updated SEC model 324 will be trained to recognize (e.g., the second set of sound classes includes N+M distinct sound classes that the updated SEC model 324 will be trained to recognize). The second set of sound classes includes the first set of sound classes (e.g., N classes) plus one or more additional sound classes (e.g., M classes). The model parameters (e.g., link weights) of the incremental model 308 are initialized to be equal to the model parameters of the SEC model 310.
[0047]
[0062] The adapter network 312 includes a neural adapter and a merger adapter. The neural adapter includes one or more adapter layers configured to receive input from the SEC model 310 and generate an output that can be merged with the output of the incremental model 308. For example, the SEC model 310 generates a first output corresponding to a first number of classes in a first set of sound classes. In a particular aspect, the first output includes one data element (e.g., N data elements) per node in the output layer of the SEC model 310. In contrast, the incremental model 308 generates a second output corresponding to a second number of classes in a second set of sound classes. For example, the second output includes one data element (e.g., N+M data elements) per node in the output layer of the incremental model 308. In this example, the adapter layer of the adapter network 312 receives the output of the SEC model 310 as input and generates an output having a second number of data elements (e.g., N+M). In a particular example, the adapter layer of the adapter network 312 includes two fully connected layers (e.g., an input layer containing N nodes, with each node in the input layer connected to every node in the output layer, and an output layer containing N+M nodes).
[0048]
[0063] The merger adapter of the adapter network 312 is configured to generate the output 314 of the updated model 306 by merging the output of the adapter layer with the output of the incremental model 308. For example, the merger adapter combines the output of the adapter layer with the output of the incremental model 308 element by element to generate a combined output, and applies an activation function (such as a sigmoid function) to the combined output to generate the output 314. The output 314 indicates the sound class to which the audio data sample 304 is assigned by the updated model 306, one or more confidence metric values determined by the updated model 306, or both.
[0049]
[0064] The model updater 152 uses the output 314 and the indications 150 corresponding to the audio data samples 304 to determine updated link weights 316, the adapter network 312, or both for the incremental model 308. The link weights of the SEC model 310 remain unchanged during training. The training process is repeated iteratively until a training termination condition is met. During training, each entry of the unknown data 146 (e.g., one entry per iteration) may be provided to the update model 306. Additionally, in some implementations, other audio data samples (e.g., audio data samples previously used to train the SEC model 310) may also be provided to the update model 306 to reduce the likelihood that the incremental model 308 will forget the previous training of the SEC model 310.
[0050]
[0065] The training termination condition may be met when all of the unknown data 146 has been presented to the updated model 306 at least once, when a certain number of training iterations have been performed, when a convergence metric meets a convergence threshold, or when some other condition indicating the end of training is met. When the training termination condition is met, the model checker 320 selects the updated SEC model 324 from between the incremental model 308 and the updated model 306 (e.g., the combination of the SEC model 310, the incremental model 308, and the adapter network 312).
[0051]
[0066] In particular aspects, model checker 320 selects updated SEC model 324 based on the accuracy of sound classes 322 assigned by incremental model 308 and the accuracy of sound classes 322 assigned by SEC model 310. For example, model checker 320 may determine an F1 score for incremental model 308 (based on sound classes 322 assigned by incremental model 308) and an F1 score for SEC model 310 (based on sound classes 322 assigned by SEC model 310). In this example, if the value of the F1 score for incremental model 308 is greater than or equal to the value of the F1 score for SEC model 310, model checker 320 selects incremental model 308 as the updated SEC model 324. In some implementations, if the F1 score value of the incremental model 308 is greater than or equal to the F1 score value of the SEC model 310 (or is less than the F1 score value of the SEC model 310 by less than a threshold amount), the model checker 320 selects the incremental model 308 as the updated SEC model 324. If the F1 score value of the incremental model 308 is less than the F1 score value for the SEC model 310 (or is less than the F1 score value for the SEC model 310 by more than a threshold amount), the model checker 320 selects the updated model 306 as the updated SEC model 324. If the incremental model 308 is selected as the updated SEC model 324, the SEC model 310, the adapter network 312, or both may be discarded.
[0052]
[0067] In some implementations, the model checker 320 is omitted or integrated with the model updater 152. For example, after training the updated model 306, the updated model 306 may be stored as the updated SEC model 324 (e.g., without selecting between the updated model 306 and the incremental model 308). As an example, while training the updated model 306, the model updater 152 may determine an accuracy metric for the incremental model 308. In this example, the training end condition may be based on the accuracy metric for the incremental model 308, such that after training, the incremental model 308 is stored as the updated SEC model 324 (e.g., without selecting between the updated model 306 and the incremental model 308).
[0053]
[0068] Utilizing the transfer learning techniques described with reference to Figure 3, the model checker 320 enables the device 100 of Figure 1 to update the SEC model to recognize previously unknown sound classes. Furthermore, the described transfer learning techniques use significantly less computational resources (e.g., memory, processing time, and power) than would be used to train a SEC model from scratch to recognize previously unknown sound classes.
[0054]
[0069] In some implementations, the operations described with reference to FIG. 2 (e.g., generating the updated SEC model 214 based on the drift data 144) are performed on the device 100 of FIG. 1, and the operations described with reference to FIG. 3 (e.g., generating the updated SEC model 324 based on the unknown data 146) are performed on a different device (such as the remote computing device 618 of FIG. 6). To illustrate, the unknown data 146 and the indications 150 may be captured at the device 100 and transmitted to a second device with more available computing resources. In this example, the second device generates the updated SEC model 324, and the device 100 downloads or receives a transmission or data representing the updated SEC model 324 from the second device. Generating the updated SEC model 324 based on the unknown data 146 is a more resource-intensive process (e.g., using more memory, power, and processor time) than generating the updated SEC model 214 based on the drift data 144. Therefore, dividing the operations described with reference to FIG. 2 and the operations described with reference to FIG. 3 between different devices may conserve resources of device 100.
[0055]
[0070] Figure 4 illustrates a specific example of the operation of the device of Figure 1. Figure 4 illustrates an implementation in which the determination of whether an active SEC model (e.g., SEC model 112) corresponds to an audio scene for which an audio data sample 110 is captured is based on comparing the current audio scene to a previous audio scene.
[0056]
[0071] 4, audio data captured by microphone 104 is used to generate audio data samples 110. The audio data samples 110 are used to perform audio classification 402. For example, one or more of the available SEC models 114 are used as active SEC models by SEC engine 120 of FIG. 1. In certain aspects, the active SEC model is selected from among the available SEC models 114 based on the audio scene indicated by scene detector 140 during a previous sampling period, also referred to as previous audio scene 408.
[0057]
[0072] The audio classification 402 generates a result 404 based on analyzing the audio data sample 110 using the active SEC model. The result 404 may indicate a sound class associated with the audio data sample 110, a probability that the audio data sample 110 corresponds to a particular sound class, or that the sound class of the audio data sample 110 is unknown. If the result 404 indicates that the audio data sample 110 corresponds to a known sound class, then a decision is made at block 406 to generate an output 124 indicating the sound class 122 associated with the audio data sample 110. For example, the SEC engine 120 of FIG. 1 may generate the output 124.
[0058]
[0073] If the result 404 indicates that the audio data sample 110 does not correspond to a known sound class, then a decision is made in block 406 to generate a trigger 126. The trigger 126 activates a drift detection scheme, which in Figure 4 includes having the scene detector 140 identify a current audio scene 407 based on data from the sensor 134.
[0059]
[0074] The current audio scene 407 is compared to the previous audio scene 408 in block 410 to determine if an audio scene change has occurred since the active SEC model was selected. In block 412, a determination is made whether the sound class of the audio data sample 110 was not recognized due to drift. For example, if the current audio scene 407 does not correspond to the previous audio scene 408, the determination in block 412 will be that drift was not the reason the sound class of the audio data sample 110 was not recognized. In this situation, the audio data sample 110 may be discarded or stored as unknown data in block 414.
[0060]
[0075] If the current audio scene 407 corresponds to the previous audio scene 408, then the determination in block 412 is that the sound class of the audio data sample 110 was not recognized due to drift because the active SEC model corresponds to the current audio scene 407. In this situation, the drifted sound class is identified in 416 and the audio data sample 110 and the sound class identifier are stored as drift data in block 418.
[0061]
[0076] Once sufficient drift data has been stored, the SEC model is updated in block 420 to generate an updated SEC model 154. The updated SEC model 154 is added to the available SEC models 114. In some implementations, the updated SEC model 154 replaces the active SEC model that generated the result 404.
[0062]
[0077] Figure 5 illustrates another specific example of the operation of the device of Figure 1. Figure 5 illustrates an implementation in which a determination of whether an active SEC model (e.g., SEC model 112) corresponds to an audio scene in which an audio data sample 110 is captured is based on comparing the current audio scene with information describing the active SEC model.
[0063]
[0078] 5, audio data captured by microphone 104 is used to generate audio data samples 110. Audio data samples 110 are used to perform audio classification 402. For example, one or more of the available SEC models 114 are used as active SEC models by SEC engine 120 of FIG. 1. In certain aspects, the active SEC models are selected from among the available SEC models 114. In some implementations, rather than selecting one or more of the available SEC models 114 as the active SEC model, an ensemble of the available SEC models 114 is used.
[0064]
[0079] The audio classification 402 generates a result 404 based on an analysis of the audio data sample 110 using one or more of the available SEC models 114. The result 404 may indicate a sound class associated with the audio data sample 110, a probability that the audio data sample 110 corresponds to a particular sound class, or that the sound class of the audio data sample 110 is unknown. If the result 404 indicates that the audio data sample 110 corresponds to a known sound class, then a decision is made at block 406 to generate an output 124 indicating the sound class 122 associated with the audio data sample 110. For example, the SEC engine 120 of FIG. 1 may generate the output 124.
[0065]
[0080] If the result 404 indicates that the audio data sample 110 does not correspond to a known sound class, then a decision is made in block 406 to generate a trigger 126. The trigger 126 activates a drift detection scheme, which in Figure 5 includes causing the scene detector 140 to identify the current audio scene based on data from the sensor 134 and determine whether the current audio scene corresponds to the SEC model that generated the result 404 that caused the trigger 126 to be sent.
[0066]
[0081] At block 412, a determination is made whether the sound class of the audio data sample 110 was not recognized due to drift. For example, if the current audio scene does not correspond to the SEC model that produced the result 404, the determination at block 412 will be that drift was not the reason the sound class of the audio data sample 110 was not recognized. In this situation, the audio data sample 110 may be discarded or stored as unknown data at block 414.
[0067]
[0082] If the current audio scene corresponds to the SEC model that produced result 404, then the determination in block 412 is that the sound class of the audio data sample 110 was not recognized due to drift. In this situation, the drifted sound class is block identified in 416, and the audio data sample 110 and the sound class identifier are stored as drift data in block 418.
[0068]
[0083] Once sufficient drift data has been stored, the SEC model is updated in block 420 to generate an updated SEC model 154. The updated SEC model 154 is added to the available SEC models 114. In some implementations, the updated SEC model 154 replaces the active SEC model that generated the result 404.
[0069]
[0084] FIG. 6 is a block diagram illustrating a specific example of the device 100 of FIG. 1. In FIG. 6, the device 100 is configured to use SEC models 114 to generate sound identification data (e.g., the output 124 of FIG. 1) in response to an input of an audio data sample (e.g., the audio data sample 110 of FIG. 1). Additionally, the device 100 of FIG. 6 is configured to update one or more of the SEC models 114 based on model update data 148. For example, the device 100 is configured to update the SEC models 114 using drift data 144 as described with reference to FIG. 2, or to update the SEC models 114 using unknown data 146 as described with reference to FIG. 3, or both. In some implementations, a remote computing device 618 updates the SEC models 114 in some situations. By way of illustration, the device 100 may update the SEC models 114 using the drift data 144, and the remote computing device 618 may update the SEC models 114 using the unknown data 146. In various implementations, device 100 may have more or fewer components than those shown in FIG.
[0070]
[0085] In particular implementations, device 100 includes a processor 604 (e.g., a central processing unit (CPU)). Device 100 may include one or more additional processors 606 (e.g., one or more digital signal processors (DSPs)). Processor 604, processor 606, or both may be configured to generate sound identification data, update SEC model 114, or both. For example, in FIG. 6, processor 606 includes SEC engine 120. SEC engine 120 is configured to analyze audio data samples using one or more of SEC models 114.
[0071]
[0086] 6, device 100 also includes memory 608 and codec 624. Memory 608 stores instructions 660 executable by processor 604 or processor 606 to implement one or more operations described with reference to FIGS. 1-7. In one example, instructions 660 include or correspond to a feature extractor 108, a SEC engine 120, a scene detector 140, a drift detector 128, a model updater 152, a model checker 320, or a combination thereof. Memory 608 may also store configuration data 130, a SEC model 114, and model update data 148.
[0072]
[0087] 6, the speaker 622 and the microphone 104 may be coupled to a codec 624. In particular aspects, the microphone 104 is configured to receive audio representative of an acoustic environment associated with the device 100 and to generate an audio signal that the feature extractor uses to generate audio data samples. In the example shown in FIG. 6, the codec 624 includes a digital-to-analog converter (DAC 626) and an analog-to-digital converter (ADC 628). In particular implementations, the codec 624 receives an analog signal from the microphone 104, converts the analog signal to a digital signal using the ADC 628, and provides the digital signal to the processor 606. In particular implementations, the processor 606 provides the digital signal to the codec 624, which converts the digital signal to an analog signal using the DAC 626, and provides the analog signal to the speaker 622.
[0073]
[0088] In FIG. 6 , the device 100 also includes an input device 136. The device 100 may also include a display 620 coupled to the display controller 610. In particular aspects, the input device 136 includes a sensor, a keyboard, a pointing device, etc. In some implementations, the input device 136 and the display 620 are combined in a touchscreen or similar touch- or motion-sensitive display. The input device 136 may be used to provide indications associated with the unknown data 146 to generate the training data 302. The input device 136 may also be used to initiate a model update operation, such as to initiate the model update process described with reference to FIG. 2 or the model update process described with reference to FIG. 3. In some implementations, the input device 136 may also or alternatively be used to select a particular SEC model of the available SEC models 114 to be used by the SEC engine 120. In certain aspects, the input device 136 may be used to configure the configuration data 130, which may be used to select a SEC model to be used by the SEC engine 120, to determine the audio scene 142, or both. The display 620 may be used to display the results of an analysis by one of the SEC models (e.g., output 124 of FIG. 1 ), to prompt the user to provide an indication associated with the unknown data 146, or both.
[0074]
[0089] In some implementations, the device 100 also includes a modem 612 coupled to a transceiver 614. In Figure 6, the transceiver 614 is coupled to an antenna 616 to enable wireless communication with other devices, such as a remote computing device 618. In other examples, the transceiver 614 is also or alternatively coupled to a communication port (e.g., an Ethernet port) to enable wired communication with other devices, such as a remote computing device 618.
[0075]
[0090] 6, device 100 includes clock 132 and sensor 134. As particular examples, sensor 134 includes one or more cameras 650, one or more position sensors 652, microphone 104, other sensors 654, or a combination thereof.
[0076]
[0091] In certain aspects, the clock 132 generates a clock signal that can be used to assign a timestamp to a particular audio data sample to indicate when the particular audio data sample was received. In this aspect, the SEC engine 120 can use the timestamp to select a particular SEC model 114 to use to analyze the particular audio data sample. Additionally or alternatively, the timestamp can be used by the scene detector 140 to determine an audio scene 142 associated with the particular audio data sample.
[0077]
[0092] In particular aspects, the camera 650 generates image data, video data, or both. The SEC engine 120 can use the image data, video data, or both to select a particular SEC model 114 to use to analyze the audio data sample. Additionally or alternatively, the image data, video data, or both can be used by the scene detector 140 to determine an audio scene 142 associated with a particular audio data sample. For example, a particular SEC model 114 can be designated for outdoor use, and the image data, video data, or both can be used to confirm that the device 100 is located in an outdoor environment.
[0078]
[0093] In particular aspects, the position sensor 652 generates position data, such as global position data, that indicates the location of the device 100. The SEC engine 120 can use the position data to select a particular SEC model 114 to use to analyze the audio data sample. Additionally or alternatively, the position data can be used by the scene detector 140 to determine an audio scene 142 associated with a particular audio data sample. For example, a particular SEC model 114 can be designated for use at home, and the position data can be used to confirm that the device 100 is located at the home location. The position sensor 652 can include a receiver for a satellite-based positioning system, a receiver for a local positioning system receiver, an inertial navigation system, a landmark-based positioning system, or a combination thereof.
[0079]
[0094] Other sensors 654 may be coupled to or included within device 100 and may include, for example, an orientation sensor, a magnetometer, a light sensor, a contact sensor, a temperature sensor, or any other sensor that may be used to generate scene data 138 useful in determining the audio scene 142 associated with device 100 at a particular time.
[0080]
[0095] In particular implementations, the device 100 is included in a system-in-package or system-on-chip device 602. In one particular implementation, the memory 608, the processor 604, the processor 606, the display controller 610, the codec 624, the modem 612, and the transceiver 614 are included in the system-in-package or system-on-chip device 602. In particular implementations, the input device 136 and the power supply 630 are coupled to the system-on-chip device 602. Further, in particular implementations, the display 620, the input device 136, the speaker 622, the sensor 134, the clock 132, the antenna 616, and the power supply 630 are external to the system-on-chip device 602, as shown in FIG. 6 . In particular implementations, each of the display 620, the input device 136, the speaker 622, the sensor 134, the clock 132, the antenna 616, and the power supply 630 may be coupled to a component of the system-on-chip device 602, such as an interface or a controller.
[0081]
[0096] Device 100 may include, correspond to, or be included within a voice-activated device, an audio device, a wireless speaker and voice-activated device, a portable electronic device, a car, a vehicle, a computing device, a communication device, an Internet of Things (IoT) device, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, a smart speaker, a mobile computing device, a mobile communication device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a game console, an appliance, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, or any combination thereof. In certain aspects, processor 604, processor 606, or a combination thereof, is included in an integrated circuit.
[0082]
[0097] 7 is an illustrative example of a vehicle 700 incorporating aspects of device 100 of FIG. 1 . According to one implementation, vehicle 700 is an autonomous vehicle. According to other implementations, vehicle 700 is a car, truck, motorcycle, aircraft, watercraft, etc. In FIG. 7 , vehicle 700 includes display 620, one or more of sensors 134, device 100, or a combination thereof. Sensors 134 and device 100 are shown using dotted lines to indicate that these components may not be visible to passengers of vehicle 700. Device 100 may be integrated into or coupled to vehicle 700.
[0083]
[0098] In particular aspects, device 100 is coupled to a display 620 and provides an output to display 620 in response to one of SEC models 114 detecting or recognizing various events (e.g., sound events) described herein. For example, device 100 provides output 124 of FIG. 1 to display 620 indicating the sound class of a sound 102 (e.g., a car horn) recognized by one of SEC models 114 in an audio data sample 110 received from microphone 104. In some implementations, device 100 can perform an action in response to recognizing a sound event, such as alerting a vehicle operator or activating one of sensors 134. In particular examples, device 100 provides an output indicating whether an action is being performed in response to the recognized sound event. In particular aspects, a user can select an option displayed on display 620 to enable or disable the performance of an action in response to a recognized sound event.
[0084]
[0099] 1, a vehicle occupancy sensor, an eye tracking sensor, a position sensor 652, or an external environmental sensor (e.g., a LIDAR sensor or a camera). In certain aspects, the sensor input of the sensor 134 indicates the location of the user. For example, the sensor 134 is associated with various locations within the vehicle 700.
[0085]
[0100] 7 includes a SEC model 114, a SEC engine 120, a drift detector 128, a scene detector 140, and a model updater 152. In other implementations, the device 100 omits the model updater 152 when installed in or used in a vehicle 700. By way of example, the model update data 148 of FIG. 6 may be sent to the remote computing device 618, which may update one of the SEC models 114 based on the model update data 148. In such implementations, the updated SEC model 114 may be downloaded to the vehicle 700 for use by the SEC engine 120. In some implementations, the device 100 further includes the model checker 320 of FIG. 3 when installed in or used in a vehicle 700.
[0086]
[0101] 1-5 thus enable a user of the vehicle 700 to generate an updated SEC model to account for drift characteristic of (and possibly inherent to) the acoustic environment in which the device 100 operates. In some implementations, the device 100 can generate an updated SEC model without user intervention. Furthermore, the techniques described with respect to FIGS. 1-5 enable a user of the vehicle 700 to generate an updated SEC model to detect one or more new sound classes. Furthermore, the SEC model may be updated without excessive use of the computing resources onboard the vehicle 700. For example, the vehicle 700 need not store in local memory all of the training data used to train the SEC model from scratch.
[0087]
[0102] FIG. 8 shows an example of device 100 coupled to or integrated into a headset 802, such as a virtual reality headset, an augmented reality headset, a mixed reality headset, an extended reality headset, a head-mounted display, or a combination thereof. A visual interface device, such as a display 620, is positioned in front of a user's eyes to enable the display of augmented reality, mixed reality, or virtual reality images or scenes to the user while headset 802 is worn. In a particular example, display 620 is configured to display an output of device 100, such as an indication of a recognized sound event (e.g., output 124 of FIG. 1 ). Headset 802 includes sensors 134, such as microphone 104 of FIG. 1 , camera 650 of FIG. 6 , position sensor 652 of FIG. 6 , other sensors 654 of FIG. 6 , or a combination thereof. Although shown in a single location, in other implementations, the sensor 134 may be located in other locations on the headset 802, such as an array of one or more microphones and one or more cameras distributed around the headset 802 to detect multimodal input.
[0088]
[0103] The sensors 134 enable detection of audio data that the device 100 uses to detect sound events or update the SEC models 114. For example, the SEC engine 120 uses one or more of the SEC models 114 to generate sound event classification data that can be provided to the display 620 to indicate that a recognized sound event, such as a car horn, was detected in an audio data sample received from the sensors 134. In some implementations, the device 100 can perform an action in response to recognizing a sound event, such as activating a camera or another one of the sensors 134 or providing haptic feedback to the user.
[0089]
[0104] 8, device 100 includes SEC model 114, SEC engine 120, drift detector 128, scene detector 140, and model updater 152. In other implementations, device 100 omits model updater 152 when installed in or used within headset 802. To illustrate, model update data 148 of FIG. 6 may be sent to remote computing device 618, which may update one of SEC models 114 based on the model update data 148. In such implementations, updated SEC model 114 may be downloaded to headset 802 for use by SEC engine 120. In some implementations, device 100 further includes model checker 320 of FIG. 3 when installed in or used within headset 802.
[0090]
[0105] FIG. 9 shows an example of device 100 integrated into a wearable electronic device 902, shown as a “smart watch,” that includes a display 620 and sensors 134. The sensors 134 enable detection of user input and audio scenes based on modalities such as location, video, speech, and gesture, for example. The sensors 134 also enable detection of sounds in the acoustic environment around the wearable electronic device 902, which the device 100 uses to detect sound events or update the SEC model 114. For example, the device 100 provides the output 124 of FIG. 1 to the display 620 indicating that a recognized sound event is detected in the audio data samples received from the sensors 134. In some implementations, the device 100 can perform an action in response to recognizing a sound event, such as activating the camera or another one of the sensors 134 or providing haptic feedback to the user.
[0091]
[0106] 9, device 100 includes SEC model 114, SEC engine 120, drift detector 128, scene detector 140, and model updater 152. In other implementations, device 100 omits model updater 152 when installed in or used within wearable electronic device 902. To illustrate, model update data 148 of FIG. 6 may be sent to remote computing device 618, which may update one of SEC models 114 based on the model update data 148. In such implementations, updated SEC model 114 may be downloaded to wearable electronic device 902 for use by SEC engine 120. In some implementations, device 100 further includes model checker 320 of FIG. 3 when installed in or used within wearable electronic device 902.
[0092]
[0107] 10 is an illustrative example of a voice-controlled speaker system 1000. The voice-controlled speaker system 1000 may have wireless network connectivity and is configured to perform auxiliary operations. In FIG. 10, the device 100 is included in the voice-controlled speaker system 1000. The voice-controlled speaker system 1000 also includes a speaker 1002 and a sensor 134. The sensor 134 includes the microphone 104 of FIG. 1 for receiving voice input or other audio input.
[0093]
[0108] During operation, in response to receiving a verbal command or a recognized sound event, the voice-controlled speaker system 1000 can perform auxiliary actions. The auxiliary actions can include adjusting the temperature, playing music, turning on lights, etc. The sensors 134 enable the device 100 to detect sound events or to detect audio data samples that are used to update one or more of the SEC models 114. Additionally, the voice-controlled speaker system 1000 can perform several actions based on sound events recognized by the device 100. For example, if the device 100 recognizes the sound of a door closing, the voice-controlled speaker system 1000 can turn on one or more lights.
[0094]
[0109] 10 , device 100 includes SEC model 114, SEC engine 120, drift detector 128, scene detector 140, and model updater 152. In other implementations, device 100 omits model updater 152 when installed in or used within a voice-controlled speaker system 1000. To illustrate, model update data 148 of FIG. 6 may be sent to remote computing device 618, which may update one of SEC models 114 based on the model update data 148. In such implementations, updated SEC model 114 may be downloaded to voice-controlled speaker system 1000 for use by SEC engine 120. In some implementations, device 100 further includes model checker 320 of FIG. 3 when installed in or used within a voice-controlled speaker system 1000.
[0095]
[0110] FIG. 11 illustrates a camera 1100 incorporating aspects of the device 100 of FIG. 1. In FIG. 11, the device 100 is incorporated into or coupled to the camera 1100. The camera 1100 includes an image sensor 1102 and one or more other sensors (e.g., sensor 134), such as the microphone 104 of FIG. 1. Additionally, the camera 1100 includes the device 100 configured to identify sound events based on audio data samples and update one or more of the SEC models 114. In certain aspects, the camera 1100 is configured to perform one or more actions in response to the recognized sound event. For example, the camera 1100 may cause the image sensor 1102 to capture an image in response to the device 100 detecting a particular sound event in the audio data samples from the sensor 134.
[0096]
[0111] 11 , device 100 includes SEC model 114, SEC engine 120, drift detector 128, scene detector 140, and model updater 152. In other implementations, device 100 omits model updater 152 when installed in or used within camera 1100. To illustrate, model update data 148 of FIG. 6 may be sent to remote computing device 618, which may update one of SEC models 114 based on the model update data 148. In such implementations, updated SEC model 114 may be downloaded to camera 1100 for use by SEC engine 120. In some implementations, device 100 further includes model checker 320 of FIG. 3 when installed in or used within camera 1100.
[0097]
[0112] FIG. 12 illustrates a mobile device 1200 incorporating aspects of the device 100 of FIG. 1. In FIG. 12, the mobile device 1200 includes or is coupled to the device 100 of FIG. 1. The mobile device 1200 includes, as illustrative, non-limiting examples, a phone or a tablet. The mobile device 1200 includes a display 620 and a sensor 134, such as the microphone 104 of FIG. 1, the camera 650 of FIG. 6, the position sensor 652 of FIG. 6, or other sensor 654 of FIG. 6. In operation, the mobile device 1200 may perform a particular action in response to the device 100 recognizing a particular sound event. For example, the action may include sending a command to another device, such as a thermostat, a home automation system, another mobile device, or the like.
[0098]
[0113] 12, the device 100 includes the SEC model 114, the SEC engine 120, the drift detector 128, the scene detector 140, and the model updater 152. In other implementations, the device 100 omits the model updater 152 when installed in or used within a mobile device 1200. To illustrate, the model update data 148 of FIG. 6 may be sent to the remote computing device 618, which may update one of the SEC models 114 based on the model update data 148. In such implementations, the updated SEC model 114 may be downloaded to the mobile device 1200 for use by the SEC engine 120. In some implementations, the device 100 further includes the model checker 320 of FIG. 3 when installed in or used within a mobile device 1200.
[0099]
[0114] FIG. 13 illustrates a flight device 1300 incorporating aspects of the device 100 of FIG. 1. In FIG. 13, the flight device 1300 includes or is coupled to the device 100 of FIG. 1. The flight device 1300 is a manned, unmanned, or remotely piloted flight device (e.g., a package delivery drone). The flight device 1300 includes a control system 1302 and sensors 134, such as the microphone 104 of FIG. 1, the camera 650 of FIG. 6, the position sensor 652 of FIG. 6, or other sensors 654 of FIG. 6. The control system 1302 controls various operations of the flight device 1300, such as cargo release, sensor activation, takeoff, navigation, landing, or a combination thereof. For example, the control system 1302 may control the flight of the flight device 1300 between a designated point and the deployment of cargo at a particular location. In certain aspects, the control system 1302 performs one or more actions in response to the detection of a particular sound event by the device 100. To illustrate, the control system 1302 may initiate a safe landing protocol in response to the device 100 detecting an aircraft engine.
[0100]
[0115] In the example shown in FIG. 13 , device 100 includes SEC model 114, SEC engine 120, drift detector 128, scene detector 140, and model updater 152. In other implementations, device 100 omits model updater 152 when installed in or used within flight device 1300. To illustrate, model update data 148 of FIG. 6 may be sent to remote computing device 618, which may update one of SEC models 114 based on the model update data 148. In such implementations, updated SEC model 114 may be downloaded to flight device 1300 for use by SEC engine 120. In some implementations, device 100 further includes model checker 320 of FIG. 3 when installed in or used within flight device 1300.
[0101]
[0116] FIG. 14 illustrates a headset 1400 incorporating aspects of the device 100 of FIG. 1. In FIG. 14, the headset 1400 includes or is coupled to the device 100 of FIG. 1. The headset 1400 includes the microphone 104 of FIG. 1 positioned to primarily capture the user's voice. The headset 1400 may also include one or more additional microphones positioned to primarily capture environmental sounds (e.g., for noise cancellation operation) and one or more of the sensors 134, such as the camera 650, the position sensor 652, or other sensors 654 of FIG. 6. In particular aspects, the headset 1400 performs one or more actions in response to the detection of a particular sound event by the device 100. By way of example, the headset 1400 may activate a noise cancellation feature in response to the device 100 detecting a gunshot. The headset 1400 may also update one or more of the SEC models 114.
[0102]
[0117] 14, device 100 includes SEC model 114, SEC engine 120, drift detector 128, scene detector 140, and model updater 152. In other implementations, device 100 omits model updater 152 when installed in or used within headset 1400. To illustrate, model update data 148 of FIG. 6 may be sent to remote computing device 618, which may update one of SEC models 114 based on the model update data 148. In such implementations, updated SEC model 114 may be downloaded to headset 1400 for use by SEC engine 120. In some implementations, device 100 further includes model checker 320 of FIG. 3 when installed in or used within headset 1400.
[0103]
[0118] FIG. 15 illustrates an appliance 1500 incorporating aspects of the device 100 of FIG. 1. In FIG. 15, the appliance 1500 is a lamp, although in other implementations, the appliance 1500 includes another Internet of Things appliance, such as a refrigerator, a coffee maker, an oven, or another household appliance. The appliance 1500 includes or is coupled to the device 100 of FIG. 1. The appliance 1500 includes a sensor 134, such as the microphone 104 of FIG. 1, the camera 650 of FIG. 6, the position sensor 652 of FIG. 6, or another sensor 654 of FIG. 6. In certain aspects, the appliance 1500 performs one or more actions in response to the detection of a particular sound event by the device 100. For illustrative purposes, the appliance 1500 may activate a light in response to the device 100 detecting a door closing. The appliance 1500 may also update one or more of the SEC models 114.
[0104]
[0119] 15, device 100 includes SEC model 114, SEC engine 120, drift detector 128, scene detector 140, and model updater 152. In other implementations, device 100 omits model updater 152 when installed in or used within appliance 1500. To illustrate, model update data 148 of FIG. 6 may be sent to remote computing device 618, which may update one of SEC models 114 based on the model update data 148. In such implementations, updated SEC model 114 may be downloaded to appliance 1500 for use by SEC engine 120. In some implementations, device 100 further includes model checker 320 of FIG. 3 when installed in or used within appliance 1500.
[0105]
[0120] Figure 16 is a flow chart illustrating an example method 1600 of operation of device 100 of Figure 1. Method 1600 may be initiated, controlled, or performed by device 100. For example, processor 604 or 606 of Figure 6 may execute instructions 660 from memory 608 to cause drift detector 128 to generate model update data 148.
[0106]
[0121] At block 1602, the method 1600 includes providing an audio data sample as an input to a sound event classification model. For example, the SEC engine 120 of FIG. 1 (or the processor 606 of FIG. 6 executing instructions 660 corresponding to the SEC engine 120) may provide the audio data sample 110 as an input to the SEC model 112. In some implementations, the method 1600 also includes capturing audio data corresponding to the audio data sample. For example, the microphone 104 of FIG. 1 may generate the audio signal 106 based on the sound 102 detected by the microphone 104. Furthermore, in some implementations, the method 1600 includes selecting a sound event classification model from among a plurality of sound event classification models stored in a memory. For example, the SEC engine 120 of FIG. 1 (or the processor 606 of FIG. 6 executing instructions 660 corresponding to the SEC engine 120) may select a SEC model 112 from among the available SEC models 114 based on sensor data associated with the audio data sample, based on input identifying an audio scene or a SEC model 112, based on when the audio data sample was received, based on configuration data, or a combination thereof.
[0107]
[0122] At block 1604, the method 1600 includes determining, based on the output of the sound event classification model in response to the audio data sample, whether the sound class of the audio data sample was recognized by the sound event classification model. For example, the SEC engine 120 of FIG. 1 (or the processor 606 of FIG. 6 executing instructions 660 corresponding to the SEC engine 120) may determine whether the sound class 122 of the audio data sample 110 was recognized by the SEC model 112. By way of example, the SEC model 112 may generate a confidence metric associated with each sound class that the SEC model 112 is trained to recognize, and the determination of whether the sound class was recognized by the SEC model may be based on the value of the confidence metric. In certain aspects, based on a determination that the sound class 122 was recognized by the SEC model 112, the SEC engine 120 generates an output 124 indicating the sound class 122 associated with the audio data sample 110.
[0108]
[0123] At block 1606, method 1600 includes determining whether the sound event classification model corresponds to an audio scene associated with the audio data sample based on the determination that the sound class was not recognized. For example, drift detector 128 of FIG. 1 (or processor 606 of FIG. 6 executing instructions 660 corresponding to drift detector 128) may determine whether SEC model 112 corresponds to an audio scene 142 associated with audio data sample 110. By way of example, scene detector 140 may determine audio scene 142 based on input received via its input device 136, based on sensor data from sensor 134, based on a timestamp associated with audio data sample 110 indicated by clock 132, based on configuration data 130, or a combination thereof. In implementations in which SEC model 112 is selected from among available SEC models, the determination of audio scene 142 may be based on information different from the information used to select SEC model 112.
[0109]
[0124] At block 1608, the method 1600 includes storing model update data based on the audio data sample based on a determination that the sound event classification model corresponds to an audio scene associated with the audio data sample. For example, the drift detector 128 of FIG. 1 (or the processor 606 of FIG. 6 executing instructions 660 corresponding to the drift detector 128) may store the model update data 148 based on the audio data sample 110. In particular aspects, if the drift detector 128 determines that the SEC model 112 corresponds to the audio scene 142 associated with the audio data sample 110, the drift detector 128 stores the drift data 144 as the model update data 148, and if the drift detector 128 determines that the SEC model 112 does not correspond to the audio scene 142 associated with the audio data sample 110, the drift detector 128 stores the unknown data 146 as the model update data 148.
[0110]
[0125] Method 1600 may also include updating the SEC model based on the model update data. For example, model updater 152 of FIG. 1 (or processor 606 of FIG. 6 executing instructions 660 corresponding to model updater 152) may update SEC model 112 based on model update data 148 as described with reference to FIG. 2, as described with reference to FIG. 3, or both.
[0111]
[0126] In certain aspects, the method 1600 includes, after storing the model update data, determining whether a threshold amount of model update data has accumulated. For example, the model updater 152 may determine when the model update data 148 of FIG. 1 includes enough data to initiate update training of the SEC model 112 (e.g., a threshold amount of model update data 148 associated with a particular SEC model and a particular sound class). The method 1600 may further include initiating an automatic update of the sound event classification model using the accumulated model update data based on determining that the threshold amount of model update data has accumulated. For example, the model updater 152 may initiate the model update without input from a user. In certain implementations, the automatic update fine-tunes the SEC model 112 to generate an updated SEC model 154. For example, prior to the automatic update, the SEC model 112 was trained to recognize multiple variants of a particular sound class, and the automatic update modifies the SEC model 112 to enable the SEC model 112 to recognize additional variants of the particular sound class as corresponding to the particular sound class.
[0112]
[0127] Due to the way training is performed, SEC models are generally closed-set. That is, the number and types of sound classes that a SEC model can recognize are fixed and limited during training. After training, a SEC model generally has a static relationship between input and output. This static relationship between input and output means that the mapping learned during training will be valid in the future (e.g., when evaluating new data) and that the relationship between input and output data will not change. However, it is difficult to collect an exhaustive set of training samples for each sound class and to properly annotate all of the available training data to train a comprehensive and sophisticated SEC model.
[0113]
[0128] In contrast, during use, SEC models face the open set problem. For example, during use, a SEC model may be presented with data samples related to both known and unknown sound events. Furthermore, the distribution of sounds or sound features in each sound class that the SEC model is trained to recognize may change over time or may not be comprehensively represented in the available training data. For example, in the case of traffic sounds, differences in sound based on location, time, busy or unbusy intersections, etc. may not be explicitly captured in the training data for the traffic sound class. For these and other reasons, there may be discrepancies between the training data used to train the SEC model and the dataset to which the SEC model is presented during use. Such discrepancies (e.g., shifts or drifts in the dataset) depend on various factors, such as location, time, and the device capturing the sound signal. Shifts in the dataset can result in poor predictions resulting from the SEC model. The disclosed systems and methods overcome these and other problems by adapting the SEC model to detect such shifted data with little or no supervision. Furthermore, in some aspects, the SEC model can be updated to recognize new sound classes without forgetting the previously trained sound classes.
[0114]
[0129] In certain aspects, no SEC model training is performed while the system is operating in inference mode. Rather, during operation in inference mode, existing knowledge, in the form of one or more previously trained SEC models, is used to analyze detected sounds. More than one SEC model may be used to analyze sounds. For example, an ensemble of SEC models may be used during operation in inference mode. A particular SEC may be selected from a set of available SEC models based on the detection of a trigger condition. For illustrative purposes, a particular SEC model will be used as the active SEC model, sometimes referred to as the “source SEC model,” whenever a trigger is activated. The trigger may be based on location, sound, camera information, other sensor data, user input, etc. For example, a particular SEC model may be trained to recognize sound events related to crowded areas such as theme parks, outdoor shopping malls, public plazas, etc. In this example, a particular SEC model may be used as the active SEC model when global positioning data indicates that the sound the device is capturing is from one of these locations. In this example, the trigger is based on the location of the sound the device is capturing, and an active SEC model is selected and loaded (e.g., in addition to or instead of a previously active SEC model) when the device is detected to be in the location.
[0115]
[0130] In conjunction with the described implementations, the apparatus includes means for providing audio data samples to the sound event classification model. For example, the means for providing audio data samples to the sound event classification model includes device 100, instructions 660, processor 604, processor 606, SEC engine 120, feature extractor 108, microphone 104, codec 624, one or more other circuits or components configured to provide audio data samples to the sound event classification model, or any combination thereof.
[0116]
[0131] The apparatus also includes means for determining, based on the output of the sound classification model, whether the sound class of the audio data sample has been recognized by the sound event classification model. For example, the means for determining whether the sound class of the audio data sample has been recognized by the sound event classification model includes device 100, instructions 660, processor 604, processor 606, SEC engine 120, one or more other circuits or components configured to determine whether the sound class of the audio data sample has been recognized by the sound event classification model, or any combination thereof.
[0117]
[0132] The apparatus also includes means for determining, in response to determining that the sound class was not recognized, whether the sound event classification model corresponds to an audio scene associated with the audio data sample. For example, the means for determining whether the sound event classification model corresponds to an audio scene associated with the audio data sample may include device 100, instructions 660, processor 604, processor 606, drift detector 128, scene detector 140, one or more other circuits or components configured to determine whether the sound event classification model corresponds to an audio scene associated with the audio data sample, or any combination thereof.
[0118]
[0133] The apparatus also includes means for storing model update data based on the audio data sample in response to determining that the sound event classification model corresponds to an audio scene associated with the audio data sample. For example, the means for storing the model update data may include the remote computing device 618, the device 100, the instructions 660, the processor 604, the processor 606, the drift detector 128, the memory 608, one or more other circuits or components configured to store the model update data, or any combination thereof.
[0119]
[0134] In some implementations, the apparatus includes means for selecting a sound event classification model from among a plurality of sound event classification models based on a selection criterion. For example, the means for selecting a sound event classification model includes device 100, instructions 660, processor 604, processor 606, SEC engine 120, one or more other circuits or components configured to select a sound event classification model, or any combination thereof.
[0120]
[0135] In some implementations, the instructions include means for updating the sound event classification model based on the model update data. For example, the means for updating the sound event classification model based on the model update data includes the remote computing device 618, the device 100, the instructions 660, the processor 604, the processor 606, the model updater 152, the model checker 320, one or more other circuits or components configured to update the sound event classification model, or any combination thereof.
[0121]
[0136] Furthermore, those skilled in the art will appreciate that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the implementations disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or a combination of both. The various illustrative components, blocks, configurations, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or as processor-executable instructions depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0122]
[0137] The steps of a method or algorithm described in connection with the implementations disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, a hard disk, a removable disk, a compact disk read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. Alternatively, the processor and the storage medium may reside as discrete components in a computing device or a user terminal.
[0123]
[0138] Certain aspects of the disclosure are described below in a first set of interrelated clauses.
[0124]
[0139] According to clause 1, the device includes one or more processors configured to provide audio data samples to a sound event classification model. The one or more processors are further configured to determine whether a sound class of the audio data sample has been recognized by the sound event classification model based on an output of the sound event classification model in response to the audio data sample. The one or more processors are also configured to determine whether the sound event classification model corresponds to an audio scene associated with the audio data sample based on a determination that the sound class has not been recognized. The one or more processors are further configured to store model update data based on the audio data sample based on a determination that the sound event classification model corresponds to an audio scene associated with the audio data sample.
[0125]
[0140] Clause 2 includes the device of clause 1, further including a microphone coupled to the one or more processors and configured to capture audio data corresponding to the audio data samples.
[0126]
[0141] Clause 3 includes the device of clause 1 or clause 2, further including a memory coupled to one or more processors and configured to store a plurality of sound event classification models, wherein the one or more processors are configured to select a sound event classification model from among the plurality of sound event classification models.
[0127]
[0142] Clause 4 includes the device of clause 3, further including one or more sensors configured to generate sensor data related to the audio data samples, wherein the one or more processors are configured to select a sound event classification model based on the sensor data.
[0128]
[0143] Clause 5 includes the device of clause 4, wherein the one or more sensors include a camera and a position sensor.
[0129]
[0144] Clause 6 includes the device of any of clauses 3 to 5, further including one or more input devices configured to receive input identifying an audio scene, wherein the one or more processors are configured to select a sound event classification model based on the audio scene.
[0130]
[0145] Clause 7 includes the device of any of clauses 3 to 6, wherein the one or more processors are configured to select a sound event classification model based on when the audio data sample is received.
[0131]
[0146] Clause 8 includes the device of any of clauses 3 to 8, wherein the memory further stores configuration data indicating one or more device settings, and the one or more processors are configured to select a sound event classification model based on the configuration data.
[0132]
[0147] Clause 9 includes the device of any of clauses 1 to 8, wherein the one or more processors are further configured to generate an output indicating a sound class associated with the audio data sample based on a determination that the sound class has been recognized.
[0133]
[0148] Clause 10 includes the device of any of clauses 1 to 9, wherein the one or more processors are further configured to store audio data corresponding to the audio data sample as training data for a new sound event classification model based on a determination that the sound event classification model does not correspond to an audio scene associated with the audio data sample.
[0134]
[0149] Clause 11 includes the device of any of clauses 1 to 10, wherein the sound event classification model is further configured to generate a confidence metric associated with the output, and the one or more processors are configured to determine whether the sound class was recognized by the sound event classification model based on the confidence metric.
[0135]
[0150] Clause 12 includes the device of any of clauses 1 to 11, wherein the one or more processors are further configured to update the sound event classification model based on the model update data.
[0136]
[0151] Clause 13 includes the device of any of clauses 1 to 12, further including one or more input devices configured to receive input identifying an audio scene, wherein the one or more processors are configured to determine whether the sound event classification model corresponds to the audio scene based on the input.
[0137]
[0152] Clause 14 includes the device of any of clauses 1 to 13, further including one or more sensors configured to generate sensor data associated with the audio data samples, wherein the one or more processors are configured to determine whether the sound event classification model corresponds to the audio scene based on the sensor data.
[0138]
[0153] Clause 15 includes the device of clause 14, wherein the one or more sensors include a camera and a position sensor.
[0139]
[0154] Clause 16 includes the device of clause 14 or clause 15, wherein the one or more processors are further configured to determine whether the sound event classification model corresponds to an audio scene based on timestamps associated with the audio data samples.
[0140]
[0155] Clause 17 includes the device of any of clauses 1 to 16, wherein the one or more processors are integrated within a mobile computing device.
[0141]
[0156] Clause 18 includes the device of any of clauses 1 to 16, wherein the one or more processors are integrated within a vehicle.
[0142]
[0157] Clause 19 includes the device of any of clauses 1 to 16, wherein the one or more processors are integrated within a wearable device.
[0143]
[0158] Clause 20 includes the device of any of clauses 1 to 16, wherein the one or more processors are integrated within an augmented reality headset, a mixed reality headset, or a virtual reality headset.
[0144]
[0159] Clause 21 includes the device of any of clauses 1 to 20, wherein the one or more processors are included in an integrated circuit.
[0145]
[0160] Clause 22 includes the device of any of clauses 1 to 21, wherein the sound event classification model is trained to recognize a particular sound class, and the model update data includes drift data representing variations in characteristics of a sound within the particular sound class that the sound event classification model is not trained to recognize as corresponding to the particular sound class.
[0146]
[0161] Certain aspects of the disclosure are described below in a second set of interrelated clauses.
[0147]
[0162] According to clause 23, the method includes providing, by one or more processors, an audio data sample as input to a sound event classification model. The method also includes determining, by the one or more processors, whether a sound class of the audio data sample was recognized by the sound event classification model based on an output of the sound event classification model responsive to the audio data sample. The method further includes determining, by the one or more processors, whether the sound event classification model corresponds to an audio scene associated with the audio data sample based on a determination that the sound class was not recognized. The method also includes storing, by the one or more processors, model update data based on the audio data sample based on a determination that the sound event classification model corresponds to an audio scene associated with the audio data sample.
[0148]
[0163] Clause 24 includes the method of clause 23, further including selecting a sound event classification model from among a plurality of sound event classification models stored in a memory coupled to the one or more processors.
[0149]
[0164] Clause 25 includes the method of clause 24, wherein the sound event classification model is selected based on user input, configuration data, location data, image data, video data, timestamps associated with the audio data samples, or a combination thereof.
[0150]
[0165] Clause 26 includes the method of any of clauses 23 to 25, wherein determining whether the sound event classification model corresponds to an audio scene is based on a confidence metric generated by the sound event classification model, user input, configuration data, location data, image data, video data, timestamps associated with the audio data samples, or a combination thereof.
[0151]
[0166] Clause 27 includes the method of any of clauses 23 to 26, further including capturing audio data corresponding to the audio data samples.
[0152]
[0167] Clause 28 includes the method of any of clauses 23 to 27, further including selecting a sound event classification model from among a plurality of available sound event classification models.
[0153]
[0168] Clause 29 includes the method of any of clauses 23 to 27, further including receiving sensor data associated with the audio data sample; and selecting a sound event classification model from among a plurality of available sound event classification models based on the sensor data.
[0154]
[0169] Clause 30 includes the method of any of clauses 23 to 27, further including receiving input identifying an audio scene and selecting a sound event classification model from among a plurality of available sound event classification models based on the audio scene.
[0155]
[0170] Clause 31 includes the method of any of clauses 23 to 27, further including selecting a sound event classification model from among a plurality of available sound event classification models based on when the audio data sample is received.
[0156]
[0171] Clause 32 includes the method of any of clauses 23-27, further including selecting a sound event classification model from among a plurality of available sound event classification models based on the configuration data.
[0157]
[0172] Clause 33 includes the method of any of clauses 23 to 32, further including generating an output indicative of a sound class associated with the audio data sample based on determining that the sound class has been recognized.
[0158]
[0173] Clause 34 includes the method of any of clauses 23 to 33, further including, based on a determination that the sound event classification model does not correspond to an audio scene associated with the audio data sample, storing audio data corresponding to the audio data sample as training data for a new sound event classification model.
[0159]
[0174] Clause 35 includes the method of any of clauses 23 to 34, wherein the output of the sound event classification model includes a confidence metric, and the method further includes determining whether the sound class was recognized by the sound event classification model based on the confidence metric.
[0160]
[0175] Clause 36 includes the method of any of clauses 23 to 35, further including updating the sound event classification model based on the model update data.
[0161]
[0176] Clause 37 includes the method of any of clauses 23 to 36, further including receiving an input identifying an audio scene, wherein determining whether the sound event classification model corresponds to an audio scene is based on the input.
[0162]
[0177] Clause 38 includes the method of any of clauses 23 to 37, further including receiving sensor data associated with the audio data samples, wherein determining whether the sound event classification model corresponds to the audio scene is based on the sensor data.
[0163]
[0178] Clause 39 includes the method of any of clauses 23 to 38, wherein determining whether the sound event classification model corresponds to an audio scene is based on timestamps associated with the audio data samples.
[0164]
[0179] Clause 40 includes the method of any of clauses 23 to 39, further including, after storing the model update data, determining whether a threshold amount of model update data has been accumulated, and based on a determination that the threshold amount of model update data has been accumulated, initiating an automatic update of the sound event classification model using the accumulated model update data.
[0165]
[0180] Clause 41 includes the method of any of clauses 23 to 40, wherein, prior to the automatic updating, the sound event classification model is trained to recognize multiple variants of a particular sound class, and the automatic updating modifies the sound event classification model to enable the sound event classification model to recognize additional variants of the particular sound class as corresponding to the particular sound class.
[0166]
[0181] Certain aspects of the disclosure are described below in a third set of interrelated clauses.
[0167]
[0182] According to clause 42, the device includes means for providing an audio data sample to a sound event classification model. The device also includes means for determining, based on an output of the sound event classification model, whether a sound class of the audio data sample has been recognized by the sound event classification model. The device further includes means for determining, in response to a determination that the sound class has not been recognized, whether the sound event classification model corresponds to an audio scene associated with the audio data sample. The device also includes means for storing model update data based on the audio data sample, in response to a determination that the sound event classification model corresponds to an audio scene associated with the audio data sample.
[0168]
[0183] Clause 43 includes the device of clause 42, further including means for selecting a sound event classification model from among the plurality of sound event classification models based on a selection criterion.
[0169]
[0184] Clause 44 includes the device of clause 42 or clause 43, further including means for updating the sound event classification model based on the model update data.
[0170]
[0185] Clause 45 includes the device of any of clauses 42 to 44, further including means for capturing audio data corresponding to the audio data samples.
[0171]
[0186] Clause 46 includes the device of any of clauses 42 to 45, further including means for storing a plurality of sound event classification models and means for selecting a sound event classification model from among the plurality of sound event classification models.
[0172]
[0187] Clause 47 includes the device of clause 46, further including means for receiving an input identifying an audio scene, wherein the sound event classification model is selected based on the input identifying the audio scene.
[0173]
[0188] Clause 48 includes the device of clause 46, further including means for determining when the audio data sample is received, wherein the sound event classification model is selected based on when the audio data sample is received.
[0174]
[0189] Clause 49 includes the device of clause 46, further including means for storing configuration data indicative of one or more device settings, wherein the sound event classification model is selected based on the configuration data.
[0175]
[0190] Clause 50 includes the device of any of clauses 42 to 49, further including means for generating an output indicative of a sound class associated with the audio data sample based on a determination that the sound class has been recognized.
[0176]
[0191] Clause 51 includes the device of any of clauses 42 to 50, further including means for storing audio data corresponding to the audio data sample as training data for a new sound event classification model based on a determination that the sound event classification model does not correspond to an audio scene associated with the audio data sample.
[0177]
[0192] Clause 52 includes the device of any of clauses 42 to 51, wherein the sound event classification model is further configured to generate a confidence metric associated with the output, and wherein the determination of whether the sound class has been recognized by the sound event classification model is based on the confidence metric.
[0178]
[0193] Clause 53 includes the device of any of clauses 42 to 52, further including means for updating the sound event classification model based on the model update data.
[0179]
[0194] Clause 54 includes the device of any of clauses 42 to 53, further including means for receiving input identifying an audio scene, wherein determining whether the sound event classification model corresponds to an audio scene is based on the input.
[0180]
[0195] Clause 55 includes the device of any of clauses 42 to 54, further including means for generating sensor data associated with the audio data samples, wherein determining whether the sound event classification model corresponds to the audio scene is based on the sensor data.
[0181]
[0196] Clause 56 includes the device of any of clauses 42 to 55, wherein determining whether the sound event classification model corresponds to an audio scene is based on timestamps associated with the audio data samples.
[0182]
[0197] Clause 57 includes the device of any of clauses 42 to 56, wherein the means for providing the audio data sample to the sound event classification model, the means for receiving an output of the sound event classification model, the means for determining whether a sound class of the audio data sample is recognized by the sound event classification model, the means for determining whether the sound event classification model corresponds to an audio scene associated with the audio data sample, and the means for storing model update data based on the audio data sample are integrated within a mobile computing device.
[0183]
[0198] Clause 58 includes the device of any of clauses 42 to 56, wherein the means for providing the audio data samples to the sound event classification model, the means for receiving an output of the sound event classification model, the means for determining whether a sound class of the audio data samples is recognized by the sound event classification model, the means for determining whether the sound event classification model corresponds to an audio scene associated with the audio data samples, and the means for storing model update data based on the audio data samples are integrated within a vehicle.
[0184]
[0199] Clause 59 includes the device of any of clauses 42 to 56, wherein the means for providing the audio data samples to the sound event classification model, the means for receiving an output of the sound event classification model, the means for determining whether a sound class of the audio data samples is recognized by the sound event classification model, the means for determining whether the sound event classification model corresponds to an audio scene associated with the audio data samples, and the means for storing model update data based on the audio data samples are integrated within the wearable device.
[0185]
[0200] Clause 60 includes the device of any of clauses 42 to 56, wherein the means for providing the audio data samples to the sound event classification model, the means for receiving an output of the sound event classification model, the means for determining whether a sound class of the audio data samples is recognized by the sound event classification model, the means for determining whether the sound event classification model corresponds to an audio scene associated with the audio data samples, and the means for storing model update data based on the audio data samples are integrated within an augmented reality headset, a mixed reality headset, or a virtual reality headset.
[0186]
[0201] Clause 61 includes the device of any of clauses 42 to 60, wherein the means for providing the audio data sample to the sound event classification model, the means for receiving the output of the sound event classification model, the means for determining whether a sound class of the audio data sample is recognized by the sound event classification model, the means for determining whether the sound event classification model corresponds to an audio scene associated with the audio data sample, and the means for storing model update data based on the audio data sample are comprised in an integrated circuit.
[0187]
[0202] Certain aspects of the present disclosure are described below in a fourth set of interrelated clauses.
[0188]
[0203] According to clause 62, a non-transitory computer-readable medium storing instructions executable by the processor to cause the processor to provide an audio data sample as input to a sound event classification model. The instructions are also executable by the processor to determine, based on an output of the sound event classification model in response to the audio data sample, whether a sound class of the audio data sample was recognized by the sound event classification model. The instructions are further executable by the processor to determine, based on a determination that the sound class was not recognized, whether the sound event classification model corresponds to an audio scene associated with the audio data sample. The instructions are further executable by the processor to store model update data based on the audio data sample, based on a determination that the sound event classification model corresponds to an audio scene associated with the audio data sample.
[0189]
[0204] Clause 63 includes the non-transitory computer-readable medium of clause 62, the instructions further causing the processor to update the sound event classification model based on the model update data.
[0190]
[0205] Clause 64 includes the non-transitory computer-readable medium of clause 62 or clause 63, the instructions further causing the processor to select a sound event classification model from among a plurality of sound event classification models stored in the memory.
[0191]
[0206] Clause 65 includes the non-transitory computer-readable medium of clause 64, the instructions causing the processor to select a sound event classification model based on the sensor data.
[0192]
[0207] Clause 66 includes the non-transitory computer-readable medium of clause 64, the instructions causing the processor to select a sound event classification model based on input identifying an audio scene associated with the audio data sample.
[0193]
[0208] Clause 67 includes the non-transitory computer-readable medium of clause 64, the instructions causing the processor to select a sound event classification model based on when the audio data sample is received.
[0194]
[0209] Clause 68 includes the non-transitory computer-readable medium of clause 64, the instructions causing the processor to select a sound event classification model based on the configuration data.
[0195]
[0210] Clause 69 includes the non-transitory computer-readable medium of any of clauses 62 to 68, wherein the instructions cause the processor to generate an output indicating a sound class associated with the audio data sample based on a determination that the sound class has been recognized.
[0196]
[0211] Clause 70 includes the non-transitory computer-readable medium of any of clauses 62 to 69, wherein the instructions cause the processor to store audio data corresponding to the audio data sample as training data for a new sound event classification model based on a determination that the sound event classification model does not correspond to an audio scene associated with the audio data sample.
[0197]
[0212] Clause 71 includes the non-transitory computer-readable medium of any of clauses 62 to 70, wherein the instructions cause the processor to generate a confidence metric associated with the output, and wherein the determination of whether the sound class was recognized by the sound event classification model is based on the confidence metric.
[0198]
[0213] Clause 72 includes the non-transitory computer-readable medium of any of clauses 62 to 71, wherein the instructions cause a processor to update the sound event classification model based on the model update data.
[0199]
[0214] Clause 73 includes the non-transitory computer-readable medium of any of clauses 62 to 72, wherein determining whether the sound event classification model corresponds to an audio scene is based on user input indicating the audio scene.
[0200]
[0215] Clause 74 includes the non-transitory computer-readable medium of any of clauses 62 to 73, wherein determining whether the sound event classification model corresponds to an audio scene is based on sensor data.
[0201]
[0216] Clause 75 includes the non-transitory computer-readable medium of any of clauses 62 to 74, wherein determining whether the sound event classification model corresponds to an audio scene is based on timestamps associated with the audio data samples.
[0202]
[0217] The above description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the disclosed embodiments. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other embodiments without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features as defined by the following claims. The inventions described in the claims of the present application as originally filed are set forth below. [C1] A device, one or more processors, the one or more processors providing audio data samples to a sound event classification model; determining whether a sound class of the audio data sample is recognized by the sound event classification model based on an output of the sound event classification model in response to the audio data sample; determining whether the sound event classification model corresponds to an audio scene associated with the audio data sample based on a determination that the sound class was not recognized; and storing model update data based on the audio data sample based on a determination that the sound event classification model corresponds to the audio scene associated with the audio data sample; and A device configured to: [C2] The device of C1, further comprising: a microphone coupled to the one or more processors and configured to capture audio data corresponding to the audio data samples. [C3] 10. The device of claim 1, further comprising: a memory coupled to the one or more processors and configured to store a plurality of sound event classification models, wherein the one or more processors are configured to select the sound event classification model from among the plurality of sound event classification models. [C4] The device of C3, further comprising: one or more sensors configured to generate sensor data associated with the audio data samples, wherein the one or more processors are configured to select the sound event classification model based on the sensor data. [C5] The device of C4, wherein the one or more sensors include a camera and a position sensor. [C6] The device of C3, further comprising: one or more input devices configured to receive input identifying the audio scene, wherein the one or more processors are configured to select the sound event classification model based on the audio scene. [C7] The device of C3, wherein the one or more processors are configured to select the sound event classification model based on when the audio data sample is received. [C8] The device of C3, wherein the memory further stores configuration data indicating one or more device settings, and the one or more processors are configured to select the sound event classification model based on the configuration data. [C9] The device of C1, wherein the one or more processors are further configured to generate an output indicating the sound class associated with the audio data sample based on a determination that the sound class is recognized. [C10] 10. The device of claim 1, wherein the one or more processors are further configured to: store audio data corresponding to the audio data sample as training data for a new sound event classification model based on a determination that the sound event classification model does not correspond to the audio scene associated with the audio data sample. [C11] 10. The device of claim 1, wherein the sound event classification model is further configured to generate a confidence metric associated with the output, and wherein the one or more processors are configured to determine whether the sound class was recognized by the sound event classification model based on the confidence metric. [C12] The device of C1, wherein the one or more processors are further configured to update the sound event classification model based on the model update data. [C13] The device of C1, further comprising: one or more input devices configured to receive input identifying the audio scene, wherein the one or more processors are configured to determine whether the sound event classification model corresponds to the audio scene based on the input. [C14] The device of C1, further comprising: one or more sensors configured to generate sensor data associated with the audio data samples, wherein the one or more processors are configured to determine whether the sound event classification model corresponds to the audio scene based on the sensor data. [C15] The device of C14, wherein the one or more sensors include a camera and a position sensor. [C16] The device of C14, wherein the one or more processors are further configured to determine whether the sound event classification model corresponds to the audio scene based on a timestamp associated with the audio data sample. [C17] 10. The device of claim 1, wherein the sound event classification model is trained to recognize a particular sound class, and the model update data includes drift data representing variations in characteristics of sounds within the particular sound class that the sound event classification model is not trained to recognize as corresponding to the particular sound class. [C18] The device of C1, wherein the one or more processors are integrated within a mobile computing device, a vehicle, a wearable device, an augmented reality headset, a mixed reality headset, or a virtual reality headset. [C19] The device of C1, wherein the one or more processors are comprised in an integrated circuit. [C20] 1. A method comprising: providing, by one or more processors, audio data samples as input to a sound event classification model; determining, by the one or more processors, whether a sound class of the audio data sample is recognized by the sound event classification model based on an output of the sound event classification model in response to the audio data sample; determining, by the one or more processors, whether the sound event classification model corresponds to an audio scene associated with the audio data sample based on a determination that the sound class was not recognized; storing, by the one or more processors, model update data based on the audio data sample based on a determination that the sound event classification model corresponds to the audio scene associated with the audio data sample; and A method for providing the above. [C21] The method of C20, further comprising selecting the sound event classification model from among a plurality of sound event classification models stored in a memory coupled to the one or more processors. [C22] The method of C21, wherein the sound event classification model is selected based on user input, configuration data, location data, image data, video data, timestamps associated with the audio data samples, or a combination thereof. [C23] The method of claim 20, wherein the determination of whether the sound event classification model corresponds to the audio scene is based on a confidence metric generated by the sound event classification model, user input, configuration data, location data, image data, video data, timestamps associated with the audio data samples, or a combination thereof. [C24] After storing the model update data, determining whether a threshold amount of model update data has been accumulated; initiating an automatic update of the sound event classification model using the accumulated model update data based on a determination that the threshold amount of model update data has been accumulated; and The method of C20, further comprising: [C25] The method of claim 24, wherein prior to the automatic updating, the sound event classification model was trained to recognize multiple variants of a particular sound class, and wherein the automatic updating modifies the sound event classification model to enable it to recognize additional variants of the particular sound class as corresponding to the particular sound class. [C26] A device, means for providing audio data samples to a sound event classification model; means for determining, based on the output of the sound event classification model, whether the sound class of the audio data sample has been recognized by the sound event classification model; means for determining, in response to determining that the sound class was not recognized, whether the sound event classification model corresponds to an audio scene associated with the audio data sample; means for storing model update data based on the audio data sample in response to determining that the sound event classification model corresponds to the audio scene associated with the audio data sample; and 1. A device comprising: [C27] The device of C26, further comprising means for selecting the sound event classification model from among a plurality of sound event classification models based on a selection criterion. [C28] The device of C26, further comprising means for updating the sound event classification model based on the model update data. [C29] A non-transitory computer-readable medium, comprising: providing an audio data sample as an input to a sound event classification model; and determining whether a sound class of the audio data sample is recognized by the sound event classification model based on an output of the sound event classification model in response to the audio data sample; determining whether the sound event classification model corresponds to an audio scene associated with the audio data sample based on a determination that the sound class was not recognized; and storing model update data based on the audio data sample based on a determination that the sound event classification model corresponds to the audio scene associated with the audio data sample; and A non-transitory computer-readable medium storing instructions executable by the processor to cause the [C30] 30. The non-transitory computer-readable medium of claim 29, wherein the instructions further cause the processor to update the sound event classification model based on the model update data.
Claims
1. A device, one or more processors, the one or more processors providing audio data samples to a sound event classification model; determining whether a sound class of the audio data sample is recognized by the sound event classification model based on an output of the sound event classification model in response to the audio data sample; determining whether the sound event classification model is trained to recognize sound events in an audio scene associated with the audio data sample based on a determination that the sound class was not recognized; storing model update data based on the audio data sample for updating the sound event classification model, the sound event classification model being trained to recognize a sound event in the audio scene associated with the audio data sample, and based on a determination that the sound class was not recognized; and A device configured to:
2. The device of claim 1 , further comprising: a microphone coupled to the one or more processors and configured to capture audio data corresponding to the audio data samples.
3. 10. The device of claim 1, further comprising: a memory coupled to the one or more processors and configured to store a plurality of sound event classification models, wherein the one or more processors are configured to select the sound event classification model from among the plurality of sound event classification models.
4. The device further comprises one or more sensors configured to generate sensor data related to the audio data samples, wherein the one or more processors are configured to select the sound event classification model based on the sensor data, and the one or more sensors include a camera and / or a position sensor; or the device further comprises one or more input devices configured to receive input identifying the audio scene, wherein the one or more processors are configured to select the sound event classification model based on the audio scene; or the one or more processors are configured to select the sound event classification model based on when the audio data sample is received; or 4. The device of claim 3, wherein the memory further stores configuration data indicative of one or more device settings, and the one or more processors are configured to select the sound event classification model based on the configuration data.
5. 10. The device of claim 1, wherein the one or more processors are further configured to generate an output indicative of the sound class associated with the audio data sample based on a determination that the sound class has been recognized.
6. 10. The device of claim 1, wherein the sound event classification model is further configured to generate a confidence metric associated with the output, and wherein the one or more processors are configured to determine whether the sound class was recognized by the sound event classification model based on the confidence metric.
7. The device of claim 1 , wherein the one or more processors are further configured to update the sound event classification model based on the model update data.
8. 10. The device of claim 1, further comprising: one or more input devices configured to receive input identifying the audio scene, wherein the one or more processors are configured to determine whether the sound event classification model is trained to recognize sound events in the audio scene based on the input.
9. one or more sensors configured to generate sensor data associated with the audio data samples, wherein the one or more processors are configured to determine whether the sound event classification model is trained to recognize sound events in the audio scene based on the sensor data, and the one or more sensors include a camera and a position sensor; or the one or more processors are further configured to determine whether the sound event classification model is trained to recognize sound events in the audio scene based on timestamps associated with the audio data samples. The device of claim 1 further comprising:
10. 2. The device of claim 1, wherein the sound event classification model is trained to recognize a particular sound class, and the model update data includes drift data representing variations in characteristics of sounds within the particular sound class that the sound event classification model is not trained to recognize as corresponding to the particular sound class.
11. The device of claim 1 , wherein the one or more processors are integrated within a mobile computing device, a vehicle, a wearable device, an augmented reality headset, a mixed reality headset, or a virtual reality headset.
12. 1. A method comprising: providing, by one or more processors, audio data samples as input to a sound event classification model; determining, by the one or more processors, whether a sound class of the audio data sample is recognized by the sound event classification model based on an output of the sound event classification model in response to the audio data sample; determining, by the one or more processors, based on a determination that the sound class was not recognized, whether the sound event classification model has been trained to recognize a sound event in an audio scene associated with the audio data sample; storing, by the one or more processors, model update data based on the audio data sample for updating the sound event classification model based on the audio data sample, wherein the sound event classification model has been trained to recognize a sound event in the audio scene associated with the audio data sample, and based on a determination that the sound class was not recognized; A method for providing the above.
13. After storing the model update data, determining whether a threshold amount of model update data has been accumulated; initiating an automatic update of the sound event classification model using the accumulated model update data based on a determination that the threshold amount of model update data has been accumulated; and The method of claim 12 , further comprising:
14. 14. The method of claim 13, wherein, prior to the automatic updating, the sound event classification model was trained to recognize multiple variants of a particular sound class, and wherein the automatic updating modifies the sound event classification model to enable it to recognize additional variants of the particular sound class as corresponding to the particular sound class.
15. 15. A non-transitory computer readable medium storing instructions executable by a processor to cause the processor to perform the method of any one of claims 12 to 14.
Citation Information
Patent Citations
Anomaly detection apparatus
JP2009259020A
Acoustic and Other Waveform Event Detection and Correction Systems and Methods
US20190103094A1
Data processing device and method, recognition device, learning data storage device, machine learning device, and program
WO2019150813A1