Method and apparatus for target sound detection
A multi-stage target sound detector with a binary classifier for initial detection and a secondary stage for refined classification addresses power consumption issues in audio context detection, ensuring efficient and accurate sound recognition.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2021-03-01
- Publication Date
- 2026-05-19
AI Technical Summary
Always-on audio context detection systems in electronic devices consume high power, reducing battery life and increasing system complexity due to the continuous scanning for multiple sound events.
A multi-stage target sound detector with a first stage using a binary classifier for low-power, low-complexity sound detection and a second stage for more powerful classification to reduce false positives, activated only when necessary.
This approach reduces power consumption while maintaining high-performance sound detection by minimizing unnecessary processing, allowing for efficient detection of target sounds with reduced average power usage.
Smart Images

Figure 0007862315000001 
Figure 0007862315000002 
Figure 0007862315000003
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the benefit of priority of U.S. Non - Provisional Patent Application No. 16 / 837,420, filed on April 1, 2020, and owned by the same applicant, which is hereby incorporated by reference in its entirety.
[0002] This disclosure generally relates to the detection of target sounds in audio data.
Background Art
[0003] Audio context detection is conventionally used to enable an electronic device to identify context information based on audio captured by the electronic device. For example, an electronic device can analyze the received sound to determine whether the received sound indicates a predetermined sound event. As another example, an electronic device can analyze the received sound to classify the surrounding environment, such as a home environment or an office environment. A "always - on" audio context detection system enables an electronic device to continuously scan an audio input to detect sound events within the audio input. However, the continuous operation of an audio context detection system results in relatively high power consumption, which shortens the battery life when implemented in a mobile device. In addition, system complexity and power consumption increase with an increase in the number of sound events that the audio context detection system is configured to detect.
Summary of the Invention
Means for Solving the Problems
[0004] According to one implementation of the present disclosure, a device for performing sound detection includes one or more processors. The one or more processors include a buffer configured to store audio data. The one or more buffers also include a target sound detector including a first stage and a second stage. The first stage includes a binary target sound classifier configured to process the audio data. The first stage is configured to activate the second stage in response to the detection of a target sound by the first stage. The second stage is configured to receive audio data from the buffer in response to the detection of a target sound.
[0005] According to another implementation of the present disclosure, a method for detecting a target sound includes the step of storing audio data in a buffer. The method also includes the step of processing the audio data in the buffer using a binary target sound classifier in a first stage of a target sound detector, and the step of activating a second stage of the target sound detector in response to the detection of a target sound by the first stage. The method further includes the step of processing the audio data from the buffer using a multi-target sound classifier in the second stage.
[0006] According to another implementation of the present disclosure, a computer-readable storage device, when executed by one or more processors, stores instructions that cause one or more processors to store audio data in a buffer and process the audio data in the buffer using a binary target sound classifier in a first stage of a target sound detector. The instructions, when executed by one or more processors, also cause one or more processors to activate a second stage of the target sound detector in response to the detection of a target sound by the first stage and process the audio data from the buffer using a multi-target sound classifier in the second stage.
[0007] According to another implementation of the present disclosure, the apparatus includes means for detecting a target sound. The means for detecting the target sound includes a first stage and a second stage. The first stage includes means for generating a binary target sound classification of audio data and for activating the second stage in response to the classification of the audio data as containing the target sound. The apparatus also includes means for buffering the audio data and for providing the audio data to the second stage in response to the classification of the audio data as containing the target sound. [Brief explanation of the drawing]
[0008] [Figure 1] This figure shows some examples of specific exemplary implementations of a system including a device that includes a multi-stage target sound detector, as illustrated by some examples of the present disclosure. [Figure 2] Figure 1 shows some examples of specific implementation configurations of the device in this disclosure. [Figure 3] Figure 1 shows another specific implementation of the device including a multi-stage audio scene detector, as illustrated by some examples of the present disclosure. [Figure 4] The diagram shows some examples of components in which a multi-stage audio scene detector may be incorporated. [Figure 5] The diagram shows another specific example of a component into which a multi-stage audio scene detector may be incorporated, as illustrated by several examples of the present disclosure. [Figure 6] Figure 1 shows another specific implementation of the device including a scene detector, as illustrated by some examples of the present disclosure. [Figure 7] Figure 6 shows some examples of components that may be incorporated into the device, according to some examples of the present disclosure. [Figure 8] Figure 6 illustrates another specific example of a component into which the device may be incorporated, as shown in some examples of the present disclosure. [Figure 9] This figure shows an example of an integrated circuit including a multi-stage target sound detector, according to some examples of the present disclosure. [Figure 10]This figure shows a first example of a vehicle including a multi-stage target sound detector, according to some examples of the present disclosure. [Figure 11] This figure shows a second example of a vehicle including a multi-stage target sound detector, according to some examples of the present disclosure. [Figure 12] Figures of headsets, such as virtual reality or augmented reality headsets, including multi-stage target sound detectors, are shown in some examples of the present disclosure. [Figure 13] This is a diagram of a wearable electronic device including a multi-stage target sound detector, as shown in some examples of the present disclosure. [Figure 14] This is a diagram of a voice-controlled speaker system including a multi-stage target sound detector, as shown in some examples of the present disclosure. [Figure 15] Figure 1 illustrates some examples of specific implementations of a target sound detection method that may be performed by the device shown in this disclosure. [Figure 16] This is a block diagram of a specific exemplary example of a device capable of performing target sound detection, as illustrated by some examples of the present disclosure. [Modes for carrying out the invention]
[0009] A device and method for using a multi-stage target sound detector to reduce power consumption are disclosed. Always-on sound detection systems that continuously scan an audio input to detect audio events within the audio input consume relatively large amounts of power, thus reducing battery life when implemented in power-constrained environments, such as in mobile devices. Power consumption can be reduced by reducing the number of audio events configured for the sound detection system to detect, but reducing the number of audio events reduces the utility of the sound detection system.
[0010] As described herein, a multistage target sound detector supports the detection of a relatively large number of target sounds of interest using relatively low power for always-on operation. The multistage target sound detector includes a first stage that supports the binary classification of audio data between all target sounds of interest (as a group) and non-target sounds. The multistage target sound detector includes a second stage for categorizing the audio data as containing one or more specific target sounds of interest for further analysis. The binary classification of the first stage enables low power consumption with low complexity and a small memory footprint to support sound event detection in always-on operation. The second stage includes a more powerful target sound classifier for distinguishing target sounds and for reducing or eliminating false positives (e.g., inaccurate detection of target sounds) that may be generated by the first stage.
[0011] In some implementations, a second stage is activated (e.g., from a sleep state) in response to the detection of one or more target sounds of interest within the audio data to enable more powerful processing of the audio data. Once processing of the audio data in the second stage is complete, the second stage can return to a low-power state. By using low-complexity binary classification in the first stage for always-on operation and selectively activating the more powerful target sound classifier in the second stage, the target sound detector enables high-performance target sound classification with reduced average power consumption for always-on operation.
[0012] In some implementations, a multi-stage environmental scene detector includes a first stage that is always on to detect whether an environmental scene change has occurred, and a more powerful second stage that is selectively activated when the first stage detects an environmental change. In some examples, the first stage includes a binary classifier configured to detect whether audio data represents an environmental scene change without identifying any specific environmental scene. In other examples, a hierarchical scene change detector includes a classifier in the first stage configured to detect a relatively small number of broad classes (e.g., indoors, outdoors, and inside a vehicle), and a more powerful classifier in the second stage is configured to detect a larger number of more specific environmental scenes (e.g., inside a car, inside a train, inside a home, inside an office, etc.). As a result, high-performance environmental scene detection can be provided with reduced average power consumption for always-on operation, in a manner similar to that of multi-stage target sound detection.
[0013] In some implementations, the target sound detector adjusts its operation based on its environment. For example, if the target sound detector is in a user's home, it may use trained data associated with household sounds such as dog barking or doorbells. If the target sound detector is in a vehicle such as a car, it may use trained data associated with vehicle sounds such as glass breaking or sirens. Various techniques may be used to determine the environment, such as using audio scene detectors, cameras, location data (e.g., from satellite-based positioning systems), or a combination of techniques. In some examples, a first stage of the target sound detector activates a camera or other component to determine the environment, and a second stage of the target sound detector is "tuned" for more accurate detection of target sounds associated with the detected environment. Using a camera or other component for environment detection allows for improved target sound detection, and keeping the camera or other component in a low-power state until activated by the first stage of the target sound detector allows for reduced power consumption.
[0014] Unless explicitly limited by the context, the term “producing” is used to indicate any of its usual meanings, such as calculating, generating, and / or providing. Unless explicitly limited by the context, the term “providing” is used to indicate any of its usual meanings, such as calculating, generating, and / or producing. Unless explicitly limited by the context, the term “coupled” is used to indicate a direct or indirect electrical or physical connection. If the connection is indirect, there may be other blocks or components between the “coupled” structures. For example, a loudspeaker may be acoustically coupled to a nearby wall via an intervening medium (e.g., air) that allows the propagation of waves (e.g., sound) from the loudspeaker to the nearby wall (or vice versa).
[0015] The term "comprising" can be used to refer to a method, apparatus, device, system, or any combination thereof, as indicated by the particular context. The term "comprising" is used in this specification and the claims and does not exclude other elements or acts. The term "based on" (such as in the case of "A is based on B") is used to indicate either (i) "at least based on" (e.g., "A is at least based on B"), and, where appropriate in the particular context, (ii) "equal to" (e.g., "A is equal to B") of its ordinary meaning. If (i) A is based on B and includes at least, this may include a configuration in which A is coupled to B. Similarly, the term "responsive to" is used to indicate either of its ordinary meanings, including "at least responsive to". The term "at least one" is used to indicate either of its ordinary meanings, including "one or more". The term "at least two" is used to indicate either of its ordinary meanings, including "two or more".
[0016] The terms "apparatus" and "device" are used generically and interchangeably unless otherwise indicated by a particular context. Unless otherwise indicated, any disclosure of the operation of an apparatus with particular features is also clearly intended to disclose a method with similar features (and vice versa), and any disclosure of the operation of an apparatus with a particular configuration is also clearly intended to disclose a method with a similar configuration (and vice versa). The terms "method", "process", "procedure", and "technique" are used generically and interchangeably unless otherwise indicated by a particular context. The terms "element" and "module" can be used to indicate a part of a larger configuration. The term "packet" can correspond to a unit of data that includes a header part and a payload part. Any incorporation by reference of a portion of a document should be understood to incorporate such definitions if the definitions of terms or variables referenced within that portion appear elsewhere in the document as well as in any figures referenced within the incorporated portion.
[0017] As used herein, the term "communication device" refers to an electronic device that can be used for voice and / or data communication over a wireless communication network. Examples of communication devices include smart speakers, speaker bars, cellular phones, personal digital assistants (PDAs), handheld devices, headsets, wireless modems, laptop computers, personal computers, and the like.
[0018] Figure 1 shows a system 100 that includes a device 102 configured to receive an input sound and process the input sound using a multi-stage target sound detector 120 to detect the presence or absence of one or more target sounds in the input sound. The device 102 includes one or more microphones represented as microphones 112 and one or more processors 160. One or more processors 160 include the target sound detector 120 and a buffer 130 configured to store audio data 132. The target sound detector 120 includes a first stage 140 and a second stage 150. In some implementations, the device 102 may include, as exemplary and non-limiting examples, a wireless speaker and voice command device with an integrated assistant application (e.g., a “smart speaker” device or home automation system), a portable communication device (e.g., a “smartphone” or headset), or a vehicle system.
[0019] The microphone 112 is configured to generate an audio signal 114 in response to a received input sound. For example, the input sound may include a target sound 106, a non-target sound 107, or both. The audio signal 114 is provided to a buffer 130 and stored as audio data 132. In an exemplary example, the buffer 130 corresponds to a pulse code modulation (PCM) buffer, and the audio data 132 corresponds to PCM data. The audio data 132 in the buffer 130 is accessible for processing by the first stage 140 and the second stage 150 of the target sound detector 120, as further described herein.
[0020] The target sound detector 120 is configured to process audio data 132 to determine whether the audio signal 114 represents one or more target sounds of interest. For example, the target sound detector 120 is configured to detect each of a set of target sounds 104, which may include an alarm 191, a doorbell 192, a siren 193, the sound of breaking glass 194, a baby crying 195, the opening and closing of a door 196, and a dog barking 197. The target sounds 191-197 included in the set of target sounds 104 are provided as illustrative examples, and it should be understood that in other implementations, the set of target sounds 104 may include fewer sounds, more sounds, or different sounds. The target sound detector 120 is further configured to detect that a non-target sound 107 originating from one or more other sound sources (represented as a non-target sound source 108) does not include any of the target sounds 191-197.
[0021] The first stage 140 of the target sound detector 120 includes a binary target sound classifier 144 configured to process audio data 132. In some implementations, the binary target sound classifier 144 includes a neural network. In some examples, the binary target sound classifier 144 includes at least one of a Bayesian classifier or a Gaussian mixture model (GMM) classifier, as an exemplary, non-limiting example. In some implementations, the binary target sound classifier 144 is trained to produce one of two outputs: a first output (e.g., 1) indicating that the audio data 132 being classified contains one or more of the target sounds 191-197, or a second output (e.g., 0) indicating that the audio data 132 does not contain any of the target sounds 191-197. In the exemplary example, the binary target sound classifier 144 is not trained to distinguish between each of the target sounds 191-197, allowing for reduced processing load and a smaller memory footprint.
[0022] The first stage 140 is configured to activate the second stage 150 in response to the detection of a target sound. For example, the binary target sound classifier 144 is configured to generate a signal 142 (also called the “activation signal” 142) for activating the second stage 150 in response to the detection of the presence of any of the multiple target sounds 104 in the audio data 132, and to refrain from generating the signal 142 in response to the detection that none of the multiple target sounds 104 are present in the audio data 132. In certain embodiments, the signal 142 is a binary signal having a first value (e.g., a first output) and a second value (e.g., a second output), and generating the signal 142 corresponds to generating a binary signal having the first value (e.g., logical 1). In this embodiment, refraining from generating the signal 142 corresponds to generating a binary signal having the second value (e.g., logical 0).
[0023] In some implementations, the second stage 150 is configured to be activated in response to a signal 142 to process audio data 132, as further described with reference to Figure 2. In an exemplary example, a specific bit in a control register represents the presence or absence of the activation signal 142, and a control circuit in or coupled to the second stage 150 is configured to read the specific bit. A value of "1" for the bit indicates the signal 142, activating the second stage 150, while a value of "0" for the bit indicates the absence of the signal 142 and that the second stage 150 may be deactivated when processing of the current portion of the audio data 132 is complete. In other implementations, the activation signal 142 may instead be implemented as a digital or analog signal on a bus or control line, an interrupt flag in an interrupt controller, or an optical or mechanical signal, in exemplary, non-limiting examples.
[0024] The second stage 150 is configured to receive audio data 132 from the buffer 130 in response to the detection of the target sound 106. In one example, the second stage 150 is configured to process one or more portions (e.g., frames) of the audio data 132 that contain the target sound 106. For example, the buffer 130 may buffer a series of frames of the audio signal 114 as audio data 132, and as a result, when an activation signal 142 is generated, the second stage 150 may process the buffered series of frames and generate a detector output 152 for each of a plurality of target sounds 104 that indicates the presence or absence of that target sound in the audio data 132.
[0025] When deactivated, the second stage 150 does not process audio data 132 and consumes less power than when activated. For example, deactivating the second stage 150 may include gating the input buffer to the second stage 150 to prevent audio data 132 from being input to the second stage 150, gating the clock signal to prevent circuit switching within the second stage 150, or both, in order to reduce dynamic power consumption. As another example, deactivating the second stage 150 may include reducing the power supply to the second stage 150 to reduce static power consumption without losing the state of the circuit elements, removing power from at least a portion of the second stage 150, or a combination thereof.
[0026] In some implementations, the target sound detector 120, buffer 130, first stage 140, second stage 150, or any combination thereof, are implemented using dedicated circuit configurations or hardware. In some implementations, the target sound detector 120, buffer 130, first stage 140, second stage 150, or any combination thereof, are implemented via firmware or software execution. For example, device 102 may include memory configured to store instructions, and one or more processors 160 are configured to execute instructions for implementing one or more of the target sound detector 120, buffer 130, first stage 140, and second stage 150.
[0027] Since the processing operations of the binary target sound classifier 144 are less complex than those performed by the second stage 150, the always-on processing of the audio data 132 in the first stage 140 uses significantly less power than the processing of the audio data 132 in the second stage 150. As a result, processing resources are saved and overall power consumption is reduced.
[0028] In some implementations, the first stage 140 is also configured to activate one or more other components of device 102. In an exemplary example, the first stage 140 may activate a camera used to detect the environment of device 102 (e.g., inside a home, outdoors, inside a car), and the second stage 150 may operate to focus on a target sound associated with the detected environment, as further described with reference to Figure 6.
[0029] Figure 2 shows an example of device 102 200 in which the binary target sound classifier 144 includes a neural network 212, and the binary target sound classifier 144 and buffer 130 are contained in a low-power domain 203, such as an always-on low-power domain of one or more processors 160. The second stage 150 is contained in another power domain 205, such as an on-demand power domain. In some implementations, the first stage 140 (e.g., the binary target sound classifier 144) and buffer 130 of the target sound detector 120 are configured to operate in always-on mode, and the second stage 150 of the target sound detector 120 is configured to operate in on-demand mode.
[0030] The power domain 205 includes a second stage 150 of the target sound detector 102, a sound context application 240, and an activation circuit configuration 230. The activation circuit configuration 230 selectively activates one or more components of the power domain 205, such as the second stage 150, in response to an activation signal 142 (e.g., a wake-up interrupt signal). For example, in some implementations, the activation circuit configuration 230 is configured to transition the second stage 150 from a low-power state 232 to an active state 234 in response to the reception of the signal 142.
[0031] For example, the activation circuit configuration 230 may include, or be combined with, a power management circuit configuration, a clock circuit configuration, a head switch or foot switch circuit configuration, a buffer control circuit configuration, or any combination thereof. The activation circuit configuration 230 may be configured to initiate power-on of the second stage 150, for example, by selectively applying or increasing the voltage of the power supply for the second stage 150, the power domain 205, or both. As another example, the activation circuit configuration 230 may be configured to selectively gate or ungate the clock signal to the second stage 150, for example, to prevent circuit operation without removing the power supply.
[0032] The second stage 150 includes a multi-target sound classifier 210 configured to generate a detector output 152 for each of the multiple target sounds 104, indicating the presence or absence of that target sound in the audio data 132. The multiple target sounds correspond to multiple classes 290 of sound events, each of which includes at least two of the following: alarm 291, doorbell 292, siren 293, glass breaking 294, baby crying 295, door opening / closing 296, or dog barking 297. It should be understood that sound event classes 291-297 are provided as illustrative examples. In other examples, the multiple classes 290 may include fewer sound events, more sound events, or different sound events. For example, in one implementation where device 102 is implemented in a vehicle (e.g., an automobile), the multiple classes 290 include, as exemplary and non-limiting examples, one or more of the following sound events more commonly encountered in a vehicle: opening and closing of vehicle doors, road noise, opening and closing of windows, radio, brakes, handbrakes being pulled or released, wipers, turn signals, or engine starting sounds. While a single set of sound event classes (e.g., multiple classes 290) is shown, in other implementations, the multi-target sound classifier 210 is configured to select from among multiple sets of sound event classes based on the environment of device 102 (e.g., one set of target sounds when device 102 is in a home, and another set of target sounds when device 102 is in a vehicle), as will be further explained with reference to Figure 6.
[0033] In some implementations, the multi-target sound classifier 210 performs "faster than real-time" processing of the audio data 132. In an exemplary, non-limiting example, the buffer 130 is sized to store approximately 2 seconds of audio data in a circular buffer configuration, where the oldest audio data in the buffer 130 is replaced by the most recently received audio data. The first stage 140 may be configured to process continuously received 20-millisecond (mS) segments (e.g., frames) of the audio data 132 periodically in real-time (e.g., the binary target sound classifier 144 processes one 20 mS segment every 20 mS) and with low power consumption. However, when the second stage 150 is activated, the multi-target sound classifier 210 processes the buffered audio data 132 at a faster rate and with higher power consumption in order to process the buffered audio data 132 more quickly and generate a detector output 152.
[0034] In some implementations, the detector output 152 includes multiple values, such as a bit or multi-bit value indicating the detection (or likelihood of detection) of each target sound. In an exemplary example, the detector output 152 includes a 7-bit value, where the first bit corresponds to the detection or non-detection of a sound classified as an alarm 291, the second bit corresponds to the detection or non-detection of a sound classified as a doorbell 292, the third bit corresponds to the detection or non-detection of a sound classified as a siren 293, the fourth bit corresponds to the detection or non-detection of a sound classified as breaking glass 294, the fifth bit corresponds to the detection or non-detection of a sound classified as a baby crying 295, the sixth bit corresponds to the detection or non-detection of a sound classified as opening and closing a door 296, and the seventh bit corresponds to the detection or non-detection of a sound classified as a dog barking 297.
[0035] The detector output 152 generated by the second stage 150 is provided to the sound context application 240. The sound context application 240 may be configured to perform one or more actions based on the detection of one or more target sounds. For example, in one implementation where device 102 is in a home automation system, the sound context application 240 may generate a user interface signal 242 to alert the user of one or more detected sound events. For example, the user interface signal 242 could cause an output device 250 (e.g., a display screen or a loudspeaker on a speech interface device) to alert the user that a dog barking and the sound of breaking glass have been detected at the back door of a building. In another example, when the user is not inside the building, the user interface signal 242 could cause an output device 250 (e.g., a transmitter coupled to a wireless network such as a cellular network or a wireless local area network) to send an alert to the user's phone or smartwatch.
[0036] In another implementation where device 102 is located inside a vehicle (e.g., an automobile), the sound context application 240 may generate a user interface signal 242 to alert the vehicle operator via an output device 250 (e.g., a display screen or voice interface) that a siren has been detected via an external microphone while the vehicle is in motion. If the vehicle is stopped and the operator has left the vehicle, the sound context application 240 may generate a user interface signal 242 to alert the vehicle owner via an output device 250 (e.g., wireless transmission to the owner's phone or smartwatch) that a baby crying has been detected via an internal microphone in the vehicle.
[0037] In another implementation where device 102 is integrated into or coupled to an audio playback device such as headphones or a headset, the sound context application 240 may, as an exemplary example, generate a user interface signal 242 to alert the user of the playback device via an output device 250 (e.g., a display screen or a loudspeaker) that a siren has been detected, or pass-through the siren for playback in the loudspeaker of the headphones or headset.
[0038] The activation circuit configuration 230 is shown as separate from the second stage 150 in the power domain 205, but in other implementations, the activation circuit configuration 230 may be included in the second stage 150. In some implementations, the output device 250 is implemented as a user interface component of device 102, such as a display screen or a loudspeaker, but in other implementations, the output device 250 may be a user interface device that is both separate from and coupled to device 102. The multi-target sound classifier 210 is configured to detect and distinguish sound events corresponding to seven classes 291-297, but in other implementations, the multi-target sound classifier 210 may be configured to detect any other sound events instead of or in addition to any one or more of the seven classes 291-297, and the multi-target sound classifier 210 may be configured to classify sound events according to any other number of classes.
[0039] Figure 3 shows one implementation configuration 300 in which device 102 includes a buffer 130 and a target sound detector 120, and also includes an audio scene detector 302. The audio scene detector 302 includes an audio scene change detector 304 and an audio scene classifier 308. The audio scene change detector 304 is configured to process audio data 132 and generate a scene change signal 306 in response to the detection of an audio scene change. In some implementation configurations, the audio scene change detector 304 is implemented in a first stage of the audio scene detector 302 (e.g., a low-power, always-on processing stage), and the audio scene classifier 308 is implemented in a second stage of the audio scene detector 302 (e.g., a more powerful, high-performance processing stage), and is activated by the scene change signal 306 in a similar manner to how the multi-target sound classifier 210 in Figure 2 is activated by the activation signal 142. Unlike target sound detection, the audio environment is always present, and the efficiency of the audio scene detector 302's operation is enhanced in the first stage by detecting changes in the audio environment without incurring the computational penalty associated with identifying an accurate audio environment.
[0040] In some implementations, the audio scene change detector 304 is configured to detect changes in the audio scene based on detecting changes in at least one of the noise statistics 310 or transient sound statistics 312. For example, the audio scene change detector 304 processes audio data 132 to determine the noise statistics 310 (e.g., the average spectral energy distribution of audio frames identified as containing noise) and transient sound statistics 312 (e.g., the average spectral energy distribution of audio frames identified as containing transient sound) time-averaged over a relatively large time window (e.g., 3 to 5 seconds). Changes between audio scenes are detected based on determining changes in the noise statistics 310, transient sound statistics 312, or both. For example, the noise and sound characteristics of an office environment are sufficiently distinct from the noise and sound characteristics in a moving car that changes from an office environment to a vehicle environment can be detected, and in some implementations, changes are detected without identifying the noise and sound characteristics as corresponding to either an office environment or a vehicle environment. In response to the detection of an audio scene change, the audio scene change detector generates a scene change signal 306 and sends it to the audio scene classifier 308.
[0041] The audio scene classifier 308 is configured to receive audio data 132 from the buffer 130 in response to the detection of an audio scene change. In some implementations, the audio scene classifier 308 is a more powerful and complex processing component than the audio scene change detector 304 and is configured to classify the audio data 132 as corresponding to one of several audio scene classes 330. In one example, the several audio scene classes 330 include indoors 332, indoors 334, indoors 336, indoors in a car 338, indoors in a train 340, on the street 342, indoors 344, and outdoors 346.
[0042] The scene detector output 352, generated by the audio scene detector 302, presents instructions for the detected audio scene, which may be provided to the sound context application 240 in Figure 2. For example, the sound context application 240 can adjust the operation of device 102 based on the detected audio scene, such as by changing the graphical user interface (GUI) on the display screen to present top-level menu items associated with the environment. For example, in an exemplary and non-limiting example, when the detected environment is inside a car, navigation and communication items (e.g., hands-free dialing) may be presented; when the detected environment is outdoors, camera and audio recording items may be presented; and when the detected environment is inside an office, note-taking and contacts items may be presented.
[0043] Although the multiple audio scene classes 330 are described as including eight classes 332-346, in other implementations, the multiple audio scene classes 330 may include at least two of the following: indoors 332, indoors 334, indoors 336, indoors in a car 338, indoors in a train 340, on the street 342, indoors 344, or outdoors 346. In other implementations, one or more of the classes 330 may be omitted, and one or more other classes may be used instead of, or in addition to, classes 332-346 or any combination thereof.
[0044] Figure 4 shows one implementation 400 of the audio scene change detector 304, which includes a scene transition classifier 414 that is trained using audio data corresponding to transitions between scenes. For example, the scene transition classifier 414 may be trained on captured audio data for transitions such as office to street, car to outdoors, and restaurant to street. In some implementations, the scene transition classifier 414 provides more robust change detection using a smaller model than the implementation of the audio scene change detector 304 described with reference to Figure 3.
[0045] Figure 5 shows one implementation configuration 500 in which the audio scene detector 302 corresponds to a hierarchical detector such that the audio scene change detector 304 classifies audio data 132 using a reduced set of audio scenes compared to the audio scene classifier 308. For example, the audio scene change detector 304 includes a hierarchical model change detector 514 configured to detect audio scene changes based on detecting changes between audio scene classes of a reduced set of classes 530. For example, the reduced set of classes 530 includes the "in-vehicle" class 502, the indoor class 344, and the outdoor class 346. In some implementation configurations, one or more (or all) of the reduced set of classes 530 include or span multiple classes used by the audio scene classifier 308. For example, the "in-vehicle" class 502 is used to classify audio scenes that the audio scene classifier 308 distinguishes as either "in-car" or "in-train". In some implementations, one or more (or all) of the reduced sets of classes 530 form a subset of classes 330 used by the audio scene classifier 308, such as indoor class 344 and outdoor class 346. In some examples, the reduced sets of classes 530 are configured to include two or three of the most likely audio scenes to be encountered in order to improve the probability of detecting audio scene changes.
[0046] The reduced class set 530 contains a reduced number of classes compared to the classes 330 of the audio scene classifier 308. For example, the first count (3) of audio scene classes in the reduced class set 530 is less than the second count (8) of audio scene classes 330. Although the reduced class set 530 is described as containing three classes, in other implementations, the reduced class set 530 may contain any number of classes (for example, at least two classes, such as two, three, four, or more classes) that are less than the number of classes supported by the audio scene classifier 308.
[0047] Because the hierarchical model change detector 514 performs detection from a smaller set of classes compared to the audio scene classifier 308, the audio scene change detector 304 can detect scene changes with reduced complexity and power consumption compared to the more powerful audio scene classifier 308. Transitions between environments that are not detected by the hierarchical model change detector 514, such as a direct transition from "indoors" to "indoors" without an intervening transition to a vehicle environment or an outdoor environment (for example, both being in the "indoors" class 344), may be less likely to occur.
[0048] Figures 3 to 5 illustrate various implementations in which both the audio scene detector 304 and the target sound detector 120 are included in device 102. In other implementations, the audio scene detector 302 may be implemented in a device that does not include the target sound detector. In an exemplary example, device 102 includes a buffer 130 and the audio scene detector 302, while omitting the first stage 140, the second stage 150, or both of the target sound detector 120.
[0049] Figure 6 shows a specific example 600 in which device 102 includes a scene detector 606 configured to detect the environment based on at least one of a camera, a location detection system, or an audio scene detector.
[0050] Device 102 includes one or more sensors 602 that generate data that the scene detector 606 can use to determine the environment 608. The one or more sensors 602 include one or more cameras and one or more sensors of a location detection system, each indicated as a camera 620 and a Global Positioning System (GPS) receiver 624. Camera 620 may include any type of image capture device and may support or include still image or video capture, visible spectrum, infrared spectrum, or ultraviolet spectrum, depth detection (e.g., structured light, time of flight), any other image capture technique, or any combination thereof.
[0051] The first stage 140 is configured to activate one or more of the sensors 602 from a low-power state in response to the detection of a target sound by the first stage 140. For example, the signal 142 may be provided to the camera 620 and the GPS receiver 624. In response to the signal 142, the camera 620 and the GPS receiver 624 transition from a low-power state (for example, when not being used by another application of device 102) to an active state.
[0052] The scene detector 606 includes an audio scene detector 302 and is configured to detect the environment 608 based on at least one of the camera 620, the GPS receiver 624, or the audio scene detector 302. As a first example, the scene detector 606 is configured to generate a first estimate of the environment 608 of device 102 based at least in part on an input signal 622 from the camera 624 (e.g., image data). For example, the scene detector 606 may be configured to process input data 622 to generate a first classification of the environment 608, such as inside a home, inside an office, inside a restaurant, inside a car, inside a train, on a street, outdoors, or indoors, based on visual features.
[0053] As a second example, the scene detector 606 is configured to generate a second estimate of the environment 608 based at least in part on location information 626 from a GPS receiver. For example, the scene detector 606 may use the location information 626 to look up map data to determine whether the location corresponds to the user's home, the user's office, a restaurant, a train route, a street, an outdoor location, or an indoor location. The scene detector 606 may be configured to determine the speed of movement of the device 102 based on the location data 626 to determine whether the device 102 is moving in a car or an airplane.
[0054] In some implementations, the scene detector 606 is configured to determine the environment 608 based on a first estimate, a second estimate, the scene detector output 352 of the audio scene detector 302, and the respective confidence levels associated with the first estimate, the second estimate, and the scene detector output 352. The indication of the environment 608 is provided to the target sound detector 120, and the operation of the multi-target sound classifier 210 is at least partially based on the classification of the environment 608 by the scene detector 606.
[0055] Figure 6 shows a device 102 including a camera 620, a GPS receiver 624, and an audio scene detector 302, but in other implementations, one or more of the camera 620, GPS receiver 624, or audio scene detector 302 may be omitted, and one or more other sensors may be added, or any combination thereof. For example, the audio scene detector 302 may be omitted or replaced with one or more other audio scene detectors. In other examples, the scene detector 606 determines the environment 608 based solely on image data 622 from the camera, solely on location data 62 from the GPS sensor 624, or solely on scene detection from the audio scene detector.
[0056] One or more sensors 602, the audio scene detector 302, and the scene detector 606 are activated in response to signal 142, but in other implementations, one or more of the scene detector 606, the audio scene detector 302, and the sensors 602, or any combination thereof, may be activated or deactivated independently of signal 142. As a non-limiting example, in environments where power is not constrained, such as in a vehicle or household electrical appliance, one or more sensors 602, the audio scene detector 302, and the scene detector 606 may remain active even if no target sound activity is detected.
[0057] Figure 7 shows an example 700 in which the multi-target sound classifier 210 is adjusted to focus on one or more specific classes 702 of several classes 290 of sound events corresponding to the environment 608. In example 700, the environment 608 is detected as "inside a car," and the multi-target sound classifier 210 is adjusted to further focus on identifying the target sound in the audio data 132 as one of the classes 290 that are more commonly encountered inside a car, namely a siren 293, the sound of breaking glass 294, a baby crying 295, or the opening and closing of a door 296, while less focusing on identifying the target sound as one of the classes that are less commonly encountered inside a car, namely an alarm 291, a doorbell 292, or a dog barking 297. As a result, target sound detection can be performed more accurately than in implementations in which environmental information is not used to focus on target sound detection.
[0058] Figure 8 shows an example 800 in which the multi-target sound classifier 210 is configured to select a specific set of sound event classes corresponding to an environment 608 from among multiple sets of sound event classes. The first set of trained data 802 includes the first set 812 of sound event classes associated with a first environment (e.g., inside a home). The second set of trained data 804 includes the second set 814 of sound event classes associated with a second environment (e.g., inside a car), and one or more additional sets of trained data including the Nth set of trained data 808, which includes the Nth set 818 of sound event classes associated with an Nth environment (e.g., inside an office), where N is an integer greater than 1. In a non-restrictive example, each of the sets of trained data 802-808 corresponds to one of the classes 330 (e.g., N=8). In some implementations, one or more of the sets of trained data 802-808 correspond to a default set of trained data to be used when the environment is undetermined. As an example, the multiple classes 290 in Figure 2 can be used as the default set of trained data.
[0059] In an exemplary implementation, a first set of sound event classes 812 corresponds to “inside the home,” and a second set of sound event classes 814 corresponds to “inside a car.” The first set of sound event classes 812 includes, as exemplary non-limiting examples, one or more sound events more commonly encountered inside a home, such as a fire alarm, a baby crying, a dog barking, a doorbell, doors opening and closing, and the sound of breaking glass. The second set of event classes 814 includes, as exemplary non-limiting examples, sound events more commonly encountered inside a car, such as doors opening and closing, road noise, windows opening and closing, the sound of a radio, brakes, handbrakes being pulled or released, wipers, turn signals, or engine starting sounds. In response to the environment 608 being detected as “inside the home,” the multi-target sound classifier 210 selects the first set of sound event classes 812 and classifies the audio data 132 based on the sound event classes of that particular set (i.e., the first set of sound event classes 812). In response to the detection of environment 608 as "inside a car," the multi-target sound classifier 210 selects a second set 814 of sound event classes and classifies the audio data 132 based on the sound event classes of that particular set (i.e., the second set 814 of sound event classes).
[0060] As a result, a larger overall number of target sounds can be detected by using a different set of sound events for each environment, without increasing the overall processing and memory usage required to perform target sound classification for any given environment. In addition, power consumption is reduced compared to the always-on operation of the sensor 602 and the scene detector 606 by using the first stage 140 to activate the sensor 602, the scene detector 606, or both.
[0061] Example 800 describes a multi-target sound classifier 210 that selects one of a set of sound event classes 812-818 based on the environment 608, but in some implementations, each of the sets of trained data 802-808 also includes trained data for a binary target sound classifier 144 to detect the presence or absence of target sounds associated with a particular environment as a group. In one example, the target sound detector 120 is configured to select a specific set of trained data from the sets of trained data 802-808 that corresponds to the detected environment 608 of device 102, and to process audio data 132 based on that specific set of trained data.
[0062] Figure 9 shows one implementation form 900 of device 102 as an integrated circuit 902 including one or more processors 160. The integrated circuit 902 also includes sensor signal inputs 910, such as one or more first bus interfaces, to enable audio signals 114 to be received from microphone 112. For example, the sensor signal input 910 receives audio signals 114 from microphone 112 and provides audio signals 114 to buffer 130. The integrated circuit 902 also includes data outputs 912, such as second bus interfaces, to enable the transmission of detector outputs 152 (for example, to a display device, memory, or transmitter, in exemplary, non-limiting examples). For example, a target sound detector 120 provides a detector output 152 to a data output 912, and the data output 912 transmits the detector output 152. The integrated circuit 902 enables the implementation of multi-stage target sound detection as a component in a system including one or more microphones, such as a vehicle shown in Figure 10 or Figure 11, a virtual reality or augmented reality headset shown in Figure 12, a wearable electronic device shown in Figure 13, a voice-controlled speaker system shown in Figure 14, or a wireless communication device shown in Figure 16.
[0063] Figure 10 shows one implementation configuration 1000 in which device 102 corresponds to or is integrated into vehicle 1002, indicated as an automobile. In some implementation configurations, multi-stage target sound detection may be performed based on audio signals received from internal microphones, such as for a baby crying inside the automobile; based on audio signals received from external microphones (e.g., microphone 112), such as for a siren; or both. The detector output 152 in Figure 1 may be provided to a display screen in vehicle 1002, to a user's mobile device, or both. For example, output device 250 includes a display screen that displays a notification indicating that a target sound (e.g., a siren) has been detected outside vehicle 1002. In another example, output device 250 includes a transmitter that sends a notification to a mobile device indicating that a target sound (e.g., a baby crying) has been detected inside vehicle 1002.
[0064] Figure 11 shows another implementation form 1100 in which device 102 corresponds to or is integrated within vehicle 1102, which is shown as a manned or unmanned aerial device (e.g., a delivery drone). Multistage target sound detection may be performed based on audio signals received from one or more microphones in vehicle 1102 (e.g., microphone 112), such as for the opening and closing of a door. For example, output device 250 includes a transmitter that sends a notification to a control device indicating that a target sound (e.g., opening and closing of a door) has been detected by vehicle 1102.
[0065] Figure 12 shows one implementation form 1200 in which device 102 is a portable electronic device corresponding to a virtual reality, augmented reality, or mixed reality headset 1202. One or more processors 160 and microphones 112 are integrated into the headset 1202. Multi-stage target sound detection may be performed based on audio signals received from the microphones 112 of the headset 1202. A visual interface device, such as an output device 250, is placed in front of the user's eyes to enable the display of augmented reality or virtual reality images or scenes to the user while the headset 1202 is being worn. In a particular example, the output device 250 is configured to display a notification indicating that a target sound (e.g., a fire alarm or doorbell) has been detected outside the headset 1202.
[0066] Figure 13 shows one implementation form 1300, where device 102 is a portable electronic device corresponding to the wearable electronic device 1302, which is labeled as a “smartwatch”. One or more processors 160 and a microphone 112 are integrated into the wearable electronic device 1302. Multistage target sound detection may be performed based on audio signals received from the microphone 112 of the wearable electronic device 1302. The wearable electronic device 1302 includes a display screen, such as an output device 250, configured to display a notification indicating that a target sound has been detected by the wearable electronic device 1302. In a particular example, the output device 250 includes a haptic device that provides haptic notification (e.g., vibrates) in response to the detection of a target sound. The haptic notification can cause the user to look at the wearable electronic device 1302 to confirm a displayed notification indicating that a target sound has been detected. Thus, the wearable electronic device 1302 can alert a user with hearing impairment or a user wearing a headset that a target sound has been detected. Figure 14 shows an exemplary example of a wireless speaker and voice-activated device 1400. The wireless speaker and voice-activated device 1400 may have wireless network connectivity and is configured to perform assistant operations. One or more processors 160, a microphone 112, and one or more cameras, such as camera 620, are included in the wireless speaker and voice-activated device 1400. Camera 620 is configured to be activated in response to an integrated assistant application 1402, such as in response to a user command to start a video conference. Camera 620 is further configured to be activated in response to the detection of the presence of any of several target sounds in audio data from microphone 112 by a binary target sound classifier 144 in a target sound detector 120, such as to function as a surveillance camera in response to the detection of a target sound. The wireless speaker and voice-activated device 1400 also includes a speaker 1404. During operation, in response to the reception of verbal commands, the wireless speaker and voice-activated device 1400 can perform assistant actions, such as through the execution of an integrated assistant application 1402. Assistant actions may include adjusting the temperature, playing music, turning on lights, starting a video conference, etc. For example, an assistant action may be performed in response to the reception of a command following a keyword (e.g., "Hello, Assistant"). Multi-stage target sound detection may be performed based on audio signals received from the microphone 142 of the wireless speaker and voice-activated device 1400. In some implementations, the integrated assistant application 1402 is activated in response to the detection by a binary target sound classifier 144 in a target sound detector 120 of the presence of any of several target sounds in the audio data from the microphone 112. An identified target sound instruction (e.g., detector output 152) is provided to the integrated assistant application 1402, which causes the wireless speaker and voice-activated device 1400 to provide a notification indicating that the target sound (e.g., opening and closing of a door) has been detected by the wireless speaker and voice-activated device 1400, such as by playing an audible speech notification via speaker 1404 or sending a notification to a mobile device.
[0067] Referring to Figure 15, a specific implementation of the multi-stage target sound detection method 1500 is shown. In a particular embodiment, one or more operations of the method 1500 are performed by at least one of the following: the binary target sound classifier 144, target sound detector 120, buffer 130, processor 160, device 102, system 100 in Figure 1, the activation signal unit 204 in Figure 2, the multi-target sound classifier 210, the activation circuit configuration 230, the sound context application 240, the output device 250, system 200, the audio scene detector 302, audio scene change detector 304, audio scene classifier 308 in Figure 3, the scene transition classifier 414 in Figure 4, the hierarchical model change detector 514 in Figure 5, the scene detector 606 in Figure 6, or a combination thereof.
[0068] Method 1500 includes storing audio data in a buffer in 1502. For example, buffer 130 in Figure 1 stores audio data 132, as described with reference to Figure 1. In a particular embodiment, the audio data 132 corresponds to an audio signal 114 received from microphone 112 in Figure 1.
[0069] Method 1500 also includes processing audio data in a buffer using a binary target sound classifier in the first stage of a target sound detector in 1504. For example, the binary target sound classifier 144 in Figure 1 processes audio data 132 stored in buffer 130, as described with reference to Figure 1. The binary target sound classifier 144 is located in the first stage 140 of the target sound detector 150 in Figure 1.
[0070] Method 1500 further includes activating a second stage of the target sound detector in response to the detection of a target sound by the first stage in 1506. For example, the first stage 140 in Figure 1 activates the second stage 150 of the target sound detector 120 in response to the detection of a target sound 106 by the first stage 140, as described with reference to Figure 1. In some implementations, the binary target sound classifier and buffer operate in an always-on mode, and activating the second stage includes sending a signal from the first stage to the second stage and transitioning the second stage from a low-power state to an active state in response to the reception of a signal in the second stage, as described with reference to Figure 2.
[0071] Method 1500 includes processing audio data from a buffer using a multi-target sound classifier in a second stage 1508. For example, the multi-target sound classifier 210 in Figure 2 processes audio data 132 from buffer 130 in the second stage 150, as described with reference to Figure 2. The multi-target sound classifier may process audio data based on multiple target sounds corresponding to multiple classes of sound events, such as class 290, or one or more of the sets of sound event classes 812-818.
[0072] Method 1500 may also include generating a detector output for each of several target sounds, such as detector output 152, indicating the presence or absence of that target sound in the audio data.
[0073] In some implementations, Method 1500 also includes processing audio data in an audio scene change detector, such as the audio scene detector 302 in Figure 3. In such implementations, in response to the detection of an audio scene change, Method 1500 includes activating an audio scene classifier, such as the audio scene classifier 308, and processing audio data from a buffer using the audio scene classifier. Method 1500 may include classifying the audio data in the audio scene classifier according to a plurality of audio scene classes, such as class 330. In an exemplary example, the plurality of audio scene classes include at least two of the following: inside a home, inside an office, inside a restaurant, inside a car, inside a train, on a street, indoors, or outdoors.
[0074] Detecting audio scene changes may be based on detecting changes in at least one of noise statistics or non-stationary sound statistics, such as those described with reference to the audio scene change detector 304 in Figure 3. Alternatively, or additionally, detecting audio scene changes may be performed using a classifier trained with audio data corresponding to transitions between scenes, such as the scene transition classifier 414 in Figure 4. Alternatively, or additionally, method 1500 may include detecting audio scene changes based on detecting changes between audio scene classes within a first set of audio scene classes (e.g., a reduced set of classes 530 in Figure 5) and classifying audio data according to a second set of audio scene classes (e.g., class 330 in Figure 3), where a first count of an audio scene class in the first set of audio scene classes (e.g., 3) is less than a second count of an audio scene class in the second set of audio scene classes (e.g., 8).
[0075] Since the processing operations of the binary target sound classifier are less complex compared to the processing operations performed by the second stage, the audio data processed in the binary target sound classifier consumes less power compared to processing audio data in the second stage. By selectively activating the second stage in response to the detection of a target sound by the first stage, method 1500 enables the saving of processing resources and a reduction in overall power consumption.
[0076] Method 1500 in Figure 15 can be implemented by a processing unit such as a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a central processing unit (CPU), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. For example, Method 1500 in Figure 15 can be implemented by an instruction-executing processor, such as those described with reference to Figure 16.
[0077] Referring to Figure 16, a block diagram of a specific exemplary implementation of the device is shown, designated as 1600 overall. In various implementations, device 1600 may have more or fewer components than shown in Figure 16. In the exemplary implementation, device 1600 may correspond to device 102. In the exemplary implementation, device 1600 may perform one or more operations described with reference to Figures 1 to 15.
[0078] In certain implementations, device 1600 includes a processor 1606 (e.g., a central processing unit (CPU)). Device 1600 may include one or more additional processors 1610 (e.g., one or more DSPs). Processor 1610 may include a speech and music coder decoder (codec) 1608, a target sound detector 120, a sound context application 240, an activation circuit configuration 230, an audio scene detector 302, or a combination thereof. Speech and music codec 1608 may include a speech coder ("vocoder") encoder 1636, a vocoder decoder 1638, or both.
[0079] Device 1600 may include memory 1686 and codec 1634. Memory 1686 may include instructions 1656 that can be executed by one or more additional processors 1610 (or processor 1606) to implement the functions described with reference to the target sound detector 120, sound context application 240, activation circuit configuration 230, audio scene detector 302, or any combination thereof. Memory 1686 may include buffer 160. Device 1600 may include wireless controller 1640 coupled to antenna 1652 via transceiver 1650.
[0080] Device 1600 may include a display 1628 coupled to a display controller 1626. Speaker 1692 and microphone 112 may be coupled to codec 1634. Codec 1634 may include a digital-to-analog converter 1602 and an analog-to-digital converter 1604. In certain implementations, codec 1634 may receive an analog signal from microphone 112, convert the analog signal to a digital signal using analog-to-digital converter 1604, and provide the digital signal to speech and music codec 1608. Speech and music codec 1608 may process the digital signal, which may be further processed by one or more of the target sound detector 120 and audio scene detector 302. In certain implementations, speech and music codec 1608 may provide the digital signal to codec 1634. Codec 1634 may convert the digital signal to an analog signal using digital-to-analog converter 1602, and provide the analog signal to speaker 1692.
[0081] In certain implementations, device 1600 may be included in a system-in-package or system-on-chip device 1622. In certain implementations, memory 1686, processor 1606, processor 1610, display controller 1626, codec 1634, and wireless controller 1640 are included in a system-in-package or system-on-chip device 1622. In certain implementations, input device 1630 and power supply 1644 are coupled to system-on-chip device 1622. Furthermore, in certain implementations, as shown in Figure 16, display 1628, input device 1630, speaker 1692, microphone 112, antenna 1652, and power supply 1644 are outside of system-on-chip device 1622. In certain implementations, each of the display 1628, input device 1630, speaker 1692, microphone 112, antenna 1652, and power supply 1644 may be coupled to a component of the system-on-chip device 1622, such as an interface or controller.
[0082] Device 1600 may include smart speakers, speaker bars, mobile communication devices, smartphones, cellular phones, laptop computers, computers, tablets, personal digital assistants, display devices, televisions, game consoles, music players, radios, digital video players, digital video disc (DVD) players, tuners, cameras, navigation devices, vehicles, headsets, augmented reality headsets, virtual reality headsets, aircraft, or any combination thereof.
[0083] In relation to the implementation described, the apparatus for processing an audio signal representing an input sound includes means for detecting a target sound. The means for detecting the target sound includes a first stage and a second stage. The first stage includes means for generating a binary target sound classification of audio data and for activating the second stage in response to the classification of the audio data as containing the target sound. For example, the means for detecting the target sound may correspond to a target sound detector 120, one or more processors 160, one or more processors 1610, one or more other circuits or components configured to detect the target sound, or any combination thereof. The means for generating a binary target sound classification and activating the second stage may correspond to a binary target sound classifier 144, one or more other circuits or components configured to generate a binary target sound classification and activate the second stage, or any combination thereof.
[0084] The apparatus also includes means for buffering audio data and providing the audio data to a second stage in response to the classification of the audio data as containing a target sound. For example, the means for buffering audio data and providing the audio data to a second stage may correspond to a buffer 160, one or more processors 160, one or more processors 1610, one or more other circuits or components configured to buffer audio data and provide the audio data to a second stage in response to the classification of the audio data as containing a target sound, or any combination thereof.
[0085] In some implementations, the device further includes means for detecting audio scenes, the means for detecting audio scenes includes means for detecting audio scene changes within audio data, and means for classifying audio data as specific audio scenes in response to the detection of audio scene changes. For example, the means for detecting audio scenes may correspond to an audio scene detector 302, one or more processors 160, one or more processors 1610, one or more other circuits or components configured to detect audio scenes, or any combination thereof. The means for detecting audio scene changes within audio data may correspond to an audio scene change detector 304, a scene transition classifier 414, a hierarchical model change detector 514, one or more other circuits or components configured to detect audio scene changes within audio data, or any combination thereof. The means for classifying audio data as specific audio scenes in response to the detection of audio scene changes may correspond to an audio scene classifier 308, one or more other circuits or components configured to classify audio data as specific audio scenes in response to the detection of audio scene changes, or any combination thereof.
[0086] In some implementations, a non-temporary computer-readable medium (e.g., memory 1686), when executed by one or more processors (e.g., one or more processors 1610 or processors 1606), includes an instruction (e.g., instruction 1656) that causes one or more processors to perform an operation to store audio data in a buffer (e.g., buffer 130) and an operation to process the audio data in the buffer using a binary target sound classifier (e.g., binary target sound classifier 144) in a first stage of a target sound detector (e.g., first stage 140 of target sound detector 120). The instruction also, when executed by one or more processors, causes one or more processors to activate a second stage of the target sound detector (e.g., second stage 150) in response to the detection of a target sound by the first stage, and to process the audio data from the buffer using a multi-target sound classifier (e.g., multi-target sound classifier 210) in the second stage.
[0087] Those skilled in the art will further understand that the various exemplary logic blocks, configurations, modules, circuits, and algorithmic steps described herein in relation to the implementations disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or a combination of both. The various exemplary components, blocks, configurations, modules, circuits, and steps have been described above in relation to their functions. Whether such functions are implemented as hardware or as processor-executable instructions depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functions in various ways for each specific application, and such decisions should not be construed as causing a departure from the scope of this disclosure.
[0088] Steps of methods or algorithms described in relation to the implementations disclosed herein may be embodied directly in hardware, in software modules executed by a processor, or in a combination of both. Software modules may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, compact disk read-only memory (CD-ROM), or any other form of non-temporary storage medium known in the art. An exemplary storage medium is coupled to a processor so that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integrated with the processor. The processor and storage medium may reside in an application-specific integrated circuit (ASIC). An ASIC may reside in a computing device or user terminal. Alternatively, the processor and storage medium may reside as separate components within a computing device or user terminal.
[0089] The foregoing description of the disclosed implementations is provided to enable those skilled in the art to create or use the disclosed implementations. Various modifications to these implementations will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other implementations without departing from the scope of this disclosure. Therefore, this disclosure is not limited to the implementations shown herein and should be given the broadest possible scope that corresponds to the principles and novel features defined by the following claims. [Explanation of symbols]
[0090] 100 Systems 102 devices 104 Set of target sounds, multiple target sounds 106 Target Sound 107 Non-target sounds 108 Non-target sound sources 112 Microphone 114 Audio signal 120 Multi-stage target sound detector, target sound detector 130 buffers 132 Audio Data 140 Section 1 142 Signal, Activation Signal 144 Binary Target Sound Classifier 150 Section 2 152 Detector output 160 processors 191-197 Target sound 191 Alarm 192 Doorbell 193 Siren 194 The sound of glass breaking 195 Baby crying 196 Door opening and closing 197 Dog barking 200 cases 203 Low-Power Domain 204 Activation Signal Unit 205 Power Domain 210 Multi-target sound classifier 212 Neural Networks 230 Activation Circuit Configuration 232 Low Power State 234 Active 240 Sound Context Application 242 User Interface Signals 250 Output Devices 290 Multiple classes, classes 291-297 Sound Event Class, Class 291 Alarm 292 Doorbell 293 Siren 294 The sound of glass breaking 295 Baby crying 296 Door opening and closing 297 Dog barking 300 Implementation Forms 302 Audio Scene Detector 304 Audio Scene Change Detector 306 Scene change signal 308 Audio Scene Classifier 310 Noise Statistics 312 Non-stationary sound statistics 330 Multiple audio scene classes, classes Classes 332-346 332 Inside the home 334 Inside the office 336 Inside the restaurant 338 Inside a car 340 Inside the train 342 Street 344 Indoor 346 Outdoor 352 Scene Detector Output 400 Implementation Forms 414 Scene Transition Classifier 500 Implementation Forms 502 "In-vehicle" class 514-Layer Model Change Detector Reduced set of 530 class 600 cases 602 Sensor 606 Scene Detector 608 Environment 620 Camera 622 Input signals, input data, image data 624 Global Positioning System (GPS) receivers, GPS receivers, GPS sensors 626 Location information, location data 700 cases 702 Specific Class 800 cases Sets of 802-808 pre-trained data 802 First set of trained data 804 Second set of trained data 808 sets of trained data Set of 812-818 sound event classes 812 Sound Event Class, First Set 814 Sound Event Class, 2nd Set 818 The Nth set of sound event classes 900 Implementation Forms 902 Integrated Circuit 910 Sensor signal input 912 Data Output 1000 Implementation Forms 1002 vehicles 1100 Implementation Forms 1102 Vehicle 1200 Implementation Forms 1202 Virtual reality, augmented reality, or mixed reality headset, headset 1300 Implementation Forms 1302 Wearable Electronic Devices 1400 Wireless Speakers and Voice-Activated Devices 1402 Integrated Assistant Application 1404 Speaker 1500 ways 1600 devices 1602 Digital-to-Analog Converter 1604 Analog-to-Digital Converter 16 06 processor 1608 Speech and music codecs, speech and music codecs 1610 Processor 1622 System-in-package or system-on-chip device, system-on-chip device 1626 Display Controller 1628 displays 1630 Input Devices 1634 codec 1636 Voice Coder ("Vocoder") Encoder 1638 Vocoder Decoder 1640 Wireless Controller 1644 Power supply 1650 Transceiver 1652 Antenna 1656 command 1686 memory 1692 Speakers
Claims
1. A device for performing sound detection, A buffer configured to store audio data, A target sound detector comprising a first stage and a second stage, The first stage includes a binary target sound classifier configured to process the audio data, The first stage is configured to activate the second stage in response to the detection of a target sound by the first stage. The second stage is configured to receive the audio data from the buffer in response to the detection of the target sound. Target sound detector and An audio scene detector configured to be activated in response to the detection of the presence of multiple target sounds in the audio data by the binary target sound classifier, An audio scene change detector configured to process the aforementioned audio data and generate a scene change signal in response to the detection of an audio scene change, The system comprises an audio scene classifier configured to receive the audio data from the buffer in response to the detection of the audio scene change, Audio scene detector and One or more processors comprising Equipped with, The binary target sound classifier, It is further configured to generate a binary signal containing a first value and a second value, The first value is set to activate the second stage by generating an activation signal in response to the detection of the presence of at least one of a plurality of target sounds in the audio data. The second value is set to refrain from generating the activation signal in response to the detection that none of the multiple target sounds are present in the audio data. The second stage includes a multi-target sound classifier configured to generate a detector output indicating the presence or absence of each of a plurality of target sounds in the audio data. device.
2. The device according to claim 1, wherein the binary target sound classifier includes a neural network.
3. The device according to claim 1, wherein the binary target sound classifier includes at least one of a Bayesian classifier or a Gaussian mixed model (GMM) classifier.
4. The plurality of target sounds correspond to a plurality of classes of sound events, The device according to claim 1.
5. The device according to claim 1, wherein the binary target sound classifier and the buffer are included in a low-power domain and are configured to operate in an always-on mode, and the second stage is configured to transition from a low-power state to an active state in response to the reception of a signal.
6. The device according to claim 1, wherein the target sound detector is configured to select a specific set of trained data from one or more sets of trained data that corresponds to the environment detected by the device, and to process the audio data based on the specific set of trained data.
7. The device according to claim 6, wherein the environment is detected based on at least one of a camera, a location detection system, or an audio scene detector.
8. The device according to claim 1, wherein the audio scene classifier is configured to classify the audio data according to a plurality of audio scene classes, the plurality of audio scene classes including at least two of the following: inside a home, inside an office, inside a restaurant, inside a car, inside a train, on a street, indoors, or outdoors.
9. The device according to claim 1, wherein the audio scene change detector is further configured to detect the audio scene change based on detecting a change in at least one of noise statistics or non-stationary sound statistics.
10. The device according to claim 1, wherein the audio scene change detector includes a classifier trained using audio data corresponding to transitions between scenes.
11. The aforementioned audio scene detector corresponds to a hierarchical detector, The audio scene change detector is configured to detect the audio scene change based on detecting changes between audio scene classes within a first set of audio scene classes. The audio scene classifier is configured to classify the audio data according to a second set of audio scene classes, wherein the first count of the audio scene class in the first set of audio scene classes is less than the second count of the audio scene class in the second set of audio scene classes. The device according to claim 1.
12. A method for detecting a target sound, The steps include storing audio data in a buffer, The steps include processing the audio data in the buffer using a binary target sound classifier in the first stage of the target sound detector, The steps include activating the second stage of the target sound detector in response to the detection of the target sound by the first stage, The steps include processing the audio data from the buffer using the multi-target sound classifier in the second stage, The steps include activating an audio scene detector in response to the detection of the presence of multiple target sounds in the audio data by the binary target sound classifier, The steps include processing the audio data in the audio scene change detector, A step of generating a scene change signal in response to the detection of an audio scene change, In an audio scene classifier, the steps include receiving the audio data from the buffer in response to the detection of the audio scene change, The binary target sound classifier includes the steps of generating a binary signal containing a first value and a second value, The multi-target sound classifier included in the second stage comprises the step of generating a detector output for each of the multiple target sounds that indicates the presence or absence of that target sound in the audio data, The first value is set to activate the second stage by generating an activation signal in response to the detection of at least one of a plurality of target sounds in the audio data, The second value is set such that the activation signal is not generated in response to the detection that none of the plurality of target sounds are present in the audio data. method.
13. The steps include processing the audio data to detect audio scene changes based on detecting changes between audio scene classes within a first set of audio scene classes, A step of classifying the audio data based on a second set of audio scene classes It further includes, The first count of the audio scene class in the first set of the audio scene class is less than the second count of the audio scene class in the second set of the audio scene class. The method according to claim 12.
14. A computer-readable storage device for storing instructions, wherein when an instruction is executed by one or more processors, the one or more processors are configured to store instructions. Storing audio data in a buffer, Processing the audio data in the buffer using a binary target sound classifier in the first stage of the target sound detector, The second stage of the target sound detector is activated in response to the detection of the target sound by the first stage, Processing the audio data from the buffer using the multi-target sound classifier in the second stage, The audio scene detector is activated in response to the detection of the presence of multiple target sounds in the audio data by the binary target sound classifier. Processing the audio data in the audio scene change detector, To generate a scene change signal in response to the detection of audio scene changes, In an audio scene classifier, the audio data is received from the buffer in response to the detection of the audio scene change, In the binary target sound classifier, a binary signal is generated that includes a first value and a second value, In the multi-target sound classifier included in the second stage, for each of the multiple target sounds, a detector output is generated indicating the presence or absence of that target sound in the audio data. The first value is set to activate the second stage by generating an activation signal in response to the detection of the presence of at least one of a plurality of target sounds in the audio data. The second value is set such that the activation signal is not generated in response to the detection that none of the plurality of target sounds are present in the audio data. Computer-readable memory device.
15. The device according to claim 4, wherein the plurality of classes of sound events include at least two of the following: alarm, doorbell, siren, glass breaking, baby crying, door opening and closing, or dog barking.
16. The device according to claim 4, wherein the plurality of classes include sound events commonly encountered in a vehicle.
17. The device according to claim 16, wherein the plurality of classes of sound events include one or more of the following: opening and closing of vehicle doors, road noise, opening and closing of windows, radio, brakes, handbrake pulling or releasing sounds, wipers, turn signals, or engine starting sounds.
18. The device according to claim 1, wherein the audio scene change detector corresponds to a hierarchical scene change detector including a classifier configured to detect a relatively small number of broad classes, the audio scene classifier corresponds to a more powerful classifier, and the more powerful classifier is configured to detect a larger number of more specific environmental scenes.