SEQUENCE MODELS FOR AUDIO SCENE RECOGNITION

DE112020004052B4Active Publication Date: 2025-07-10NEC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
DE112020004052
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-08-19
Filing Date
2020-08-20
Publication Date
2025-07-10
Estimated Expiration
2040-08-20

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Computer-implemented method for classifying audio scenes, comprising: Generating (610) intermediate audio features from an input acoustic sequence; and Classifying (620), using a nearest neighbor search, segments of the input acoustic sequence based on the intermediate audio features to generate a final intermediate feature as a classification for the input acoustic sequence, each of the segments corresponding to a respective different one of different acoustic windows; wherein the generating step comprises: Learning (610A) the intermediate audio features from Mel-Frequency Cepstrum Coefficient (MFCC) features extracted from the input acoustic sequence; Dividing (610B) a scene into the different acoustic windows with different MFCC characteristics; and Feeding (610E) the MFCC features from each of the different acoustic windows into respective LSTM units such that a hidden state from each of the respective LSTM units is passed through an attention layer to identify feature correlations between hidden states at different time steps corresponding to different ones of the different acoustic windows.
Need to check novelty before this filing date? Find Prior Art

Description

INFORMATION REGARDING ASSOCIATED APPLICATIONS

[0001] This application claims priority to U.S. Non-Provisional Patent Application No. 16 / 997,314, filed August 19, 2020, U.S. Provisional Patent Application No. 62 / 915,022, filed August 27, 2019, and U.S. Provisional Patent Application No. 62 / 915,668, filed October 16, 2019, each of which is incorporated herein by reference in its entirety. BACKGROUNDTechnical field

[0002] The present invention relates to scene recognition and, in particular, to sequence models for audio scene recognition. Description of the related prior art

[0003] Audio (or acoustic) scene analysis is the task of identifying the category (or categories) of an environment using acoustic signals. The task of audio scene analysis can be formulated in two ways: (1) scene recognition, where the goal is to link or associate a single category with an entire scene (e.g., park, restaurant, train, etc.), and (2) event detection, where the goal is to detect shorter sound events in an audio scene (e.g., door knocking, laughter, keyboard click, etc.). Audio scene analysis has several important applications, some of which include, for example, multimedia retrieval (automatically identifying or tagging sports or music scenes); intelligent surveillance systems (identifying specific sounds in the environment); acoustic monitoring; searching audio archives; cataloging and indexing.An important step in audio scene analysis is processing the raw audio data with the goal of computing representative audio features that can be used to identify the correct categories (also known as the feature selection process). GUO, Jinxi [et al.]: Attention based CLDNNs for short-duration acoustic scene classification. In: Interspeech 2017, August 20-24, 2017, Stockholm, Sweden. pp. 469-473, DOI: http: / / dx.doi.org / 10.21437 / Interspeech.2017-440 , discloses a method for classifying audio scenes in which an audio sequence or scene is divided into different acoustic windows and a hidden state of LSTM units is passed through an attention layer. REN, Zhao [et al.]: Attention-based convolutional neural networks for acoustic scene classification. In: Detection and classification of Acoustic Scenes and Events 2018, 19-20, November 2018, Surrey, UK. 5 p. URL: https: / / openresearch.surrey.ac.uk / esploro / outputs / conferencePresentation / Attention-based-Convolutional-Neural-Networks-for-Acoustic / 99513850202346 discloses a method for classifying audio scenes in which MFCC features are extracted from an acoustic sequence and attention weighting is performed. SUMMARY

[0004] The invention is defined in the independent claims. The dependent claims define embodiments of the invention.

[0005] According to aspects of the present invention, a computer-implemented method for classifying audio scenes is provided. The method includes generating intermediate audio features from an input acoustic sequence. The method further includes classifying segments of the input acoustic sequence using a nearest neighbor search based on the intermediate audio features to generate a final intermediate feature as a classification for the input acoustic sequence. Each of the segments corresponds to a respective different one of different acoustic windows. The generating step includes learning the intermediate audio features from Mel-Frequency Cepstrum Coefficient (MFCC) features extracted from the input acoustic sequence. The generating step further includes splitting the same scene into the different acoustic windows with varying ones of the MFCC features.The generation step also includes feeding the MFCC features from each of the different acoustic windows into respective LSTM units such that a hidden state from each of the respective LSTM units is passed through an attention layer to identify feature correlations between hidden states at different time steps corresponding to different ones of the different acoustic windows.

[0006] According to other aspects of the present invention, a computer program product for classifying audio scenes is provided. The computer program product includes a non-transitory computer-readable storage medium having program instructions embodied thereon. The program instructions are computer-executable to cause the computer to perform a method. The method includes generating intermediate audio features from an input acoustic sequence. The method further includes classifying segments of the input acoustic sequence using a nearest neighbor search based on the intermediate audio features to generate a final intermediate feature as a classification for the input acoustic sequence, each of the segments corresponding to a respective different one of different acoustic windows.The generation step includes learning the intermediate audio features from mel-frequency cepstrum coefficient (MFCC) features extracted from the input acoustic sequence. The generation step further includes dividing the same scene into the different acoustic windows with varying MFCC features. The generation step also includes feeding the MFCC features from each of the different acoustic windows into respective LSTM units, such that a hidden state from each of the respective LSTM units is passed through an attention layer to identify feature correlations between hidden states at different time steps corresponding to different ones of the different acoustic windows.

[0007] According to still other aspects of the present invention, a computer processing system for classifying audio scenes is provided. The system includes a storage device for storing program code. The system further includes a hardware processor operatively coupled to the storage device for running the program code to generate intermediate audio features from an input acoustic sequence and, using a nearest neighbor search, classify segments of the input acoustic sequence based on the intermediate audio features to generate a final intermediate feature as a classification for the input acoustic sequence. Each of the segments corresponds to a respective different one of different acoustic windows.The hardware processor runs the program code to generate the intermediate audio features, to learn the intermediate audio features from mel-frequency cepstrum coefficient (MFCC) features extracted from the input acoustic sequence, to divide the same scene into the different acoustic windows with varying ones of the MFCC features, and to feed the MFCC features from each of the different acoustic windows into respective LSTM units, such that a hidden state from each of the respective LSTM units is passed through an attention layer to identify feature correlations between hidden states at different time steps corresponding to different ones of the different acoustic windows.

[0008] These and other features and advantages will become apparent from the following detailed description of illustrative embodiments thereof, which should be read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The disclosure will provide details in the following description of preferred embodiments with reference to the following figures, wherein: Fig. 1 is a block diagram illustrating an exemplary computing device according to an embodiment of the present invention; Fig. 2 is a flowchart illustrating an exemplary method for detecting audio scenes according to an embodiment of the present invention; Fig. 3 is a high-level diagram illustrating an exemplary audio scene detection architecture according to an embodiment of the present invention; Fig. 4 is a block diagram showing the intermediate audio feature learning portion of the Fig. 3 further shows, according to an embodiment of the present invention; Fig. 5 is a flowchart illustrating an exemplary method for the intermediate audio feature learning portion of the Fig. 3 further shows, according to an embodiment of the present invention; Fig. 6-7 are flowcharts illustrating an exemplary method for classifying audio scenes, according to an embodiment of the present invention; Fig. 8 is a flowchart illustrating an exemplary method for time series-based audio scene classification according to an embodiment of the present invention; Fig. 9 is a block diagram showing an exemplary triplet loss according to an embodiment of the present invention; Fig. 10 is a block diagram illustrating an exemplary scene-based precision evaluation approach according to an embodiment of the present invention; Fig. 11 is a block diagram illustrating another scene-based precision evaluation approach according to an embodiment of the present invention; and Fig. 12 is a block diagram illustrating an exemplary computing environment according to an embodiment of the present invention. DETAILED DESCRIPTION OF PREFERRED EMBODIMENTS

[0010] According to embodiments of the present invention, systems and methods for sequence models for audio scene recognition are provided.

[0011] Time series analysis is an important branch of data science that deals with the analysis of data collected from one or more sensors over time. Based on the observation that audio data is a time series, the present invention provides an end-to-end architecture for using time series data analysis to analyze audio data.

[0012] The observation underlying various embodiments of the present invention is that the basic audio features of an audio scene (obtained after signal processing) form a multiply varied time series, where each feature corresponds to a sensor and its value represents the sensor readings over time.

[0013] According to one or more embodiments of the present invention, a multivariate time series analysis tool called Data2Data (D2D) is provided. D2D learns representations (or embeddings) of time series data and uses them to perform fast retrieval, i.e., given a query time series segment, to identify the most similar historical time series segment. Retrieval is an important building block for time series classification.

[0014] One or more embodiments of the present invention provide audio scene analysis for time series analysis. To interpret audio scenes as time series data, the audio scenes can be fed into the D2D platform for rapid retrieval for classification and anomaly detection.

[0015] Thus, one or more embodiments of the present invention present a deep learning framework for accurately classifying an audio environment after listening for less than one second. The framework relies on a combination of recurrent neural networks and attention to learn embeddings for each audio segment. A key feature of the learning process is an optimization mechanism that minimizes an audio loss function. This function is designed to promote embeddings to preserve segment similarity (through a distance-based component) and penalize indeterminate segments while capturing the importance of more relevant ones (through an importance-based component).

[0016] One or more embodiments of the present invention generate intermediate audio features and classify them using a nearest neighbor classifier. The intermediate audio features attempt to both capture correlations between different acoustic windows in the same scene and isolate and mitigate the effect of "uninteresting" features / portions, such as silence or noise. To learn the intermediate audio features, basic mel-frequency cepstrum coefficient (MFCC) audio features are first generated. Then, the entire scene is divided into (possibly overlapping) windows, and the basic features from each window are fed into LSTM units.The hidden state of each LSTM unit (there are as many hidden states as there are time steps in the current window) is taken and passed through an attention layer to identify correlations between the states at different time steps. To generate the final intermediate feature for each window, the triplet loss function is optimized, to which a regularization parameter calculated on the last element of each intermediate feature is added. The goal of the regularization parameter is to reduce the importance of the silent segments.

[0017] Thus, one or more embodiments of the present invention address audio scene classification (ASC), i.e., the task of identifying the category of the environment using acoustic signals.

[0018] To achieve the goal of early detection, ASC is formulated as a retrieval problem. This allows us to divide the audio data into short segments (less than one second), learn embeddings for each segment, and use the embeddings to classify each segment as it is "heard." Given a query segment (e.g., a short ambient sound), the query segment is classified into the class of the most similar historical segment according to an embedding similarity function, such as the Euclidean distance.

[0019] A natural question is how to find embeddings that enable fast and accurate retrieval of short audio segments. Good embeddings must meet two criteria. First, they must maintain similarity: segments belonging to the same audio scene category should have similar embeddings. Second, they must capture the importance of each segment within a scene. For example, in a playground scene, segments containing children's laughter are more relevant to the scene; in contrast, segments containing silence or white noise are less important, as they are found in many other types of scenes.

[0020] Fig. Figure 1 is a block diagram illustrating an exemplary computing device 100 according to an embodiment of the present invention. Computing device 100 is configured to perform audio scene recognition.

[0021] Computing device 100 may be embodied as any type of computing device capable of performing the functions described herein, including, without limitation, a computer, a server, a rack-based server, a blade server, a workstation, a desktop computer, a laptop computer, a notebook computer, a tablet computer, a mobile computing device, a portable computing device, a network device, a web device, a distributed computing system, a processor-based system, and / or a consumer electronics device. Additionally or alternatively, computing device 100 may be embodied as one or more compute sleds, storage sleds, or other racks, sleds, computing enclosures, or other components of a physically disaggregated computing device. As described in Fig. 1, computing device 100 illustratively includes processor 110, an input / output subsystem 120, memory 130, a data storage device 140, and a communications subsystem 150, and / or other components and devices typically found in a server or similar computing device. Of course, in other embodiments, computing device 100 may include other or additional components, such as those typically found in a server computer (e.g., various input / output devices). Additionally, in some embodiments, one or more of the illustrative components may be incorporated into or otherwise form a portion of another component. For example, memory 130, or portions thereof, may be incorporated into processor 110 in some embodiments.

[0022] Processor 110 may be embodied as any type of processor capable of performing the functions described herein. Processor 110 may be embodied as a single processor, multiple processors, central processing unit(s) (CPU(s)), graphics processing unit(s) (GPU(s)), single- or multi-core processor(s), digital signal processor(s), microcontroller(s), or other processor(s) or processing / control circuit(s).

[0023] The memory 130 may be embodied as any type of volatile or non-volatile memory or data storage capable of performing the functions described herein. In operation, the memory 130 may store various data and software used during operation of the computing device 100, such as operating systems, applications, programs, libraries, and drivers. The memory 130 is communicatively coupled to the processor 110 via the I / O subsystem 120, which may be embodied as circuitry and / or components to enable input / output operations with the processor 110, the memory 130, and other components of the computing device 100. For example, the I / O subsystem 120 may be embodied as memory control hubs, input / output control hubs, platform control hubs, integrated control circuitry, firmware devices, communication links (e.g.,Point-to-point connections, bus connections, wires, cables, optical fibers, circuit board traces, etc.) and / or other components and subsystems are embodied or otherwise included to enable or facilitate the input / output operations. In some embodiments, the I / O subsystem 120 may form a portion of a system-on-a-chip (SOC) and may be incorporated onto a single integrated circuit chip along with the processor 110, the memory 130, and other components of the computing device 100.

[0024] The data storage device 140 may be embodied as any type of device or devices configured for the short-term or long-term storage of data, such as memory devices and circuits, memory cards, hard disk drives, solid-state drives, or other data storage devices. The data storage device 140 may store program code for audio scene recognition / classification. The program code may control a hardware-based device in response to a recognition / classification result. The communications subsystem 150 of the computing device 100 may be embodied as any network interface controller or other communications circuitry, device, or collection thereof that can enable communications between the computing device 100 and other remote devices over a network.The communication subsystem 150 may be configured to use any one or more communication technologies (e.g., wired or wireless communications) and associated protocols (e.g., Ethernet, InfiniBand®, Bluetooth®, Wi-Fi®, WiMAX, etc.) to effect such communication.

[0025] As shown, computing device 100 may also include one or more peripheral devices 160. Peripheral devices 160 may include any number of additional input / output devices, interface devices, and / or other peripheral devices. For example, in some embodiments, peripheral devices 160 may include a display, a touch screen, graphics circuitry, a keyboard, a mouse, a speaker system, a microphone, a network interface, and / or other input / output devices, interface devices, and / or peripheral devices.

[0026] Of course, computing device 100 may also include other elements (not shown), as would be readily contemplated by one of ordinary skill in the art, as well as omit certain elements. For example, various other input devices and / or output devices may be included in computing device 100, depending on the particular implementation thereof, as would be readily understood by one of ordinary skill in the art. For example, various types of wireless and / or wired input and / or output devices may be used. Furthermore, additional processors, controllers, memory, and so forth may also be used in various configurations. These and other variations of processing system 100 are readily contemplated by one of ordinary skill in the art in light of the teachings of the present invention provided herein.

[0027] As used herein, the term "hardware processor subsystem" or "hardware processor" may refer to a processor, memory (including RAM, cache(s) and so on), software (including memory management software), or combinations thereof that work together to perform one or more specific tasks. In useful embodiments, the hardware processor subsystem may include one or more data processing elements (e.g., logic circuits, processing circuits, instruction execution devices, etc.). The one or more data processing elements may be included in a central processing unit, a graphics processing unit, and / or a separate processor- or computer element-based controller (e.g., logic gates, etc.). The hardware processor subsystem may include one or more onboard memories (e.g., caches, specific ordedicated memory arrays, read-only memory, etc.). In some embodiments, the hardware processor subsystem may include one or more memories that may be onboard or offboard, or that may be dedicated or dedicated to use by the hardware processor subsystem (e.g., ROM, RAM, BIOS (Basic Input / Output System), etc.).

[0028] In some embodiments, the hardware processor subsystem may include and execute one or more software elements. The one or more software elements may include an operating system and / or one or more applications and / or specific code to achieve a specified result.

[0029] In other embodiments, the hardware processor subsystem may include dedicated, specialized circuitry that performs one or more electronic processing functions to achieve a particular result. Such circuitry may include one or more application-specific integrated circuits (ASICs), FPGAs, and / or PLAs.

[0030] These and other variations of a hardware processor subsystem are also contemplated according to embodiments of the present invention.

[0031] Fig. 2 is a flowchart illustrating an exemplary method 200 for audio scene recognition, according to an embodiment of the present invention.

[0032] At a block 210, raw audio data is input.

[0033] At a block 220, the raw audio data is processed to extract basic audio features therefrom.

[0034] At a block 230, time series processing is performed to obtain audio segments.

[0035] At a block 240, a time series analysis is performed to obtain audio segment representations.

[0036] In a block 250, the audio segment representations are stored in a database.

[0037] At a block 260, an action is performed in response to the audio segment representations.

[0038] Now, various of the blocks of method 200 will be described in more detail according to various embodiments of the present invention.

[0039] Raw audio processing (block 220). The input is the raw audio data, and the output is represented by basic audio features obtained after applying signal processing techniques. The audio data is processed by applying several transforms as follows. First, the signal is broken down or split into several overlapping frames, each frame being 25 ms in size. Then, the fast Fourier transform is applied to each frame to extract the energy levels for each frequency present in the sound. The frequency levels are then mapped to the mel scale to better match the hearing capabilities of the human ear. Finally, the cosine transform is applied to the logs of the mel powers to obtain the mel frequency cepstrum coefficients (MFCCs). MFCCs are powerful basic audio features for scene recognition.Alternatively, the procedure can be terminated after applying the FFT and the frequency spectrum powers can be used as basic audio features.

[0040] Time series-based processing (block 230). The entire training data is now represented as baseline audio feature vectors over time. If each feature is considered equivalent to a sensor and the values of the feature over time are considered as values collected by the sensor, the entire training data can be viewed as a multiply varied time series. The data is divided into several, possibly overlapping segments. Each segment contains all baseline audio feature vectors over a user-defined period of time. Dividing the data into overlapping short-range windows is typical for time series analysis and allows for better capture of short-range dependencies and correlations in sound.

[0041] Time series analysis (block 240). Each audio segment is fed into our Data2Data (D2D) framework. Each base audio feature vector in a segment is the input to an LSTM unit. The unit continuously updates its state as it reads more and more audio features. The final output of the LSTM is the representation of the segment, capturing dependencies between the audio feature vectors that are part of the segment. All representations are saved to a database and used for later retrieval.

[0042] Fig. 3 is a high-level diagram illustrating an exemplary audio scene detection architecture 300, according to an embodiment of the present invention.

[0043] The audio scene recognition architecture 300 includes a raw audio data loading section 310, a raw audio processing section 320, a basic audio feature segmentation section 330, and an intermediate audio feature learning section 340.

[0044] The basic audio feature segmentation portion 330 includes an audio segment 331.

[0045] The intermediate audio feature learning sub-area 340 includes an LSTM sub-area 341, an attention sub-area 342, and a final representation (feature) sub-area 343.

[0046] Raw Audio Data Load 310. This element loads the set of audio scenes used for training and their labels from a file. In one embodiment, the data is in wav format. Of course, other formats can be used. All training data is concatenated so that it appears as one long audio scene.

[0047] Raw audio processing subsection 320. The audio data is processed by applying several transforms as follows. First, the signal is split into several overlapping frames, each frame being 25 ms long. The fast Fourier transform is applied to each frame to extract the energy levels for each frequency present in the sound. The frequency levels are mapped onto the mel scale to better match the hearing capabilities of the human ear. Finally, the cosine transform is applied to the logs of the mel powers to obtain the mel frequency cepstrum coefficients (MFCCs). Previous research has shown that MFCCs are strong baseline audio features for scene recognition. Alternatively, the procedure can be terminated after applying the FFT, and the frequency spectrum powers can be used as baseline audio features.

[0048] Base audio feature segmentation sub-area 330. The entire training data is now represented as a vector of base audio feature vectors. To capture dependencies among different base audio feature vectors, the data is divided into several, possibly overlapping segments. Each segment contains all base audio feature vectors over a user-defined time period.

[0049] Intermediate audio feature learning sub-area 340. Each audio segment is fed into a deep architecture composed of a recurrent layer and an attention layer.

[0050] LSTM sub-area 341. Each base audio feature vector in a segment is the input to an LSTM unit. The unit continuously updates its state as it reads more and more audio features. The final output of the LSTM unit can be thought of as a representation of the segment, capturing long-term dependencies between the audio feature vectors that are part of the segment. A bidirectional LSTM is used, meaning that each segment is fed in time order and in reverse time order, yielding two final representations.

[0051] Attention sub-area 342. The two final representations obtained from the recurrent layer may not be sufficient to capture all correlations between base feature vectors of the same segment. An attention layer is used to identify correlations between LSTM states at different times. The input to the attention layer is represented by the hidden states of the LSTM across all time steps of the segment.

[0052] Subdomain for a final representation (feature) 343. To obtain the final intermediate feature 350, the two final outputs of LSTM are concatenated and the results are multiplied by the attention weights.

[0053] Fig. 4 is a block diagram illustrating the intermediate audio feature learning portion 340 of the Fig. 3 further shows, according to an embodiment of the present invention.

[0054] Fig. Figure 4 shows the optimization step used to learn the intermediate audio features. Each iteration of learning attempts to minimize a loss function calculated using the current intermediate features of a randomly selected batch of segments. The weights and perceptual biases of the deep network (block 340) are backpropagated and updated. The loss function 410 to be minimized is composed of two different quantities, as follows: Loss = Audio Triplet Loss + Silence Regularization

[0055] Audio Triplet Loss 410 is based on the classic triplet loss. To compute the triplet loss 410, two segments that are part of the same class and one that is part of a different class are selected, and an attempt is made to bring the intermediate features 405 of the segments of the same class closer together and those of the other class farther apart. The silence weight is defined as the last element in the representation of each segment. The silence weight is likely to be low if the segment is silence. The audio Triplet Loss 410 is computed by multiplying the triplet loss by the silence weights of each of the segments in the triplet. The idea behind this is that silence segments, even if they are part of different classes, are similar and should not contribute to learning (i.e., their representations should not be pushed apart or away from each other by the optimization).

[0056] In addition to a triplet loss, a new term is added, called silence regularization. Silence regularization is the sum of the silence weights and is intended to prevent the silence weights from all becoming 0 at the same time.

[0057] Fig. 5 is a flowchart illustrating intermediate audio feature learning portion 340 of the Fig. 3 further shows, according to an embodiment of the present invention.

[0058] In a block 510, the Fourier transform of the audio scene is calculated.

[0059] In a block 520, the powers of the spectrum obtained above are mapped to the Mel scale.

[0060] At a block 530, the logs of the powers at each of the MEL frequencies are calculated.

[0061] At a block 540, the discrete cosine transform of the list of MEL protocol powers is calculated.

[0062] At block 550, the MFCCs are calculated as the amplitudes of the resulting spectrum. Three components of our audio scene classification architecture will now be described as follows: raw audio processing to generate baseline audio features; the encoder to compute high-level audio segment representations; and loss function optimization to guide the computation of good embeddings. Some of the contributions of the present invention lie in encoder and loss optimization.

[0063] A description will now be given regarding raw audio processing according to an embodiment of the present invention.

[0064] Each audio scene is decomposed using windowed FFT and 20 mel-frequency cepstrum coefficients are extracted. Their first derivatives are added, along with 12 harmonic and percussive features known to augment the raw feature set, to obtain 52 baseline audio features for each FFT window.

[0065] One should X = (x 1 , x 2 , ..., x n ) T ∈ R n×TLet an audio segment of length T (e.g., T consecutive FFT windows) be represented with n base features (where n = 52). Each segment is associated with the label of the scene to which it belongs. One goal is audio segment retrieval: given a query segment, one finds the most similar historical segments using a similarity measurement function, such as the Euclidean distance. The query segment is then classified into the same category as the most similar historical segment.

[0066] A description will now be given regarding a learning embedding according to an embodiment of the present invention.

[0067] To enable fast and efficient retrieval, compact representations are learned for each historical audio segment and compared to the representations rather than the base audio features. The embedding is assumed to be given by the following mapping function: h=F(X) where X ∈ R n×T is an audio segment of n basis features over T time steps and h ∈ R d is an embedding vector of size d. F is a nonlinear mapping function.

[0068] A combination of bidirectional LSTM and attention is used to F An LSTM is selected to capture long-term temporal dependencies and attention to highlight the more important audio parts in a segment. To capture correlations between audio at different time steps in a segment, all hidden states of the LSTM from each time step are fed into an attention layer, which evaluates the importance of each time step using a nonlinear score function attn. score(ht = tanh(h t V + b). V and b are learned together with F. The scores are normalized using Softmax as follows: at=exp(attnscore(ht))∑i=1Texp(attnscore(ht)) and the embedding of the segment is calculated as a weighted average of each hidden state: h=∑i=1Tatht

[0069] Our coding architecture is reminiscent of neural machine translation in that it combines LSTM and attention. However, it computes self-attention between the encoder's hidden states rather than attention between the decoder's current state and the encoder's hidden states.

[0070] Other deep encoders that preserve similarity between audio segments can be used to compute embeddings; recurrent networks and an attention mechanism are efficient at identifying important features in audio. The present invention focuses on providing accurate early detection given a reasonably accurate encoder.

[0071] A description will now be given regarding the loss according to an embodiment of the present invention.

[0072] The loss function is designed to meet two criteria. First, it must promote embeddings to reflect class membership. In other words, segments that are part of the same class should have similar embeddings, and segments that are part of different classes should have different embeddings. This goal is achieved by using a distance-based component, such as the triplet loss: Lsimilarity=max(‖F(a)−F(p)‖2−‖F(a)−F(n)‖2+α,0) where a, p, and n ∈ X are audio segments such that a and p have the same label and a and n have different labels. The second criterion is shaped by our goal of quickly classifying scenes. It is desirable to be able to detect or recognize as few ambient sounds as possible after listening. Thus, it is desirable to emphasize the segments that can distinguish a scene (e.g., children's laughter in a playground scene) and downplay those that are less descriptive (e.g., silence, white noise). To capture the importance of each segment, an audio importance score is defined. The importance score is a linear projection of the segment embedding that is learned jointly with the encoder. The score is normalized using softmax, similar to Equation 2, to determine the importance weight, w iof each segment and use them to calculate the total loss: L=(∏wi)similarity+αaudio(−∑wi) where w i represents the weightings of the segments used to Similarity to calculate, e.g. a, p and n from equation 4, and α audio is a regularization parameter. The first term of the equation ensures that only important segments are used in the triplet loss calculation, while the second term attempts to maximize the weights of such segments.

[0073] Attention and importance assessments complement each other in highlighting different segments within an audio scene. Attention assessment helps identify useful time steps within a segment, while importance assessment helps retrieve relevant segments within a scene.

[0074] Fig. 6-7 are flowcharts illustrating an exemplary method 600 for classifying audio scenes, according to an embodiment of the present invention.

[0075] At block 610, intermediate audio features are generated that both capture correlations between different acoustic windows in a same scene and isolate and mitigate the effect of uninteresting features in the same scene. In one embodiment, the uninteresting features may include silence and / or noise. The intermediate audio features are generated to isolate and mitigate the effect of uninteresting features in the same scene by using a triplet loss that pushes different classes further apart than similar classes in a classification space.

[0076] In one embodiment, block 610 may include one or more of blocks 510A through 510C.

[0077] At a block 610A, the intermediate audio features learn from Mel-Frequency Cepstrum Coefficient (MFCC) features extracted from an input acoustic sequence.

[0078] At block 610B, the same scene is divided into different acoustic windows with varying MFCC features. In one embodiment, the entire scene may be divided into overlapping windows to exploit dependencies between windows.

[0079] At block 610C, the input acoustic sequence is preprocessed by applying a fast Fourier transform (FFT) to each of the different acoustic windows to extract respective acoustic frequency energy levels therefor. In one embodiment, the respective acoustic frequency energy levels may be used as the intermediate audio features.

[0080] At a block 610D (in the case where the respective acoustic frequency energy levels are not used as the intermediate audio features), mapping the respective acoustic frequency energy levels to a mel scale to match human hearing capabilities and applying a cosine transform to logs of the respective acoustic frequency energy levels to obtain the MFCC features.

[0081] At block 610E, the MFCC features from each of the different acoustic windows are fed into LSTM units such that a hidden state from each of the LSTM units is passed through an attention layer to identify correlations between hidden states at different time steps corresponding to different ones of the different acoustic windows. In one embodiment, the LSTM units may contain as many hidden states as there are time steps in a given current one of the windows.

[0082] At a block 620, segments of an input acoustic sequence are classified using a nearest neighbor search based on the intermediate audio features to generate a final intermediate feature as a classification for the input acoustic sequence.

[0083] In one embodiment, block 620 may include one or more of blocks 620A and 620B.

[0084] At a block 620A, the final intermediate feature is generated for each of the different acoustic windows by optimizing a triplet loss function to which a regularization parameter calculated for each of the intermediate audio features is added to reduce an importance of the uninteresting features.

[0085] At a block 620B, the final intermediate classification is determined by majority vote on classifications for the segments forming the input acoustic sequence.

[0086] In one embodiment, the regularization parameter is computed on a last element of each of the intermediate audio features, where the last element is a silence weight.

[0087] At a block 630, a hardware device is controlled to perform an action in response to a classification.

[0088] Example actions may include, but are not limited to, detecting anomalies in computer processing systems and controlling the system in which an anomaly is detected. For example, a query in the form of acoustic time series data from a hardware sensor or a sensor network (e.g., a mesh) may be characterized as anomalous behavior (dangerous or otherwise excessive operating speed (e.g., motor, gear linkage), dangerous or otherwise excessive operating heat (e.g., motor, gear linkage), dangerous or otherwise out-of-tolerance alignment (e.g., motor, gear linkage, etc.)) using a text message as a label / classification compared to historical sequences. Accordingly, a potentially faulty device may be turned off, its operating speed reduced, an alignment (e.g.,hardware-based) procedure and so on, based on the implementation.

[0089] Another example action may be operating parameter tracing, where a history of parameter change over time may be logged, as used to perform other functions, such as hardware machine control functions including turning on or off, slowing down, speeding up, positional adjustment, and so on, upon detection of a given operating condition that matches a given output classification.

[0090] Example environments where the present invention may be employed include, but are not limited to, power plants, information technology systems, manufacturing facilities, computer processing systems (e.g., server farms, storage pools, etc.), multimedia retrieval (automatic tagging of sports or music scenes), intelligent surveillance systems (identifying specific sounds in the environment), acoustic monitoring, audio archive searching, cataloging and indexing, and so forth. These and other environments are readily contemplated by one of ordinary skill in the art in light of the teachings of the present invention provided herein.

[0091] Fig. 8 is a flowchart illustrating an exemplary method 800 for time series-based classification of audio scenes, according to an embodiment of the present invention.

[0092] At a block 810, intermediate audio features are generated from respective segments of an input acoustic time series for a same scene captured by a sensor device.

[0093] In one embodiment, block 810 includes one or more of blocks 810A through 810C.

[0094] At a block 810A, the intermediate audio features are learned from Mel-Frequency Cepstrum Coefficient (MFCC) features extracted from the input acoustic time series.

[0095] In a block 810B, the same scene is divided into the different acoustic windows with different MFCC characteristics.

[0096] At a block 810C, the MFCC features from each of the different acoustic windows are fed into respective LSTM units such that a hidden state from each of the respective LSTM units is passed through an attention layer to identify feature correlations between hidden states at different time steps corresponding to different ones of the different acoustic windows.

[0097] At block 820, the respective segments of the input acoustic time series are classified using a nearest neighbor search based on the intermediate audio features to generate a final intermediate feature as a classification for the input acoustic time series. Each of the respective segments corresponds to a respective different one of the different acoustic windows.

[0098] At block 830, replacing a hardware device monitored by the sensor occurs in response to the final intermediate feature. Or performing another action, such as any of the example actions described herein, with respect to a resulting classification.

[0099] Fig. 9 is a block diagram showing an exemplary triplet loss 900, according to an embodiment of the present invention.

[0100] The triplet loss contains a sampling triplets formed from samples of anchors, positives and negatives.

[0101] The following applies to scanning negatives.

[0102] Random: Random sampling from another class.

[0103] Semi-hard negatives 901: Sampling negatives that are no closer to the anchor than the positive from another class, i.e. d(a,p) < d(a,n) < d(a, p) + remainder.

[0104] Hard Negatives 902: Sampling negatives that are closer to the anchor than the positive from another class, i.e. d(a,n) < d(a, p).

[0105] Fig. 10 is a block diagram illustrating an exemplary scene-based precision evaluation approach 1000, according to an embodiment of the present invention.

[0106] The approach 1000 includes predicted scene labels 1001, predicted segment labels 1002, and true scene labels 1003.

[0107] Approach 1 (scene-based precision): If more than half of the segments are correctly predicted for each audio scene, this scene is considered to be correctly predicted. Precision=True PositiveTrue Positive+False Positive

[0108] Fig. 11 is a block diagram illustrating another scene-based precision evaluation approach 1100, according to an embodiment of the present invention.

[0109] The approach 1100 includes predicted scene labels 1101, predicted segment labels 1102, and true scene labels 1103.

[0110] Approach 2 (scene-based precision): For each audio scene, the two most frequently predicted segment labels are counted. If a correct label for that audio scene falls within these two labels, then that scene is determined to be correctly predicted.

[0111] Fig. 12 is a block diagram illustrating an exemplary computing environment 1200, according to an embodiment of the present invention.

[0112] The environment 1200 includes a server 1210, a plurality of client devices (collectively designated by reference numeral 1220), a controlled system A 1241, a controlled system B 1242.

[0113] Communication between the entities of environment 1200 may be performed via one or more networks 1230. For illustrative purposes, a wireless network 1230 is shown. In other embodiments, any of wired, wireless, and / or a combination thereof may be used to enable or facilitate communication between the entities.

[0114] Server 1210 receives time series data from client devices 1220. Server 1210 may control one of systems 1241 and / or 1242 based on a prediction thereby made. In one embodiment, the time series data may be data related to the controlled systems 1241 and / or 1242, such as, but not limited to, sensor data.

[0115] Embodiments described herein may be entirely hardware, entirely software, or include both hardware and software elements. In a preferred embodiment, the present invention is implemented in software, including, but not limited to, firmware, resident software, microcode, etc.

[0116] Embodiments may include a computer program product accessible from a computer-usable or computer-readable medium that provides program code for use by, or in connection with, a computer or any instruction execution system. A computer-usable or computer-readable medium may include any device that stores, communicates, propagates, or transports the program for use by, or in connection with, the instruction execution system, apparatus, or device. The medium may be a magnetic, optical, electronic, electromagnetic, infrared, or semiconductor system (or apparatus or device) or propagation medium.The medium may include a computer-readable storage medium such as semiconductor or solid-state memory, magnetic tape, removable computer diskette, random access memory (RAM), read-only memory (ROM), fixed magnetic disk, and optical disk, etc.

[0117] Each computer program may be tangibly stored in a machine-readable storage medium or device (e.g., a program memory or a magnetic disk) readable by a general or special-purpose programmable computer, for configuring and controlling operation of a computer when the storage medium or device is read by the computer to perform the procedures described herein. The inventive system may also be considered to be embodied in a computer-readable storage medium configured with a computer program, where the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.

[0118] A data processing system capable of storing and / or executing program code may include at least one processor coupled directly or indirectly to storage elements through a system bus. The storage elements may include local memory used during actual execution of the program code, mass storage, and cache memory that provides temporary storage of at least some program code to reduce the number of times code is retrieved from mass storage during execution. Input / output or I / O devices (including, but not limited to, keyboards, displays, pointing devices, etc.) may be coupled to the system either directly or through intervening I / O controllers.

[0119] Network adapters may also be coupled to the system to enable the computing system to interface with other computing systems or remote printers or storage devices via intervening private or public networks. Modems, cable modems, and Ethernet cards are just a few of the types of network adapters currently available.

[0120] Reference in the specification to "a single embodiment" or "an embodiment" of the present invention, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the appearances of the phrase "in a single embodiment" or "in an embodiment," as well as any other variations that appear at various points throughout the specification, are not necessarily all referring to the same embodiment. However, it is to be understood that features from one or more embodiments may be combined with the given teachings of the present invention provided herein.

[0121] It should be understood that the use of any of the following " / ", "and / or", and "at least one of", such as in the cases of "A / B", "A and / or B", and "at least one of A and B", is intended to encompass only the selection of the first listed option (A), or the selection of the second listed option (B), or the selection of both options (A and B). As another example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such language is intended to encompass only the selection of the first listed option (A), or only the selection of the second listed option (B), or only the selection of the third listed option (C), or only the selection of the first and second listed options (A and B), or only the selection of the first and third listed options (A and C), or only the selection of the second and third listed options (B and C), or only the selection of all three options (A, B, and C).This can be expanded for as many items as are listed.

[0122] The foregoing is to be considered in all respects as illustrative and exemplary, but not restrictive, and the scope of the invention disclosed herein is to be determined not from the detailed description, but from the claims as interpreted according to the full breadth allowed by the patent laws. It is to be understood that the embodiments shown and described herein are merely illustrative of the present invention, and that those skilled in the art may implement various modifications without departing from the scope and spirit of the invention. Those skilled in the art could implement various other combinations of features without departing from the scope and spirit of the invention.Having thus described aspects of the invention with the detail and particularity required by the patent laws, what is claimed and desired to be protected by the patent is set forth in the appended claims.

Claims

[1] Computer-implemented method for classifying audio scenes, comprising: Generating (610) intermediate audio features from an input acoustic sequence; and Classifying (620), using a nearest neighbor search, segments of the input acoustic sequence based on the intermediate audio features to generate a final intermediate feature as a classification for the input acoustic sequence, each of the segments corresponding to a respective different one of different acoustic windows; wherein the generating step comprises: Learning (610A) the intermediate audio features from Mel-Frequency Cepstrum Coefficient (MFCC) features extracted from the input acoustic sequence; Dividing (610B) a scene into the different acoustic windows with different MFCC characteristics; and Feeding (610E) the MFCC features from each of the different acoustic windows into respective LSTM units such that a hidden state from each of the respective LSTM units is passed through an attention layer to identify feature correlations between hidden states at different time steps corresponding to different ones of the different acoustic windows. [2] The computer-implemented method of claim 1, wherein the intermediate audio features both capture feature correlations between different acoustic windows in a scene and isolate and mitigate an effect of uninteresting features in the same scene. [3] The computer-implemented method of claim 2, wherein the classifying step comprises generating the final intermediate feature for each of the different acoustic windows by optimizing a triplet loss function to which a regularization parameter calculated on each of the intermediate audio features is added to reduce an importance of the uninteresting features, and wherein the uninteresting features include silence. [4] The computer-implemented method of claim 3, wherein the triplet loss function adjusts a triplet selection algorithm to avoid using the uninteresting features as silence and noise by using a silence and noise perceptual distortion. [5] The computer-implemented method of claim 3, wherein the regularization parameter is calculated on a last element of each of the intermediate audio features, the last element being a silence weight. [6] The computer-implemented method of claim 3, wherein the regularization parameter comprises a sum of silence weights and prevents all silence weights from reaching a value of zero at the same time. [7] The computer-implemented method of claim 1, wherein an entirety of a scene is divided into overlapping windows to exploit dependencies between windows. [8] The computer-implemented method of claim 1, further comprising controlling a hardware device to perform an action response to a classification of a scene. [9] The computer-implemented method of claim 1, wherein the intermediate audio features are generated to isolate and mitigate the effect of uninteresting features in a scene using a triplet loss that pushes different classes farther apart than similar classes in a classification space. [10] The computer-implemented method of claim 1, further comprising calculating an embedding of the input acoustic sequence as a weighted average of each of the hidden states. [11] The computer-implemented method of claim 10, wherein the embedding is the final intermediate feature. [12] The computer-implemented method of claim 1, further comprising receiving a query segment and finding a most similar historical segment using a nearest neighbor. [13] The computer-implemented method of claim 1, wherein the respective LSTMs are bidirectional and feed segments of the input acoustic sequence in time order and in reverse time order to provide two final representations. [14] The computer-implemented method of claim 13, wherein the final intermediate feature for a given one of the segments is obtained by concatenating the two final representations multiplied by the attention weights determined in the attention layer. [15] The computer-implemented method of claim 1, wherein the learning step learns the intermediate audio features by minimizing a loss function calculated using the intermediate audio features of a randomly selected batch of segments from the input acoustic sequence. [16] The computer-implemented method of claim 1, wherein the final intermediate feature is determined by majority vote on classifications for the segments forming the input acoustic sequence. [17] The computer-implemented method of claim 1, wherein each intermediate audio feature represents a sensor. [18] A computer program product for classifying audio scenes, the computer program product comprising a non-transitory computer-readable storage medium having program instructions embodied thereon, the program instructions being executable by a computer to cause the computer to perform a method comprising: Generating (610) intermediate audio features from an input acoustic sequence; and Classifying (620), using a nearest neighbor search, segments of the input acoustic sequence based on the intermediate audio features to generate a final intermediate feature as a classification for the input acoustic sequence, each of the segments corresponding to a respective different one of different acoustic windows; wherein the generating step comprises: Learning (610A) the intermediate audio features from Mel-Frequency Cepstrum Coefficient (MFCC) features extracted from the input acoustic sequence; Dividing (610B) a scene into the different acoustic windows with different MFCC characteristics; and Feeding (610E) the MFCC features from each of the different acoustic windows into respective LSTM units such that a hidden state from each of the respective LSTM units is passed through an attention layer to identify feature correlations between hidden states at different time steps corresponding to different ones of the different acoustic windows. [19] The computer program product of claim 18, wherein the intermediate audio features both capture feature correlations between different acoustic windows in a scene and isolate and mitigate an effect of uninteresting features in the same scene. [20] Computer processing system for the classification of audio scenes, comprising: a storage device (130) that stores a program code; and a hardware processor (110) operatively coupled to the storage device that executes the program code to to generate intermediate audio features from an input acoustic sequence; and, using a nearest neighbor search, classify segments of the input acoustic sequence based on the intermediate audio features to generate a final intermediate feature as a classification for the input acoustic sequence, each of the segments corresponding to a respective different one of different acoustic windows; wherein the hardware processor executes the program code to generate the intermediate audio features to learn the intermediate audio features from Mel-Frequency Cepstrum Coefficient (MFCC) features extracted from the input acoustic sequence; to divide a scene into different acoustic windows with varying MFCC characteristics; and feeding the MFCC features from each of the different acoustic windows into respective LSTM units, such that a hidden state from each of the respective LSTM units is passed through an attention layer to identify feature correlations between hidden states at different time steps corresponding to different ones of the different acoustic windows.

Citation Information

Patent Citations

  • US-PATENTANMELDUNGNR.62/915,022

  • US-PATENTANMELDUNGNR.16/997,314

  • US-PATENTANMELDUNGNR.62/915,668