Method and system for detecting abnormal sounds
By employing an attentive neural process architecture to analyze masked regions of the spectrogram and calculate anomaly scores, the method effectively addresses inefficiencies and inaccuracies in existing abnormal sound detection techniques, achieving improved detection efficiency and accuracy.
Patent Information
- Application Number
- JP2024531746
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-19
- Filing Date
- 2022-05-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-05-12
AI Technical Summary
Existing methods for detecting abnormal sounds in audio signals are inefficient and often produce inaccurate results due to insufficient frequency variation and the time-consuming nature of processing long audio signals with few abnormal sounds.
The use of an attentive neural process architecture within a neural network to detect abnormal sounds by dividing the spectrogram into context and target regions, processing masked areas, and calculating reconstruction errors to determine anomaly scores.
This approach enables efficient and accurate detection of abnormal sounds in audio signals, even in cases with limited frequency variation and long signal processing times, by focusing on specific regions of interest within the spectrogram.
Smart Images

Figure 0007675936000012 
Figure 0007675936000013 
Figure 0007675936000014
Abstract
Description
[Technical field]
[0001] The present disclosure relates generally to anomaly detection, and more particularly to a method and system for detecting abnormal sounds. [Background technology]
[0002] Diagnosis and monitoring of machine operation performance is important for various applications. Diagnosis and monitoring tasks are usually performed manually by skilled technicians. For example, skilled technicians can determine abnormal sounds by listening to and analyzing sounds generated by the machine. By automating the manual process for analyzing the sounds, the voice signal generated by the machine can be processed and abnormal sounds in the voice signal can be detected. The automated voice diagnosis can be trained based on deep learning techniques to detect abnormal sounds. Typically, the automated voice diagnosis can be trained to detect abnormal sounds using training data corresponding to normal operating conditions of the voice diagnosis. Such abnormal sound detection based on training data is an unsupervised method. Unsupervised abnormal sound detection may be suitable for detecting certain types of anomalies, such as sudden transient disturbances or impulse sounds that can be detected based on sudden temporary changes.
[0003] However, sudden temporary changes may result in insufficient change information in the audio frequency domain to detect abnormal sounds. Lack of frequency changes to detect abnormal sounds may result in undesirable inaccurate results. In some cases, abnormal sounds in an audio signal can be detected by processing the entirety of the audio signal for non-stationary sounds. However, audio signals may have fewer occurrences of abnormal sounds. Processing such audio signals with fewer occurrences of abnormal sounds for a long time may consume time and computing resources and is infeasible. In some cases, fewer occurrences of abnormal sounds may not be detected due to the long processing time.
[0004] Therefore, there is a need to solve the above problems, and more specifically, to develop a method and system for detecting abnormal sounds in an audio signal in an efficient and feasible manner. Summary of the Invention
[0005] Various embodiments of the present disclosure disclose systems and methods for detecting abnormal sounds in an audio signal. Some embodiments aim to detect abnormal sounds using deep learning techniques.
[0006] Conventionally, abnormal sounds in an audio signal may be detected based on an autoencoder or a variational autoencoder. The autoencoder may compress an audio signal and reconstruct the original audio signal from the compressed data. The variational autoencoder may reconstruct the original audio signal by determining parameters of a probability distribution (e.g., a Gaussian distribution) of the audio signal. By comparing the reconstructed audio signal with the original audio signal, a reconstruction error may be determined to detect abnormal sounds in the audio signal. More specifically, the audio signal may be represented as a spectrogram, which includes a visual representation of the signal strength or volume of the audio signal at various frequencies versus time.
[0007] In some embodiments, a spectrogram of a particular region of the time-frequency domain of the audio signal can be masked. The masked region may be pre-specified during training of the neural network. The neural network processes the unmasked region to reconstruct a spectrogram of the masked region of the audio signal. The reconstructed spectrogram is compared with the original spectrogram to determine a reconstruction error. The reconstruction error is the difference between the original spectrogram and the reconstructed spectrogram. The reconstruction error can be used to detect abnormal sounds.
[0008] Some embodiments of the present disclosure are based on the understanding that an autoencoder for detecting abnormal sounds can be trained based on training data corresponding to non-abnormal sound data, such as normal sounds of a machine operating normally. An autoencoder trained with non-abnormal data can model the data distribution of "normal" (non-abnormal) data samples. However, the reconstruction error may be large because an autoencoder that learns to reconstruct normal data can detect abnormal sounds. In some cases, an autoencoder may be trained on a specific region at a pre-specified fixed time-frequency location to reconstruct the region of the audio signal. However, an autoencoder may not be suitable for performing a dynamic search to determine a region that is distinguishable from normal sounds. A distinguishable region is a region that corresponds to a potential abnormal sound in the audio signal.
[0009] However, an autoencoder trained on normal sound data may not be able to reconstruct abnormal sounds that are different from normal sounds. During inference, some sounds may vary in time and exhibit highly non-stationary behavior. For example, non-stationary sounds generated by a machine (e.g., a valve or slider) may exhibit time-varying or non-stationary behavior. An autoencoder has difficulty reconstructing abnormal sounds based on time-varying or non-stationary sounds. In such a situation, the reconstruction error corresponding to the time-varying or non-stationary sounds determined by the autoencoder may be inaccurate. The reconstruction error may be high even for normal operating conditions of the machine, making it difficult to detect abnormal sounds.
[0010] Some embodiments of the present disclosure are based on the recognition that a partial time signal from ambient information in an audio signal can be processed. This partial time signal processing approach does not require processing the full length of the audio signal to generate a reconstructed spectrogram. Also, processing the partial audio signal can improve the performance of non-stationary sounds, including speech signals and sound waves with various frequencies.
[0011] To that end, a certain region of the spectrogram of the audio signal can be masked based on some time signal. The autoencoder can reconstruct the spectrogram by processing the masked region of the spectrogram. The reconstructed spectrogram can be compared with the spectrogram to obtain a reconstruction error. The reconstruction error may be used for the masked region of the spectrogram as an anomaly score. However, the autoencoder may not be accurate in detecting abnormal sounds because it may present frequency information of the audio signal. Also, the autoencoder may not be able to incorporate prior information about time and / or frequency regions where abnormal sounds may occur in the spectrogram.
[0012] Some embodiments are based on the recognition that the difficulty of detecting anomalies in a non-stationary audio signal may correspond to the variability and diversity of the time and frequency locations of the anomalous regions of the corresponding spectrogram. In particular, reconstructing a region of the non-stationary audio signal (e.g., speech, electrocardiogram (ECG) signal, mechanical sounds, etc.) and inspecting the region for anomalies can exclude the region from the variability of the remaining regions of the audio signal and focus the anomaly detection on the region of interest. However, the versatility of the non-stationary audio signal may result in a versatility of the time and frequency locations that may contain anomalous sounds. Thus, during an online mode of anomaly detection, i.e., during online abnormal sound detection, a specific region of the audio signal may be inspected. Additionally or alternatively, potentially anomalous regions in the audio signal may also be inspected during online abnormal sound detection.
[0013] To that end, some embodiments of the present disclosure disclose a neural network for detecting abnormal sounds in a non-stationary audio signal using an attentional neural process architecture. The attentional neural process architecture is a meta-learning framework for estimating the distribution of a signal. Some embodiments are based on the understanding that the attentional neural process architecture can be used to restore missing parts of an image. For example, if a finger accidentally occludes a part of a camera taking a photo, a person's face in the photo may not be completely captured. The captured photo may have the person's face partially covered by an occlusion, for example, the forehead part of the person's face may be covered by the occlusion. Since the occlusion is known, the covered part of the person's face can be restored. To that end, in some embodiments, the attentional neural process architecture may be configured to search for and restore different regions in a spectrogram of the audio signal. The different regions may include regions of the spectrogram that may correspond to potential abnormal sounds. In some embodiments, the regions of potential abnormal sounds may be determined based on signal characteristics or prior knowledge, such as known abnormal sound behaviors. Using signal characteristics or prior knowledge does not require pre-defined data of the regions at training time.
[0014] Thus, to detect abnormal sounds, the spectrogram of the audio signal can be divided into regions, such as a context region and a target region. The context region includes a time-frequency portion selected from the spectrogram. The target region corresponds to a time-frequency portion predicted from the spectrogram to detect abnormal sounds. In some embodiments, the neural network may be trained by randomly or pseudo-randomly selecting different partitions of the training spectrogram into the context region and the target region. The trained spectrogram corresponds to abnormal sounds that can be used to create an abnormal spectrogram library. The abnormal spectrogram library can be used to identify difficult-to-predict target regions of the spectrogram during testing of the neural network. In some embodiments, the identified target regions can be utilized as one or more hypotheses to determine a maximum anomaly score. The maximum anomaly score corresponds to a likely abnormal region (i.e., an abnormal sound) in the spectrogram. In some embodiments, the one or more hypotheses may include an intermediate frame hypothesis procedure for recovering a temporally intermediate portion of the spectrogram; a frequency masking hypothesis procedure for recovering specific frequency regions of the spectrogram from high or low frequency regions of the spectrogram; a frequency masking hypothesis procedure for recovering individual frequency regions from adjacent or harmonically related frequency regions of the spectrogram; an energy based hypothesis procedure for recovering high energy time-frequency portions of the spectrogram; a procedure for recovering randomly selected subsets of the masked frequency regions and time frames of the spectrogram; a likelihood bootstrapping procedure for running different context regions of the spectrogram and recovering the entire spectrogram with a high likelihood of reconstruction; and an ensemble procedure for finding the maximum anomaly score by combining the above hypothesis generation procedures.
[0015] Further, during testing of the neural network, multiple sections of the spectrogram can be generated and corresponding anomaly scores can be determined based on a predetermined protocol, for example, the mean square error, Gaussian log-likelihood, or any other statistical expression of the reconstruction error can be calculated. A maximum anomaly score can be determined from the anomaly scores, and the maximum anomaly score can be used to detect an abnormal sound. After detecting the abnormal sound, a control action can be performed.
[0016] Some embodiments disclose an iterative approach to determine regions that are difficult to reconstruct from a spectrogram. To that end, a set of context regions and a corresponding set of target regions can be generated by dividing the spectrogram into different combinations of context regions and target regions. The set of context regions is provided to a neural network. The set of context regions can be processed by running the neural network multiple times. Specifically, the target regions can be reconstructed by running the neural network once for each context region in the set of context regions. The set of target regions can be reconstructed by summing each reconstructed target region from each run of the neural network. A set of anomaly scores can be obtained by comparing the reconstructed set of target regions with the set of target regions. More specifically, each target region of the reconstructed set of target regions is compared with each target region of the corresponding set of target regions. This comparison determines a reconstruction error between each target region of the reconstructed set of target regions and each target region of the set of target regions. The reconstruction error may be used for all target regions as an anomaly score. In some embodiments, the anomaly score may correspond to an average or total anomaly score, which may be determined based on a pooling operation on the set of anomaly scores. The pooling operation may include an average pooling operation, a weighted average pooling operation, a max pooling operation, a median pooling operation, etc.
[0017] In some embodiments, the total anomaly score can be used as a first anomaly score to further split the spectrogram into separate context and target regions. The neural network can recover the target region by processing the context region. A second anomaly score can be obtained by comparing the recovered target region with the split target region. A final anomaly score can be obtained by summing the first anomaly score and the second anomaly score using a pooling operation. The final anomaly score can be used to detect abnormal sounds and a control action can be performed based on the final anomaly score. The neural network can process the context region using an attentional neural network architecture.
[0018] In some embodiments, the attention-based neural process architecture may include an encoder neural network, a cross attention module, and a decoder neural network. The encoder neural network may be trained to accommodate an input set of any size. Each element of the input set may include a value and coordinates of the element in the context domain. The encoder neural network may also output an embedding vector for each element of the input set. In some exemplary embodiments, the encoder neural network may jointly encode all elements in the context domain using a self-attention mechanism. The self-attention mechanism corresponds to the attention mechanism for computing the encoded representation of the elements in the context domain by allowing each element to interact or associate with each other.
[0019] The cross-attention module may be trained to compute a unique embedding vector for each element in the target region by processing embedding vectors of elements in the context region located at adjacent coordinates. In some exemplary embodiments, the cross-attention module may compute the unique embedding vectors using multi-head attention. Multi-head attention may compute the embedding vectors in parallel by implementing an attention mechanism. The decoder neural network outputs a probability distribution for each element in the target region. The probability distribution may be obtained from the coordinates of the target region and the embedding vectors of the corresponding elements in the target region. In some exemplary embodiments, the decoder neural network may output parameters of the probability distribution. The probability distribution may correspond to a conditionally independent Gaussian distribution. In some other exemplary embodiments, the decoder neural network may output parameters of the probability distribution that may correspond to a mixture of conditionally independent Gaussian distributions.
[0020] Additionally or alternatively, a sliding window on the spectrogram of the audio signal can be used to determine abnormal sounds in the audio signal. The neural network can use an attention-based neural network architecture to process the sliding window to determine an anomaly score for detecting abnormal sounds. Since the detection of abnormal sounds is completed within the sliding window, the detection speed of abnormal sounds can be improved.
[0021] Accordingly, one embodiment discloses a computer-implemented method for detecting abnormal sounds. The method includes receiving a spectrogram of an audio signal having elements defined by values in a time-frequency domain. The value of each element of the spectrogram is identified by a coordinate in the time-frequency domain. The method includes partitioning the time-frequency domain of the spectrogram into a context domain and a target domain. The method includes recovering values of elements having coordinates in the target domain of the spectrogram by providing values of the elements in the context domain and the coordinates of the elements in the context domain to a neural network including an attention-based neural processing architecture. The method includes determining an anomaly score for detecting abnormal sounds in the audio signal based on a comparison of the recovered values of the elements in the target domain and the values of the elements in the partitioned target domain. The method includes performing a control action based on the anomaly score.
[0022] Accordingly, another embodiment discloses a system for detecting abnormal sounds. The system comprises at least one processor and a memory storing instructions that, when executed by the at least one processor, cause the system to receive a spectrogram of an audio signal having elements defined by values in a time-frequency domain of the spectrogram. A value of each element of the spectrogram is identified by a coordinate in the time-frequency domain. The at least one processor can cause the system to divide the time-frequency domain of the spectrogram into a context domain and a target domain. The at least one processor can cause the system to provide the values of the elements in the context domain and the coordinates of the elements in the context domain to a neural network including an attention-based neural processing architecture to recover spectrogram values of elements having coordinates in the target domain. The at least one processor can cause the system to determine an anomaly score for detecting abnormal sounds in the audio signal based on a comparison of the recovered values of the elements in the target domain and the values of the elements in the divided target domain. The at least one processor can further cause the system to perform a control action based on the anomaly score.
[0023] Further features and advantages will become more readily apparent from the following detailed description taken in conjunction with the accompanying drawings.
[0024] The present disclosure is further described in the following detailed description with reference to a number of drawings as non-limiting examples of exemplary embodiments of the present disclosure. Like reference numerals represent like parts in the several drawings. The drawings are not necessarily drawn to scale. Instead, emphasis may be placed on illustrating the principles of embodiments of the present disclosure. [Brief description of the drawings]
[0025] [Figure 1] 1 is a schematic block diagram illustrating a system for detecting abnormal sounds in an audio input signal according to an embodiment of the present disclosure. [Diagram 2]FIG. 2 illustrates a step-by-step process for detecting abnormal sounds in an audio signal according to an embodiment of the present disclosure. [Diagram 3] FIG. 13 illustrates a step-by-step process for detecting abnormal sounds in an audio signal, according to some other embodiments of the present disclosure. [Figure 4A] 3A-3C are diagrams illustrating example representations of a context region and a corresponding set of target regions of a spectrogram of an audio signal according to an embodiment of the present disclosure. [Figure 4B] 1A-1C are diagrams illustrating example representations of a sliding window on a spectrogram of an audio signal in accordance with some embodiments of the present disclosure. [Figure 5A] 3A-3C are diagrams illustrating example representations of context regions and corresponding target regions of a spectrogram of an audio input signal in accordance with some embodiments of the present disclosure. [Figure 5B] 5A-5C are diagrams illustrating example representations of context regions and corresponding target regions of a spectrogram of an audio input signal in accordance with certain other embodiments of the present disclosure. [Figure 5C] 5A-5C are diagrams illustrating example representations of context regions and corresponding target regions of a spectrogram of an audio input signal in accordance with certain other embodiments of the present disclosure. [Figure 5D] 5A-5C are diagrams illustrating example representations of context regions and corresponding target regions of a spectrogram of an audio input signal in accordance with certain other embodiments of the present disclosure. [Figure 5E] 5A-5C are diagrams illustrating example representations of context regions and corresponding target regions of a spectrogram of an audio input signal in accordance with certain other embodiments of the present disclosure. [Figure 6] FIG. 2 is a schematic diagram illustrating an anomaly library for detecting abnormal sounds in an audio signal, according to some embodiments of the present disclosure. [Figure 7] FIG. 1 is a schematic diagram illustrating an architecture for detecting abnormal sounds in an audio signal in accordance with some embodiments of the present disclosure. [Figure 8]FIG. 1 is a flow diagram illustrating a method for detecting abnormal sounds according to an embodiment of the present disclosure. [Figure 9] FIG. 1 is a block diagram illustrating a system for detecting abnormal sounds in accordance with an embodiment of the present disclosure. [Figure 10] FIG. 10 illustrates a use case for detecting abnormal sounds using the system of FIG. 9 in accordance with an embodiment of the present disclosure. [Figure 11] 10 illustrates a use case for detecting abnormal sounds using the system of FIG. 9 according to another embodiment of the present disclosure. [Figure 12] 10 illustrates a use case for detecting abnormal sounds using the system of FIG. 9 in accordance with some other embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0026] Although the above drawings illustrate embodiments of the present disclosure, other embodiments are contemplated, as discussed above. The present disclosure provides exemplary embodiments by way of example and not by way of limitation. Those skilled in the art can devise numerous other variations and implementations that fall within the scope and spirit of the principles of the embodiments of the present disclosure.
[0027] In the following description, for the purpose of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent to those skilled in the art that one or more embodiments can be practiced without these specific details. In addition, the apparatus and method are shown as block diagrams in order not to obscure the present disclosure. Various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the subject matter described in the appended claims.
[0028] As used in the present specification and claims, the terms "for example," "for example," "such as," and the verbs "comprise," "have," "include," and other verb forms thereof, when used with a list of one or more components or other items, should be construed as open ended, meaning not to exclude other additional components or items from this list. The term "based on" means based at least in part on. Furthermore, it should be understood that the phrases and terms used herein are for the purpose of description and should not be considered as limiting unless specifically defined as such. Any headings used herein are for convenience only and have no legal or limiting effect.
[0029] In the following description, specific details are given to provide a thorough understanding of the embodiments. However, those skilled in the art will understand that the embodiments can be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in block diagrams so as not to obscure the embodiments in unnecessary detail. Also, well-known processes, structures, and techniques may be shown without unnecessary detail so as not to obscure the embodiments. Moreover, like reference numbers and names in the various drawings refer to like elements.
[0030] Although most of the explanations are given using mechanical sounds as the target sound source, the same methods can be applied to other types of audio signals. System Overview
[0031] FIG. 1 is a block diagram illustrating a system 100 for detecting abnormal sounds in an audio input signal 108, according to an embodiment of the present disclosure. The audio processing system 100 (hereinafter referred to as system 100) includes a processor 102 and a memory 104. The memory 104 is configured to store instructions for detecting abnormal sounds. In some embodiments, the memory 104 is configured to store a neural network 106 for detecting abnormal sounds. In some exemplary embodiments, the audio input signal 108 may correspond to a non-stationary sound, such as a human voice, the sound of a running machine, etc. The audio input signal 108 may be represented in a spectrogram. In some cases, the spectrogram may correspond to a log-mel spectrogram that represents an acoustic time-frequency representation of the audio input signal 108.
[0032] The processor 102 is configured to execute the stored instructions to cause the system 100 to receive a spectrogram of an audio signal. The spectrogram includes elements defined by values in a time-frequency domain of the spectrogram. The value of each element of the spectrogram 108 is identified by a coordinate in the time-frequency domain. The time-frequency domain of the spectrogram is divided into a context domain and a target domain. The context domain corresponds to one or more subsets of the time-frequency domain, e.g., a time-frequency portion in the spectrogram. The target domain corresponds to a predicted time-frequency portion in the spectrogram that may be used for anomaly detection.
[0033] The values of the elements in the context region and the coordinates of the elements in the context region are provided to a neural network 106. The neural network 106 includes an attention-based neural process architecture 106A for recovering values of the elements having coordinates in the target region. The recovered values may correspond to abnormal sounds in the spectrogram. Recover the target region based on the recovered values. Determine an anomaly score for detecting abnormal sounds by comparing the recovered target region with the segmented target region. The anomaly score is a reconstruction error for determining the difference between the recovered target region and the segmented target region.
[0034] In some demonstrative embodiments, the attention-based neural processing architecture 106A can encode the coordinates of each of the elements in the context region with an observation value, which may correspond to known abnormal sound behavior, such as a human scream or a scraping sound during machine operation.
[0035] In some exemplary embodiments, the neural network 106 may be trained by randomly or pseudo-randomly selecting different sections of the training spectrogram into the context region and the target region. Additionally or alternatively, the neural network 106 may be trained based on signal characteristics or prior knowledge, such as known abnormal sound behavior. For example, the known abnormal sound behavior may correspond to an abnormal sound of a damaged part of a machine. Furthermore, the neural network 106 generates multiple sections of the spectrogram and corresponding anomaly scores during execution according to a predetermined protocol. An anomaly score for detecting abnormal sounds can be obtained by averaging multiple sections of one or more random context regions or target regions.
[0036] In one exemplary embodiment, the neural network 106 may be trained to segment the spectrogram using a random RowCol selection method. The random RowCol method trains the neural network 106 by randomly selecting one or two time frames in the columns of the spectrogram and up to two frequency bands in the rows of the spectrogram as a set of target regions. The remaining time-frequency portions of the spectrogram are used as a set of context regions.
[0037] In another exemplary embodiment, the neural network 106 may be trained using an intermediate frame selection method. The intermediate frame selection method selects an intermediate frame (frame (L+1) / 2) as a target region of the L frames of the spectrogram. In another exemplary embodiment, the neural network 106 may be trained using a likelihood bootstrapping method. The likelihood bootstrapping method divides the spectrogram into multiple different combinations of context regions and target regions by performing multiple forward passes. By processing different combinations of context regions and target regions, the attention-based neural processing architecture 106A can restore target regions with frame and frequency values that are difficult to reconstruct potential abnormal sounds in the audio input signal 108.
[0038] In some embodiments, a set of context regions and a corresponding set of target regions can be generated by dividing the spectrogram into different combinations of context regions and target regions. The set of context regions is provided to the neural network 106. The neural network 106 may be run multiple times to process the set of context regions. Specifically, the target regions are reconstructed by running the neural network 106 once for each context region in the set of context regions. The attention-based neural processing architecture 106A outputs the reconstructed target regions. Each reconstructed target region from each run of the neural network 106 may be summed into a reconstructed set of target regions. A set of anomaly scores can be determined by comparing the reconstructed set of target regions to a set of target regions. More specifically, each target region of the reconstructed set of target regions is compared to each target region of the corresponding set of target regions. The comparison determines a reconstruction error between each target region of the reconstructed set of target regions and each target region of the set of target regions. The reconstruction error may be used for all target regions as an anomaly score, such as an anomaly score 110. In some embodiments, the anomaly score 110 may correspond to an average anomaly score, which may be determined based on a weighted combination of a set of anomaly scores.
[0039]
number
[0040] In some cases, it may be difficult to detect some sounds in the audio input signal 108 that correspond to abnormal sounds. For example, sounds generated by a damaged or broken part of a machine may be abnormal. However, if the sound of the damaged part is smaller than other sounds of the machine, it may be difficult to detect the sound of the damaged part as an abnormal sound. In some other cases, training the abnormal sound detection based on training data of normal sounds may not be able to accurately detect abnormal sounds in the audio input signal 108. Such abnormal sounds may be difficult to reconstruct from the corresponding regions of the spectrogram (i.e., the time and frequency values of the abnormal sounds). Therefore, the system 100 may divide the spectrogram of the audio input signal 108 into different combinations of context regions and target regions to detect abnormal sounds. This will be further explained with reference to FIG. 2 below.
[0041] FIG. 2 illustrates a step-by-step process 200 for detecting abnormal sounds in an audio signal 108 according to an embodiment of the present disclosure. The process 200 is performed by the system 100 of FIG. 1. In step 202, the system 100 receives an audio input signal 108. The audio input signal 108 may correspond to sounds generated by a machine, such as CNC machining of a workpiece. The CNC machine may include multiple actuators (e.g., motors) that assist one or more tools in performing one or more tasks, such as soldering or assembling the workpiece. Each of the multiple actuators may generate vibrations due to deformation of the workpiece during machining of the workpiece. These vibrations are mixed with a signal from a motor that drives a cutting tool of the CNC. The mixed signal may include sounds due to a failure of a part of the CNC, such as a failure of the cutting tool.
[0042] In step 204, the audio input signal 108 is processed to extract a spectrogram of the audio input signal 108. A spectrogram is an acoustic time-frequency domain of the audio input signal 108. A spectrogram includes elements defined by values (e.g., pixels in the time-frequency domain). The value of each element is identified by its coordinate in the time-frequency domain. For example, time frames in the time-frequency domain are represented as columns and frequency bands in the time-frequency domain are represented as rows.
[0043] In step 206, the time-frequency domain of the spectrogram is divided into different combinations of context domains and target domains. The context domains correspond to one or more subsets of the time-frequency domain, e.g., a time-frequency portion of the spectrogram. The target domains correspond to a predicted time-frequency portion from the spectrogram that may be used for anomaly detection. A set of context domains and a corresponding set of target domains are generated by masking each of the different combinations of context domains and target domains. For example, the set of context domains and the corresponding set of target domains are masked into a set of context-target domain masks, e.g., context-target domain mask 1 206A, context-target domain mask 2 206B, and context-target domain mask N 206N (hereinafter referred to as context-target domain masks 206A-206N).
[0044] In an example embodiment, the set of context regions and the corresponding set of target regions may be masked based on a random RowCol selection method. The random RowCol selection method randomly selects one or two time frames and up to two frequency band values of the spectrogram as the set of target regions. The set of context regions corresponds to the remaining time-frequency portion in the spectrogram. In another example embodiment, the set of context regions and the corresponding set of target regions may be masked based on an intermediate frame selection method. The intermediate frame selection method selects an intermediate frame (frame (L+1) / 2) of the spectrogram as the set of target regions corresponding to L frames of the spectrogram. In yet another example embodiment, the set of context regions and the corresponding set of target regions may be masked based on a likelihood bootstrapping method. The likelihood bootstrapping method divides the spectrogram into multiple different combinations of context regions and target regions, e.g., context-target region masks 206A-206N, by performing multiple forward passes of the values of the spectrogram.
[0045] The context-target region masks 206A-206N are input to the neural network 106. The neural network 106 processes the context-target region masks 206A-206N using an attention-based neural processing architecture 106A.
[0046] In step 208, the attention-based neural processing architecture 106A is run multiple times to process the context regions in the context-target region masks 206A-206N, one for each context region in the set of context-target region masks 206A-206N, to reconstruct the corresponding target region. The reconstructed target regions are summed to form a reconstructed set of target regions. Each target region in the reconstructed set of target regions is compared to its corresponding target region in the set of context-target region masks 206A-206N.
[0047] At step 210, a set of anomaly scores is determined based on the comparison. The set of anomaly scores may be represented as an anomaly score vector that summarizes information in the set of context regions that may be most relevant to each frequency bin in the set of target regions.
[0048] In step 212, the attention-based neural processing architecture 106A is used to obtain a composite region of the target regions by concatenating the summarized anomaly score vector with the corresponding target region vector positions of the set of target regions. Specifically, the composite region can be obtained by performing a composite of regions on the anomaly score vector. Each element of the anomaly score vector corresponds to an anomaly score of a restored target region. In some exemplary embodiments, the composite of regions may be performed using a pooling operation such as an average pooling operation, a weighted average pooling operation, a maximum pooling operation, a median pooling operation, etc.
[0049] In step 214, a final anomaly score is obtained, e.g., anomaly score 110. The final anomaly score corresponds to an anomalous region in the spectrogram. In some exemplary embodiments, the final anomaly score may be determined based on a weighted combination of the set of anomaly scores. For example, the set of anomaly scores may be averaged to obtain the final anomaly score.
[0050] The final anomaly score can be used to further split the spectrogram to determine regions that may be difficult to reconstruct from the spectrogram. Splitting the spectrogram using the final anomaly score is described in more detail below with reference to FIG.
[0051] 3 shows a step-by-step process 300 for detecting abnormal sounds in an audio input signal 108, according to some other embodiments of the present disclosure. The process 300 is performed by the system 100. In step 302, the audio input signal 108 is received. In step 304, a spectrogram is extracted from the audio input signal 108. Steps 302 and 304 are similar to steps 202 and 204 of the process 200.
[0052] In step 306, the spectrogram is divided into a first context region and a corresponding first target region. A first context-target region mask is generated by masking the first context region and the first target region. The first context region and the first target region may be masked based on one of a random RowCol selection method, an intermediate frame selection method, and a likelihood bootstrapping method, etc. The first context-target region mask is input to the neural network 106. The neural network 106 processes the first context-target region mask using an attention-based neural processing architecture 106A.
[0053] In step 308, the attention-based neural process architecture 106A is executed to recover the target region from the context region by processing the context region in the first context-target region mask. The recovered target region is compared to the corresponding target region of the first context-target region mask. The comparison of the target region and the recovered target region determines a first anomaly score. In some exemplary embodiments, the second target region can be recovered by repeatedly running the neural network with values and coordinates of the second context region. The values and coordinates of the second context region may correspond to a time-frequency portion of the second context region of the spectrogram.
[0054] In step 310, the time-frequency portion of the second context region may be sampled based on the complete spectrogram reconstructed by the attention-based neural process. In some exemplary embodiments, the reconstructed spectrogram may include time-frequency portions of the original spectrogram that have a low reconstruction likelihood.
[0055] In step 312, the first anomaly score is used to identify a second partition of the spectrogram, i.e., a reconstructed spectrogram. Specifically, a second partition can be performed by comparing the reconstructed spectrogram in step 310 with the original spectrogram obtained in step 304. To this end, the reconstructed spectrogram is divided into a second context region and a second target region based on the second partition. The second target region may include a region in the original spectrogram with high reconstructibility, and the second target region may include a region in the original spectrogram with low reconstructibility. The second context region and the second target region are also input to the neural network 106. The neural network 106 processes the second context region using an attention-based neural processing architecture 106A.
[0056] In step 314, the second target region may be reconstructed by repeatedly executing the attention-based neural processing architecture 106A with the values and coordinates of the second context region. A second anomaly score is determined by comparing the reconstructed second target region with the segmented second target region.
[0057] In step 316, the second anomaly score is output as a final anomaly score. The final anomaly score can be used to detect an abnormal sound and to perform a control action on the detected abnormal sound. In some other embodiments, the control action may be performed based on a combination of the first anomaly score and the second anomaly score, or both.
[0058] In some cases, the restored target region can be utilized as one or more hypotheses to determine the maximum anomaly score. The one or more hypotheses are described in corresponding Figures 4A and 4B, Figures 5A, 5B, 5C, 5D, and 5E. Determining the maximum anomaly score based on the restored target region is further described with reference to Figure 6.
[0059] FIG. 4A illustrates an example representation 400A of a set of context regions 406 and a corresponding set of target regions 408 of a spectrogram 404 of an audio input signal 402 according to an embodiment of the present disclosure. The audio input signal 402 is an example of an audio input signal 108. The system 100 extracts a spectrogram 404 from the audio input signal 402. The spectrogram 404 includes elements defined by values in a time-frequency domain. The value of each element of the spectrogram 404 corresponds to a coordinate in the time-frequency domain. The time-frequency domain of the spectrogram 404 is divided into a set of context regions 406 and a set of target regions 408. The set of context regions 406 corresponds to one or more time-frequency portions of the spectrogram 404. The set of target regions 408 corresponds to time-frequency portions predicted from the spectrogram 404.
[0060]
number
[0061] The set of context regions 406 is provided to the neural network 106. The neural network 106 executes the attention-based neural processing architecture 106A to recover the target region from the spectrogram 404 by processing the set of context regions 406.
[0062] In some cases, the audio input signal 402 may correspond to a long audio signal that may contain abnormal sounds, such as transient disturbances in the audio input signal 402. In such cases, a sliding window procedure can be used to determine the abnormal sounds, as shown in FIG.
[0063] 4B illustrates an example representation 400B of a sliding window 410 on a spectrogram 404 of an audio input signal 402 according to some embodiments of the present disclosure. The audio input signal 402 may include a long audio signal. For example, the spectrogram 404 of the audio input signal 402 may include a frame length of 1024 samples with a 512 hop length between successive frames, e.g., a train of spectrograms 404 and 128 mel bands. The sliding window 410 may be input to the neural network 106. The neural network 106 may determine an anomaly score by processing the sliding window 410 using an attention-based neural processing architecture 106A.
[0064] In some exemplary embodiments, multiple sliding windows of five frames with one frame hop can be used as input to the neural network 106. By averaging multiple sliding windows, an anomaly score for each sample of the spectrogram 404 can be obtained. A pooling operation can be used to combine the anomaly scores of corresponding samples to detect abnormal sounds in the audio input signal 402. By using the sliding window 410, it is not necessary to process the entire length of the audio input signal 402 since the abnormal sound detection is completed within the sliding window 410.
[0065] In some exemplary embodiments, the spectrogram 404 may be divided into a set of context regions and a corresponding set of target regions that include masking of frequency bands of the spectrogram 404. Such a set of context regions and a corresponding set of target regions are shown in Figures 5A and 5B.
[0066] FIG. 5A illustrates an example representation 500A of a context region 502 and a corresponding target region 504 of a spectrogram 404 of an audio input signal 402 according to some other embodiments of the present disclosure. The context region 502 and the target region 504 may be obtained by an intermediate frame selection method (as shown in 400A). For example, an intermediate frame of the spectrogram 404 may be selected from L frames in the time-frequency domain of the spectrogram 404. The intermediate frame may be determined as a frame ((L+1) / 2) for dividing the spectrogram 404 into the context region 502 and the target region 504. As shown in FIG. 5A, the context region 502 may mask a range of continuous frequency bands by adding a horizontal bar to the top of the spectrogram 404. As shown in FIG. 5A, the target region 504 may mask a range of continuous frequency bands by adding a horizontal bar to the bottom of the spectrogram 404.
[0067] The attention-based neural processing architecture 106A recovers a frequency band (i.e., a target region) from the context region 502. The recovered frequency band may correspond to a high-frequency band reconstructed from a lower portion of the context region 502. For example, the high-frequency band may be recovered from a lower portion of the context region 502, which may include a low-frequency band of the spectrogram 404. The recovered frequency band is compared with the frequency band of the target region 504 to obtain an anomaly score (e.g., anomaly score 110) for detecting an abnormal sound in the audio input signal 402.
[0068] In some cases, corresponding context and target regions can be obtained by masking each frequency band of the spectrogram 404, as shown in FIG. 5B.
[0069] 5B illustrates an example representation 500B of a context region 506 and a corresponding target region 508 of a spectrogram 404 according to some other embodiments of the present disclosure. In some example embodiments, the context region 506 and the target region 508 are formed by dividing the spectrogram 404 into individual frequency regions. The context region 506 and the target region 508 may be provided to the neural network 106. The neural network 106 may process the context region 506 using an attention-based neural processing architecture 106A. The attention-based neural processing architecture 106A recovers the individual frequency bands from adjacent and harmonically related frequency bands in the context region 506. The recovered individual frequency bands are compared to the frequency bands of the target region 508 to determine an anomaly score for detecting an abnormal sound in the audio input signal 402.
[0070] 5C illustrates an example representation 500C of a context region 510 and a corresponding target region 512 of a spectrogram 404 according to some other embodiments of the present disclosure. In some example embodiments, the context region 510 and the target region 512 can be formed by dividing the spectrogram 404 into multiple regions, as shown in FIG. 5C. The context region 510 and the target region 512 are provided to the neural network 106. The neural network 106 recovers the high-energy time-frequency portion from the unmasked time-frequency portion of the context region 510 by processing the context region 510 using the attention-based neural processing architecture 106A. The recovered high-energy time-frequency portion is compared with the high-energy time-frequency portion of the target region 512 to determine an anomaly score for detecting an abnormal sound in the audio input signal 402.
[0071] FIG. 5D illustrates an example representation 500D of a context region 514 and a corresponding target region 516 of a spectrogram 404 in accordance with some other embodiments of the present disclosure.
[0072] In some exemplary embodiments, as shown in FIG. 5D, a set of masked frequency bands and time frames are randomly selected and the spectrogram 404 is partitioned to form a context region 514 and a target region 516. The context region 514 and the target region 516 are provided to the neural network 106. The context region 514 is processed using the attention-based neural processing architecture 106A to recover the set of randomly selected masked frequency bands and time frames from the context region 514. The recovered set of masked frequency bands and time frames are compared with the set of masked frequency bands and time frames in the target region 516 to determine an anomaly score for detecting abnormal sounds in the audio input signal 402.
[0073] FIG. 5E illustrates an example representation 500E of multiple partitions of a spectrogram 404 according to some other embodiments of the present disclosure. In some exemplary embodiments, the spectrogram 404 can be divided into different combinations of context regions and target regions. The different combinations of context regions and target regions may correspond to time-frequency portions of the spectrogram 404 that are distinguishable from normal sound time-frequency portions of the spectrogram 404. The division of the spectrogram in this case is referred to as stage 518. In stage 518, the spectrogram 404 may be sampled as context regions, such as context region 520, with different percentages of the time-frequency portion. In some exemplary embodiments, the time-frequency portions of the spectrogram 404 can be downsampled uniformly. By downsampling the spectrogram 404, samples can be removed from the audio input signal 108 while maintaining the length of the audio input signal 108 with respect to time. For example, the context region 520 can be obtained by sampling the time-frequency portions of the spectrogram 404 at nC=62.5%. The time-frequency portions sampled from the context region 520 may be run in multiple forward passes with multiple sections of the spectrogram 404. For example, as shown in FIG. 5E, all spectrograms, such as reconstructed spectrogram 522, can be reconstructed by processing the time-frequency portions sampled from the context region multiple times.
[0074] A first anomaly score can be determined by comparing the reconstructed spectrogram 522 to the spectrogram 404. The first anomaly score can be used to identify a second partition in the time-frequency domain of the spectrogram 404. In some exemplary embodiments, a dynamic search can be performed using the first anomaly score to determine a region that is distinguishable from normal sounds. The distinguishable region may correspond to a potential abnormal sound in the audio signal. The dynamic search allows the system 100 to process a portion of the audio signal without having to process the entire length of the audio signal. As shown in FIG. 5E, the second partition of the spectrogram 404 is referred to as stage 524.
[0075] In stage 524, the first anomaly score is used to divide the spectrogram 522 into a second context region (e.g., context region 526) and a second target region (e.g., target region 528). The context region 526 includes a time-frequency portion of the spectrogram 404 that has a high likelihood of reconstruction. The remaining time-frequency portion of the spectrogram 522 may correspond to the target region 528. The context region 526 is provided to the neural network 106. The neural network 106 can process values and coordinates of the time-frequency portion of the context region 526 using an attention-based neural processing architecture 106A. The attention-based neural processing architecture 106A reconstructs the target region that has a time-frequency portion that has a low likelihood of reconstruction. A second anomaly score is determined by comparing the reconstructed target region with the target region 528. The second anomaly score is used to detect an abnormal sound and perform a control action based on the detected abnormal sound. In some embodiments, a control action may be performed based on a combination of the first anomaly score and the second anomaly score, or both.
[0076] In some cases, a library of abnormal sound data can be created that includes abnormal sound behaviors, such as vibration sounds during machine operation, as will be further described with reference to FIG.
[0077] FIG. 6 is a schematic diagram 600 of an abnormal spectrogram library 602 according to some other embodiments of the present disclosure. In some exemplary embodiments, the abnormal spectrogram library 602 is created based on known abnormal sound behavior. For example, the library 602 may include abnormal data 604A, abnormal data 604N, etc. Each of the abnormal data 604A and the abnormal data 604N may include a context region having a context index, a target region having a target index and a threshold, and a threshold that may be compared to a corresponding anomaly score 606 to determine whether an anomaly has occurred. In some exemplary embodiments, the threshold may be determined based on a previous observation of an abnormal sound detection. For example, the anomaly score may be determined based on an anomaly score observed when previously dividing a spectrogram (e.g., spectrogram 404) into a context region (e.g., context region 406) and a target region (e.g., target region 408). The determined anomaly score may be used as a threshold. The threshold may be stored in the library 602 that may be used to detect anomalies in any audio sample within the same spectrogram partition.
[0078] In some embodiments, the attention-based neural processing architecture 106A can use the library 602 to identify target regions, i.e., time-frequency portions, from the context regions of the spectrogram 404. It is difficult to predict the time-frequency portions in the spectrogram 404 of the audio input signal 402. The target regions thus identified can be used as one or more hypotheses to detect the maximum anomaly scores 606.
[0079] In some embodiments, the target region with the maximum anomaly score 606 can be found by testing one or more hypotheses, including an intermediate frame hypothesis procedure, a frequency masking hypothesis procedure that aims to recover a specific frequency region, a frequency masking hypothesis procedure that aims to recover individual frequency bands, an energy-based hypothesis procedure that aims to recover high energy time-frequency parts from a context region, a procedure that aims to recover a randomly selected subset of masked frequency bands, a likelihood bootstrapping procedure (described in FIG. 5E), and an ensemble procedure. The ensemble procedure can be used to find the maximum anomaly score 606 by combining multiple hypothesis generation procedures described above.
[0080] In some embodiments, a mid-frame hypothesis procedure (as described in Figures 4A and 4B) can be used to recover the mid-temporal portion of the spectrogram from the side portions of the spectrogram that bracket the mid-frame portion from both sides of the frame of the spectrogram. A frequency masking hypothesis procedure can be used to recover a particular frequency region of the spectrogram from the unmasked surrounding region, e.g., the spectrogram context region 502 (described in Figure 5A). The unmasked surrounding region corresponds to the upper portion of the spectrogram. The recovery corresponds at least to reconstructing the high frequencies at the top of the spectrogram from the low frequencies at the bottom of the spectrogram, or reconstructing the low frequencies at the bottom from the high frequencies at the top of the spectrogram. A frequency masking hypothesis procedure can be used to recover individual frequency bands from adjacent and / or harmonically related frequency bands (described in Figure 5A). An energy-based hypothesis procedure (as described in Figure 5C) can be used to recover high energy time-frequency portions of the spectrogram from the remaining unmasked time-frequency portions of the spectrogram. A procedure (described in FIG. 5D) can be used to recover a randomly selected subset of the masked frequency bands and time frames from the remaining unmasked regions of the spectrogram. A likelihood bootstrapping procedure can be used to perform multiple passes with different context regions determined by first sampling different percentages of the time-frequency portions as context regions, e.g., context regions 520. The attention-based neural processing architecture 106A reconstructs all spectrograms (e.g., spectrogram 522) by processing the context regions (as described in FIG. 5E) and determines the time-frequency portions of the reconstructed spectrograms with high reconstructed likelihood. The time-frequency portions are used as context to reconstruct the time-frequency regions with low reconstructed likelihood.
[0081] The attention-based neural processing architecture 106A for recovering target regions for detecting abnormal sounds is further described with reference to FIG.
[0082] 7 is a schematic diagram illustrating a network architecture 700 for detecting abnormal sounds according to some embodiments of the present disclosure. The network architecture 700 corresponds to the attention-based neural process architecture 106A. The network architecture 700 includes an encoder neural network 702, a cross-attention module 704, and a decoder neural network 706.
[0083]
number
[0084]
number
[0085]
number
[0086] Self-attention can estimate parameters corresponding to the conditionally independent Gaussian distribution, such as parameters 716. To that end, parameters 716 for each point in the target region 408 can be obtained by inputting the concatenated values and coordinates 708 of the context region 406 through the encoder neural network 702.
[0087]
number
[0088]
number
[0089]
number
[0090]
number
[0091]
number
[0092]
number
[0093] The anomaly scores are used across the target region 408 to detect anomalous sounds. Control actions may be taken based on the detection of anomalous sounds. This is further described with reference to Figures 10, 11 and 12.
[0094] 8 is a flow diagram illustrating a method 800 for detecting abnormal sounds according to an embodiment of the disclosure. Method 800 is performed by system 200. At operation 802, method 800 includes receiving a spectrogram (e.g., spectrogram 404) of an audio signal (e.g., audio input signal 402) having elements defined by values in the time-frequency domain. The value of each element of the spectrogram is identified by a coordinate in the time-frequency domain.
[0095] At operation 804, the method 800 includes partitioning the time-frequency domain of the spectrogram into context regions and target regions. In some embodiments, a set of context regions and a corresponding set of target regions are generated by partitioning the spectrogram into different combinations of context regions and target regions (see FIG. 2).
[0096] At operation 806, the method 800 includes recovering values of elements having coordinates in the target region of the spectrogram by providing values of elements in the context regions and coordinates of elements in the context regions to a neural network including an attention-based neural processing architecture. In some exemplary embodiments, a set of context regions can be processed by running the neural network multiple times. A set of target regions can be recovered by running the neural network once for each context region in the set of context regions. In some embodiments, the neural network is trained by randomly or pseudo-randomly selecting different partitions of the training spectrogram for the context regions and the target regions.
[0097] At operation 808, the method 800 includes determining an anomaly score for detecting an abnormal sound in the audio signal based on a comparison of the reconstructed values of the elements in the target regions and the values of the elements in the segmented target regions. In some embodiments, a set of anomaly scores may be determined from the reconstructed set of target regions. For example, the set of anomaly scores may be determined by comparing each target region of the reconstructed set of target regions with a corresponding target region. The anomaly score may be determined based on a weighted combination of the set of anomaly scores.
[0098] At operation 810, the method 800 includes performing a control action based on the anomaly score. In some exemplary embodiments, the anomaly score can be used as a first anomaly score to identify a second partition of the spectrogram. The first anomaly score can be used to partition the spectrogram into a second context region (e.g., second context region 526) and a second target region (e.g., target region 528). The first context region corresponds to a time-frequency region of the spectrogram with a high likelihood of reconstruction. The target region of the spectrogram can be restored by providing the context region 526 to the neural network 106. The restored target region corresponds to a time-frequency region of the spectrogram with a low likelihood of reconstruction. A second anomaly score can be determined by comparing the restored target region to the partitioned target region. The second anomaly score can be used to perform a control action. In some embodiments, the control action can be performed based on a combination of the first anomaly score and the second anomaly score, or both.
[0099] 9 is a block diagram illustrating an abnormal sound detection system 900 according to an embodiment of the present disclosure. The abnormal sound detection system 900 is an example of the system 100. The abnormal sound detection system 900 includes a processor 902 configured to execute stored instructions and a memory 904 that stores instructions for a neural network 906. The neural network 906 includes an attention-based neural processing architecture (e.g., attention-based neural processing architecture 106A). In some embodiments, the attention-based neural processing architecture corresponds to an encoder-decoder model (e.g., network architecture 700). The encoder-decoder model includes an encoder neural network (e.g., encoder neural network 702), a cross-attention module (e.g., cross-attention module 704), and a decoder neural network (e.g., decoder neural network 706).
[0100] The processor 902 may be a single-core processor, a multi-core processor, a graphics processing unit (GPU), a computing cluster, or any number of other configurations. The memory 904 may include random access memory (RAM), read-only memory (ROM), flash memory, or any other suitable memory system. The memory 904 may also include a hard drive, an optical drive, a thumb drive, a drive array, or any combination thereof. The processor 902 is connected to one or more input and output interfaces / devices via a bus 912.
[0101] The abnormal sound detection system 900 may also include an input interface 918. The input interface 918 is configured to receive the audio data 910. In some embodiments, the abnormal sound detection system 900 may receive the audio data 910 over a network 916 using a network interface controller (NIC) 914. The NIC 914 may be configured to connect the abnormal sound detection system 900 to the network 916 over the bus 106. In some cases, the audio data 910 may be online data, such as an online audio stream received over the network 916. In some other cases, the audio data 910 may be recorded data stored in the storage device 908. In some embodiments, the storage device 908 is configured to store a training data set for training the neural network 906. The storage device 908 may also be configured to store an abnormal sound library, such as the library 602.
[0102] In some demonstrative embodiments, the abnormal sound detection system 900 can receive audio data from one or more sensors, collectively referred to as sensors 924. The sensors 924 can include a camera that captures an audio signal, an audio receiver, etc. For example, the camera can capture a video of a scene that includes audio information of the scene. The scene may correspond to an indoor or outdoor environment that includes audio information of one or more objects or people in the scene.
[0103] The abnormal sound detection system 900 may also include an output interface 920. The output interface 920 is configured to output the abnormality score via an output device 926. The output device 922 may output the detected abnormal sound based on the abnormality score. The output device 922 may include a display screen (e.g., a monitor) of a computer, laptop, mobile phone, smart watch, etc. The output device 922 may also include an audio output device (e.g., a speaker) of a computer, laptop, mobile phone, smart watch, etc. The abnormality score is used to perform a control action. For example, the control action may include sending a notification, such as an alarm, to a machine operator upon detection of the abnormal sound.
[0104] FIG. 10 illustrates a use case 1000 for detecting abnormal sounds using the abnormal sound detection system 900 according to an embodiment of the present disclosure.
[0105] In the illustrated exemplary scenario, a machine 1004, such as an ultrasound machine, a heart rate monitor, or the like, is used to diagnose or monitor a patient 1002. For example, the machine 1004 monitors the heart rate of the patient 1002. The machine 1004 is connected to the system 100. In some exemplary embodiments, the machine 1004 may be connected to the system 100 via a network. In some other exemplary embodiments, the system 100 may be provided inside the machine 1004. The machine 1004 transmits monitoring data to the system 100. The monitoring data may correspond to an audio recording or a live audio stream corresponding to the heart rate of the patient 1002. The system 100 determines the abnormal sound by processing the monitoring data and calculating an abnormality score. If an abnormal sound is detected, the detected abnormal sound is reported to an operating room, such as an emergency room 1008, to assist the patient.
[0106] In some other cases, the patient 1002 may be monitored by a camera, such as the camera 1006. The camera 1006 is connected to the system 100. The camera 1006 may capture a video of the patient 1002, including audio information of the patient 1002. The audio information may be processed by the system 100 for abnormal sound detection. For example, the patient 1002 may cough violently while sleeping. Audio data including the corresponding coughing sound may be transmitted to the system 100. The system 100 detects abnormal sounds, such as coughing sounds of the patient 1002, by processing the audio data and calculating an anomaly score. In some cases, the system 100 may utilize the library 602 including the abnormal coughing sound data to detect whether the patient's cough is abnormal or not. The emergency room 1008 may be alerted according to the detected anomaly and notify a doctor or nurse to assist the patient 1002.
[0107] FIG. 11 illustrates a use case 1100 for detecting abnormal sounds in a machine 1102 using the abnormal sound detection system 900 according to an embodiment of the present disclosure. The machine 1102 may include one or more elements (e.g., actuators) collectively referred to as machine elements 1104. Each of the machine elements 1104 may perform a unique task and may be connected to an adjustment device 1106. Examples of tasks performed by the machine elements 1104 may include machining, soldering, or assembly of the machine 1102. In some cases, the machine elements 1104 may operate simultaneously and the adjustment device 1106 may control each of the machine elements 1104 individually. The adjustment device 1106 is, for example, a tool for performing a task.
[0108] The machine 1102 may be connected to a sensor 1110, which may include an audio device such as a microphone or an array of microphones. The sensor 1110 may capture vibrations generated by each of the mechanical elements 1104 during operation of the machine 1108. Furthermore, some of the mechanical elements 1104 may be co-located in the same spatial region, such that the mechanical elements 1104 are not captured individually by the sensor 1110. The vibrations captured by the sensor 1110 may be recorded as an acoustic mixture signal 1112. The acoustic mixture signal 1112 includes the sum of the vibration signals generated by each of the mechanical elements 1104.
[0109] The acoustic mixture signal 1112 is transmitted to the abnormal sound detection system 900. In some embodiments, the abnormal sound detection system 900 extracts a spectrogram (e.g., spectrogram 404) of the acoustic mixture signal 1112. The spectrogram may include at least some of the sound sources that may occupy the same time, space, and frequency spectrum in the acoustic mixture signal 1112. The spectrogram of the acoustic mixture signal 1112 is divided into a context region of the spectrogram and a corresponding predicted target region. The context region is processed by the neural network 906 of the abnormal sound detection system 900. The neural network 906 uses an attention-based neural processing architecture (e.g., attention-based neural processing architecture 106A) to restore the time frame and frequency domain of the spectrogram. The restored time frame and frequency domain are obtained as the restored target region of the spectrogram. The restored target region may include the lowest reconstruction likelihood value of the spectrogram. Further, the restored target area and the segmented target area are compared to determine an anomaly score. The abnormal sound detection system 900 outputs an anomaly score that is used to detect abnormal sounds in the acoustic mixture signal 111. By detecting abnormal sounds, smooth operation of the machine 1102 can be maintained and malfunctions in performing tasks can be avoided. The detected abnormal sounds may be notified to the operator 1114 to perform control actions such as terminating the operation of the machine, notifying manual intervention, etc. For example, the operator 1114 may be an automatic operator that may be programmed to terminate the operation of the machine 1102 if an abnormal sound is detected. In other cases, the operator 1114 may be a manual operator that may intervene to perform actions such as replacing one of the machine elements 1104, repairing the machine element 1104, etc.
[0110] 12 is a diagram illustrating a use case 1200 for detecting abnormal sounds using the abnormal sound detection system 900 according to an embodiment of the present disclosure. In some example embodiments, the abnormal sound detection system 900 may be used for an infant monitoring application. As an example, an infant monitoring device 1202 may be connected to the abnormal sound detection system 900 via a network, such as the network 916. Alternatively, the abnormal sound detection system 900 may be provided internally to the infant monitoring device 1202.
[0111] In the illustrated example scenario, an infant monitor 1202 is monitoring an infant 1204 in a room 1206. The infant monitor 1202 can capture audio signals that can include the infant's crying, white noise or music being played in the room, etc. In some examples, the white noise or music can be louder than the crying. If the white noise or music is loud, a caregiver 1208 in another room may not be able to hear the crying. In such cases, the infant monitor 1202 can transmit an audio signal to the abnormal sound detection system 900.
[0112] The abnormal sound detection system 900 may receive a spectrogram of the audio signal. The spectrogram is divided into a context region and a target region. The neural network 106 recovers the target region by processing the context region with the attention-based neural processing architecture 106A. The recovered target region may include values and coordinates in the time-frequency domain of the spectrogram corresponding to the crying sound. An abnormality score is determined by comparing the recovered target region with the divided target region. The abnormality score may correspond to the crying sound detected as an abnormal sound by the abnormal sound detection system 900. The detected abnormal sound may be notified to the caregiver 1208 via the user device 1210. As an example, the user device 1210 may include an application interface of the infant monitoring device 1202. The caregiver 1208 may perform an action if an abnormal sound is detected.
[0113] In this manner, the system 100 can be used to detect abnormal sounds in an audio signal in an efficient and feasible manner. More specifically, the system 100 can detect abnormal sounds based on training data that includes only normal data. The system 100 can improve the accuracy for detecting abnormal sounds by processing audio data that includes both time and frequency information. Furthermore, the system 100 can process audio data that exhibits time variations and non-stationary behavior and is versatile. Furthermore, the system 100 can detect abnormal sounds by performing a dynamic search for abnormal regions in the audio data. The dynamic search allows the system 100 to process a portion of the audio signal rather than having to process the entire length of the audio signal. This can improve the overall computation speed.
[0114] Also, each embodiment may be described as a process that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart describes operations as a sequential process, many operations may be performed in parallel or simultaneously. Also, the order of operations may be changed. A process may be terminated when the operations of a process are completed, but the process may include additional steps not discussed or shown. Moreover, not all operations within a specifically described process need be included in every embodiment. A process may be a method, a function, a procedure, a subroutine, a subprogram, etc. When a process is a function, the termination of the function corresponds to the return of the function to the calling function or to the main function.
[0115] Furthermore, embodiments of the disclosed subject matter may be implemented, at least in part, manually or automatically. The manual or automated implementation may be implemented, or at least assisted, by machine, hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented by software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks may be stored in a machine-readable medium. A processor may perform the necessary tasks.
[0116] The embodiments of the present disclosure described above may be implemented in many ways. For example, the embodiments may be implemented in hardware, software, or a combination thereof. When implemented in software, the software code may be executed on any suitable processor or collection of processors, whether located on a single computer or distributed across multiple computers. Such a processor may be implemented as an integrated circuit. An integrated circuit element may include one or more processors. However, the processor may be implemented in any suitable circuit.
[0117] Also, the various methods or steps outlined herein may be coded as software executable on one or more processors employing any one of a variety of operating systems or platforms. Moreover, such software may be written using any of a number of suitable programming languages and / or programming or scripting tools, and compiled as executable machine language code or intermediate code that runs on a framework or virtual machine. Typically, the functionality of the program modules may be combined or distributed in various embodiments as desired.
[0118] The embodiments of the present disclosure may be embodied as methods, which are provided as examples. The operations performed as part of the method may be ordered in any suitable manner. Thus, embodiments may be constructed that perform operations in a different order than the operations performed sequentially in the exemplary embodiments, and that may include performing some operations simultaneously. Accordingly, the appended claims are intended to cover all such variations and modifications that are within the true spirit and scope of the present disclosure.
[0119] Although the present disclosure has been described with reference to certain preferred embodiments, it should be understood that various other modifications and alterations can be made within the spirit and scope of the disclosure. It is therefore intended in the appended claims to cover all such variations and modifications that are within the true spirit and scope of the present disclosure.
Claims
1. 1. A sound processing system for detecting abnormal sounds, comprising: At least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the system to perform the following operations: The operation includes: receiving a spectrogram of an audio signal having elements defined by values in a time-frequency domain of the spectrogram, the value of each element of the spectrogram being identified by a coordinate in the time-frequency domain; The operation includes: partitioning the time-frequency domain of the spectrogram into a context domain and a target domain; recovering values of elements having coordinates in the target region of the spectrogram by providing values of elements in the context region and coordinates of the elements in the context region to a neural network including an attention-based neural processing architecture; determining an anomaly score for detecting the abnormal sound in the audio signal based on a comparison of the restored values of the elements in the target region and values of the elements in the segmented target region; and performing a control action based on the anomaly score.
2. The at least one processor generating a set of context regions and a corresponding set of target regions by dividing the spectrogram into different combinations of context regions and target regions; reconstructing a set of target regions by running the neural network multiple times, once for each context region in the set of context regions; determining a set of anomaly scores by comparing each target region in the set of reconstructed target regions to a corresponding target region in the set of target regions; The speech processing system of claim 1 , configured to determine the anomaly score based on a pooling operation on the set of anomaly scores.
3. the context area is a first context area, the target area is a first target area, the anomaly score is a first anomaly score, The processor, identifying a second partition of the time-frequency domain based on the first anomaly score; Dividing the second partition of the spectrogram into a second context region and a second target region; reconstructing the second target region by repeatedly running the neural network with values and coordinates of the second context region, and generating a second anomaly score based on a comparison of the reconstructed second target region to the segmented second target region; 3. The speech processing system of claim 2, configured to perform a second control action based on the second anomaly score, a combination of the first anomaly score and the second anomaly score, or both.
4. The neural network is trained by randomly or pseudo-randomly selecting different sections of a training spectrogram for the context and target regions; 2. The speech processing system of claim 1, wherein during execution of the neural network, the processor is configured to generate a plurality of partitions of the spectrogram and a corresponding plurality of anomaly scores according to a predetermined protocol, and to perform a control action based on a maximum anomaly score.
5. The at least one processor Create an anomalous spectrogram library based on known anomalous behavior, Using the abnormal spectrogram library to identify difficult-to-predict target regions; The speech processing system of claim 4 , further configured to utilize the identified target regions as one or more hypotheses to find the maximum anomaly score.
6. the at least one processor is configured to determine a target region having the highest anomaly score by testing the one or more hypotheses; The one or more hypotheses include: an intermediate frame hypothesis procedure aimed at recovering a temporally intermediate portion of the spectrogram from lateral portions of the spectrogram that bracket the intermediate portion of the frame from either side of the frame of the spectrogram; a frequency masking hypothesis procedure aimed at recovering a specific frequency region of the spectrogram from a surrounding unmasked region of the spectrogram, the recovery of the specific frequency region corresponding at least to reconstructing high frequencies of the spectrogram from low frequencies of the spectrogram or reconstructing low frequencies from high frequencies of the spectrogram; a frequency masking hypothesis procedure aimed at recovering individual frequency bands from adjacent and / or harmonically related frequency bands of said spectrogram; an energy-based hypothesis procedure that aims to recover high-energy time-frequency parts of the spectrogram from the remaining unmasked time-frequency parts of the spectrogram; a procedure aimed at recovering a randomly selected subset of masked frequency bands and time frames from the remaining unmasked regions of the spectrogram; a likelihood bootstrapping procedure that performs multiple passes using different context regions of the spectrogram determined by first reconstructing all spectrograms sampling different percentages of the time-frequency portions as context of the spectrogram, determining time-frequency regions of the reconstructed spectrograms that are more likely to be reconstructed, and reconstructing time-frequency regions that are less likely to be reconstructed using the time-frequency regions of the reconstructed spectrograms that are more likely to be reconstructed as context; and an ensemble procedure for determining the maximum anomaly score by combining a plurality of the hypothesis generation procedures described above.
7. The attention-based neural processing architecture comprises: an encoder neural network trained to receive an input set of any size, the input set corresponding to the values and coordinates of elements in the context region, the encoder neural network generating an embedding vector for each element of the input set; a cross-attention module trained to compute a unique embedding vector for each element in the target region by processing the embedding vectors of the elements in the context region located at adjacent coordinates; and a decoder neural network that outputs a probability distribution for each element in the target region based on its coordinates in the target region and the unique embedding vector of the element in the target region.
8. The speech processing system of claim 7 , wherein the encoder neural network jointly encodes all elements in the context region using a self-attention mechanism.
9. The speech processing system of claim 7 , wherein the cross-attention module uses multi-head attention.
10. The speech processing system of claim 7 , wherein the decoder neural network outputs at least one of a plurality of parameters of conditionally independent Gaussian distributions and a plurality of parameters of a mixture of conditionally independent Gaussian distributions.
11. the at least one processor is configured to implement a sliding window on the spectrogram; The voice processing system of claim 1 , wherein the neural network using an attention-based neural network architecture determines the anomaly score for detecting the abnormal sound by processing the sliding window.
12. 1. A computer-implemented method for detecting abnormal sounds, comprising: receiving a spectrogram of an audio signal having elements defined by values in a time-frequency domain, the value of each element of the spectrogram being identified by a coordinate in the time-frequency domain; Partitioning the time-frequency region of the spectrogram into a context region and a target region; Recovering values of elements having coordinates within the target region of the spectrogram by providing the values of elements within the context region and the coordinates of the elements within the context region to a neural network including an attention-based neural processing architecture. To do, determining an anomaly score for detecting the abnormal sound of the audio signal based on a comparison of the restored values of the elements in the target region and the values of the elements in the segmented target region; and performing a control action based on the anomaly score.
13. generating a set of context regions and a corresponding set of target regions by dividing the spectrogram into different combinations of context regions and target regions; reconstructing a set of target regions by running the neural network multiple times, once for each context region in the set of context regions; determining a set of anomaly scores by comparing each target region in the set of reconstructed target regions to a corresponding target region; and determining the anomaly score based on a pooling operation on the set of anomaly scores.
14. the context area is a first context area, the target area is a first target area, the anomaly score is a first anomaly score, The method comprises: identifying a second partition of the time-frequency domain based on the first anomaly score; and Dividing the second partition of the spectrogram into a second context region and a second target region; reconstructing the second target region by repeatedly running the neural network with values and coordinates of the second context region, and generating a second anomaly score based on a comparison of the reconstructed second target region and the segmented second target region; 14. The method of claim 13, further comprising: performing a second control action based on the second anomaly score, a combination of the first anomaly score and the second anomaly score, or both.
15. training the neural network by randomly or pseudo-randomly selecting different sections of a training spectrogram for context and target regions; generating, during execution of the neural network, a plurality of partitions of the spectrogram and a corresponding plurality of anomaly scores according to a predetermined protocol; and performing a control action based on the maximum anomaly score.
16. creating an anomalous spectrogram library based on known anomalous behavior; using the abnormal spectrogram library to identify difficult-to-predict target regions; 20. The method of claim 15, further comprising: utilizing the identified target regions as one or more hypotheses to find the maximum anomaly score.
17. The method further includes determining a target region having the highest anomaly score by testing the one or more hypotheses; The one or more hypotheses include: an intermediate frame hypothesis procedure aimed at reconstructing a temporally intermediate portion of the spectrogram from the flanking portions of the spectrogram which bracket the intermediate portion from either side; a frequency masking hypothesis procedure aimed at recovering a specific frequency region of the spectrogram from a surrounding unmasked region of the spectrogram, the recovery of the specific frequency region corresponding at least to reconstructing high frequencies of the spectrogram from low frequencies of the spectrogram or reconstructing low frequencies from high frequencies of the spectrogram; a frequency masking hypothesis procedure aimed at recovering individual frequency bands from adjacent and / or harmonically related frequency bands; an energy-based hypothesis procedure that aims to recover high-energy time-frequency parts of the spectrogram from the remaining unmasked time-frequency parts of the spectrogram; a procedure aimed at recovering a randomly selected subset of masked frequency bands and time frames from the remaining unmasked regions of the spectrogram; a likelihood bootstrapping procedure, which first reconstructs all spectrograms by sampling different proportions of time-frequency parts as the context of the spectrograms, and then from the reconstructed spectrograms, select only the time-frequency regions with a high likelihood of reconstruction and use them as context to reconstruct the time-frequency regions with a low likelihood of reconstruction; and an ensemble procedure that determines the maximum anomaly score by combining a number of the hypothesis generation procedures described above.
18. The attention-based neural processing architecture comprises: receiving an input set of any size using a trained encoder neural network of the attention-based neural processing architecture, the input set corresponding to the values and coordinates of elements in the context region, and an embedding vector for each element of the input set being output by the encoder neural network; The attention-based neural processing architecture comprises: computing a unique embedding vector for each element in the target region by processing the embedding vectors of the elements in the context region located at adjacent coordinates using a trained cross-attention module of the attention-based neural processing architecture; and outputting, using a trained decoder neural network, a probability distribution for each element in the target region based on its coordinates in the target region and the unique embedding vector of the element in the target region.
19. 20. The method of claim 18, further comprising the encoder neural network jointly encoding all elements in the context region using a self-attention mechanism.
20. 20. The method of claim 18, further comprising the decoder neural network outputting at least one of a plurality of parameters of conditionally independent Gaussian distributions and a plurality of parameters of a mixture of conditionally independent Gaussian distributions.
Citation Information
Patent Citations
Determination system, separation device, determination method, and program
JP2020177021A