System and method for teacherless abnormal sound detection

The system employs a multi-head neural network to generate interdependent embedding vectors for multiple surrogate tasks, enhancing resilience against domain shift and improving the robustness of anomaly sound detection in varying acoustic environments.

JP2025516388AActive Publication Date: 2025-05-27MITSUBISHI ELECTRIC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025512399
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-23
Filing Date
2023-06-22
Publication Date
2025-05-27
Estimated Expiration
2043-06-22

AI Technical Summary

Technical Problem

Existing anomaly detection systems fail when presented with acoustic characteristics that vary between training and inference due to domain shift, unable to distinguish between unexpected signal changes from anomalies and expected changes from domain shift.

Method used

A system and method using a multi-head neural network to generate interdependent embedding vectors for multiple surrogate tasks, which are then analyzed individually to enhance resilience against domain shift, allowing for robust unsupervised anomaly sound detection.

Benefits of technology

The approach effectively addresses the domain shift problem, improving the robustness of anomaly sound detection by learning rich embedding vectors that are resilient to variations in acoustic characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516388000001_ABST
    Figure 2025516388000001_ABST
Patent Text Reader

Abstract

A system and method for detecting abnormal sounds are disclosed. The method comprises receiving an audio signal from a sound source in a recording environment. The sound source and the recording environment are characterized by a set of attributes including a first attribute related to a first attribute type and a second attribute related to a second attribute type. A multi-head neural network is trained to extract from the received audio signal a first embedding vector indicating the first attribute type and a second embedding vector indicating the second attribute type. To classify the attributes of the first attribute type, the first embedding vector is compared with a first set of embedding vectors, and to classify the attributes of the second attribute type, the second embedding vector is compared with a second set of embedding vectors to determine an anomaly detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to audio processing, and more specifically, to the detection of anomalies in audio signals and the explanation of detected anomalies.

Background Art

[0002] The diagnosis and monitoring of the operating performance of machines are important in various applications. The diagnosis and monitoring tasks are often performed manually by skilled technicians. For example, a skilled technician may listen to and analyze the sounds generated by a machine to determine abnormal sounds. The manual process of analyzing sounds may be automated to process the sound signals generated by the machine and detect abnormal sounds within the sound signals.

[0003]

[0004] In some scenarios, automated sound diagnosis is performed to detect abnormal sounds based on deep learning-based techniques. Typically, automated sound diagnosis can be used to detect abnormal sounds using training data corresponding to the normal operating conditions of sound diagnosis. The detection of abnormal sounds based on such training data is an unsupervised approach. Unsupervised detection of abnormal sounds may be suitable for the detection of certain types of anomalies, such as impact sounds that may be detected based on sudden transient disturbances or sudden temporal changes.Anomaly detection without a teacher is the problem of learning a model that can detect anomalies when only data in a normal operating state can be used for training model parameters. Typical application examples are the state monitoring and diagnosis of machine sounds in applications such as predictive maintenance and factory automation. Typical approaches to anomaly detection without a teacher include those based on architectures such as autoencoders, where a model trained only on normal data for input reconstruction should show a large reconstruction error when presented with abnormal examples during inference. Another class of approaches called surrogate task models learn a normality model using another supervised training task and measure the deviation from normal values to predict anomalies. Examples of surrogate tasks include (1) OE (outlier exposure) where sounds known to be completely different from the target machine are used as synthetic anomalies, (2) prediction of metadata (such as machine instances) or attributes (such as operating load), or (3) learning to predict what kind of augmentation (such as time stretching or pitch shifting) is applied to an audio clip.

[0005] One major drawback of existing anomaly detection systems is that they fail when presented with acoustic characteristics that vary between the normal data collected for training and the normal data collected during inference due to factors such as domain shift, i.e., different background noises, different operating voltages, etc. Such failures are usually caused by algorithms that cannot distinguish between unexpected signal changes due to anomaly sounds and expected signal changes due to domain shift.

[0006] Therefore, it is necessary to overcome the above-mentioned problems related to anomaly detection under domain shift conditions. More specifically, it is necessary to develop a method and system for detecting anomaly sounds in an audio signal in an efficient and feasible manner. SUMMARY OF THE INVENTION

[0007] Various embodiments of the present disclosure disclose a system and method for detecting abnormal sounds in an audio signal. The purpose of some embodiments is to perform abnormal sound detection using deep learning techniques.

[0008] The purpose of some embodiments is to provide unsupervised abnormal sound detection by learning a model capable of detecting abnormalities when only data in a normal operating state is available for training model parameters. Typical application examples are state monitoring and diagnosis of machine sounds in applications such as predictive maintenance and factory automation.

[0009] One major drawback of existing abnormal sound detection algorithms is that they fail when presented with acoustic characteristics that vary between the normal data collected for training and the normal data collected during inference due to factors such as domain shift, i.e., different background noises, different operating voltages, etc. Such failures are usually caused by algorithms that cannot distinguish between unexpected signal changes due to abnormal sounds and expected signal changes due to domain shift.

[0010] Typical approaches for unsupervised abnormal sound detection include those based on architectures such as autoencoders, where a model trained only on normal data for input reconstruction should exhibit a large reconstruction error when an abnormal example is presented during inference. Another class of approaches, referred to herein as surrogate task models, learn a model of normality using an alternative supervised training task and then measure the deviation from the normal value to predict abnormalities. Examples of surrogate tasks include (1) OE that uses a sound known to be completely different from the target machine as a synthetic abnormality, (2) prediction of metadata (such as machine instance) or attributes (such as operating load), or (3) learning to predict what extensions (such as time stretching or pitch shifting) are applied to an audio clip.

[0011] Some embodiments are based on the recognition that unsupervised anomaly sound detection using a surrogate task model can be adapted to a specific operation scenario by selecting a specific surrogate task, and thus can be more robust in a specific application compared to an approach based on an architecture such as an autoencoder. However, selecting a specific surrogate task for a specific application may not be realistic in some scenarios. Moreover, even if a surrogate task model is developed for a specific advantageous task, this model still has the domain shift problem.

[0012] Therefore, an object of some embodiments is to provide a system and method for unsupervised anomaly sound detection that is robust to the domain shift problem. Further, or alternatively, an object of some embodiments is to provide a surrogate task model approach for learning a normality model that is strong against the domain shift problem. Further, or alternatively, an object of some embodiments is to provide an anomaly sound detection system configured to perform anomaly detection, and an application that can benefit from such detection.

[0013] Some embodiments are based on the understanding that the surrogate task model approach can be extended to consider multiple surrogate tasks instead of just one surrogate task. In other words, this approach can be extended to learn the normality models of multiple surrogate tasks. Theoretically, since two tasks may be better than one task, this approach can be stronger against domain shift. However, some embodiments are based on the experimentally verified recognition that simply extending the surrogate task model approach to consider multiple surrogate tasks does not necessarily make the learned normality model stronger against the domain shift problem.

[0014] Some embodiments are based on the recognition that in order to enhance the resilience of a normality model against domain shift, two conditions need to be met: (1) the embedding vectors generated to classify multiple surrogate tasks need to be generated together, i.e., they need to have a dependency relationship with each other, and (2) these embedding vectors need to be analyzed individually to detect anomalies. Since certain surrogate tasks do not provide sufficient information on their own to learn a powerful model, the requirement for interdependent generation enables the learning of rich embedding vectors. In individual tests, the surrogate tasks are separated to obtain resilience against domain shift.

[0015] Some embodiments are based on the recognition that these two conditions can be met when the embedding vectors of different surrogate tasks are generated by a multi-head neural network having one input and multiple outputs, and the outputs of the multi-head neural network are analyzed individually. For example, in one embodiment, an audio signal is processed using a multi-head neural network trained to extract a first embedding vector indicating a first attribute type and a second embedding vector indicating a second attribute type from the received audio signal. Since the first attribute type and the second attribute type are different from each other, they can be adapted by different embodiments to classify different surrogate tasks, and thus the first condition can be met. Examples of multi-head neural networks include convolutional neural network modules connected to multiple thin output layers including a first output layer for outputting the attributes of the first attribute type and a second output layer for outputting the attributes of the second attribute type.

[0016] Next, in some embodiments, a first anomaly score is generated by comparing a first embedding vector with a first set of normal embedding vectors, and a second anomaly score is generated by comparing a second embedding vector with a second set of normal embedding vectors. The anomaly detection result is determined based on one or a combination of the first anomaly score and the second anomaly score and satisfies a second condition for independent evaluation.

[0017] Accordingly, an anomaly sound detection system is disclosed in one embodiment. The anomaly sound detection system includes at least one processor and a memory storing instructions. When the instructions are executed by the at least one processor, the anomaly sound detection system is caused to receive an audio signal generated by a sound source in a recording environment, where the sound source and the recording environment are characterized by a set of attributes including a first attribute regarding a first attribute type and a second attribute regarding a second attribute type. The received audio signal is processed using a multi-head neural network trained to extract from the received audio signal a first embedding vector indicating the first attribute type and a second embedding vector indicating the second attribute type. The first embedding vector is compared with a first set of normal embedding vectors previously generated by the multi-head neural network to classify the attributes of the first attribute type, and the second embedding vector is compared with a second set of normal embedding vectors previously generated by the multi-head neural network to classify the attributes of the second attribute type, and an anomaly detection result is determined. Then, the anomaly sound detection system is configured to render the anomaly detection result.

[0018] Another embodiment discloses a method implemented by a computer for detecting abnormal sounds. The method comprises receiving an audio signal generated by a sound source in a recording environment, where the sound source and the recording environment are characterized by a set of attributes including a first attribute related to a first attribute type and a second attribute related to a second attribute type. The method implemented by the computer further comprises processing the received audio signal using a multi-head neural network trained to extract from the received audio signal a first embedding vector indicating the first attribute type and a second embedding vector indicating the second attribute type. The method implemented by the computer further comprises comparing the first embedding vector with a first set of normal embedding vectors previously generated by the multi-head neural network to classify the attributes of the first attribute type, and comparing the second embedding vector with a second set of normal embedding vectors previously generated by the multi-head neural network to classify the attributes of the second attribute type, determining an anomaly detection result, and rendering the anomaly detection result.

[0019] Still other features and advantages will become more readily apparent from the following detailed description when considered in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0020]

Fig. 1A

Fig. 1B

Fig. 2

Fig. 3A

Fig. 3B

Fig. 4

Fig. 5A

Fig. 5B

Fig. 6

Fig. 7

Fig. 8

MODE FOR CARRYING OUT THE INVENTION

[0021] [Description of Embodiments] In the following description, for the purpose of thorough understanding of the present disclosure, numerous specific details are set forth. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, in order to avoid obscuring the present disclosure, the apparatus and methods are shown only in block diagram form. Various changes may be made in the functions and arrangements of the elements without departing from the spirit and scope of the disclosed subject matter as set forth in the appended claims.

[0022] As used in this specification and the claims, terms such as "for example," "by way of example," and "such as," and the verbs "comprising," "having," "including," and each of their other verb forms, when used in conjunction with a listing of one or more components or other items, are to be construed as open-ended, meaning that such listing should not be regarded as excluding other additional components or items. The term "based on" means at least in part based on. Further, it should be understood that the phrases and terms used in this specification are for purposes of explanation and should not be regarded as limiting. Any headings utilized within this specification are for convenience only and have no legal or limiting effect.

[0023] Specific details are provided in the following description so that the embodiments may be fully understood. However, one skilled in the art will understand that the embodiments may be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may sometimes be shown as components in block diagram form in order to avoid obscuring the embodiments with unnecessary detail. In other instances, well-known processes, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments. Further, like reference numbers and designations in the various drawings indicate like elements. Overview of the System

[0024] FIG. 1A shows a recording environment 100 for detecting abnormal sounds according to an embodiment of the present disclosure. The recording environment 100 includes an abnormal sound detection system 102 coupled to a sound source 108 that generates an audio signal 110. The abnormal sound detection system 102 includes a processor 104 and a memory 106. The memory 106 is configured to store instructions for generating an abnormal detection result 114 for detecting abnormal sounds. Abnormal sounds include abnormal sound signals that are not normally expected to occur during operation. For example, in a manufacturing process, the operating sounds related to a machine can include normal sounds and abnormal sounds, i.e., non-normal sounds. For example, the sound that occurs during the operation of a machine may be regarded as an abnormal sound because it indicates a failure or wear of machine parts. The purpose of some embodiments is to detect such abnormal sounds using the abnormal sound detection system 102. Furthermore, the detection of abnormal sounds is used to predict failures related to the machine. For example, the accurate sound source 108 of the abnormal detection or abnormal sound is the faulty part of the machine. The operation of the abnormal sound detection system 102 may be embodied as instructions stored in the memory 106 and executed by the processor 104.

[0025] In some embodiments, the memory 106 is configured to store instructions for implementing a multi-head neural network 112 to facilitate the detection of abnormal sounds. The memory 106 corresponds to at least one of RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices, or any other storage medium that can be used to store desired information and is accessible by the abnormal sound detection system 102. The memory 106 includes a non-transitory computer storage medium in the form of volatile memory and / or non-volatile memory. The memory 106 may be removable, non-removable, or a combination thereof. Exemplary memory devices include solid state memory, hard drives, optical disk drives, and the like.

[0026] The abnormal sound detection system 102 is configured to receive an audio signal 110 generated by a sound source 108 in a recording environment 100. The recording environment 110 can correspond to an environment specific to an application, such as a manufacturing factory, a vehicle, or a studio. The audio signal 110 can correspond to non-steady sounds such as the sound of an operating machine or the sound of an operating engine. The audio signal 110 may be converted into a representation in the time-frequency domain, such as a spectrogram. Generally, a spectrogram includes elements defined by values, such as pixels in the time-frequency domain. Each value of each element is identified by coordinates in the time-frequency domain. For example, a time frame in the time-frequency domain is represented as a column, and a frequency band in the time-frequency domain is represented as a row.

[0027] The audio signal 110 is associated with the sound generated from the sound source 108. In the case of a manufacturing setup, the sound source 108 corresponds to at least one of a machine, an electrical device, or an engine. For example, the sound source 108 can include a ball mill, a grinding machine, a packaging machine, and the like.

[0028] The sound source 108 and the recording environment 100 are characterized by a set of attributes. FIG. 1B is a schematic diagram showing a set of attributes 116 of the sound source 108 and the recording environment 100.

[0029] The set of attributes 116 includes, for the sound source 108, at least a first attribute 118 related to a first attribute type 118a. Also, the set of attributes 116 includes, for the recording environment 100, at least a second attribute 120 related to a second attribute type 120a.

[0030] Therefore, in one embodiment, a first attribute type 118a and a second attribute type are selected such that the classification of the attributes of the first attribute type does not depend on one or a combination of the recording environment 100, the type of the sound source 108, or the operating state of the sound source 108. On the other hand, the classification of the attributes of the second attribute type 120a depends on one or a combination of the recording environment 100, the type of the sound source 108, or the operating state of the sound source 108. That is, the first attribute type 118a includes attributes such that the classification of the attributes of the first attribute type 118a does not depend on one or a combination of the recording environment 100, the type of the sound source 108, and the operating state of the sound source 108, and the second attribute type 120a includes other attributes such that the classification of the other attributes of the second attribute type 120a depends on one or a combination of the recording environment 100, the type of the sound source 108, and the operating state of the sound source 108.

[0031] For example, the first attribute 118 of the first attribute type 118a is associated with the sound source 108 and classifies, for example, a motor from a sound signal. In fact, the motor and / or the type of the motor remains the same regardless of the environment in which the operating sound of the motor is recorded. Therefore, the first attribute 118 is the sound of the motor, and the first attribute type 118a can be normal sound or abnormal sound.

[0032] Furthermore, in the above example, the second attribute 120 of the second attribute type 120a can indicate the operating state of the motor. Therefore, the motor includes different characteristics related to the operating state of the motor. The characteristics include the characteristics of the input provided to the motor, such as the operating voltage required for the motor. The characteristics also include the characteristics of the output of the motor, such as the torque generated by the motor. For example, the second attribute 120 can be the state s of the motor when the motor is operating at a low speed. i This may be the case, and the second attribute type may be the voltage value for supplying power to the motor in this state s. i In another example, the second attribute type 120a is the torque generated by the motor that can change during the operation of the motor from the state s to another state. i ​

[0033] Thus, the audio signal 110 is a combination of a set 116 of attributes including a first attribute 118 of a first attribute type 118a and a second attribute 120 of a second attribute type 120a. The abnormal sound detection system 102 is configured to process the audio signal 110 characterized by the set 116 of attributes including the first attribute 118 of the first attribute type 118a and the second attribute 120 of the second attribute type 120a using the multi-head neural network 112.

[0034] FIG. 2 is a block diagram 200 showing the operation of the multi-head neural network 112 according to an embodiment of the present disclosure. The multi-head neural network 112 is configured to process an audio signal 110 (which can be a test audio signal or an audio signal at the time of received inference) with one or more linear layers that are thin output layers, i.e., a first output layer 112a and a second output layer 112b. To that end, in some implementations, the multi-head neural network 112 includes a convolutional neural network module connected to a plurality of thin output layers such as the first output layer 112a and the second output layer 112b. In FIG. 2, it can be understood that, for the purpose of illustration only, the number of thin output layers is shown as 2. In practice, any same number of output layers may be used as thin output layers without departing from the scope of the present disclosure.

[0035] The multi-head neural network 112 is configured to process the audio signal 110 to extract a first embedding vector 202a indicating a first attribute type 118a and a second embedding vector 202b indicating a second attribute type 120a. To that end, the first output layer 112a outputs the extracted first embedding vector 202a, and the second output layer 202b outputs the extracted second embedding vector 202b. In one example, the multi-head neural network 112 provides a 512-D global depthwise convolution (GDC) output, which is connected to each of a linear output layer, i.e., the first output layer 112a and the second output layer 112b. Both the first output layer 112a and the second output layer 112b are 1×1 convolutions in the example.

[0036] Thereafter, the first embedding vector 202a and the second embedding vector 202b are provided to a classification module 204 including a first classifier 204a for classifying the attributes of the first attribute type 118a, which are embedded in the first embedding vector 202a and the second embedding vector 202b respectively, and a second classifier for classifying the attributes of the second attribute type 120a. As already understood, the classification of the attributes of the first attribute type 118a is independent of the recording environment 100, while the classification of the attributes of the second attribute type 120a depends on the recording environment 100. In one example, each of the first classifier 204a and the second classifier 204b uses a softmax activation function.

[0037] Some embodiments are based on the recognition that in the abnormal sound detection system 102, when surrogate tasks are separated from each other, the type of task can be selected to further improve the resistance to the domain shift problem. For example, there are tasks that depend on the domain and tasks that do not depend on the domain. Therefore, while classifying the attributes of the first attribute type 118a as independent of the recording environment, by classifying the attributes of the second attribute type 120a as dependent on the recording environment, the first surrogate task based on the classification of the first attribute type 118a is independent of the domain, while the second surrogate task based on the classification of the second attribute type 120a depends on the domain.

[0038] As an example of the surrogate task of the attributes of the first attribute type 118a, the determination of the motor from the sound signal is included. In fact, the motor and / or the type of motor are the same regardless of the environment in which the operating sound of the motor is recorded. In other words, a classifier trained to classify the embedded vector indicating the first type of attribute should produce the same result for classifying the outputs of multiple executions of the multi-head neural network 112.

[0039] Conversely, in the example of using a motor, the second type of attribute can indicate the operating state of the motor and can vary between different executions of the multi-head neural network 112. For example, the second type of attribute can indicate the voltage supplying power to the motor or the torque generated by the motor that can change during the operation of the motor.

[0040] In an embodiment, the abnormal sound detection system 102 includes a multi-head neural network 112 and includes a machine learning module (not shown) configured to perform different operations during training and execution. Further, the abnormal sound detection system 102 also includes an output interface (not shown) for outputting an abnormal detection result.

[0041] FIG. 3A is a schematic diagram showing the operation of the training stage of the abnormal sound detection system 102 according to an embodiment of the present disclosure. As shown in FIG. 3A, the training stage includes receiving a training audio signal 301 including a first set 118 of attributes of a first attribute type 118a and a second set 120 of attributes of a second attribute type 120a. The first attribute type 118a includes attributes such that the classification of the attributes of the first attribute type 118a does not depend on one or a combination of the recording environment 100, the type of sound source that generates the training audio signal 301, and the operating state of the sound source that generates the training audio signal 301. The second attribute type 120a includes other attributes such that the classification of the other attributes of the second attribute type 120a depends on one or a combination of the recording environment 100, the type of sound source that generates the training audio signal 301, and the operating state of the sound source that generates the training audio signal 301.

[0042] The multi-head neural network 112 is trained with training data including training audio signals regarding different values of the attributes of the first attribute type 118a and different values of the attributes of the second attribute type 120a. By doing so, the accuracy of the normality model of the abnormal sound detection system 102 is improved.

[0043] For example, for samples of different training audio signals 301, different values of the first attribute type 118a may exist, such as a first value 118a1, an nth value 118an (n is an arbitrary number), etc. Also, samples of different training audio signals 301 may have the same value of the first attribute type 118a.

[0044] Similarly, for samples of different training audio signals 301, different values of the second attribute type 120a may exist, such as a first value 120a1, an mth value 120am (m is an arbitrary number), etc. Also, samples of different training audio signals 301 may have the same value of the second attribute type 120a.

[0045] During training, different values of the first attribute type 118a and the second attribute type 120a form ground truth values associated with different tasks. The multi-head neural network is trained jointly with the classification module to generate different encodings that are advantageous for classification so as to match the ground truth values associated with different tasks. In particular, in some implementations, the classification module is used only during the training phase for training the multi-head neural network and is not used during the test phase.

[0046] Accordingly, each training audio signal 301 is associated with a corresponding ground truth value of an attribute of the first attribute type, for example, a first value 118a1 which is one of the different values of the first attribute type 118a. Similarly, a corresponding ground truth value of an attribute of the second attribute type 120a, for example a first value 120a1, is one of the different values of the second attribute type 120a. These ground truth values are used to train the multi-head neural network 112 in combination with the classification module 204. The multi-head neural network 112 processes each training audio signal sample 301, and the first output layer 112a extracts a first embedding vector such as an embedding vector 302a1 for the ground truth value 118a1. The second output layer 112b extracts a second embedding vector, for example, an embedding vector 302b1 for the ground truth value 120a. To that end, there can be many other different embedding vectors obtained from different training audio signal samples that are associated with the same ground truth value 118a1.

[0047] For example, for another training audio signal sample 301n, another set of embedding vectors such as embedding vector 302an and embedding vector 302bm can be extracted. And the different embedding vectors extracted for different audio signal samples form a first set 302a of embedding vectors extracted by the first output layer 112a and a second set 302b of embedding vectors extracted by the second output layer 112b. In the diagram of FIG. 3A, only two training audio signal samples are shown for the sake of brevity of explanation. However, any same number of training audio signal samples may be illustrated without departing from the scope of the present disclosure.

[0048] For example, next, the first embedding vector 302a1 and the second embedding vector 302b1 are passed to the classification module 204, where class labels are assigned to the first embedding vector 302a1 and the second embedding vector 302b1. For this purpose, the first classifier 204a classifies the first embedding vector 302a1, and the second classifier 204b classifies the second embedding vector 302b1. The parameters of the multi-head neural network 112 and the classification module 204 are jointly optimized by minimizing a classification loss function based on the classification output of the first classifier 204a and the ground truth value 118a1 of the attribute of the first attribute type 118a, and the classification output of the second classifier 204b and the ground truth value 120a of the attribute of the second attribute type. During training, a first set 302a of embedding vectors consisting of all the first embedding vectors obtained for each training audio signal 301 and a second set 302b of embedding vectors consisting of all the second embedding vectors obtained for each training audio signal 301 are associated with corresponding ground truth values such as the first value 118a1, the nth value 118an, the first value 120a1, and the mth value 120am of the attributes associated with the training audio signals for which these embedding vectors are generated, and are stored as metadata in the anomaly sound detection system 102, for example, the memory 106 of the anomaly sound detection system 102.

[0049] After such optimization is performed in the first training stage, the parameters of the multi-head neural network 112 are fixed. The first set 302a of embedded vectors and the second set 302b of embedded vectors stored in the abnormal sound detection system 102 during the last epoch of training can be used for the test stage, or, using the final multi-head neural network 112 with the fixed parameters obtained at the end of training, in the same way as these sets of embedded vectors were obtained during training, by reprocessing all the training audio signals, these can be replaced.

[0050] During the test stage, the multi-head neural network 112 is configured to receive a test audio signal and process it to obtain a first embedded vector and a second embedded vector. Next, anomaly detection is performed by comparing the first embedded vector with the embedded vectors from the first set 302a of embedded vectors and the second embedded vector with the embedded vectors from the second set 302b of embedded vectors. In some embodiments, the first embedded vector obtained from the test audio signal is compared with all the embedded vectors of the first set 302a of embedded vectors, and the second embedded vector obtained from the test audio signal is compared with all the embedded vectors of the second set 302a of embedded vectors.

[0051] FIG. 3B is a schematic diagram showing the operation of the test stage of the abnormal sound detection system according to an embodiment of the present disclosure.

[0052] In an embodiment, the multi-head neural network 112 accepts a value 306 of an attribute of a first attribute type 118a for a test audio signal, and selects, from a first set 302a of embedding vectors, a first subset 302aa' of embedding vectors having a matching first attribute, where the attribute value matches the accepted value 306 of the attribute for the test audio signal.

[0053] In another embodiment, the multi-head neural network 112 accepts a set of acceptable values of an attribute of the first attribute type 118a, and is configured to place the embedding vector for a training signal into the first subset 302aa' of embedding vectors only if the value of the first attribute type for that training signal belongs to the set of acceptable values of the first attribute type 118a. For example, if the first attribute type 118a corresponds to the operating voltage of a motor, embedding vectors corresponding only to acceptable values of the operating voltage of the motor, where it is known that no damage or harm is done to the motor, are placed into the first subset 302aa' of embedding vectors.

[0054] In an embodiment, the multi-head neural network 112 accepts a value 308 of an attribute of a second attribute type 120a for a test audio signal, and selects, from a second set 302a of embedding vectors, a second subset 302bb' of embedding vectors having a matching second attribute, where the attribute value matches the accepted value 308 of the attribute of the second attribute type 120a for the test audio signal.

[0055] In another embodiment, the multi-head neural network 112 accepts a set of acceptable values of an attribute of the second attribute type 120a, and is configured to place the embedding vector for a training signal into the second subset 302bb' of embedding vectors only if the value of the second attribute type for that training signal belongs to the set of acceptable values of the second attribute type 120a.

[0056] The abnormal sound detection system 102 is configured to store in the memory 106 a training data set 302 of embedded vectors including a first set 302a of embedded vectors and a second set 302b of embedded vectors, and associated values of a first attribute type and a second attribute type for the training audio signals from which they were obtained. Thus, the anomaly detection is based on the first set 302a of embedded vectors and the second set 302b of embedded vectors generated and stored during training, or the first subset 302aa' of embedded vectors and the second subset 302bb' of embedded vectors extracted therefrom based on one of the above criteria, and the first embedded vector 202a and the second embedded vector 202b that are generated at runtime and may be different from the first set 302a of embedded vectors and the second set 302b of embedded vectors previously generated and previously classified by the multi-head neural network 112. It may be possible to perform this by comparing them respectively. For this purpose, the reference to "previously" may be understood to be equivalent to "during training" within the scope of the considerations in the present disclosure.

[0057] Furthermore, during the training phase of the multi-head neural network 112, the training dataset 302 of the embedding vectors is passed to a classification module such as the classification module 204 shown in FIG. 2. The classification module 204 is a surrogate task classifier (typically implemented as a multinomial logistic regression (aka softmax) layer) for predicting the labels of the received audio signal embeddings generated by the multi-head neural network 112. This prediction is based on surrogate tasks such as (1) the task of predicting the exact model of the machine that generated the sound corresponding to the embedding (e.g., there are many different types of valve machines, all of which may have different model numbers), or (2) the task of predicting the attributes of the machine based on the operating principle of the machine, such as the fan speed or the voltage of the power supply, (3) the task of predicting the background characteristics of the location where the normal sound was recorded (e.g., whether it was recorded in factory A, factory B, a laboratory, etc.), or (4) the task of creating synthetic anomalies by corrupting the sound signal or presenting examples from another machine and enabling the comparison of the predicted label of the surrogate task with the true label, but is not limited to these. Furthermore, during training, the predicted label of the surrogate task can be compared with the true label using a cross-entropy loss function.

[0058] Some embodiments recognize that since the values of the attributes during the execution of the anomaly detection system 102 are unknown, this comparison is advantageous for domain-dependent attributes, but the value of the first attribute type 118a may also be known during execution. Comparing the first embedding vector 202a with the embedding vectors within the first set 302a of embedding vectors corresponding to different values of the first attribute type 118a is unnecessary because only the normal first embedding vectors having the same value of the first attribute 118 of the first attribute type 118a are relevant when calculating the overall anomaly score during execution. Thus, in some embodiments, the first embedding vector 202a is only compared with the first subset 302aa' of embedding vectors corresponding to the same value of the first attribute type 118a.

[0059] At runtime or inference time, the abnormal sound detection system 102 is configured to compare a first embedded vector 202a with a first set 302a of embedded vectors previously generated by the multi-head neural network 112 and stored in the memory 106 or a first subset 302aa' of the embedded vectors to classify the attributes of the first attribute type 118a of the received audio signal 110 at inference time. The second embedded vector 202b is compared with a second set 302b of embedded vectors previously generated by the multi-head neural network 112 or a second subset 302bb' of the embedded vectors to classify the attributes of the second attribute type 120a of the received audio signal 110 at inference time. The comparison result is used to determine the abnormal detection result of the audio signal 110 received at inference time.

[0060] In an embodiment, the comparison of the first embedded vector 202a with the first set 302a of embedded vectors or the first subset 302aa' of the embedded vectors, and the comparison of the second embedded vector 202b with the second set 302b of embedded vectors or the second subset 302bb' of the embedded vectors are performed using a nearest neighbor distance metric, which is one of the Euclidean distance or the cosine distance.

[0061] Furthermore, the abnormal sound detection system 102 is configured to determine an abnormality score of the received audio signal 110 by using the received audio signal 110 and the abnormality detection result. The abnormal sound detection system 102 determines whether the received audio signal 110 is a normal audio signal or an abnormal audio signal based on the abnormality score. When the abnormality score exceeds a preset threshold, the received audio signal 110 is determined to be abnormal. In one example, both the abnormality score and the preset threshold may be numerical values between 0 and 1. For example, the preset threshold may be 0.5, and the determined abnormality score may be 0.4. Since the abnormality score 0.4 is smaller than the preset threshold 0.5, in this example, the received audio signal 110 is determined to be normal or not abnormal. The preset threshold is set based on experimental data in some embodiments.

[0062] FIG. 4 is a schematic diagram 400 showing a process of determining an abnormality score according to an exemplary embodiment of the present disclosure. For the purpose of facilitating the description in the present disclosure, the training data set 302 of the embedded vectors shown in FIG. 4 includes the first set 302a of the embedded vectors, the first subset 302aa' of the embedded vectors, the second set 302b of the embedded vectors, and the second subset 302bb' of the embedded vectors, which were described above in FIGS. 3A and 3B.

[0063] As shown in FIG. 4, the first embedded vector 202a generated at runtime is compared with a training data set of embedded vectors, which includes the first set 302a of the embedded vectors and the first subset 302aa' of the embedded vectors, previously generated (such as during training) by the multi-head neural network 112 to generate a first abnormality score 402.

[0064] The second embedded vector 202b generated at runtime is compared with a training data set of embedded vectors, including a second set 302b of embedded vectors and a second subset 302bb' of embedded vectors, previously generated by the multi-head neural network 112, to generate a second anomaly score 404. Next, the first anomaly score 402 and the second anomaly score 404 are combined (406) to determine an anomaly detection result 114 based on a combination 406 of the first anomaly score 402 and the second anomaly score 404.

[0065] In one example, the combination 406 of the first anomaly score 402 and the second anomaly score 404 is a weighted combination where the weight of the first anomaly score 402 is smaller than the weight of the second anomaly score 404. For this purpose, since the first attribute 118 is less likely to cause a failure in the machine that generates the audio signal 110 compared to the second attribute 120, a smaller weight of the first anomaly score 402 is used. In embodiments, if it is proven that there is an attribute with low reliability in anomaly prediction over time, different weights are considered for different anomaly scores.

[0066] FIG. 5A is a schematic diagram 500a showing an example method of anomaly score generation according to an exemplary embodiment of the present disclosure.

[0067] The anomaly sound detection system 102 is configured to determine an anomaly score, such as a combined anomaly score 406, by concatenating a first embedded vector 202a and a second embedded vector 202b to generate a concatenated embedded vector 502. Further, a nearest neighbor distance algorithm 504 is used to compare the concatenated embedded vector 502 with the embedded vectors of the training data set 302 by calculating the minimum distance between the concatenated embedded vector 502 and the embedded vectors of the training data set 302.

[0068] To that end, the abnormal sound detection system 102 is configured to compare the generated concatenated embedding vector 502 with each of the embedding vectors of the training dataset 302 using a distance measurement method for calculating the minimum distance between the concatenated embedding vector 502 and the embedding vectors of the training dataset 302. The distance measurement method includes at least one of a Euclidean distance method, a cosine distance method, or a weighted Euclidean distance. For example, the first embedding vector 202a is compared with the first set 302a of embedding vectors and the first subset 302aa' of embedding vectors to determine the first distance measure. Similarly, the second embedding vector 202b is compared with the second set 302b of embedding vectors and the second subset 302bb' of embedding vectors to determine the second distance measure. The first distance measure and the second distance measure are then combined by a distance measurement algorithm 504 to generate a combined anomaly score 406. The combination may include a sum, an average, a weighted average, etc.

[0069] FIG. 5B is a block diagram 500b showing another method of determining the combined anomaly score 406 with the aid of a learned weighting method according to an embodiment of the present disclosure. The abnormal sound detection system 102 is configured to determine the combined anomaly score 406 individually for the first embedding vector 202a and the second embedding vector 202b of separated dimensions by applying the distance measurement algorithm 504 individually to the first embedding vector 202a and the second embedding vector 202b to determine the first anomaly score 504a and the second anomaly score 504b, respectively. The separation of dimensions corresponds to the first embedding vector 202a and the second embedding vector 202b having different dimensions, and the dimensions refer to a quantitative or qualitative measure of the features included in the dimensions.

[0070] Furthermore, the separate anomaly scores for the separated dimensions are combined using the learned weighting method 506 to determine the combined anomaly score 406. To that end, each of the first anomaly score 504a and the second anomaly score 504b is weighted before being combined by a known combination method such as addition, averaging. The weight of the first anomaly score 504a generated for the first set 118 of attributes (independent of the recording environment 100) is, in one example, smaller than the weight of the second anomaly score 504b generated for the second set 120 of attributes (dependent on the recording environment 100). This is because the values of the first set 118 of attributes may be known at runtime, and comparing the first embedding vector 202a with the first set 302a of embedding vectors corresponding to different values of the first type of attribute and the first subset 302aa' of embedding vectors is either unnecessary or less relevant because only the first subset 302aa' of embedding vectors having the same value of the first type of attribute 118 is relevant in calculating the overall anomaly score during execution. Such weighted combination improves the accuracy and performance of the multi-head neural network 112 by not overly focusing on redundant processing operations.

[0071] Furthermore, based on the contribution of the combined anomaly score 406 to a particular combination of separated dimensions, the abnormal sound detection system 102 generates an anomaly detection result 114 to predict the reason behind the abnormal sound. For example, if the separated dimension corresponding to the speed prediction during execution contributes the most to the combined anomaly score 406, the abnormal sound detection system 102 further causes the processor 104 to generate a control command to investigate only the part of the machine that controls the speed as being likely the cause of the detected abnormal sound.

[0072] In one example, during inference, a database of prototype normal sound embeddings in the form of a training dataset of normal embedding vectors is used to compare one or more embedding vectors 202 of the received audio signal 110 with the training dataset of normal embedding vectors. If the training dataset of normal embedding vectors is small enough, all training samples are used as prototypes. However, in some embodiments, the training dataset of normal embedding vectors is further reduced by using an algorithm such as K-means clustering.

[0073] FIG. 6 is a schematic diagram showing a use case 600 for collecting a training data set 302 of embedding vectors used to train a multi-head neural network 112 according to an embodiment of the present disclosure. The abnormal sound detection system 102 detects abnormal sounds using the collected training data set 302 of embedding vectors. The use case 600 shows different models of the ball board, namely, the ball board model 602, the ball board model 604, and the ball board model 606 in the factory 608 (hereinafter, for simplicity of explanation, different ball board models are referred to as ball boards 602 to 606). To determine when an abnormal sound occurs, the ball boards 602 to 606 are monitored by microphones 610. In one example, the ball board model 602 is used for a plate material. The ball board model 604 is used for an iron plate, and the ball board model 606 is used for a steel plate. The training data set 302 of embedding vectors is created using only the recordings of the sounds of the ball boards 602 to 606 collected during normal operation. The ball boards 602 to 606 are complex machines and are controlled by many different operating parameters. Examples of operating parameters include the speed at which the motor of the ball board rotates, the type of material to be drilled, and the like. By changing exemplary operating parameters and recording the sounds they generate, the abnormal sound detection system 102 collects a training data set 302 of embedding vectors corresponding to normal training data (through the collection of corresponding audio signals and the extraction of embedding vectors from them). Further, a surrogate task is defined as a classifier aimed at predicting specific operating parameters from the sound signals generated by the ball boards 602 to 606 (for example, predicting the rotation speed). Since the ball boards 602 to 606 have a plurality of operating parameters, a plurality of different classifiers can be trained to classify the plurality of operating parameters.

[0074] However, once the classifier is trained, the abnormal sound detection system 102 may not know the specific environment of an individual pachinko machine during inference. This specific environment may include, for example, a specific factory 608 where the pachinko machines 602-606 are operating (and the associated background noise). These specific environments are called domains. To make the abnormal sound detection system 102 operate well in as many domains as possible, different classifiers perform predictions on attributes that do not change based on the domain, i.e., predictions of domain-shared attribute features (e.g., prediction of the model numbers of the pachinko machines 602-606, prediction of the rotational speed of a drill, etc.), and are trained to perform predictions of features that change based on the domain, which are called domain-specific attributes in this specification (e.g., prediction of the environment in which the pachinko machine is operating, or prediction of the type of material of the pachinko machine).

[0075] By calculating a separate anomaly score for each classifier during inference, the abnormal sound detection system 102 can determine anomaly scores in different ways from the embedded vectors trained by different classifiers (e.g., the first classifier 204a and the second classifier 204b). For example, if any pachinko machine is made of plastic and the abnormal sound detection system 102 has never obtained a recording of the sound of a pachinko machine drilled into plastic before, the abnormal sound detection system 102 can weight the anomaly score from the domain-specific classifier. Alternatively, the abnormal sound detection system 102 can determine the weights between different classifiers based on the accuracy with which the classifier is trained (i.e., the more accurate classifier has a higher weight). As an additional approach for determining the weights of classifiers in different dimensions, after several known anomalies are observed, the abnormal sound detection system 102 can adjust the weights in different dimensions so that the observed anomalies are detected in the future with high reliability without requiring retraining of the classifier.

[0076] Implementation Example

[0077] [Number]

[0078]

Number

[0079]

Number

[0080]

Number

[0081]

Number

[0082] wherein, w S and w A are scalar weights and are optimized after the training is completed.

[0083] In the case of a separated section, only the section embedding z S is used in (4).

[0084] In the case of a separated attribute, only the attribute embedding z A is used in (4).

[0085] At the time of testing, since the section label of the test sample is known, the restriction on the training set sample when calculating the NN distance is set from D and only samples belonging to the appropriate section are obtained.

[0086] An implementation example will be further described using the flowchart shown in FIG. 7.

[0087] FIG. 7 is a flowchart 700 showing a method for detecting an abnormal sound signal according to an embodiment of the present disclosure. The method is executed by an abnormal sound detection system 102. The method starts at operation 702.

[0088] At 704, the method comprises receiving an audio signal (e.g., audio signal 110 of FIG. 1) from a sound source 108 within a recording environment 100. The sound source 108 corresponds to at least one of a machine, an electrical device, or an engine. In one example, the sound source 108 may include a ball mill, a grinding machine, a packaging machine, etc. Further, the abnormal sound detection system 102 is configured to utilize the received audio signal 110 to generate one or more embedding vectors with the aid of a multi-head neural network such as the multi-head neural network 112. The sound source 108 and the recording environment 100 are characterized by an attribute set 116 including a first attribute 118 regarding a first attribute type 118a and a second attribute 120 regarding a second attribute type 120a. The first attribute 118 is independent of the recording environment 100, while the second attribute is dependent on the recording environment 100. For example, in the case of a motor operating in an industrial automation environment, the attribute regarding the type of the motor is a first attribute type 118a that is independent of the recording environment 100 of the industrial automation setup, and the attributes regarding the operating state of the motor such as the operating voltage and the operating torque are dependent on the type of the industrial automation setup. Thus, these attributes related to the operating state of the motor are of the second attribute type 120a. In this way, the received audio signal 110 is formed as a combination of one or more of the first attribute 118 of the first attribute type 118a and the second attribute 120 of the second attribute type 120a.

[0089] At 706, the received audio signal 110 is processed using a multi-head neural network, such as the aforementioned multi-head neural network 112, that is trained to extract from the received audio signal 110 a first embedding vector indicative of a first attribute type 118a and a second embedding vector indicative of a second attribute type 120a. As shown in FIG. 2, the multi-head neural network 112 passes the received audio signal 110 through a first output layer 112a to extract a first embedding vector 202a of the first attribute type 118a and through a second output layer 112b to extract a second embedding vector 202b of the second attribute type 120a.

[0090] In 708, the first embedding vector is compared with a first set of embedding vectors previously generated by a multi-head neural network to classify attributes of a first attribute type, and the second embedding vector is compared with a second set of embedding vectors previously generated by a multi-head neural network to classify attributes of a second attribute type, and an anomaly detection result is determined. As previously described in conjunction with FIG. 4, the first embedding vector 202a is compared with a first set of embedding vectors 302a, and this first set of embedding vectors 302a was previously generated by the multi-head neural network 112 during training and is also included in a first subset 302aa' of the embedding vectors of the training data set 302 of the embedding vectors stored in the memory 106 of the abnormal sound detection system 102. Similarly, the second embedding vector 202b is compared with a second set of embedding vectors 302b, and this second set of embedding vectors 302b was previously generated by the multi-head neural network 112 during training and is also included in a second subset 302bb' of the embedding vectors of the training data set 302 of the embedding vectors stored in the memory 106 of the abnormal sound detection system 102. By the comparison, a first anomaly score 402 and a second anomaly score 404 are generated, which are combined to generate a combined anomaly score 406. This combination is performed using any of the techniques previously described in connection with FIGS. 5A and 5B, i.e., using a concatenated combination of embedding vectors, or using a learned weighting method for anomaly scores, or using any other equivalent technique. Next, the generated combined anomaly score 406 is used to determine the anomaly detection result 114. For example, the anomaly detection result 114 is the detection of an abnormal sound when the determined combined anomaly score exceeds a predefined threshold. In one example, the predefined threshold can be set by the machine learning module of the processor 104, for example, by using 95 percent of the anomaly scores calculated over the entire non-abnormal training set.

[0091] At 710, the anomaly detection result 114 is rendered. The rendering can be performed on, for example, one or more of a display, a user interface, an audio interface, or a combination thereof associated with the abnormal sound detection system 102. For example, the abnormal sound detection system 102 displays an abnormal sound signal on a display interface showing spectrograms of different audio signals received in the recording environment 100. The spectrogram of the abnormal sound signal may be highlighted in a color different from that of the non-abnormal sound signal. For example, the spectrogram of the abnormal sound signal may be highlighted in red, and the spectrogram of the non-abnormal sound signal may be displayed in green. Further, the display may include more information regarding the source of the abnormal sound signal that can be obtained from the first embedded vector 202a or the second embedded vector 202b generated for the received audio signal 110. The method ends at 712.

[0092] The method shown in FIG. 7 provides more accurate and efficient detection of abnormal sound signals, which can be used in various applications such as machines, engines, and numerical control parts.

[0093] FIG. 8 is a block diagram 800 showing the hardware framework of the abnormal sound detection system 102 according to an embodiment of the present disclosure. In some exemplary embodiments, the block diagram 800 includes one or more microphones 802a for collecting the audio signal 110 and a training data set 806 (which corresponds to the training data set 302 shown in FIG. 3).

[0094] The abnormal sound detection system 102 includes a hardware processor 808. The hardware processor 808 communicates with a computer storage memory such as a memory 810. The memory 810 includes stored data including algorithms, instructions, and other data implemented by the hardware processor 808. The hardware processor 808 may be considered to include two or more hardware processors according to specific application requirements. The two or more hardware processors may be either internal processors or external processors. The abnormal sound detection system 102 is combined with other components including an output interface and a transceiver, among several devices.

[0095] In some alternative embodiments, the hardware processor 808 is connected to a network 804 that communicates with an audio signal 110 source. The network 804 includes, but is not limited to, one or more local area networks (LANs) and / or wide area networks (WANs) as non-limiting examples. Also, the network 804 includes enterprise-scale computer networks, intranets, and the Internet. The abnormal sound detection system 102 includes one or more numbers of client devices, storage components, and data sources. Each of the one or more client devices, storage components, and data sources includes one or more devices that cooperate in the distributed environment of the network 804.

[0096] In some other alternative embodiments, the hardware processor 808 is connected to a network-enabled server 814 that is connected to a client device 816. The network-enabled server 814 corresponds to a dedicated computer connected to the network that processes client requests received from the client device 816 and runs software intended to provide an appropriate response on the client device 816. The hardware processor 808 is connected to an external memory device 818 that stores all the necessary data used for the detection of abnormal sound signals, and a transmitter 820. The transmitter 820 facilitates data transmission between the network-enabled server 814 and the client device 816. Further, an output 822 associated with the detection of abnormal sound signals is generated.

[0097] The audio signal 110 and the training data set 806 are further processed by a multi-head neural network 112. The multi-head neural network 112 is trained using the training data set 806 of normal embedding vectors. (As described previously).

[0098] The abnormal sound detection system 102 is configured to detect malfunctioning parts in a manufacturing setup based on abnormal sound detection disclosed in various embodiments described herein.

[0099] Many changes and other embodiments of the present disclosure described herein will come to mind to those skilled in the art to which these disclosures pertain, having the benefit of the teachings presented in the foregoing description and the related drawings. It is to be understood that the present disclosure is not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Further, although the foregoing description and the related drawings illustrate embodiments in the context of examples of particular combinations of elements and / or functions, it is to be understood that alternative embodiments may be provided by different combinations of elements and / or functions without departing from the scope of the appended claims. In this regard, for example, different combinations of elements and / or functions than those explicitly described above may be considered, as may be described in some of the appended claims. Although specific terms are used herein, they are used only in a general and descriptive sense and not for purposes of limitation.

Claims

**Claim 1** An abnormal sound detection system, comprising at least one processor, and a memory storing instructions which, when executed by the at least one processor, cause the abnormal sound detection system to, receive an audio signal generated by a sound source in a recording environment, the sound source and the recording environment being characterized by a set of attributes including a first attribute regarding a first attribute type and a second attribute regarding a second attribute type, and the instructions further cause the abnormal sound detection system to, process the received audio signal using a multi-head neural network trained to extract from the received audio signal a first embedding vector indicating the first attribute type and a second embedding vector indicating the second attribute type, compare the first embedding vector with a first set of normal embedding vectors previously generated by the multi-head neural network to classify the attribute of the first attribute type, and compare the second embedding vector with a second set of embedding vectors previously generated by the multi-head neural network to classify the attribute of the second attribute type, thereby determining an abnormality detection result, and cause rendering of the abnormality detection result. An abnormal sound detection system. **Claim 2** The multi-head neural network includes a convolutional neural network module connected to a plurality of thin output layers including a first output layer for outputting the attribute of the first attribute type and a second output layer for outputting the attribute of the second attribute type. The abnormal sound detection system according to claim 1. **Claim 3** The first attribute type includes the attribute such that classification of the attribute of the first attribute type does not depend on one or a combination of the recording environment, the type of the sound source, and the operating state of the sound source, The second attribute type includes the other attribute such that classification of the other attribute of the second attribute type depends on one or a combination of the recording environment, the type of the sound source, and the operating state of the sound source. The abnormal sound detection system according to claim 1. **Claim 4** The sound source is a machine that generates the audio signal during operation, the first embedded vector indicates the type of the machine, and the second embedded vector indicates the operating state of the machine. The abnormal sound detection system according to claim 3.

5. The machine includes a motor, the type of the machine includes the type of the motor, and the operating state includes one or a combination of the characteristics of the input to the motor and the characteristics of the output of the motor. The abnormal sound detection system according to claim 4.

6. The machine is one or a combination of a drone, a robot, a numerically controlled machine, and an engine. The abnormal sound detection system according to claim 4.

7. The classification of the attributes of the first attribute type is managed by a first classifier used to train the multi-head neural network to generate a first set of the normal embedded vectors. The abnormal sound detection system according to claim 1.

8. The classification of the attributes of the second attribute type is managed by a second classifier used to train the multi-head neural network to generate a second set of the normal embedded vectors. The abnormal sound detection system according to claim 1.

9. The multi-head neural network is trained with training data including audio signals related to different values of the attributes of the first attribute type and different values of the attributes of the second attribute type. During processing of the received audio signal using the multi-head neural network, the value of the attribute of the first type of the received audio signal is known, and the first set of the normal embedded vectors includes a first subset of the embedded vectors indicating the known values of the attribute of the first type. The abnormal sound detection system according to claim 1.

10. The multi-head neural network is trained with training data including audio signals related to different values of the attributes of the first attribute type and different values of the attributes of the second attribute type. During processing of the received audio signal using the multi-head neural network, the processor receives the value of the attribute of the first type of the received audio signal, The anomaly sound detection system according to claim 1, wherein only a first subset of the embedding vectors indicating the accepted values of the attributes of the first type is configured to be arranged in a first set of the embedding vectors.

11. The comparison between the first embedding vector and the first set of the embedding vectors previously generated by the multi-head neural network, and the comparison between the second embedding vector and the second set of the embedding vectors previously generated by the multi-head neural network are performed using a nearest neighbor distance metric, and the nearest neighbor distance metric is one of a Euclidean distance, a cosine distance, and a weighted Euclidean distance. The anomaly sound detection system according to claim 1.

12. The at least one processor causes the anomaly sound detection system to determine the anomaly detection result based on a combined anomaly score for detecting an anomaly sound signal, and the anomaly sound signal is detected when the combined detection score exceeds a preset threshold. The anomaly sound detection system according to claim 1.

13. To determine the anomaly detection result, the processor compares the first embedding vector with the first set of the embedding vectors to generate a first anomaly score, compares the second embedding vector with the second set of the embedding vectors to generate a second anomaly score, is configured to determine the anomaly detection result based on a combination of the first anomaly score and the second anomaly score, and the combination of the first anomaly score and the second anomaly score generates the combined anomaly score. The anomaly sound detection system according to claim 1.

14. The combination of the first anomaly score and the second anomaly score is a weighted combination with a weight of the first anomaly score smaller than a weight of the second anomaly score. The anomaly sound detection system according to claim 13.

15. To determine the anomaly detection result, the processor combines the first embedding vector and the second embedding vector to obtain a combined embedding vector. The combined embedding vector is compared with a set of previously obtained normal combined embedding vectors by combining a first embedding vector and a second embedding vector generated by the multi-head neural network for a normal audio signal, to generate a combination anomaly score. The abnormal sound detection system according to claim 1, which is configured to determine an abnormality detection result based on the combination anomaly score.

16. To obtain the combined embedding vector by combining the first embedding vector and the second embedding vector, the processor is configured to concatenate the first embedding vector and the second embedding vector. The abnormal sound detection system according to claim 15.

17. The rendered abnormality detection result is used to predict a failure associated with at least one of the sound source and the recording environment. The abnormal sound detection system according to claim 1.

18. A method implemented by a computer for detecting abnormal sounds, comprising: Receiving an audio signal generated by a sound source in a recording environment, wherein the sound source and the recording environment are characterized by a set of attributes including a first attribute regarding a first attribute type and a second attribute regarding a second attribute type. The method further includes: Processing the received audio signal using a multi-head neural network trained to extract a first embedding vector indicating the first attribute type and a second embedding vector indicating the second attribute type from the received audio signal; Comparing the first embedding vector with a first set of embedding vectors previously generated by the multi-head neural network to classify the attributes of the first attribute type, and comparing the second embedding vector with a second set of embedding vectors previously generated by the multi-head neural network to classify the attributes of the second attribute type, and determining an abnormality detection result; And rendering the abnormality detection result. A method implemented by a computer.

19. The method realized by a computer according to claim 18, wherein the multi-head neural network includes a convolutional neural network module connected to a plurality of thin output layers including a first output layer for outputting an attribute of the first attribute type and a second output layer for outputting an attribute of the second attribute type.

20. The first attribute type includes the attribute such that classification of the attribute of the first attribute type does not depend on one or a combination of the recording environment, the type of the sound source, and the operating state of the sound source. The second attribute type includes the other attribute such that classification of the other attribute of the second attribute type depends on one or a combination of the recording environment, the type of the sound source, and the operating state of the sound source. The method realized by a computer according to claim 18.

Citation Information

Patent Citations

  • System and Method for Producing Metadata of an Audio Signal

    US20220108698A1

  • Learning device, abnormality detection device, learning method, and abnormality detection method

    WO2021214833A1