System and method for classifying noise by using multi-feature extraction

The noise classification system enhances accuracy and consistency in noisy environments by employing multi-feature extraction and a transformer model, addressing the limitations of conventional methods in reflecting temporal variations and maintaining high accuracy.

WO2026084121A1PCT designated stage Publication Date: 2026-04-23PAICHAI UNIV IND ACADEMIC COOPERATION FOUND
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
PAICHAI UNIV IND ACADEMIC COOPERATION FOUND
Filing Date
2024-11-19
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Conventional audio signal processing methods, particularly those using frequency domain analysis and machine learning algorithms like SVM or k-NN, struggle to maintain high accuracy in complex noisy environments due to sensitivity to noise and inability to reflect temporal variations effectively.

Method used

A noise classification system utilizing multi-feature extraction and a transformer model, incorporating MFCC, delta MFCC, and double delta MFCC vectors, along with a preprocessing step for noise removal and normalization, to generate integrated feature vectors for training and classification.

Benefits of technology

Improves noise classification accuracy by reflecting temporal characteristics and learning complex dependencies, enabling real-time identification and management of noise sources in diverse environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024018277_23042026_PF_FP_ABST
    Figure KR2024018277_23042026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a method for classifying noise by using multi-feature extraction, comprising: a step of loading an audio signal; a preprocessing step of removing noise from the audio signal and normalizing same; a step of extracting features of the preprocessed audio signal so as to generate feature vectors and integrating the feature vectors so as to generate an integrated feature vector; a step of training a transformer model for noise classification by using the integrated feature vector as an input; a step of evaluating the performance of the transformer model with verification data; and a step of outputting the type of noise classified using the transformer model if it is determined that the performance of the transformer model is not poor.
Need to check novelty before this filing date? Find Prior Art

Description

Noise classification system and method through multiple feature extraction

[0001] The present invention relates to a noise classification system and method through multi-feature extraction, and more specifically, to a noise classification technology utilizing multi-feature extraction and a transformer model to provide high accuracy and consistent performance in various noise environments.

[0002] Conventional audio signal processing methods have primarily relied on frequency domain analysis techniques such as the Fast Fourier Transform (FFT). While these techniques are effective for analyzing the frequency components of a signal, they fail to adequately reflect temporal variations. Frequency domain analysis methods are sensitive to noise, which can lead to performance degradation in noisy environments. Recently, there have been increasing attempts to incorporate temporal characteristics through time-frequency analysis techniques, and there is a demand for technologies that can better reflect the temporal characteristics of noise by combining these dynamic features of temporal variation.

[0003] Furthermore, existing noise classification systems have primarily utilized machine learning algorithms such as Support Vector Machines (SVM) or k-Nearest Neighbors (k-NN). However, these algorithms struggle to maintain high accuracy in complex noisy environments. To address this issue, there is a demand for machine learning algorithms that demonstrate superior performance in learning the complex dependencies of data, enabling high accuracy even in noisy environments.

[0004] The present invention aims to provide a noise classification system and method through multiple feature extraction.

[0005] As a technical means for achieving the above-mentioned task, one embodiment may provide a noise classification method through multiple feature extraction comprising: a step of loading an audio signal; a preprocessing step of removing noise from the audio signal and normalizing it; a step of extracting features of the preprocessed audio signal to generate feature vectors and integrating the feature vectors to generate an integrated feature vector; a step of training a transformer model for noise classification using the integrated feature vector as input; a step of evaluating the performance of the transformer model using validation data; and a step of outputting the type of noise classified using the transformer model if it is determined that the performance of the transformer model is not poor.

[0006] There is a technical effect of significantly improving the accuracy of noise classification by analyzing noise data through the combination of various acoustic features and learning the complex dependencies of sequence data through artificial intelligence.

[0007] By monitoring and analyzing noise in real time, it is possible to identify immediate noise sources and respond quickly. This offers the advantage of improving users' living environments and minimizing discomfort caused by noise. Accordingly, it can be utilized in various application fields, such as urban environment monitoring, smart homes, and security systems. In each field, real-time response and management are possible through noise detection and analysis.

[0008] By systematically recording and analyzing noise data, noise problems can be prevented and managed. This has the advantageous effect of reducing conflicts caused by noise and creating a pleasant living environment.

[0009] FIG. 1 is a diagram schematically illustrating a noise classification system through multiple feature extraction according to one embodiment.

[0010] FIGS. 2 and 3 are flowcharts illustrating a noise classification method through multiple feature extraction according to one embodiment.

[0011] FIG. 4 is a diagram showing an example of preprocessing an audio signal according to one embodiment.

[0012] FIG. 5 is a diagram illustrating an example of extracting and integrating features from an audio signal according to one embodiment.

[0013] FIG. 6 is a diagram showing the configuration of a transformer model according to one embodiment.

[0014] FIG. 7 is a diagram showing an example of utilizing a noise classification system according to one embodiment.

[0015] FIG. 8 is a block diagram illustrating the configuration of a noise classification system according to one embodiment.

[0016] A first aspect of the present invention provides a noise classification method through multiple feature extraction, comprising the steps of: loading an audio signal; a preprocessing step of removing noise from the audio signal and normalizing it; extracting features of the preprocessed audio signal to generate feature vectors and integrating the feature vectors to generate an integrated feature vector; training a transformer model for noise classification using the integrated feature vector as input; evaluating the performance of the transformer model using validation data; and, if it is determined that the performance of the transformer model is not poor, outputting the type of noise classified using the transformer model.

[0017] Additionally, the preprocessing step may include a step of removing noise from the audio signal, a step of padding if it is determined that the length of the audio signal is insufficient, and a step of normalizing the noise-removed signal so that the mean is 0 and the standard deviation is 1 if it is determined that the length of the audio signal is not insufficient.

[0018] In addition, MFCC vectors, delta MFCC vectors, and double delta MFCC vectors can be extracted to reflect temporal changes in the preprocessed audio signal, and MFCC vectors, delta MFCC vectors, and double delta MFCC vectors can be combined to generate an integrated feature vector.

[0019] Additionally, the transformer model may include an input layer that receives the integrated feature vector, which is a 39-dimensional vector; an embedding layer that converts the integrated feature vector into an embedding vector and maps it to 512 dimensions; a positional encoding that adds positional information to the embedding vector to provide input order information; an encoder block in which six encoding blocks are sequentially connected; and an output layer that includes a softmax activation function that classifies 10 noise classes.

[0020] Additionally, the method may include an additional step of retraining by adjusting hyperparameters if it is determined that the performance of the Transformer model is poor.

[0021] In addition, noise can be removed from an audio signal using a Wiener filter, and the Wiener filter has Equation 1, and Equation 1 is (step, is the power spectral density of the audio signal, can be the power spectral density of the noise.

[0022] A second aspect of the present invention may provide a computer-readable storage medium having a program recorded thereon for executing the method of the first aspect on a computer.

[0023] The invention will be described in detail below by way of exemplary embodiments with reference to the attached drawings. The following embodiments are intended only to embody the invention and do not limit or restrict the scope of the invention. Anything that can be easily inferred by a person skilled in the art to which the invention pertains from the detailed description and embodiments shall be interpreted as falling within the scope of the invention.

[0024] Terms such as 'composed' or 'comprising' as used in this specification should not be interpreted as necessarily including all of the various components or steps, and should be interpreted as some of the components or steps may not be included, or additional components or steps may be included.

[0025] The terms used in this specification are described in their current, general form considering the functions mentioned herein; however, they may refer to various other terms depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Therefore, the terms used in this specification should not be interpreted solely by their names, but should be interpreted based on the meaning of the terms and the overall content of this specification. Furthermore, singular expressions include a plural meaning unless the context clearly indicates a singular meaning.

[0026] The terms “above” and similar designations used in this specification (particularly in the claims) may indicate both singular and plural forms. Furthermore, unless there is a description explicitly specifying the order of the steps describing the method according to this specification, the described steps may be performed in a suitable order. The present invention is not limited by the order in which the described steps are described.

[0027] Phrases such as "in one embodiment" appearing in various places in this specification do not necessarily refer to the same embodiment.

[0028] Some embodiments of this specification may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various numbers of hardware and / or software configurations that execute specific functions.

[0029] The embodiments relate to a noise classification system and method through multiple feature extraction, and detailed descriptions of matters widely known to those skilled in the art to which the following embodiments belong are omitted. The present invention will be described in detail below with reference to the attached drawings.

[0030] FIG. 1 is a diagram schematically illustrating a noise classification system through multiple feature extraction according to one embodiment. A noise classification system (100) according to one embodiment may include a data collection device (110), a computing device (120), a user mobile device (130), and an AI model storage device (140).

[0031] The data collection device (110) can collect audio data. For example, the audio data collected may be audio signals related to noise. The data collection device (110) can generate feature vectors through preprocessing after filtering out invalid audio files. The computing device (120) can train and verify an artificial intelligence model for noise classification. The artificial intelligence model can be trained using the integrated feature vectors received from the data collection device (110). The performance of the trained model can be verified, and if the performance is poor, it can be retrained. The trained model can then be stored in the AI ​​model storage device (140), and the AI ​​model storage device can transmit the AI ​​model to the user's mobile device (130). For example, the trained or retrained AI model can be stored for real-time prediction.

[0032] Accordingly, the user mobile device (130) can receive audio data in real time from the data collection device (110) and predict the type of noise through an AI model. The user mobile device (130) can immediately provide the prediction result to the user.

[0033] In this specification, the learning and prediction performance of noise classification can be maximized by utilizing a Transformer model as an AI model. Through the Transformer model's multi-head self-attention mechanism, complex dependencies in sequence data can be learned, and the characteristics of noise signals can be effectively learned. This provides a high-performance noise classification model and enables real-time prediction.

[0034] This noise classification technology can play a crucial role in various application fields, such as urban environment monitoring, smart homes, and security systems. In urban environment monitoring, it offers the advantage of detecting and responding to traffic and construction noise in real time. In smart home systems, it can analyze appliance noise to detect abnormal sounds and warn users. In security systems, it can detect abnormal noise in real time to identify intrusion attempts. In other words, because it detects and analyzes noise in real time to identify its source, it helps users recognize problems early and assists them in resolving them by collaborating with management offices or neighbors.

[0035] Such noise classification technology requires real-time processing and high accuracy, and this invention aims to meet these requirements by utilizing multi-feature extraction and transformer models. This enables high classification accuracy and consistent performance even in diverse noise environments, while also allowing for real-time data processing and analysis.

[0036] Existing frequency domain analysis methods had the problem of not being able to sufficiently reflect temporal changes. The present invention combines MFCC, delta MFCC (Δ-MFCC), and double delta MFCC (ΔΔ-MFCC) to effectively reflect the temporal characteristics of noise signals. Through this, dynamic changes occurring in various noise environments can be captured, and more sophisticated noise classification becomes possible.

[0037] Furthermore, conventional noise classification systems have struggled to maintain high classification accuracy in complex noise environments. This specification utilizes multi-feature extraction and a transformer model to improve noise classification accuracy and provide consistent performance in various environments, thereby maximizing the analysis and prediction capabilities of noise data.

[0038] Furthermore, to address the issue of signal quality degradation in noisy environments, signal quality can be improved through noise removal and standard normalization. This enables the preprocessed signal to provide higher accuracy during the feature extraction and model training processes.

[0039] FIGS. 2 and 3 are flowcharts illustrating a noise classification method through multiple feature extraction according to one embodiment.

[0040] In step S210, an audio signal can be loaded. In one embodiment, an analog audio signal input through a microphone and an audio interface can be collected. The analog signal can then be converted into a digital signal and saved as an audio file format (e.g., WAV and MP3 files). An audio signal detected in real time by the microphone can be acquired, or a pre-recorded audio file can be loaded. For example, a user can install a microphone near the ceiling of a house to detect and collect noise originating from an upstairs floor.

[0041] As shown in FIG. 3, when loading an audio signal, the validity of the data can be checked, and if it is not valid, an error message can be output and the audio signal loading can be stopped.

[0042] In step S220, preprocessing can be performed to remove noise from the audio signal and normalize it.

[0043] As illustrated in FIG. 3, the preprocessing process according to one embodiment removes noise from an audio signal using a Wiener filter, and if it is determined that the length of the audio signal is insufficient, padding can be applied. Additionally, if it is determined that the length of the audio signal is not insufficient, the noise-removed signal can be normalized so that the mean is 0 and the standard deviation is 1. However, if the intensity of the audio signal is too weak, noise removal may be omitted.

[0044] For an example of preprocessing, we will look into it in more detail by referring to Fig. 4.

[0045] In step S230, features of the preprocessed audio signal can be extracted to generate feature vectors, and the feature vectors can be integrated to generate an integrated feature vector. In one embodiment, an MFCC vector, a delta MFCC vector, and a double delta MFCC vector can be extracted to reflect temporal changes in the preprocessed audio signal, and an integrated feature vector can be generated by combining the MFCC vector, the delta MFCC vector, and the double delta MFCC vector. By extracting these features, important information, dynamic characteristics of the signal, and acceleration can be reflected.

[0046] In step S240, a transformer model for noise classification can be trained using the integrated feature vector as input.

[0047] As illustrated in FIG. 3, in one embodiment, class imbalance in the training data can be checked before inputting the integrated feature vector into the transformer model. If there is too much or too little data of a specific class, the transformer model may be trained in a skewed manner; therefore, if the imbalance is severe, it can be balanced through under / oversampling.

[0048] In one embodiment, the transformer model may include an input layer that receives the integrated feature vector, which is a 39-dimensional vector; an embedding layer that converts the integrated feature vector into an embedding vector and maps it to 512 dimensions; a positional encoding that adds positional information to the embedding vector to provide input order information; an encoder block in which six encoding blocks are sequentially connected; and an output layer that includes a softmax activation function that classifies 10 noise classes. An example of the configuration of the transformer model will be examined with reference to FIG. 6.

[0049] In step S250, the performance of the Transformer model can be evaluated using validation data.

[0050] In one embodiment, the performance of the transformer model is evaluated using validation data, and if it is determined that the performance of the transformer model is poor, the hyperparameters can be adjusted to retrain the model.

[0051] In step S260, if it is determined that the performance of the transformer model is not poor, the types of noise classified using the transformer model can be output. At this time, the results can be fed back to the user, and the types of noise classified can be provided through a user interface. For example, the noise occurrence time, type, and intensity can be displayed in graph and text formats. Additionally, if the noise exceeds a certain level, a real-time warning message can be sent to the user.

[0052] FIG. 4 is a diagram illustrating an example of preprocessing an audio signal according to one embodiment. For example, noise can be removed from an audio signal using a Wiener filter. The Wiener filter is a minimum mean squared error technique and is a method that minimizes the difference between the original speech signal and the filtered signal. The Wiener filter maximizes the signal-to-noise ratio (SNR) while remaining faithful to the original waveform. It generates an optimal filter using the power spectrum of the audio signal x(t) and the noise n(t), and the recognition rate is higher than when using the spectrum subtraction method.

[0053] The following formula is used to remove noise.

[0054]

[0055] In mathematical formula 1 is the power spectral density of the audio signal, and is the power spectral density of the noise.

[0056] Also, y(t) is obtained by applying a Wiener filter to the audio signal x(t).

[0057]

[0058] In mathematical equation 2, F is the Fourier transform and is the inverse Fourier transform. After calculating the Fourier transform of a filtered signal, it can be converted back to the time domain by performing the inverse Fourier transform.

[0059] Preprocessing is divided into two types: noise removal and normalization. Normalization is a method that maintains signal consistency by converting the mean to 0 and the standard deviation to 1. This can improve the performance of model training.

[0060] The normalized signal is calculated as follows.

[0061]

[0062] In mathematical formula 3 is a normalized signal, and is the average of the signal y(t), and is the standard deviation of the signal y(t).

[0063] Looking at Figure 4, the original signal waveforms of street music, a dog barking sound, and an air conditioner sound are shown from top to bottom on the left, and the preprocessed signal waveforms of street music, a dog barking sound, and an air conditioner sound are shown from top to bottom on the right. As shown in Figure 4, there is little change in the shape of the waveforms before and after preprocessing. However, it can be seen that the numerical range has increased after preprocessing. Part of the data changes before and after preprocessing are displayed below the graph.

[0064] FIG. 5 is a diagram illustrating an example of extracting and integrating features from an audio signal according to one embodiment. The feature extraction and combination process will be visually explained with reference to FIG. 5.

[0065] First, MFCC is extracted from the normalized audio signal. MFCC effectively analyzes the frequency components of the audio signal to match human hearing, allowing it to express the timbre and characteristics of the noise. To achieve this, the audio signal is divided into short frames, and a Fourier transform is performed on each frame to obtain the frequency spectrum S(f). Then, a Mel filter bank is applied to the frequency spectrum to obtain the Mel spectrum. This is done by applying the following formula.

[0066]

[0067] As shown in Equation 4, the Mel filter bank converts frequency components into a Mel scale, applies weights, and then sums them.

[0068] Then, to reflect human auditory characteristics, each value of the Mel spectrum is converted to a logarithm. Next, an inverse Fourier transform is performed on the log Mel spectrum to obtain the MFCC.

[0069]

[0070] Next, the delta MFCC (Δ-MFCC) is extracted. The delta MFCC can be extracted by calculating the first-order rate of change of the MFCC. This provides dynamic features that reflect the temporal change of the MFCC. In other words, it can represent the change pattern of the sound. The following formula is used to extract the delta MFCC.

[0071]

[0072] In mathematical formula 6 is the MFCC coefficient at time t, and is the Δ-MFCC coefficient at time t. N is the number of frames to calculate the difference.

[0073] Subsequently, the double delta MFCC (ΔΔ-MFCC) is extracted. As the time rate of change of the delta MFCC, the double delta MFCC can reflect more precise temporal changes. Through this, the acceleration of the signal can be calculated to represent more complex dynamic information. The following formula is used to extract the delta MFCC.

[0074]

[0075] In mathematical formula 7 is the Δ-MFCC coefficient at time t, and ΔΔ-MFCC coefficients at time t. N is the number of frames to calculate the difference.

[0076] Then, three feature vectors (MFCC, Δ-MFCC, and ΔΔ-MFCC) are combined into a single vector to generate a feature vector containing rich information. When the three feature vectors are combined, the combined feature provides richer information to the model, enabling better predictive performance. Figure 5 shows an example of the MFCC feature vector, delta MFCC feature vector, double delta MFCC feature vector, and the combined feature vector represented as a Mel spectrogram.

[0077] FIG. 6 is a diagram showing the configuration of a transformer model according to one embodiment.

[0078] In one embodiment, the transformer model may include an input layer that receives an integrated feature vector which is a 39-dimensional vector, an embedding layer that converts the integrated feature vector into an embedding vector and maps it to 512 dimensions, a positional encoding that adds positional information to the embedding vector to provide input order information, an encoder block in which six encoding blocks are sequentially connected, and an output layer that includes a softmax activation function that classifies 10 noise classes.

[0079] The input layer can receive an integrated feature vector (e.g., a combination of an MFCC vector, a delta MFCC vector, and a double delta MFCC vector) as input data for the transformer model. In this case, the integrated feature vector can be a 39-dimensional vector.

[0080] The embedding layer can convert the integrated feature vector into a high-dimensional embedding vector so that the Transformer model can understand it. To do this, the input data can be mapped to 512 dimensions.

[0081] Positional encoding can preserve input order information by adding positional information to embedding vectors. It is a process of adding positional information to embeddings to provide information about the input order; in other words, it is an embedding vector in which positional information has been added to each vector.

[0082] Each encoding block of an encoder block having 6 encoding blocks consists of Multi-Head Attention, Feed Forward Network, layer normalization, and dropout, and can be connected sequentially.

[0083] Multi-head attention can learn the relationships between input data through multiple attention mechanisms. Seven attention heads each process a 63-dimensional vector, allowing data to be processed at various points in time. It can extract important features by learning complex interactions between inputs.

[0084] Feedforward networks can perform non-linear transformations by passing the output of each attention through two dense layers. This can improve the network's expressiveness by learning more complex features. The first layer has a ReLu activation function and can be extended to 2048 dimensions.

[0085] Add & Layer Normalization can normalize the result of an attention function by adding the input to it. Additionally, the output of a feedforward network can be added to the input once again and normalized. This helps maintain the stability of the audio signal and training efficiency, while also aiding in network stability and convergence.

[0086] The output layer is a softmax layer for classification that can predict various noise classes. The output layer receives the output of the encoder and ultimately predicts the noise class; for example, it may include a Dense layer containing a softmax activation function that classifies 10 noise classes.

[0087] The reason Transformer models contain only encoders and no decoders is that they perform classification tasks rather than sequence generation. Decoders are generally used to generate new sequences based on input sequences. In other words, the structure processes input data using only encoder blocks and classifies the results directly in the output layer. This simplifies the model and allows it to be optimized for classification tasks.

[0088] Transformer models are models in which information from all locations is directly connected, enabling them to efficiently learn very long sequences of data. For example, CRNNs are effective at processing spatial and temporal information, while LSTMs are strong at processing temporal sequence data. CNNs are strong at processing visual information such as spectrograms, but are relatively weak at processing information regarding temporal order. RNNs exhibit different performance depending on the characteristics of the data or the problem.

[0089] FIG. 7 is a diagram showing an example of utilizing a noise classification system according to one embodiment.

[0090] As illustrated in Fig. 7, User A, who resides in an apartment and experiences discomfort due to noise from the upstairs neighbor, can install a noise classification system to resolve this issue. Specifically, a microphone can be installed near the ceiling to measure the noise. The noise classification system can monitor noise from the upstairs neighbor in real time, analyze the noise type and intensity, and send notifications to User A's mobile device. Based on the system's analysis, it may be predicted that footsteps and furniture moving noises occurred late at night. User A can check the time, type, and intensity of the noise through a web interface and, based on this information, cooperate with the management office to raise the issue with the upstairs neighbor. Through this process, the upstairs neighbor installs a noise-reducing mat, and User A is able to resolve the sleep disturbance problem.

[0091] In another embodiment, User B resides in a house equipped with a smart home system. One day, User B wants to check if the refrigerator is operating normally. A noise classification system embedded in the smart home system can monitor noise generated by the refrigerator in real time. The noise classification system can analyze the noise emitted by the refrigerator to determine whether it is normal operating noise or abnormal noise. If abnormal noise is detected, the noise classification system can send a notification to User B's mobile device. User B can check the results of the noise analysis through a web interface and schedule a refrigerator repair if necessary. This offers User B the advantage of detecting abnormal appliance operation early and resolving the problem quickly.

[0092] In another embodiment, User C, a city manager, may use a noise classification system to manage noise pollution within the city. The noise classification system can collect noise in real time through microphones installed in various locations, such as major roads, construction sites, and parks. The system can analyze the collected noise to classify various noise types, such as traffic noise, construction noise, or crowd noise. For example, if traffic noise in a specific area exceeds a certain threshold, the noise classification system can send a warning notification to User C's mobile device. User C can check the type and location of the noise through a web interface and take prompt response measures to alleviate traffic congestion. This provides the city manager with the advantageous effect of effectively managing noise pollution and improving the living environment of residents.

[0093] As mentioned above, noise problems can be prevented and managed by systematically recording and analyzing noise data using a noise classification system. For example, analyzing the frequency and patterns of noise occurrence can be used to implement noise reduction measures or serve as mediation data for resolving noise issues. This offers the advantage of reducing conflicts caused by noise and creating a pleasant living environment.

[0094] FIG. 8 is a block diagram illustrating the configuration of a noise classification system according to one embodiment. As illustrated in FIG. 8, a noise classification system (100) according to one embodiment may include a data collection device (110), a computing device (120), a user mobile device (130), and an AI model storage device (140). However, not all illustrated components are essential components. The system (100) may be implemented with more components than illustrated, or with fewer components. The above components will be examined in turn below.

[0095] The data collection device (110) can collect audio data through the microphone sensor module (111). Then, the analog signal can be converted into a digital signal and the digital signal can be saved in an audio file format (e.g., WAV and MP3 files). At this time, invalid audio files can be filtered out, and then a feature vector can be generated through preprocessing.

[0096] A feature vector can be generated by preprocessing a pre-recorded audio file through a pre-data preprocessing and conversion module (113), or by preprocessing an audio signal detected in real time by a microphone sensor module (111) through a real-time data preprocessing and conversion module (115).

[0097] In the computing device (120), a transformer model for noise classification can be trained and validated. First, the integrated feature vector received from the data collection device (110) can be labeled in the data labeling module (121) to train the transformer model in the model training module (123). In the model evaluation module (125), the performance of the trained transformer model can be evaluated using validation data. If the performance of the trained model is poor after validation, it can be retrained in the model training module (123). Then, the trained model can be stored in the model storage and load module (127) and transmitted to the AI ​​model storage device (140).

[0098] The model load module (131) of the user mobile device (130) can load a transformer model received from the AI ​​model storage device (140), and the prediction module (133) can predict the type of noise by inputting real-time audio data received from the data collection device (110) into the transformer model. The user notification module (135) can immediately provide the prediction result to the user.

[0099] Meanwhile, the embodiments of the present invention described above can be written as a program executable on a computer and can be implemented on a general-purpose digital computer that operates the program using a computer-readable recording medium.

[0100] The above computer-readable recording media includes storage media such as magnetic storage media (e.g., ROM, floppy disk, hard disk, etc.) and optical reading media (e.g., CD-ROM, DVD, etc.).

[0101] Although embodiments of the present invention have been described above with reference to the attached drawings, those skilled in the art will understand that the present invention may be implemented in other specific forms without changing its technical concept or essential features. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive.

Claims

1. In a noise classification method through multiple feature extraction, Step of loading the audio signal; A preprocessing step for removing noise from and normalizing the above audio signal; A step of extracting features of the preprocessed audio signal to generate feature vectors, and integrating the feature vectors to generate an integrated feature vector; A step of training a transformer model for noise classification using the above integrated feature vector as input; A step of evaluating the performance of the above transformer model using verification data; and If it is determined that the performance of the above transformer model is not poor, a step of outputting the type of noise classified using the above transformer model; A method including 2. In Paragraph 1, The above preprocessing step is, A step of removing noise from the above audio signal; If it is determined that the length of the above audio signal is insufficient, a step of padding; and If it is determined that the length of the audio signal is not insufficient, a step of normalizing the noise-removed signal so that the mean is 0 and the standard deviation is 1; A method including 3. In Paragraph 1, The step of generating the above integrated feature vector is, A method comprising the step of extracting an MFCC vector, a delta MFCC vector, and a double delta MFCC vector to reflect temporal changes of the preprocessed audio signal, and combining the MFCC vector, the delta MFCC vector, and the double delta MFCC vector to generate an integrated feature vector.

4. In Paragraph 1, The above Transformer model is, An input layer that receives the above-mentioned integrated feature vector, which is a 39-dimensional vector; An embedding layer that converts the above integrated feature vector into an embedding vector and maps it to 512 dimensions; Positional encoding that provides input order information by adding position information to the above embedding vector; An encoder block in which six encoding blocks are connected sequentially; and An output layer containing a softmax activation function that classifies 10 noise classes; A method including 5. In Paragraph 1, A method comprising an additional step of retraining by adjusting hyperparameters if it is determined that the performance of the above-mentioned transformer model is poor.

6. In Paragraph 2, The step of removing the above noise is, The method includes the step of removing noise from the above audio signal using a Wiener filter, and The above Wiener filter has Equation 1, The above mathematical formula 1 is (step, is the power spectral density of the audio signal, is the power spectral density of the noise, method.

7. A computer-readable storage medium having a program recorded thereon for executing the method of any one of paragraphs 1 through 6 on a computer.