Coal gangue identification method based on audio feature fusion
By combining the Mel-frequency cepstrum and gammatone frequency cepstrum features and simulating the auditory characteristics of the human ear, the problems of low efficiency and insufficient accuracy in coal gangue detection in existing technologies are solved, higher recognition accuracy and robustness are achieved, adaptation to complex acoustic environments is achieved, and the false recognition rate in noisy environments is reduced.
Patent Information
- Application Number
- CN202510783054.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-12
AI Technical Summary
Existing gangue detection methods rely on manual detection or image recognition, which have the disadvantages of low detection efficiency, large errors, and great environmental influence. In addition, MFCC features have poor robustness and limited accuracy in complex environments, making it difficult to accurately capture subtle changes in coal gangue audio signals.
A method based on audio feature fusion is adopted, combining Mel frequency cepstrum and gammatone frequency cepstrum features, extracting audio features through Mel filter banks and gammatone filter banks, and combining with deep learning network for coal gangue recognition. The advantages of Mel frequency cepstrum and gammatone frequency cepstrum are utilized to simulate the auditory characteristics of the human ear and improve recognition accuracy and robustness.
It significantly improves the accuracy and robustness of coal gangue audio recognition, reduces the misrecognition rate in noisy environments, can better capture the steady-state and transient characteristics of coal gangue audio, and improves the recognition accuracy and stability of the system in complex environments.
Smart Images

Figure CN120636458A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of coal mine safety monitoring, and in particular to a coal gangue audio recognition method based on audio feature fusion. Background Art
[0002] Gangue is waste rock generated during the mining process. Its presence can cause gas and coal dust explosions, leading to coal mine accidents and posing a potential threat to mine safety. Therefore, timely detection and identification of the presence of gangue is crucial for mine safety.
[0003] Traditional coal gangue detection methods mainly rely on manual detection or image recognition technology, which has problems such as low detection efficiency, large errors, and great environmental impact.
[0004] Audio recognition, as a new detection method, has been widely researched in recent years. As mentioned earlier, in coal mines, the sounds produced by coal gangue and other materials in the mine differ during mining operations. These differences can be detected by the human ear and can be used for audio recognition. By collecting the audio signals generated by coal gangue during mining operations and analyzing them using audio signal analysis techniques, efficient coal gangue detection can be achieved. Existing audio recognition methods often use Mel-Frequency Cepstrum Coefficient (MFCC) features for coal gangue audio recognition. MFCCs are audio features based on the human auditory perception of sounds of different frequencies and perform well in low-noise environments. However, MFCC features have poor robustness, making them sensitive to ambient noise and less adaptable to non-stationary audio signals. Furthermore, the limited resolution of MFCC features makes it difficult to accurately capture subtle spectral variations in coal gangue audio signals. In complex environments, MFCC features suffer from poor robustness, limited accuracy, and information loss, leaving room for improvement in recognition accuracy. Summary of the Invention
[0005] In order to improve the accuracy and robustness of coal gangue audio recognition, the present invention proposes a coal gangue audio recognition method based on audio feature fusion, which includes the following steps:
[0006] Collect coal gangue audio signals;
[0007] Preprocessing the coal gangue audio signal to obtain a frequency spectrum of the audio signal;
[0008] Using a Mel filter bank to extract features from the spectrogram to obtain a Mel frequency cepstrum;
[0009] Using a gammatone filter bank to extract features from the spectrum to obtain a gammatone frequency cepstrum;
[0010] Based on the Mel frequency cepstrum and the gammatone frequency cepstrum, a classification network is used to obtain a coal gangue recognition result.
[0011] In some embodiments, the preprocessing of the coal gangue audio signal specifically includes:
[0012] The audio signal is divided into a plurality of overlapping short time frames, and each frame of the audio signal is windowed; each frame of the windowed audio signal is Fourier transformed to obtain a spectrum of each frame of the audio signal; and a spectrum diagram is drawn according to the spectrum of each frame of the audio signal.
[0013] In some embodiments, extracting features from the spectrogram using a Mel filter bank to obtain a Mel-frequency cepstrum specifically includes:
[0014] Calculate a power spectrum based on the spectrum graph;
[0015] Filtering the power spectrum using the Mel filter bank to obtain the short-time energy output by each Mel filter;
[0016] Take the logarithm of the short-time energy output by each Mel filter and arrange the results into a two-dimensional matrix in time order to obtain the logarithmic Mel spectrum;
[0017] Perform discrete cosine transform on the logarithmic Mel spectrum to obtain Mel frequency cepstral coefficients;
[0018] The Mel-frequency cepstrum coefficients of each frame signal are arranged into a two-dimensional matrix in time sequence to obtain the Mel-frequency cepstrum.
[0019] In some more specific embodiments, the Mel filter bank is pre-constructed according to a Mel frequency scale, and the Mel filter bank is a triangular filter bank; in the Mel frequency scale, the relationship between the Mel frequency M(f) and the normal frequency f is: M(f)=1125ln(1+f / 700).
[0020] In some embodiments, extracting features from the spectrogram using a gammatone filter bank to obtain a gammatone frequency cepstrum specifically includes:
[0021] filtering the spectrogram using a gammatone filter bank to obtain short-time energy output by each gammatone filter; taking the logarithm of the short-time energy output by each gammatone filter, and arranging the results into a two-dimensional matrix in time order to obtain a logarithmic gammatone spectrum;
[0022] Perform discrete cosine transform on the logarithmic gammatone spectrum to obtain the gammatone frequency cepstral coefficients;
[0023] The gamma-tone frequency cepstrum coefficients of each frame signal are arranged in time sequence into a two-dimensional matrix to obtain a gamma-tone frequency cepstrum.
[0024] In some embodiments, the classification network is one of a support vector machine, a deep neural network, and a random forest.
[0025] The coal gangue audio recognition method provided by the present invention combines two audio features, the Mel frequency cepstrum and the gammatone frequency cepstrum, and fully utilizes the advantages of these two audio features in simulating the auditory characteristics of the human ear, thereby significantly improving the accuracy and robustness of coal gangue audio recognition. Combining MFC and GFC can make the auditory perception characteristics of the two complement each other, improve adaptability to the complex acoustic environment under coal mines, improve robustness, and significantly reduce the error recognition rate in noisy environments. Combining MFC and GFC can better capture the wideband characteristics of coal gangue audio, which is conducive to simultaneously capturing the steady-state and transient characteristics of coal gangue audio, thereby significantly improving the accuracy of coal gangue audio recognition.
[0026] In addition, the present invention uses cepstral images instead of cepstral coefficients as the input of the classification network, making full use of the two-dimensional spatiotemporal relationship between the coefficients and reducing information loss. The method also uses the deep learning network's powerful information mining capabilities for graphic data to process cepstral images, thereby improving the accuracy of coal gangue audio recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 A technical roadmap for the coal gangue audio recognition method provided by the present invention;
[0029] Figure 2 A frequency response diagram of a Mel filter bank provided by an embodiment of the present invention;
[0030] Figure 3 A time-frequency diagram after being filtered by a Mel filter provided by an embodiment of the present invention;
[0031] Figure 4 A Mel-frequency cepstrum provided by an embodiment of the present invention;
[0032] Figure 5 A frequency response diagram of a gammatone filter bank provided by an embodiment of the present invention;
[0033] Figure 6A gammatone frequency cepstrum diagram is provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0034] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0035] In order to improve the accuracy and robustness of coal gangue audio recognition, the present invention proposes a coal gangue audio recognition method based on audio feature fusion. The method fuses two audio features, Mel-Frequency Cepstral (MFC) and Gammatone Frequency Cepstral (GFC), to fully utilize the advantages of these two audio features in simulating the auditory characteristics of the human ear and improve the coal gangue audio recognition effect.
[0036] The gamma-tone frequency cepstrum is also an audio feature based on the human auditory characteristics. It uses a gamma-tone filter bank to simulate the human auditory characteristics. Compared with MFC, GFC has higher recognition accuracy, can more sensitively capture subtle changes in the audio signal's spectrum, and has higher noise robustness. Combining GFC and MFC features can overcome the poor robustness and limited precision of existing MFCC features, improving the accuracy and robustness of coal gangue audio recognition. In addition, since both GFC and MFC simulate the spectral characteristics of the human ear, fusing GFC and MFC features can fully utilize the advantages of both, better approximating the human ear's auditory perception, enhancing the recognition ability of coal gangue audio signals, and improving recognition accuracy.
[0037] The coal gangue audio recognition method provided by the present invention is described in detail below with reference to the accompanying drawings.
[0038] The technical route of the coal gangue audio recognition method provided by the present invention is as follows Figure 1 As shown, the method is divided into the following four stages:
[0039] Audio signal acquisition stage: collecting the coal gangue audio signal generated at the tail beam of the hydraulic support;
[0040] Preprocessing stage: preprocess the collected audio signal;
[0041] Feature extraction stage: MFC and GFC feature extraction are performed on the preprocessed audio signal to obtain Mel frequency cepstrum and gammatone frequency cepstrum;
[0042] Classification and recognition stage: The Mel frequency cepstrum and gammatone frequency cepstrum are input into the classification network to obtain the coal gangue recognition results.
[0043] The specific implementation process of each stage of the coal gangue audio recognition method provided by the present invention is described in detail below with reference to the accompanying drawings.
[0044] S1, audio signal acquisition stage
[0045] This stage is used to collect coal gangue audio signals.
[0046] During top coal caving, a high-sensitivity microphone array can be placed near the tail beam of the hydraulic support in a fully mechanized coal mining face to collect the audio signals generated by the gangue at the tail beam. The gangue directly impacts the tail beam during caving, and the audio signals generated here can more directly reflect the characteristics of the gangue. Therefore, the tail beam is selected as the gangue audio signal collection point.
[0047] The audio acquisition system should have high anti-interference and real-time performance to ensure timely acquisition of real and effective coal gangue audio signals, so as to perform real-time coal gangue audio recognition and link the coal gangue recognition results with relevant alarm or control mechanisms in real time to ensure safe operation of mines, improve the level of coal mine safety management, and reduce the incidence of coal mine accidents.
[0048] S2, preprocessing stage
[0049] This stage is used to preprocess the collected audio signals to provide a basis for subsequent feature extraction. The preprocessing stage mainly includes the following steps:
[0050] S21, framing: dividing the audio signal collected in S1 into multiple overlapping short time frames, each frame containing 20 to 40 milliseconds of data, and the overlap between adjacent frames is about 50%.
[0051] S22, windowing: applying a window function (such as a Hamming window) to each frame signal to reduce spectrum leakage.
[0052] S23. Short-time Fourier transform: performing short-time Fourier transform on each frame of the windowed signal to obtain a frequency spectrum of each frame of the audio signal; and drawing a spectrum graph based on the frequency spectrum of each frame of the audio signal.
[0053] The short-time Fourier transform converts a signal from the time domain to the frequency domain, representing it as a two-dimensional array of frequency and time. The short-time Fourier transform is shown in formula (1).
[0054]
[0055] Among them, s i (n) is the time domain signal of the data of the i-th frame, k is the discrete frequency point, n is the discrete time point, the range of n is 1-400, N is the number of Fourier transform points, h(n) is the window function, S i(k) represents the complex coefficient (spectrum) at frequency k of the i-th frame signal; e -j2πkn / N is a complex exponential function, where j is the imaginary unit.
[0056] The above introduces the preprocessing process of audio signals. The following introduces the MFC and GFC feature extraction processes respectively.
[0057] S3, feature extraction stage
[0058] S31, MFC feature extraction
[0059] S311, calculate power spectrum
[0060] To extract MFC features, we first need to calculate the power spectrum based on the spectrogram of the audio signal. To calculate the power spectrum, we need to square the amplitude of each frequency component in the spectrogram to obtain the energy of the frequency component, as shown in formula (2).
[0061]
[0062] Among them, P i (k) is the power spectrum at frequency k in the i-th frame.
[0063] The power spectrum represents the power distribution of a signal at different frequencies. For a time domain signal, its power spectrum is a representation of the power of the signal at each frequency component.
[0064] S312, constructing a Mel filter bank
[0065] This step requires constructing a set of triangular filters according to the Mel frequency scale as a Mel filter bank, which is used to perform weighted averaging on the spectrum and map the spectrum to the Mel frequency scale.
[0066] The human ear has a higher resolution for low-frequency sounds than for high-frequency sounds. The Mel frequency scale maps linear frequencies to a scale that is more consistent with human ear perception through nonlinear transformation, which can better simulate the human ear's perception of frequency.
[0067] Exemplarily, 26 Mel filters are selected to form a Mel filter bank, and the relationship between the Mel frequency M(f) and the normal frequency f is shown in formula (3).
[0068] M(f)=1125ln(1+f / 700) (3)
[0069] To construct a Mel filter bank, we first need to convert the normal frequency into Mel frequency according to formula (3); then we need to take 26 Mel frequency points at equal intervals on the Mel frequency scale as the center frequency position of each Mel filter.
[0070] The frequency response of the Mel filter bank is shown in Equation (4).
[0071]
[0072] In the formula Where H_m(k) is the frequency response of the mth filter, f(m) is the center frequency of the mth filter, f(m+1) and f(m-1) are the edge frequencies of the mth filter, and M is the number of triangular filters.
[0073] The frequency response of the Mel filter bank is as follows Figure 2 As shown in the figure, each solid line represents the frequency response of a Mel filter in the filter bank. The filter frequency response is triangular on the frequency axis, with the highest response at the center frequency and linearly decreasing to zero toward either side. The filters in the filter bank are densely distributed at low frequencies and relatively sparse at high frequencies to simulate the human ear's frequency perception. The frequency ranges of adjacent filters overlap by 50%. This filter bank performs a weighted average of the spectrum, with the highest weight at the center frequency of each filter and decreasing linearly toward the edges.
[0074] The filter bank consists of 26 filters, each of which is composed of a vector of length 257, where only the vector values in the frequency range to be collected are non-zero. The input signal of the filter bank is a 257-point power spectrum. After the power spectrum is filtered by 26 filters, the energy of each Mel filter output needs to be calculated separately.
[0075] S313, filter the power spectrum using a Mel filter bank
[0076] The energy value of each frequency point in the power spectrum is multiplied by the response value of each filter at that frequency point, and each Mel filter obtains a set of Mel filter signals; for each filter, the energy values corresponding to all frequency points in the Mel filter signal are added together to obtain the short-time energy of each Mel filter output signal. This process is the process of weighted summation of the spectrum within each filter range. The short-time energy output by these Mel filters is filled into the frequency band corresponding to the filter, and the following can be obtained: Figure 3 The time-frequency diagram shown, where the vertical axis represents the frequency based on the Mel frequency scale, the horizontal axis represents the time, and the color represents the signal energy at that frequency and time point. The brighter the color, the stronger the signal energy at that frequency and time point.
[0077] S314, take logarithm
[0078] The logarithm of the short-time energy output by each Mel filter is taken to obtain the logarithmic energy of each frame; the logarithmic energy of each frame is arranged into a two-dimensional matrix in time sequence and visualized to obtain the logarithmic Mel spectrum.
[0079] Taking the logarithm of the output of the Mel filter can compress the dynamic range of the spectrum, enhance the low-frequency details of the spectrum, and better simulate the human ear's perception of sound intensity.
[0080] S315, Discrete Cosine Transform
[0081] Perform discrete cosine transform (DCT) on the logarithmic Mel spectrum to obtain the Mel frequency cepstrum. The formula of discrete cosine transform is shown in formula (5).
[0082]
[0083] Where L is the order of the MFCC coefficients, that is, the number of MFCC coefficients obtained by discrete cosine transform. C(n) is the transformed Mel-frequency cepstral coefficient, and s(m) is the logarithmic energy value in the log Mel-frequency spectrum.
[0084] Perform DCT on the logarithmic energy value of each frame in the logarithmic Mel spectrum obtained in step S314 to obtain L Mel-frequency cepstral coefficients of each frame; arrange the Mel-frequency cepstral coefficients of each frame in chronological order to form a time × coefficient two-dimensional matrix, and visualize the matrix using a heat map, with the horizontal axis being time and the vertical axis being the cepstral coefficient number (1-L), to obtain the following: Figure 4 Mel-frequency cepstrum shown.
[0085] The above introduces the specific steps of Mel feature extraction. Through the above steps, we can get Figure 4 The Mel-frequency cepstrum is shown in Figure 2. Next, we will introduce the specific process of gamma-tone feature extraction.
[0086] S32, Gammatone feature extraction
[0087] S321, constructing a gammatone filter bank
[0088] A gammatone filter bank consisting of 32 gammatone filters is selected. The time domain expression of the gammatone filter is shown in formula (6):
[0089] g(t)=at n-1 e -2πbt cos(2πf c t+φ0) (6)
[0090] Where fc represents the center frequency of the gammatone filter, φ0 represents the initial phase, a is the filter amplitude, n represents the filter order, which is related to the filter shape, b represents the filter bandwidth, b = 3dB, and t represents time.
[0091] Among them, the center frequency fc of the gammatone filter is evenly distributed according to the equivalent rectangular bandwidth ERB scale to accurately simulate the frequency response characteristics of the basilar membrane of the human ear.
[0092] The frequency response of the gammatone filter bank is as follows Figure 5 As shown in Figure 2, the gammatone filter has the highest response at the center frequency and gradually decreases to zero in a cosine curve toward either side. The gammatone filter bank has high resolution in the low-frequency range and low resolution in the high-frequency range, which is consistent with the frequency resolution characteristics of the human ear. The gammatone filter bank effectively simulates the frequency response characteristics of the basilar membrane of the human ear, thereby extracting audio features relevant to human ear perception.
[0093] S322. Filter the spectrum using a gammatone filter bank
[0094] The spectrum diagram is obtained by preprocessing in step S2.
[0095] The amplitude value of each frequency point in the spectrum is multiplied by the response value of each gammatone filter at that frequency point. Each gammatone filter generates a set of gammatone filtered signals. For each gammatone filter, the amplitude values corresponding to all frequency points in the filtered signal are added together to obtain the short-term amplitude of the output signal of each gammatone filter. The amplitudes output by these filters are combined to obtain the filtered time-frequency diagram.
[0096] S323, take logarithm
[0097] The logarithm of the short-time amplitude output by each gammatone filter is taken to obtain the logarithmic amplitude of each frame; the logarithmic amplitude of each frame is arranged into a two-dimensional matrix in time sequence and visualized to obtain the logarithmic gammatone spectrum.
[0098] Taking the logarithm of the Mel filter output can better simulate the nonlinear perception of sound intensity by the human ear.
[0099] S324, Discrete Cosine Transform
[0100] Perform DCT on the logarithmic amplitude value of each frame in the log gammatone spectrum obtained in step S323 according to formula (5) to obtain L gammatone frequency cepstral coefficients of each frame; arrange the gammatone frequency cepstral coefficients of each frame in chronological order to form a time × coefficient two-dimensional matrix, and visualize the matrix using a heat map, with the horizontal axis being time and the vertical axis being the cepstral coefficient number (1-L), to obtain the following: Figure 6 Gammatone frequency cepstrum plot shown.
[0101] The above describes the specific process of gammatone feature extraction. Through these steps, we can obtain the gammatone frequency cepstrum of the audio signal. Next, we will describe the specific process of coal gangue identification using a classification network.
[0102] S4, classification and recognition stage
[0103] In this stage, it is necessary to use a classification network to obtain the coal gangue recognition results based on the Mel frequency cepstrum and gammatone frequency cepstrum obtained in S3.
[0104] The classification network can be a support vector machine (SVM), a deep neural network, a random forest, etc. Before using the classification network for gangue recognition, it is necessary to use samples to perform necessary training and optimization on the classification network so that it can achieve high-precision recognition of gangue audio.
[0105] The gangue identification result can indicate the presence or absence of gangue, or the type of gangue present. The gangue identification result can be presented in a visual format, such as a list or image. The gangue identification result can also be linked to relevant alarm or control mechanisms to ensure safe mine operations.
[0106] The present invention uses cepstral images instead of cepstral coefficients as the input of the classification network, thereby avoiding information loss in the coefficient extraction process and improving the accuracy of coal gangue recognition.
[0107] The following is a specific step to complete the identification of coal gangue using a classification network:
[0108] S41. Constructing a training set and a test set for gangue identification. The features of the training set and the test set are Mel-frequency cepstrum and Gammatone frequency cepstrum, and the labels are the gangue identification results;
[0109] S42, using the Histogram of Oriented Gradients (HOG) algorithm to extract feature vectors from the Mel-frequency cepstral map and the gamma-tone frequency cepstral map provided by the training set;
[0110] S43, inputting the feature vector extracted by HOG and the corresponding label into the linear kernel SVM model to train the SVM model;
[0111] S44. After the training is completed, the accuracy of the model in identifying coal gangue is evaluated using the test set; if the accuracy reaches the preset standard, the model training is considered complete;
[0112] S45. Use the HOG algorithm to extract feature vectors from the Mel-frequency cepstrum and the gammatone frequency cepstrum obtained in step S3, and input them into the trained SVM model to obtain coal gangue recognition results.
[0113] The above is an embodiment of using SVM to identify coal gangue. In actual application, other image processing algorithms or classification networks can be selected as needed to complete spectral feature extraction and coal gangue identification.
[0114] The above describes the specific steps of the coal gangue audio recognition method provided by the present invention. As can be seen from the above, compared with the existing technology, the coal gangue audio recognition method provided by the present invention has the following beneficial effects:
[0115] 1. The present invention combines the two audio features of Mel frequency cepstrum and gammatone frequency cepstrum, fully utilizing their advantages in simulating the auditory characteristics of the human ear, enhancing the recognition ability of coal gangue audio signals and improving the recognition accuracy.
[0116] The advantages of combining MFC and GFC audio features for coal gangue audio recognition mainly include:
[0117] (1) Combining MFC and GFC allows their auditory perception characteristics to complement each other, improving adaptability to the complex acoustic environment of coal mines. The combined audio features of MFC and GFC are more robust in noisy environments and can significantly reduce the error recognition rate in noisy environments. In a noisy environment (SNR = 5dB), MFC + GFC reduces the error rate by 15-20% compared to single feature.
[0118] (2) MFC has a high sensitivity to mid- and low-frequency audio, but poor resolution of high frequencies, while GFC is more evenly distributed across the entire frequency band and can retain more high-frequency components. Combining the two can better capture the wide-band characteristics of coal gangue audio while ensuring the resolution of mid- and low-frequency audio, thereby significantly improving the accuracy of coal gangue audio recognition.
[0119] (3) MFC focuses on static spectral features, while GFC can better characterize transient changes in audio. The combination of the two can simultaneously capture the steady-state and transient characteristics of coal gangue audio, thereby significantly improving the accuracy of coal gangue audio recognition. On the TIMIT speech database, the recognition rate of MFC alone is 89.2%, and that of GFC is 87.5%. The fusion of the two can reach 91.8% (DNN-HMM model).
[0120] 2. The present invention uses cepstral images instead of cepstral coefficients as the input of the classification network, making full use of the two-dimensional spatiotemporal relationship between the coefficients and reducing information loss. In addition, this method uses the powerful information mining ability of deep learning networks for graphic data to process cepstral images, thereby improving the accuracy of coal gangue identification.
[0121] 3. The present invention improves the audio signal preprocessing, feature extraction and fusion methods, thereby enhancing the robustness and stability of the system in complex environments.
[0122] The coal gangue audio recognition method provided by the present invention can accurately identify coal gangue in real time, has good application prospects, can greatly improve the level of coal mine safety management, and reduce the incidence of coal mine accidents.
[0123] In the description of the embodiments of the present application, words such as "exemplary," "for example," or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary," "for example," or "for example" in the embodiments of the present application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.
[0124] In the description of the embodiments of this application, the term "and / or" is merely a description of an association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can represent the following three situations: A exists alone, B exists alone, and A and B exist at the same time. In addition, unless otherwise specified, the term "plurality" means two or more.
[0125] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly identifying the technical features being referred to. Thus, features specified as "first" or "second" may explicitly or implicitly include one or more of such features. The terms "include," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.
[0126] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for identifying coal gangue, characterized in that: The following steps are involved: Collect coal gangue audio signals; Preprocessing the coal gangue audio signal to obtain a frequency spectrum of the audio signal; Using a Mel filter bank to extract features from the spectrogram to obtain a Mel frequency cepstrum; Using a gammatone filter bank to extract features from the spectrum to obtain a gammatone frequency cepstrum; Based on the Mel frequency cepstrum and the gammatone frequency cepstrum, a classification network is used to obtain a coal gangue recognition result.
2. The coal gangue identification method according to claim 1, characterized in that: The preprocessing of the coal gangue audio signal specifically includes: The audio signal is divided into a plurality of overlapping short time frames, and each frame of the audio signal is windowed; each frame of the windowed audio signal is Fourier transformed to obtain a spectrum of each frame of the audio signal; and a spectrum diagram is drawn according to the spectrum of each frame of the audio signal.
3. The coal gangue identification method according to claim 1, characterized in that: The feature extraction of the spectrogram using the Mel filter bank to obtain the Mel frequency cepstrum specifically includes: Calculate a power spectrum based on the spectrum graph; Filtering the power spectrum using the Mel filter bank to obtain the short-time energy output by each Mel filter; Take the logarithm of the short-time energy output by each Mel filter and arrange the results into a two-dimensional matrix in time order to obtain the logarithmic Mel spectrum; Perform discrete cosine transform on the logarithmic Mel spectrum to obtain Mel frequency cepstral coefficients; The Mel-frequency cepstrum coefficients of each frame signal are arranged into a two-dimensional matrix in time sequence to obtain the Mel-frequency cepstrum.
4. The coal gangue identification method according to claim 3, characterized in that: The Mel filter bank is pre-constructed according to the Mel frequency scale, and the Mel filter bank is a triangular filter bank; in the Mel frequency scale, the relationship between the Mel frequency M(f) and the normal frequency f is: M(f)=1125ln(1+f / 700).
5. The coal gangue identification method according to claim 1, characterized in that: The method of extracting features from the spectrum using a gammatone filter bank to obtain a gammatone frequency cepstrum specifically includes: Filtering the spectrogram using a gammatone filter bank to obtain a short-time energy output by each gammatone filter; Take the logarithm of the short-time energy output by each gammatone filter and arrange the results into a two-dimensional matrix in time order to obtain the logarithmic gammatone spectrum; Perform discrete cosine transform on the logarithmic gammatone spectrum to obtain the gammatone frequency cepstral coefficients; The gamma-tone frequency cepstrum coefficients of each frame signal are arranged in time sequence into a two-dimensional matrix to obtain a gamma-tone frequency cepstrum.
6. The coal gangue identification method according to claim 1, characterized in that: The classification network is one of a support vector machine, a deep neural network, and a random forest.