A method for extracting abnormal features of idlers based on voiceprint spectrum separation
The abnormal characteristics of the rollers are extracted through the soundprint separation method, which solves the problems of low manual inspection efficiency and difficulty in identifying rollers in complex underground noise environments, and realizes accurate identification and timely alarms of rollers, ensuring safe production of coal mines.
Patent Information
- Application Number
- CN202211375620.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-04
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-11-04
AI Technical Summary
In the prior art, the abnormal detection of rollers in belt conveyors relies on manual inspection, which has high labor intensity, low efficiency and easy to miss inspection. In complex noise environments underground in coal mines, it is difficult to extract abnormal soundprint characteristics of rollers, low recognition accuracy, and high false alarm rate.
Using a method based on voiceprint spectrum separation, the sound signal of the roller is obtained, and the time and frequency domain characteristics are extracted after preprocessing. The soundprint spectrum is separated by wavelet threshold denoising and HPSS. Combined with MFCC transformation, the roller abnormality is identified and the abnormality is confirmed through the voiceprint characteristics.
It realizes accurate and rapid identification of roller abnormalities and prompt alarms, instead of manual inspection, avoid major accidents, ensure safe production of coal mines, and expands the abnormal identification and disaster alarms applied to other rotating machinery equipment.
Smart Images

Figure CN115753984B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of voiceprint feature extraction, and relates to a method for extracting abnormal features of idlers based on voiceprint spectrum separation. Background Art
[0002] Belt conveyors are the main means of material transportation in industries such as coal mines, non-coal mines, power, and docks, and contain a large number of idler groups. As important load-bearing components, idlers are numerous and easily damaged, and are extremely likely to cause accidents such as belt conveyor fires, tape tears, and deviation, resulting in casualties and economic losses, and seriously threatening the safe production of coal mines. At present, the abnormal detection of idlers on belt conveyors mainly relies on manual inspections. Inspectors rely on physical senses such as hearing, vision, and touch to judge the operation of equipment, with high labor intensity, high risk, low efficiency, and easy to miss inspections. When an idler operates abnormally, obvious noises will be generated. Online monitoring of idler abnormalities through sound can replace manual inspections, liberate inspectors from harsh environments and heavy labor, and achieve the purpose of reducing personnel, improving efficiency, and enhancing safety.
[0003] In coal mines, the sound sources underground are complex, the low-frequency noise interference is large, and the signal-to-noise ratio is low. The abnormal voiceprint features of idlers are not obvious and are difficult to extract. The method of identifying idler abnormalities based on a single sound feature has low accuracy and a high false alarm rate, and is limited in on-site applications. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method for extracting abnormal features of idlers based on voiceprint spectrum separation.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A method for extracting abnormal features of idlers based on voiceprint spectrum separation, the method comprising the following steps:
[0007] S1: Obtain the idler sound signal;
[0008] S2: Preprocess the idler sound signal;
[0009] S3: After preprocessing, obtain time-domain feature 1, compare the value of time-domain feature 1 with a preset threshold 1, and if it is greater than the preset threshold 1, issue warning 1;
[0010] S4: After preprocessing, perform frequency-domain transformation to obtain frequency-domain feature 2, compare the value of frequency-domain feature 2 with a preset threshold 2, and if it is greater than the preset threshold 2, issue warning 2;
[0011] S5: When both warning 1 and warning 2 occur, generate a voiceprint spectrum, and separate the voiceprint spectrum through HPSS to obtain a harmonic component and a shock wave component;
[0012] S6: Perform MFCC transformation on the shock wave component to obtain voiceprint feature 3. When the value of voiceprint feature 3 exceeds the preset threshold, confirm that the idler is abnormal and send an idler abnormality alarm.
[0013] Optionally, in S2, for the preprocessing of the idler sound signal, wavelet threshold denoising is used to perform wavelet transform, wavelet decomposition, threshold processing, and signal reconstruction on the original sound signal to remove Gaussian white noise; the decomposition level is selected as 3, the wavelet basis is selected as db6, and the threshold processing function introduces an exponential function and a parameter to smooth and denoise the signal; the expression of the threshold processing function is as follows:
[0014]
[0015] where: 1 / exp(w j,k -λ) is the reciprocal of the introduced exponential function, and a is the introduced parameter.
[0016] Optionally, S3 is specifically as follows:
[0017] S3-1: The pre-emphasis processing is to pass the preprocessed sound signal through a high-pass filter to enhance the high-frequency part of the sound signal and prevent the loss of high-frequency signal information; the first-order high-pass filter is defined as follows:
[0018] H(Z) = 1 - μz -1
[0019] where μ is between 0.9 and 1.0;
[0020] S3-2: Frame processing: Divide several sound sampling points into one frame. Within this frame, the characteristics of the sound signal are stable. The frame length is taken as 1 s, and the sampling frequency is 22050 Hz; to avoid frequency aliasing after frame processing, windowing is performed on the framed sound signal, that is, each framed sound signal is multiplied by a Hamming window, and the Hamming window function is defined as follows:
[0021]
[0022] where different values of a produce different Hamming windows, and a is taken as 0.46;
[0023] S3-3: The idler sound signal has time-varying and short-term stationary characteristics. Use short-time average energy and amplitude to characterize a feature of the current signal. The short-time average energy is defined as:
[0024]
[0025] The short-time average amplitude is defined as:
[0026]
[0027] Where N is the window length, and w(n,a) is the Hamming window function;
[0028] S3-4: The short-time zero-crossing rate characterizes the number of times a sound signal x(n) crosses zero or a set threshold. The number of zero crossings reflects the low-frequency and high-frequency characteristics of the abnormal sound of the idler. The short-time zero-crossing rate is defined as:
[0029]
[0030] Where N is the window length, m is the zero point or the specified threshold, w(n,a) is the Hamming window function, and the sgn(x) function is the sign function. The formula is:
[0031]
[0032] S3-5: The peak-to-peak value indicates the strength of the vibration when the idler is running. This value often increases sensitively when a fault occurs. The peak-to-peak value is defined as follows:
[0033] F n = max(x(n)) - min(x(n))
[0034] S3-6: Kurtosis represents the impact pulse generated when the idler fails. The more serious the fault, the greater the amplitude of the impact response. It is sensitive to early faults in rotating equipment. Kurtosis is defined as follows:
[0035]
[0036] Where P is the average value, and the formula is as follows:
[0037]
[0038] S3-7: Normalize the short-time average energy, short-time average amplitude, short-time zero-crossing rate, peak-to-peak value, and kurtosis characteristics. After processing, they are E‘, M‘, Z‘, F‘, and C‘ respectively. The normalization processing expression is as follows:
[0039]
[0040] Where G‘ represents the normalized eigenvalue, and its range is (0,1). G is the eigenvalue before normalization;
[0041] S3-8: Multiply E‘, M‘, Z‘, F‘, and C‘ by the weight coefficients a(a1,a2,a3,a4,a5) respectively and then add them to obtain the eigenvalue T1 of the time-domain feature 1. The expression is as follows:
[0042] T1 = E‘a1 + M‘a2 + Z‘a3 + F‘a4 + C‘a5
[0043] Among them, the value ranges of a1, a2, a3, a4, and a5 are (0, 1), which are preset.
[0044] S3-9: If the time-domain eigenvalue T1 is greater than the preset threshold 1, then warning 1 is issued.
[0045] Optionally, in the step S4, the steps are as follows:
[0046] S4-1: After the preprocessed signal is subjected to Fourier transform, the frequency-domain energy is calculated, which is an effective feature for distinguishing non-silent and silent states.
[0047]
[0048] Among them, ω0 is half of the sampling frequency, and F i (ω) represents the Fourier transform of the i-th frame signal.
[0049] S4-2: The frequency domain is divided into 4 sub-bands, and the sub-band energy ratio represents the ratio of the energy of the j-th sub-band in the i-th frame to the total energy of the frequency domain, which is expressed as:
[0050]
[0051] Among them, represents the upper boundary frequency of the j-th sub-band, represents the lower boundary frequency of the j-th sub-band; the abnormal sound of the idler mainly concentrates on the primary sub-band.
[0052] S4-3: Extract the resonance peaks in the spectral envelope, and calculate the sharpness and 1 / 3 octave eigenvalue; the sharpness is expressed as follows:
[0053]
[0054] In the formula, N'(z) is the loudness spectrum on the critical band Z, the integral of the loudness spectrum over the critical band is the loudness, g(z) is the additional coefficient, Z is the critical band, and d(z) is the differential of the critical band.
[0055] S4-4: Normalize the frequency-domain energy, sub-band energy ratio, resonance peak features, sharpness, and 1 / 3 octave features. The normalization processing expression is shown in S3-7; multiply the normalized features by the weight coefficients b (b1, b2, b3, b4, b5) respectively and then add them to obtain the eigenvalue F2 of the frequency-domain feature 2. The expression is shown in S3-8. Among them, the value ranges of b1, b2, b3, b4, and b5 are (0, 1), which are preset.
[0056] S4-5: If the frequency-domain eigenvalue F2 is greater than the preset threshold 2, then warning 2 is issued.
[0057] Optionally, in the step S5, the steps are as follows:
[0058] S5-1: When both Warning 1 and Warning 2 occur, generate a voiceprint spectrum, separate the voiceprint spectrum through HPSS to obtain a harmonic component and a shock wave component. The HPSS steps are as follows:
[0059] (1) Perform a short-time Fourier transform (STFT) χ on the preprocessed sound signal:
[0060]
[0061] Further obtain the energy spectrogram y:
[0062] y(m, k) = |χ(m, k)| 2
[0063] (2) Generate a harmonic component spectrogram and a shock wave component spectrogram from the energy spectrogram through a set of horizontal median filters and a set of vertical median filters respectively:
[0064]
[0065] (3) Construct a masking matrix by means of binary masking:
[0066]
[0067] Use the masking matrix M h and M p to divide all the time-domain frequency-domain bins in the STFT into either the harmonic component or the shock wave component, and then divide the original STFT result into the harmonic component χ h and the shock wave component χ p :
[0068] χ h = (χ ⊙ Mh)
[0069] χ p = (χ ⊙ M p )
[0070] (4) Finally, perform an inverse short-time Fourier transform iSTFT on χ h and χ p to obtain the harmonic component χ h and the shock wave component χ p of the original signal, thus completing the HPSS.
[0071] Optionally, in step S6, the steps are as follows:
[0072] S6-1: Pass the energy spectrogram y obtained in S5-1 through a Mel filter bank to obtain a Mel spectrogram;
[0073] S6-2: The frequency response H of the triangular filter m(k) is defined as:
[0074]
[0075] where m is the number of filters, and f(m) is the center frequency corresponding to the m-th filter;
[0076] S6-3: Calculate the logarithmic energy s(m) output by each filter bank, and convert the Mel spectrogram into a logarithmic Mel spectrogram:
[0077]
[0078] where M is the maximum number of filters;
[0079] S6-4: Calculate the logarithmic energy of a frame of signal based on the single logarithmic energy, and perform discrete cosine transform (DCT) on the logarithmic energy output by each filter. The expression is as follows, to obtain MFCC, that is, voiceprint feature 3. When its value exceeds the preset threshold, it is confirmed that the idler is abnormal, and an idler abnormality alarm is issued;
[0080]
[0081] where L refers to the order of MFCC, taking values from 12 to 16.
[0082] The beneficial effects of the present invention are as follows: A method for extracting idler abnormal features based on voiceprint spectrum separation disclosed by the present invention collects idler sounds in real time by arranging mining acoustic sensors at certain intervals along the belt conveyor. The mining acoustic sensors output electrical signals to the mining acoustic analyzer, which has a built-in idler abnormality recognition algorithm, can accurately and quickly identify abnormalities and give alarms in a timely manner. Through the mining repeater, the alarm information and sound data are connected to the underground Ethernet ring network and transmitted to the ground equipment abnormal monitoring software, replacing manual inspection to achieve 24*7 online monitoring, avoiding major accidents such as fires and belt tearing, reducing unplanned downtime, and ensuring the safe and efficient operation of the main coal flow transportation. At the same time, the present invention can be extended and applied to the abnormal recognition of rotating mechanical equipment such as winches, main ventilation fans, air compressors, and pumps in coal mines and non-coal mines, as well as the alarm of disasters such as coal mine rock bursts, water inrusions, and roof falls, promoting the development of intelligent diagnosis of coal mine equipment failures and intelligent early warning technologies for coal mine disasters.
[0083] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. Brief Description of the Drawings
[0084] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail with reference to the accompanying drawings. Among them:
[0085] Figure 1 It is a schematic flow diagram of a method for extracting abnormal features of idlers based on voiceprint spectrum separation according to the present invention;
[0086] Figure 2 It is a schematic connection diagram of the equipment abnormal monitoring system according to the present invention. Specific embodiments
[0087] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0088] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be construed as limitations on the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0089] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be construed as limitations on the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0090] The method for extracting abnormal features of idlers based on voiceprint spectrum separation according to the present invention has a process as Figure 1 shown. The main steps of the present invention are as follows:
[0091] By arranging mine voice sensors along the belt conveyor to collect the idler sounds in real time and converting them into electrical signals for transmission to a mine voice analyzer, as Figure 2 shown.
[0092] The mine-used sound analyzer preprocesses the sound signal of the idler, that is, wavelet transform, wavelet decomposition, threshold processing, and signal reconstruction are performed on the original sound signal by using wavelet threshold denoising to remove Gaussian white noise. In the present invention, the decomposition level is selected as 3, the wavelet basis is selected as db6, and the exponential function and parameters are introduced into the threshold processing function to make the signal smoother and the denoising effect better. The expression of the threshold processing function is as follows:
[0093]
[0094] Where: 1 / exp(w j,k -λ) is the reciprocal of the introduced exponential function, and a is the introduced parameter.
[0095] After pre-emphasis and framing processing, the short-time average energy, short-time average amplitude, short-time zero-crossing rate, peak-to-peak value, and kurtosis feature are calculated, and normalization processing is performed to obtain the time-domain feature T1. If it is greater than the preset threshold 1, an early warning 1 is issued. The detailed steps are as follows:
[0096] S3-1: The pre-emphasis processing is to pass the preprocessed sound signal through a high-pass filter to enhance the high-frequency part of the sound signal and prevent the loss of high-frequency signal information. The first-order high-pass filter is defined as follows:
[0097] H(Z) = 1 - μz -1
[0098] Where, μ is between 0.9 and 1.0, and generally 0.97 is taken.
[0099] S3-2: Framing processing: Several sound sampling points are divided into one frame. Within this frame, the characteristics of the sound signal can be regarded as stable. The frame length is taken as 1 s, and the sampling frequency is 22050 Hz. To avoid frequency aliasing after framing, a window is added to the framed sound signal, that is, each frame of the sound signal is multiplied by a Hamming window. The definition of the Hamming window function is as follows:
[0100]
[0101] Where, different a values can generate different Hamming windows, and generally a is taken as 0.46.
[0102] The idler sound signal has time-varying and short-time stationary characteristics, and the short-time average energy and amplitude can be used to characterize a feature of the current signal. The short-time average energy is defined as:
[0103]
[0104] The short-time average amplitude is defined as:
[0105]
[0106] Among them, N is the window length, and w(n,a) is the Hamming window function.
[0107] S3-4: The short-time zero-crossing rate characterizes the number of times a sound signal x(n) crosses zero or a set threshold. The number of zero crossings reflects the low-frequency and high-frequency characteristics of the abnormal sound of the idler. The short-time zero-crossing rate is defined as:
[0108]
[0109] Among them, N is the window length, m is the zero point or the specified threshold, w(n,a) is the Hamming window function, and the sgn(x) function is the sign function. The formula is:
[0110]
[0111] S3-5: The peak-to-peak value indicates the strength of the vibration when the idler is running. This value often increases sensitively when a fault occurs. The peak-to-peak value is defined as follows:
[0112] F n = max(x(n)) - min(x(n))
[0113] S3-6: Kurtosis represents the impact pulse generated when the idler fails. The more serious the fault, the greater the amplitude of the impact response. It is more sensitive to early faults of rotating equipment. The kurtosis is defined as follows:
[0114]
[0115] Among them, P is the average value. The formula is as follows:
[0116]
[0117] S3-7: Normalize the short-time average energy, short-time average amplitude, short-time zero-crossing rate, peak-to-peak value, and kurtosis characteristics. After processing, they are E‘, M‘, Z‘, F‘, and C‘ respectively. The normalization processing expression is as follows:
[0118]
[0119] Among them, G‘ represents the normalized eigenvalue, and its range is (0,1). G is the eigenvalue before normalization.
[0120] S3-8: Multiply E‘, M‘, Z‘, F‘, and C‘ by the weight coefficients a(a1,a2,a3,a4,a5) respectively and then add them to obtain the eigenvalue T1 of the time-domain feature 1. The expression is as follows:
[0121] T1 = E‘a1 + M‘a2 + Z‘a3 + F‘a4 + C‘a5
[0122] Among them, the value ranges of a1, a2, a3, a4, and a5 are (0, 1) and can be preset.
[0123] S3-9: If the time-domain eigenvalue T1 is greater than the preset threshold 1, warning 1 is issued.
[0124] Perform frequency-domain transformation on the preprocessed idler sound signal to obtain frequency-domain feature 2. Compare the value of frequency-domain feature 2 with the preset threshold 2. If it is greater than the preset threshold 2, warning 2 is issued. The detailed steps are as follows:
[0125] S4-1: After performing Fourier transform on the preprocessed signal, calculate the frequency-domain energy, which is an effective feature for distinguishing non-silent and silent.
[0126]
[0127] Among them, ω0 is half of the sampling frequency, and Fi(ω) represents the Fourier transform of the i-th frame signal.
[0128] S4-2: Divide the frequency domain into 4 subbands. The subband energy ratio represents the ratio of the energy of the j-th subband in the i-th frame to the total frequency-domain energy, and can be expressed as:
[0129]
[0130] Among them, represents the upper boundary frequency of the j-th subband, represents the lower boundary frequency of the j-th subband. The abnormal sound of the idler mainly concentrates on the primary subband.
[0131] S4-3: Extract the resonance peaks in the spectral envelope, and calculate the sharpness and 1 / 3 octave eigenvalue. The sharpness is expressed as follows:
[0132]
[0133] In the formula, N'(z) is the loudness spectrum on the critical band Z. The integral of the loudness spectrum over the critical band is the loudness, g(z) is the additional coefficient, Z is the critical band, and d(z) is the differential of the critical band.
[0134] S4-4: Normalize the frequency-domain energy, subband energy ratio, resonance peak features, sharpness, and 1 / 3 octave features. The normalization processing expression is shown in S3-7. Multiply the normalized features by the weight coefficients b (b1, b2, b3, b4, b5) respectively and then add them to obtain the eigenvalue F2 of frequency-domain feature 2. The expression is shown in S3-8. Among them, the value ranges of b1, b2, b3, b4, and b5 are (0, 1) and can be preset.
[0135] S4-5: If the frequency-domain eigenvalue F2 is greater than the preset threshold 2, warning 2 is issued.
[0136] When both Warning 1 and Warning 2 occur, a voiceprint spectrum is generated. The voiceprint spectrum is separated by HPSS to obtain a harmonic component and a shock wave component. The detailed steps are as follows:
[0137] S5-1: When both Warning 1 and Warning 2 occur, a voiceprint spectrum is generated. The voiceprint spectrum is separated by HPSS to obtain a harmonic component and a shock wave component. The HPSS steps are as follows:
[0138] (1) Perform a short-time Fourier transform (STFT) χ on the preprocessed sound signal:
[0139]
[0140] Further obtain the energy spectrogram y:
[0141] y(m, n) = |χ(m, k)| 2
[0142] (2) Generate a harmonic component spectrogram and a shock wave component spectrogram for the energy spectrogram through a set of horizontal median filters and a set of vertical median filters respectively:
[0143]
[0144] (3) Construct a masking matrix by means of binary masking:
[0145]
[0146] Use the masking matrices Mh and Mp to divide all the time-frequency bins in the STFT into either the harmonic component or the shock wave component, and then divide the original STFT result into the harmonic component χ h and the shock wave component χ p :
[0147] χ h = (χ ⊙ M h )
[0148] χ p = (χ ⊙ M p )
[0149] (4) Finally, perform an inverse short-time Fourier transform iSTFT on χ h and χ p to obtain the harmonic component χ h and the shock wave component χ p of the original signal, thus completing the HPSS.
[0150] Perform an MFCC transform on the shock wave component. When the value of the voiceprint feature 3 exceeds the preset threshold, it is confirmed that the idler is abnormal, and an idler abnormality alarm is issued. The detailed steps are as follows:
[0151] S6-1: Pass the energy spectrogram y obtained in S5-1 through a Mel filter bank to obtain a Mel spectrogram;
[0152] S6-2: The frequency response H m (k) is defined as:
[0153]
[0154] where m is the number of filters, and f(m) is the center frequency corresponding to the m-th filter.
[0155] S6-3: Calculate the logarithmic energy s(m) output by each filter bank, and convert the Mel spectrogram into a logarithmic Mel spectrogram:
[0156]
[0157] where M is the maximum number of filters.
[0158] S6-4: Calculate the logarithmic energy of a frame of signal based on the individual logarithmic energy, and perform a discrete cosine transform (DCT) on the logarithmic energy output by each filter. The expression is as follows to obtain MFCC, that is, voiceprint feature 3. When its value exceeds a preset threshold, it is confirmed that the idler is abnormal and an idler abnormality alarm is issued.
[0159]
[0160] where L refers to the order of MFCC, generally taking 12 to 16.
[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for extracting abnormal features of idlers based on voiceprint spectrum separation, characterized in that: The method includes the following steps: S1: Obtain the idler sound signal; after pre-emphasis and frame processing, calculate the short-time average energy, short-time average amplitude, short-time zero-crossing rate, peak-to-peak value, and kurtosis feature, and perform normalization processing. After processing, they are E’, M’, Z’, F’, and C’ respectively; the normalization processing expressions are as follows: Among them, G’ represents the normalized eigenvalue, and its range is (0,1), and G is the eigenvalue before normalization; multiply E’, M’, Z’, F’, and C’ by the weight coefficients a, that is, a1, a2, a3, a4, a5 respectively, and then add them to obtain the eigenvalue T1 of time-domain feature 1; the value ranges of a1, a2, a3, a4, a5 are (0,1) and can be preset; if the time-domain eigenvalue T1 is greater than the preset threshold 1, then issue warning 1; S2: Preprocessing of the idler sound signal; if the time-domain eigenvalue T1 is greater than the preset threshold 1, warning 1 is issued; calculate the frequency-domain energy, sub-band energy ratio, formant features and sharpness, octave band features, and perform normalization processing; multiply the normalized features by the weight coefficients b, namely b1, b2, b3, b4, b5, respectively, and then add them up to obtain the eigenvalue F2 of the frequency-domain feature 2, where the value ranges of b1, b2, b3, b4, b5 are (0, 1) and can be preset; if the frequency-domain eigenvalue F2 is greater than the preset threshold 2, warning 2 is issued; S3: After preprocessing, obtain time-domain feature 1, compare the value of time-domain feature 1 with the preset threshold 1, and if it is greater than the preset threshold 1, issue warning 1; S4: After preprocessing, perform frequency-domain transformation to obtain frequency-domain feature 2, compare the value of frequency-domain feature 2 with the preset threshold 2, and if it is greater than the preset threshold 2, issue warning 2; S5: When both warning 1 and warning 2 occur, generate a voiceprint spectrum, and separate the voiceprint spectrum through HPSS to obtain a harmonic component and a shock wave component; S6: Perform MFCC transformation on the shock wave component to obtain voiceprint feature 3. When the value of voiceprint feature 3 exceeds the preset threshold, confirm that the idler is abnormal and issue an idler abnormality alarm.
2. The method for extracting abnormal features of idlers based on voiceprint spectrum separation according to claim 1, characterized in that: In S2, for the preprocessing of the idler sound signal, wavelet threshold denoising is used to perform wavelet transformation, wavelet decomposition, threshold processing, and signal reconstruction on the original sound signal to remove Gaussian white noise; the decomposition level is selected as 3, the wavelet basis is selected as db6, and the threshold processing function introduces an exponential function and parameters to smooth and denoise the signal; the threshold processing function expression is as follows: where: 1 / exp(w j,k -λ) is the reciprocal of the introduced exponential function, and a is the introduced parameter.
3. The method for extracting abnormal features of idlers based on voiceprint spectrum separation according to claim 1, characterized in that: The specific content of S3 is as follows: S3-1: The pre-emphasis process is to pass the preprocessed sound signal through a high-pass filter to enhance the high-frequency part of the sound signal and prevent the loss of high-frequency signal information; the first-order high-pass filter is defined as follows: H(Z) = 1 - μz -1 Among them, u is between 0.9 and 1.0; S3-2: Frame processing: Divide several sound sampling points into one frame. Within this frame, the characteristics of the sound signal are stable. The frame length is taken as 1s, and the sampling frequency is 22050Hz; to avoid frequency aliasing after frame processing, windowing is performed on the framed sound signal, that is, each framed sound signal is multiplied by a Hamming window, and the Hamming window function is defined as follows: Among them, different a values generate different Hamming windows, and a is taken as 0.46; S3-3: The idler sound signal has time-varying and short-term stationary characteristics. The short-time average energy and amplitude are used to characterize a feature of the current signal. The short-time average energy is defined as: The short-time average amplitude is defined as: Among them, N is the window length, and w(n,a) is the Hamming window function; S3-4: The short-time zero-crossing rate characterizes the number of times a section of sound signal x(n) crosses the zero point or a set threshold. The number of zero-crossing times reflects the low-frequency and high-frequency characteristics of the abnormal sound of the idler. The short-time zero-crossing rate is defined as: Where N is the window length, m is the zero point or specified threshold, w(n,a) is the Hamming window function, and sgn(x) is the sign function. The formula is as follows: S3-5: The peak-to-peak value indicates the intensity of the vibration during the operation of the idler. This value often increases sensitively when a fault occurs. The peak-to-peak value is defined as follows: F n = max(x(n)) - min(x(n)) S3-6: Kurtosis represents the impact pulse generated during idler failure. The more severe the failure, the greater the amplitude of the impact response. It is sensitive to early faults in rotating equipment. The definition of kurtosis is as follows: Where P is the average value, and the formula is as follows: S3-7: Normalize the short-time average energy, short-time average amplitude, short-time zero-crossing rate, peak-to-peak value, and kurtosis feature. After processing, they are E‘, M’, Z‘, F’, and C‘ respectively. The normalization processing expressions are as follows: Where G‘ represents the normalized eigenvalue, and its range is (0,1). G is the eigenvalue before normalization. S3-8: Multiply E‘, M’, Z‘, F’, and C‘ by the weight coefficients a(a1,a2,a3,a4,a5) respectively and then sum them to obtain the eigenvalue T1 of the time-domain feature 1. The expression is as follows: T1 = E'a1 + M'a2 + Z'a3 + F'a4 + C'a5 Where a1,a2,a3,a4,a5 have a preset value range of (0,1). S3-9: If the time-domain eigenvalue T1 is greater than the preset threshold 1, then issue warning 1.
4. A method for extracting abnormal features of idlers based on voiceprint spectrum separation according to claim 1, characterized in that: In the step S4, the steps are as follows: S4-1: After performing Fourier transform on the preprocessed signal, calculate the frequency-domain energy, which is an effective feature for distinguishing non-silent and silent states. where ω0 is half of the sampling frequency, and F i (ω) represents the Fourier transform of the i-th frame signal; S4-2: Divide the frequency domain into 4 subbands. The subband energy ratio represents the ratio of the energy of the j-th subband in the i-th frame to the total frequency-domain energy, which is expressed as: Among them, represents the upper boundary frequency of the j-th sub-band, represents the lower boundary frequency of the j-th sub-band; the abnormal sound of the idler mainly concentrates on the primary sub-band; S4-3: Extract the resonance peaks in the spectral envelope and calculate the sharpness and 1 / 3 octave eigenvalue. The sharpness is expressed as follows: In the formula, N'(z) is the loudness spectrum on the critical band Z. The integral of the loudness spectrum over the critical band is the loudness. g(z) is the additional coefficient, Z is the critical band, and d(z) is the differential of the critical band. S4-4: Normalize the frequency-domain energy, subband energy ratio, resonance peak feature, sharpness, and 1 / 3 octave feature. The normalization processing expressions are shown in S3-7. Multiply the normalized features by the weight coefficients b, namely b1,b2,b3,b4,b5 respectively, and then sum them to obtain the eigenvalue F2 of the frequency-domain feature 2. Where b1,b2,b3,b4,b5 have a preset value range of (0,1). S4-5: If the frequency-domain eigenvalue F2 is greater than the preset threshold 2, then issue warning 2.
5. A method for extracting abnormal features of idlers based on voiceprint spectrum separation according to claim 1, characterized in that: In the step S5, the steps are as follows: S5-1: When both warning 1 and warning 2 occur, generate a voiceprint spectrum. Separate the voiceprint spectrum through HPSS to obtain the harmonic component and shock wave component. The HPSS steps are as follows: (1) Perform short-time Fourier transform STFTχ on the preprocessed sound signal: Further obtain the energy spectrogram y: y(m, k) = |χ(m, k)| 2 (2) Generate the harmonic component spectrogram and shock wave component spectrogram respectively through a group of horizontal median filters and a group of vertical median filters on the energy spectrogram: (3) Construct a masking matrix by means of binary masking: Using the masking matrix M h and M p divide all time-frequency bins in the STFT into harmonic components or shock wave components, and then partition the harmonic components χ h and shock wave components χ p : χ h = (χ ☉ M h ) χ p = (χ ⊙ M p ) (4) Finally, perform the inverse short-time Fourier transform iSTFT on χ h and χ p to obtain the harmonic component χ h and the shock wave component χ p , thus completing HPSS.
6. A method for extracting abnormal features of idlers based on voiceprint spectrum separation according to claim 1, characterized in that: In the step S6, the steps are as follows: S6-1: Pass the energy spectrogram y obtained in S5-1 through the Mel filter bank to obtain the Mel spectrogram. S6-2: The frequency response H of the triangular filter m (k) is defined as: Among them, m is the number of filters, and f(m) is the center frequency corresponding to the m-th filter; S6-3: Calculate the logarithmic energy s(m) output by each filter bank, and convert the Mel spectrogram into a logarithmic Mel spectrogram: Among them, M is the maximum number of filters; S6-4: Calculate the logarithmic energy of a frame of signal based on the single logarithmic energy, and perform discrete cosine transform DCT on the logarithmic energy output by each filter. The expression is as follows to obtain MFCC, that is, voiceprint feature 3. When its value exceeds the preset threshold, it is confirmed that the idler is abnormal and an idler abnormality alarm is issued; Among them, L refers to the order of MFCC, taking 12 to 16.
Citation Information
Patent Citations
Carrier roller fault diagnosis method and system based on machine learning and storage medium
CN112504673A
Diagnostic method and equipment for abnormal noise of machine
JP1994300619A