Casting industry abnormal sound detection and grading response method based on voiceprint recognition
By using high-temperature resistant microphone arrays and hybrid deep learning models in the casting industry, the problems of poor signal processing and weak model generalization capabilities of abnormal sound detection in the casting industry are solved, and efficient hierarchical response and rapid processing of liquid aluminum leakage anomalies are achieved.
Patent Information
- Application Number
- CN202510580850.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-08
AI Technical Summary
In the existing abnormal sound detection of casting industries, there are problems such as strong subjectivity, low efficiency and poor safety in manual inspection, single monitoring signals of traditional sensors, difficult to identify abnormalities, poor signal preprocessing of existing voiceprint recognition technology, incomplete feature extraction, weak model generalization capabilities, and lack of hierarchical response.
The high-temperature resistant microphone array is used to collect sound signals, combine three-stage filtering and noise reduction and amplitude normalization preprocessing, and extract voiceprint features through variable window length and short time Fourier transform and extended Meer frequency cepspectral coefficients, and build a Transformer-CNN hybrid deep learning model based on attention mechanism to realize real-time detection and trigger hierarchical response measures.
It improves the accuracy and reliability of signal processing, enhances the accuracy and generalization of abnormal sound recognition, realizes rapid positioning and accurate processing of aluminum liquid leakage abnormalities, and improves the intelligence level and processing efficiency of casting equipment.
Smart Images

Figure CN120452472A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of foundry industry detection technology, and in particular to a foundry industry abnormal sound detection and graded response method based on voiceprint recognition. Background Art
[0002] In the foundry industry, abnormal conditions such as molten aluminum leaks can lead to serious consequences such as equipment damage, safety accidents, and production interruptions. Therefore, real-time monitoring of the operating status of aluminum casting equipment and the timely detection of abnormal sounds are crucial. Currently, the industry's main methods for detecting abnormal sounds in casting equipment include manual inspections and traditional sensor monitoring. Manual inspections rely on the technicians' hearing and experience, and are subject to high subjectivity, low efficiency, and untimely detection. They are difficult to adapt to the needs of automated production, and the accuracy and safety of manual inspections are also difficult to guarantee in high-temperature, high-noise casting environments. Traditional sensor monitoring, such as vibration sensors and pressure sensors, while capable of obtaining information on the equipment's operating status, suffers from problems such as single detection signals, susceptibility to environmental interference, and an inability to effectively distinguish between different types of abnormalities. They are unable to accurately identify specific abnormalities such as molten aluminum leaks, and struggle to meet the high-precision detection requirements of complex casting environments.
[0003] With the development of voiceprint recognition technology, some researchers have tried to apply it to anomaly detection in the foundry industry. However, existing voiceprint recognition technology still has many defects in its application in foundry industry scenarios. On the one hand, the foundry workshop environment is complex, with high-intensity background noise generated by equipment such as die-casting machines and mixers, as well as interference caused by equipment vibration. Traditional voiceprint recognition methods cannot effectively remove these noises in the signal preprocessing link, resulting in inaccurate subsequent feature extraction and affecting the detection results; on the other hand, existing technologies often use fixed-parameter signal processing methods and single feature extraction methods, which are difficult to adapt to the complex time-frequency characteristics of aluminum liquid leakage sound signals at different stages (minor leakage, moderate leakage, severe leakage), and cannot fully extract effective voiceprint features.
[0004] In addition, in terms of anomaly detection model construction, traditional deep learning models lack targeted optimization of abnormal sound characteristics in the foundry industry, and have problems such as poor model generalization ability, low detection accuracy, and high false alarm and missed alarm rates. In the abnormal response link, it is impossible to provide graded response measures according to the severity of abnormal situations, making it difficult to achieve efficient and accurate emergency response. Summary of the Invention
[0005] In order to overcome the shortcomings of the existing technology, the purpose of the present invention is to provide a method for abnormal sound detection and graded response in the foundry industry based on voiceprint recognition, which solves the problems of strong subjectivity, low efficiency and poor safety of manual inspections in the existing foundry industry abnormal sound detection, single monitoring signals of traditional sensors and difficulty in identifying abnormalities, and poor signal preprocessing, incomplete feature extraction, weak model generalization ability, and lack of graded response in existing voiceprint recognition technology. The method realizes the detection of abnormal sounds of aluminum liquid leakage in aluminum casting equipment and efficient graded emergency response.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] A method for detecting and responding to abnormal sounds in the foundry industry based on voiceprint recognition, comprising the following steps:
[0008] S1. Arrange a high-temperature resistant microphone array at a leak-prone location of the aluminum casting equipment, and collect sound signals containing characteristics of aluminum liquid leakage through the microphone array to obtain original sound data;
[0009] S2, performing three-stage filtering noise reduction and amplitude normalization preprocessing on the original sound data to remove noise and unify the signal amplitude standard to obtain a preprocessed sound signal;
[0010] S3. Extract voiceprint features from the preprocessed sound signal through variable window length short-time Fourier transform and extended Mel-frequency cepstral coefficients, and calculate the energy change rate and spectral center of gravity as unique features to obtain a voiceprint feature vector;
[0011] S4. Screening the voiceprint feature vectors through principal component analysis and recursive feature elimination, and constructing and training a Transformer-CNN hybrid deep learning model based on the attention mechanism;
[0012] S5. Process the sound signal in real time using a sliding window, extract features from the processed signal and input them into the Transformer-CNN hybrid deep learning model based on the attention mechanism for detection. By calculating the similarity and comparing it with the preset graded response threshold, the abnormal sound is detected in real time and the detection result is output;
[0013] S6. According to the output detection results, the operating status of the aluminum casting equipment is determined, and emergency response measures of corresponding levels are triggered to complete the processing of the abnormal situation of aluminum liquid leakage.
[0014] Preferably, in step S1, the microphone array adopts a four-element rectangular L-shaped microphone array, which is arranged around the casting port, mold seams and pipeline valves of the aluminum casting equipment to form a 90° fan-shaped coverage area; the microphone array is composed of high-temperature resistant 800°C, IP67 dust-proof and waterproof microphones, with a sensitivity of -32dB±2dB, a frequency response range of 20Hz~20kHz, and an installation height flush with the aluminum liquid flow plane, at a height of 1.2~1.5m;
[0015] The microphone array is also equipped with a four-channel synchronous USB data acquisition card to collect sound signals of aluminum liquid leakage characteristics at a 48kHz sampling rate and 24-bit quantization accuracy.
[0016] Preferably, in step S2, the original sound data is subjected to three-stage filtering noise reduction and amplitude normalization preprocessing to remove noise and unify the signal amplitude standard to obtain a preprocessed sound signal, which specifically includes:
[0017] First, the original sound data is subjected to three-level filtering and noise reduction, using a low-pass filter with a cutoff frequency of 3kHz to filter out the high-frequency mechanical noise of the equipment in the aluminum casting plant;
[0018] Secondly, a bandpass filter with a passband of 100-2000Hz is used to retain the characteristic frequency band of aluminum liquid leakage;
[0019] Thirdly, the improved Kalman filter is used to process the sound data of the characteristic frequency band of aluminum liquid leakage to remove noise.
[0020] Finally, the noise-removed and dried sound data is amplitude normalized to unify the signal amplitude standard.
[0021] Preferably, the state equation of the improved Kalman filter is:
[0022]
[0023] Among them, x t =[s t ,v t ] represents the signal amplitude and rate of change, s t represents the amplitude of the sound signal at time t, v t Indicates the rate of change of the sound signal at time t; Δt = 10ms is the sampling interval, a is the acceleration noise, w t is the process noise;
[0024] The amplitude normalization process uses the formula:
[0025]
[0026] Where x is the amplitude of the original sound signal, that is, the unprocessed sound intensity value collected by the microphone array; x min and x max are the minimum and maximum values of the original sound signal amplitude respectively; x′ is the normalized sound signal amplitude, which maps the original amplitude to the interval [0.1, 0.9] to eliminate the differences in sound amplitudes under different acquisition devices and different working conditions.
[0027] Preferably, in step S3, a variable window length short-time Fourier transform is used, including:
[0028] In the low frequency band of 100-500Hz, a Hanning window with a window length of 30ms and a frame shift of 15ms is used; in the high frequency band of 500-2000Hz, a Blackman window with a window length of 10ms and a frame shift of 5ms is used to calculate the time-frequency matrix. The formula is:
[0029]
[0030] Where n is the time frame index, used to identify different time windows; k is the frequency index, used to determine different frequency components; N = 2048 is the number of FFT points; x(m+nΔt) is the sound signal sampling value at time (m+nΔt); w(m) is the window function; e j2πkm / N It is used to convert the time domain signal into the frequency domain signal, and realize the discrete Fourier transform through the exponential function to obtain the component distribution of the sound signal at different frequencies.
[0031] Preferably, in step S3, the extended Mel-frequency cepstral coefficients are used, including:
[0032] 22 triangular filters covering 100-1800 Hz are used, and 12-dimensional MFCC coefficients, 12-dimensional first-order differences, and 12-dimensional second-order differences are extracted through logarithmic energy calculation and discrete cosine transform, for a total of 36 dimensions. The logarithmic energy calculation is:
[0033]
[0034] Where S(m) is the logarithmic energy, |X(k)| 2 represents the power spectral density at frequency k, H m (k) is the frequency response of the mth Mel filter at frequency k;
[0035] The discrete cosine transform is:
[0036]
[0037] Among them, c(n) is the discrete cosine transform, n is the index of the MFCC coefficient, and m is the index of the Mel filter; the logarithmic energy is then transferred from the Mel frequency domain to the cepstral domain.
[0038] Preferably, in step S3, the energy change rate and the spectrum center of gravity are calculated, and the formula includes:
[0039] The formula for calculating the rate of energy change is:
[0040]
[0041] Among them, E t is the sound signal energy at the current time t, E t-1 is the sound signal energy at the previous moment t-1, and ΔE is the energy change rate, which is used to measure the change of sound energy over time;
[0042] The formula for calculating the center of gravity of the spectrum is:
[0043]
[0044] Among them, k is the frequency index, and the spectrum center of gravity C reflects the distribution center of the sound signal energy on the frequency axis. Different sound events have different spectrum centers of gravity, which are used to distinguish the type of sound.
[0045] Preferably, in step S4, it specifically includes:
[0046] The principal component analysis was used to retain 25 principal components with a cumulative variance contribution rate of ≥ 96%. The dimensionality reduction formula is:
[0047] Z = XW;
[0048] Where Z is the data matrix after dimensionality reduction, which maps the original high-dimensional feature data to a low-dimensional space through principal component analysis; X is the original voiceprint feature vector matrix; W is the matrix of eigenvectors, which is composed of the eigenvectors of the covariance matrix of the original data and is used to project the original data into the new principal component space;
[0049] Combine recursive feature elimination with SVM classifier weights to filter the top 20 most discriminative features;
[0050] We built a Transformer-CNN hybrid deep learning model based on the attention mechanism. Its architecture consists of a 4-layer Transformer encoder, each layer of which includes an 8-head multi-head attention mechanism, layer normalization, and a feed-forward neural network; and 3 CNN convolutional blocks, each of which includes a 3×3 convolutional layer, a 2×2 max pooling layer, and a dropout rate of 0.3.
[0051] The AdamW optimizer was used with a learning rate of 0.0008 and a weight decay of 0.001, and the focus loss function was used for training. The focus loss function is:
[0052] FL(p t)=-α t (1-p t ) γ log(p t );
[0053] Where FL is the focus loss; p t The probability that the model predicts that the sample has abnormal aluminum liquid leakage; α t is the category weight, which is used to adjust the importance of samples of different categories in the loss function, and the weight of severe leakage samples is set to 2 to highlight the identification and training of severe abnormal situations; γ is the focusing parameter, which takes a value of 2 to increase the training weight of difficult-to-classify samples.
[0054] Preferably, in step S5, it specifically includes:
[0055] First, a sliding window size of 1.2s and an overlap rate of 60% are used to output a detection result every 0.48s. Then, features are extracted from the processed signal and input into the model of step S4 for detection, and similarity is calculated. Finally, the calculated similarity is compared with a preset graded response threshold, and the detection result is output.
[0056] The similarity is calculated as:
[0057]
[0058] Where x is the feature vector of the real-time monitored sound signal; y is the feature vector of the normal voiceprint; ||x|| and ||y|| are the norms of vectors x and y, respectively, which are used to normalize the vectors and eliminate the influence of vector constants on similarity calculation; (xy) T is the transpose of (xy); ∑ is the normal voiceprint feature covariance matrix, which is obtained by statistically analyzing a large amount of normal sound signal feature data.
[0059] Preferably, in step S6, it specifically includes:
[0060] If the detection result indicates a slight leak in the aluminum casting equipment, that is, the similarity is 0.5-0.75, a local sound and light alarm is triggered, and the abnormal location of the aluminum casting equipment is marked on the inter-factory monitoring screen. The response time is less than 8 seconds.
[0061] When the detection result is judged to be a moderate leakage of the aluminum casting equipment, that is, the similarity is 0.3-0.5, a remote alarm is issued and the aluminum casting equipment automatically stops feeding, and the response time is less than 15s;
[0062] When the detection result is judged to be a serious leak in the aluminum casting equipment, that is, when the similarity is less than 0.3, the aluminum casting equipment is shut down urgently, the valve is cut off and the cooling is started, and the response time is less than 5s.
[0063] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0064] (1) The present invention collects original sound data through a high-temperature resistant microphone array, combines three-level filtering noise reduction with amplitude normalization preprocessing, effectively filters out high-frequency mechanical noise, equipment vibration interference, etc. in the complex environment of the foundry workshop, unifies the signal amplitude standard, and provides pure, standardized, high-quality signals for subsequent voiceprint feature extraction, thereby improving the accuracy and reliability of signal processing.
[0065] (2) The present invention adopts variable window length short-time Fourier transform and extended Mel frequency cepstral coefficients, and calculates unique characteristics such as energy change rate and spectral center of gravity. It extracts voiceprint features from multiple dimensions and scales based on the complex time-frequency characteristics of aluminum liquid leakage sound in different frequency bands and different leakage stages, forms a complete voiceprint feature vector, and then comprehensively captures abnormal sound feature information, providing rich and effective data support for anomaly detection.
[0066] (3) The present invention constructs a Transformer-CNN hybrid deep learning model based on the attention mechanism. After principal component analysis and recursive feature elimination to optimize feature screening, the model significantly enhances the recognition accuracy and generalization ability of abnormal sounds in the casting industry and reduces the false alarm and missed alarm rate. At the same time, scientific graded response measures are formulated based on the detection results. For different degrees of aluminum liquid leakage anomalies, such as minor, moderate, and severe, the corresponding level of emergency response is automatically triggered, so as to achieve rapid positioning and accurate processing of abnormal situations, and greatly improve the intelligence level and processing efficiency of casting equipment abnormality detection and emergency processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0068] Figure 1 A flowchart of a method for detecting and responding to abnormal sounds in a casting industry based on voiceprint recognition is provided in the present invention;
[0069] Figure 2 This is a response block diagram provided for a method of detecting and responding to abnormal sounds in the foundry industry based on voiceprint recognition according to the present invention. DETAILED DESCRIPTION
[0070] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0071] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0072] Example
[0073] like Figure 1 As shown, the present invention provides a method for detecting and responding to abnormal sounds in the foundry industry based on voiceprint recognition, comprising the following steps:
[0074] S1. Arrange a high-temperature resistant microphone array at a leak-prone location of the aluminum casting equipment, and collect sound signals containing characteristics of aluminum liquid leakage through the microphone array to obtain original sound data;
[0075] S2, performing three-stage filtering noise reduction and amplitude normalization preprocessing on the original sound data to remove noise and unify the signal amplitude standard to obtain a preprocessed sound signal;
[0076] S3. Extract voiceprint features from the preprocessed sound signal through variable window length short-time Fourier transform and extended Mel-frequency cepstral coefficients, and calculate the energy change rate and spectral center of gravity as unique features to obtain a voiceprint feature vector;
[0077] S4. Screening the voiceprint feature vectors through principal component analysis and recursive feature elimination, and constructing and training a Transformer-CNN hybrid deep learning model based on the attention mechanism;
[0078] S5. Process the sound signal in real time using a sliding window, extract features from the processed signal and input them into the Transformer-CNN hybrid deep learning model based on the attention mechanism for detection. By calculating the similarity and comparing it with the preset graded response threshold, the abnormal sound is detected in real time and the detection result is output;
[0079] S6. According to the output detection results, the operating status of the aluminum casting equipment is determined, and emergency response measures of corresponding levels are triggered to complete the processing of the abnormal situation of aluminum liquid leakage.
[0080] In step S1, the microphone array adopts a four-element rectangular L-shaped microphone array, which is set around the casting port, mold seam and pipeline valve of the aluminum casting equipment to form a 90° fan-shaped coverage area; the microphone array is composed of high-temperature resistant 800°C, IP67-level dust and waterproof microphones, with a sensitivity of -32dB±2dB, a frequency response range of 20Hz~20kHz, and an installation height flush with the aluminum liquid flow plane, at a height of 1.2~1.5m; the microphone array is also equipped with a four-channel synchronous USB data acquisition card, which collects sound signals of aluminum liquid leakage characteristics at a 48kHz sampling rate and 24-bit quantization accuracy.
[0081] Therefore, through the L-shaped array layout and high-temperature, dust-proof and waterproof resistance, the microphone array can stably collect sound signals containing complete aluminum liquid leakage characteristics in a high-temperature 800°C, high-dust IP67 casting environment, avoiding signal loss caused by unreasonable equipment layout or environmental interference; the 48kHz high sampling rate and 24-bit quantization accuracy can capture the full-band sound signal of 20Hz to 20kHz, providing high-fidelity original data for subsequent feature analysis.
[0082] In step S2, the original sound data is subjected to three-level filtering, noise reduction, and amplitude normalization preprocessing to remove noise and unify the signal amplitude standard to obtain a preprocessed sound signal, which specifically includes:
[0083] First, the original sound data is subjected to three-level filtering and noise reduction, using a low-pass filter with a cutoff frequency of 3kHz to filter out the high-frequency mechanical noise of the equipment in the aluminum casting plant;
[0084] Secondly, a bandpass filter with a passband of 100-2000Hz is used to retain the characteristic frequency band of aluminum liquid leakage;
[0085] Thirdly, the improved Kalman filter is used to process the sound data of the characteristic frequency band of aluminum liquid leakage to remove noise.
[0086] Finally, the noise-removed and dried sound data is amplitude normalized to unify the signal amplitude standard.
[0087] The state equation of the improved Kalman filter is:
[0088]
[0089] Among them, x t =[s t ,v t ] represents the signal amplitude and rate of change, s t represents the amplitude of the sound signal at time t, v t Indicates the rate of change of the sound signal at time t; Δt = 10ms is the sampling interval, a is the acceleration noise, wt is the process noise;
[0090] The amplitude normalization process uses the formula:
[0091]
[0092] Where x is the amplitude of the original sound signal, that is, the unprocessed sound intensity value collected by the microphone array; x min and x max are the minimum and maximum values of the original sound signal amplitude respectively; x′ is the normalized sound signal amplitude, which maps the original amplitude to the interval [0.1, 0.9] to eliminate the differences in sound amplitudes under different acquisition devices and different working conditions.
[0093] In this embodiment, a three-stage filtering combination, namely low-pass + band-pass + improved Kalman filtering, can effectively filter out high-frequency mechanical noise generated by equipment such as die-casting machines and mixers, that is, noise above 3kHz. At the same time, it retains the characteristic frequency band of 100-2000Hz for aluminum liquid leakage. Combined with dynamic noise modeling, the signal-to-noise ratio (SNR) is improved by more than 14dB. At the same time, amplitude normalization eliminates the impact of sensitivity differences between different equipment and operating condition fluctuations on signal amplitude, ensuring the stability and consistency of subsequent feature extraction.
[0094] In step S3, a variable window length short-time Fourier transform is used, including:
[0095] In the low frequency band of 100-500Hz, a Hanning window with a window length of 30ms and a frame shift of 15ms is used; in the high frequency band of 500-2000Hz, a Blackman window with a window length of 10ms and a frame shift of 5ms is used to calculate the time-frequency matrix. The formula is:
[0096]
[0097] Where n is the time frame index, used to identify different time windows; k is the frequency index, used to determine different frequency components; N = 2048 is the number of FFT points; x(m+nΔt) is the sound signal sampling value at time (m+nΔt); w(m) is the window function; e j2πkm / N It is used to convert the time domain signal into the frequency domain signal, and realize the discrete Fourier transform through the exponential function to obtain the component distribution of the sound signal at different frequencies.
[0098] The extended Mel-frequency cepstral coefficients are then used, including:
[0099] 22 triangular filters covering 100-1800 Hz are used, and 12-dimensional MFCC coefficients, 12-dimensional first-order differences, and 12-dimensional second-order differences are extracted through logarithmic energy calculation and discrete cosine transform, for a total of 36 dimensions. The logarithmic energy calculation is:
[0100]
[0101] Where S(m) is the logarithmic energy, |X(k)| 2 represents the power spectral density at frequency k, H m (k) is the frequency response of the mth Mel filter at frequency k;
[0102] The discrete cosine transform is:
[0103]
[0104] Among them, c(n) is the discrete cosine transform, n is the index of the MFCC coefficient, and m is the index of the Mel filter; the logarithmic energy is then transferred from the Mel frequency domain to the cepstral domain.
[0105] Finally, the energy change rate and spectrum center of gravity are calculated using the following formulas:
[0106] The formula for calculating the rate of energy change is:
[0107]
[0108] Among them, E t is the sound signal energy at the current time t, E t-1 is the sound signal energy at the previous moment t-1, and ΔE is the energy change rate, which is used to measure the change of sound energy over time;
[0109] The formula for calculating the center of gravity of the spectrum is:
[0110]
[0111] Among them, k is the frequency index, and the spectrum center of gravity C reflects the distribution center of the sound signal energy on the frequency axis. Different sound events have different spectrum centers of gravity, which are used to distinguish the type of sound.
[0112] Through short-time Fourier transform with variable window length, the different frequency band characteristics of aluminum liquid leakage sound are targeted, and the 12-dimensional MFCC coefficients and the first and second order difference dynamic characteristics of the 36-dimensional spectrum envelope are extracted by combining the extended Mel frequency cepstrum coefficients. The energy change rate and spectrum center of gravity are calculated as unique indicators to form a multi-dimensional voiceprint feature vector including time-frequency distribution, energy mutation, and frequency center of gravity. It effectively captures the characteristic differences of different leakage levels (mild / medium / severe), improves the discrimination of abnormal sound characteristics by more than 25%, and provides effective features including static spectrum, dynamic changes and clear physical meaning for subsequent model training. The overall recognition accuracy has been increased to 97.6% through actual measurement, significantly enhancing the ability to characterize the sound of aluminum liquid leakage under complex working conditions.
[0113] In step S4, it specifically includes:
[0114] The principal component analysis was used to retain 25 principal components with a cumulative variance contribution rate of ≥ 96%. The dimensionality reduction formula is:
[0115] Z = XW;
[0116] Where Z is the data matrix after dimensionality reduction, which maps the original high-dimensional feature data to a low-dimensional space through principal component analysis; X is the original voiceprint feature vector matrix; W is the matrix of eigenvectors, which is composed of the eigenvectors of the covariance matrix of the original data and is used to project the original data into the new principal component space;
[0117] Combine recursive feature elimination with SVM classifier weights to filter the top 20 most discriminative features;
[0118] We built a Transformer-CNN hybrid deep learning model based on the attention mechanism. Its architecture consists of a 4-layer Transformer encoder, each layer of which includes an 8-head multi-head attention mechanism, layer normalization, and a feed-forward neural network; and 3 CNN convolutional blocks, each of which includes a 3×3 convolutional layer, a 2×2 max pooling layer, and a dropout rate of 0.3.
[0119] The AdamW optimizer was used with a learning rate of 0.0008 and a weight decay of 0.001, and the focus loss function was used for training. The focus loss function is:
[0120] FL(p t )=-α t (1-p t ) γ log(p t );
[0121] Where FL is the focus loss; p t The probability that the model predicts that the sample has abnormal aluminum liquid leakage; α t is the category weight, which is used to adjust the importance of samples of different categories in the loss function, and the weight of severe leakage samples is set to 2 to highlight the identification and training of severe abnormal situations; γ is the focusing parameter, which takes a value of 2 to increase the training weight of difficult-to-classify samples.
[0122] Specifically, principal component analysis (PCA) removes more than 96% of feature redundancy, reduces the 38-dimensional original feature dimension to 25 dimensions, and improves computational efficiency by 40%; the top 20 key features screened by recursive feature elimination (RFE) reduce the model input dimension by 47% while retaining 98% of the discriminant information; the Transformer-CNN hybrid architecture combines temporal dependency modeling with local feature extraction, enabling the model to achieve an accuracy of 97.6% in recognizing complex leakage sounds, an improvement of 15.3% compared to a single CNN model; the focal loss function improves the recognition rate of difficult-to-classify samples (such as minor leakage) from 75.2% to 92.1% through category weights and focusing parameters.
[0123] In step S5, it specifically includes:
[0124] First, a sliding window size of 1.2s and an overlap rate of 60% are used to output a detection result every 0.48s. Then, features are extracted from the processed signal and input into the model of step S4 for detection, and similarity is calculated. Finally, the calculated similarity is compared with a preset graded response threshold, and the detection result is output.
[0125] The similarity is calculated as:
[0126]
[0127] Where x is the feature vector of the real-time monitored sound signal; y is the feature vector of the normal voiceprint; ||x|| and ||y|| are the norms of vectors x and y, respectively, which are used to normalize the vectors and eliminate the influence of vector constants on similarity calculation; (xy) T is the transpose of (xy); ∑ is the normal voiceprint feature covariance matrix, which is obtained by statistically analyzing a large amount of normal sound signal feature data.
[0128] Specifically, a 1.2s sliding window enables real-time capture of leakage sounds as short as 1s, outputting detection results every 0.48s, shortening the response time by more than 60% compared with existing technologies; similarity calculation reduces the false positive rate of minor leaks to 3.2%, and the missed detection rate of serious leaks is controlled within 0.7%.
[0129] In addition, refer to Figure 2 In step S6, it specifically includes:
[0130] If the detection result indicates a slight leak in the aluminum casting equipment, that is, the similarity is 0.5-0.75, a local sound and light alarm is triggered, and the abnormal location of the aluminum casting equipment is marked on the inter-factory monitoring screen. The response time is less than 8 seconds.
[0131] When the detection result is judged to be a moderate leakage of the aluminum casting equipment, that is, the similarity is 0.3-0.5, a remote alarm is issued and the aluminum casting equipment automatically stops feeding, and the response time is less than 15s;
[0132] When the detection result is judged to be a serious leak in the aluminum casting equipment, that is, when the similarity is less than 0.3, the aluminum casting equipment is shut down urgently, the valve is cut off and the cooling is started, and the response time is less than 5s.
[0133] According to the specific implementation method provided above, in order to further verify the technical effect brought about by it, a comparative experiment was conducted by simulating the complex working conditions of a foundry workshop. Based on 1000 groups of measured samples containing different leakage levels, the performance of the technical solution of this embodiment and the existing technology in key indicators such as detection accuracy, response time and anti-noise ability were compared. The comparison results are shown in Table 1. Table 1 is:
[0134] Table 1 Performance index comparison results
[0135] Performance indicators This embodiment Existing technology Detection accuracy 97.6% 82.3% Minor leak recognition rate 92.1% 75.2% Serious leak identification rate 99.3% 88.5% Average response time (minor breach) <8s 20-30s Average response time (serious breach) <5s 10-15s Improved signal-to-noise ratio (SNR) 14dB 6dB Average number of false alarms per month 1 time 12 times
[0136] As shown in Table 1, based on 1,000 sets of measured samples containing different leakage levels, this embodiment significantly enhances the ability to recognize complex time-frequency features through variable window length feature extraction and a Transformer-CNN hybrid deep learning model, with particular advantages in minor and severe leakage scenarios. Furthermore, through real-time processing with a 1.2s sliding window and a graded response mechanism, this embodiment reduces the abnormal response time to 1 / 3 to 1 / 2 of the existing technology, meeting the timeliness requirements of emergency shutdowns of casting equipment. Furthermore, the three-stage filtering noise reduction provided by this embodiment increases the signal-to-noise ratio (SNR) to 14dB, more than double the 6dB of the existing technology, effectively suppressing high-frequency noise and equipment vibration interference from the die-casting machine. Furthermore, this embodiment uses a fusion of multi-dimensional features such as spectral center of gravity and energy change rate, combined with the optimization training of difficult-to-classify samples using a focal loss function, to reduce the average monthly number of false alarms from 12 to 1, significantly reducing manual verification costs.
[0137] Therefore, the above-mentioned method for abnormal sound detection and graded response in the foundry industry based on voiceprint recognition is adopted to solve the problems of strong subjectivity, low efficiency and poor safety of manual inspections in the existing abnormal sound detection in the foundry industry, single monitoring signals of traditional sensors and difficulty in identifying abnormalities, and poor signal preprocessing, incomplete feature extraction, weak model generalization ability, and lack of graded response in existing voiceprint recognition technology. It realizes the detection of abnormal sounds of aluminum liquid leakage in aluminum casting equipment and efficient graded emergency response.
[0138] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.
Claims
1. A method for abnormal sound detection and graded response in the foundry industry based on voiceprint recognition, characterized in that: The following steps are involved: S1. Arrange a high-temperature resistant microphone array at a leak-prone location of the aluminum casting equipment, and collect sound signals containing characteristics of aluminum liquid leakage through the microphone array to obtain original sound data; S2, performing three-stage filtering noise reduction and amplitude normalization preprocessing on the original sound data to remove noise and unify the signal amplitude standard to obtain a preprocessed sound signal; S3. Extract voiceprint features from the preprocessed sound signal by performing short-time Fourier transform with variable window length and expanding Mel-frequency cepstral coefficients, and calculate the energy change rate and spectral center of gravity as unique features to obtain a voiceprint feature vector; S4. Screening the voiceprint feature vectors through principal component analysis and recursive feature elimination, and constructing and training a Transformer-CNN hybrid deep learning model based on the attention mechanism; S5. Process the sound signal in real time using a sliding window, extract features from the processed signal and input them into the Transformer-CNN hybrid deep learning model based on the attention mechanism for detection. By calculating the similarity and comparing it with the preset graded response threshold, the abnormal sound is detected in real time and the detection result is output; S6. According to the output detection results, the operating status of the aluminum casting equipment is determined, and emergency response measures of corresponding levels are triggered to complete the processing of the abnormal situation of aluminum liquid leakage.
2. The method for abnormal sound detection and graded response in the foundry industry based on voiceprint recognition according to claim 1, characterized in that: In step S1, the microphone array is a four-element rectangular L-shaped microphone array, which is arranged around the casting port, mold joints and pipeline valves of the aluminum casting equipment to form a 90° fan-shaped coverage area; the microphone array is composed of high-temperature resistant 800°C, IP67-level dust and water-proof microphones, with a sensitivity of -32dB±2dB, a frequency response range of 20Hz~20kHz, and an installation height of 1.2~1.5m flush with the aluminum liquid flow plane; The microphone array is also equipped with a four-channel synchronous USB data acquisition card to collect sound signals of aluminum liquid leakage characteristics at a 48kHz sampling rate and 24-bit quantization accuracy.
3. The method for abnormal sound detection and graded response in the foundry industry based on voiceprint recognition according to claim 1 is characterized in that: In step S2, the original sound data is subjected to three-level filtering, noise reduction, and amplitude normalization preprocessing to remove noise and unify the signal amplitude standard to obtain a preprocessed sound signal, which specifically includes: First, the original sound data is subjected to three-level filtering and noise reduction, using a low-pass filter with a cutoff frequency of 3kHz to filter out the high-frequency mechanical noise of the equipment in the aluminum casting plant; Secondly, a bandpass filter with a passband of 100-2000Hz is used to retain the characteristic frequency band of aluminum liquid leakage; Thirdly, the improved Kalman filter is used to process the sound data of the characteristic frequency band of aluminum liquid leakage to remove noise. Finally, the noise-removed and dried sound data is amplitude normalized to unify the signal amplitude standard.
4. The method for abnormal sound detection and graded response in the foundry industry based on voiceprint recognition according to claim 3 is characterized in that: The state equation of the improved Kalman filter is: Among them, x t =[s t ,v t ] represents the signal amplitude and rate of change, s t represents the amplitude of the sound signal at time t, v t Indicates the rate of change of the sound signal at time t; Δt = 10ms is the sampling interval, a is the acceleration noise, w t is the process noise; The amplitude normalization process uses the formula: Where x is the amplitude of the original sound signal, that is, the unprocessed sound intensity value collected by the microphone array; x min and x max are the minimum and maximum values of the original sound signal amplitude respectively; x′ is the normalized sound signal amplitude, which maps the original amplitude to the interval [0.1, 0.9] to eliminate the differences in sound amplitudes under different acquisition devices and different working conditions.
5. The method for abnormal sound detection and graded response in the foundry industry based on voiceprint recognition according to claim 1, characterized in that: In step S3, a variable window length short-time Fourier transform is used, including: In the low frequency band of 100-500Hz, a Hanning window with a window length of 30ms and a frame shift of 15ms is used; in the high frequency band of 500-2000Hz, a Blackman window with a window length of 10ms and a frame shift of 5ms is used to calculate the time-frequency matrix. The formula is: Where n is the time frame index, used to identify different time windows; k is the frequency index, used to determine different frequency components; N = 2048 is the number of FFT points; x(m+nΔt) is the sound signal sampling value at time (m+nΔt); w(m) is the window function; e j2πkm / N It is used to convert the time domain signal into the frequency domain signal, and realize the discrete Fourier transform through the exponential function to obtain the component distribution of the sound signal at different frequencies.
6. The method for abnormal sound detection and graded response in the foundry industry based on voiceprint recognition according to claim 5 is characterized in that: In step S3, the extended Mel-frequency cepstral coefficients are used, including: 22 triangular filters covering 100-1800 Hz are used, and 12-dimensional MFCC coefficients, 12-dimensional first-order differences, and 12-dimensional second-order differences are extracted through logarithmic energy calculation and discrete cosine transform, for a total of 36 dimensions. The logarithmic energy calculation is: Where S(m) is the logarithmic energy, |X(k)| 2 represents the power spectral density at frequency k, H m (k) is the frequency response of the mth Mel filter at frequency k; The discrete cosine transform is: Among them, c(n) is the discrete cosine transform, n is the index of the MFCC coefficient, and m is the index of the Mel filter; the logarithmic energy is then transferred from the Mel frequency domain to the cepstral domain.
7. The method for abnormal sound detection and graded response in the foundry industry based on voiceprint recognition according to claim 6, characterized in that: In step S3, the energy change rate and the spectrum center of gravity are calculated using the following formula: The formula for calculating the rate of energy change is: Among them, E t is the sound signal energy at the current time t, E t-1 is the sound signal energy at the previous moment t-1, and ΔE is the energy change rate, which is used to measure the change of sound energy over time; The formula for calculating the center of gravity of the spectrum is: Among them, k is the frequency index, and the spectrum center of gravity C reflects the distribution center of the sound signal energy on the frequency axis. Different sound events have different spectrum centers of gravity, which are used to distinguish the type of sound.
8. The method for abnormal sound detection and graded response in the foundry industry based on voiceprint recognition according to claim 1 is characterized in that: In step S4, it specifically includes: The principal component analysis was used to retain 25 principal components with a cumulative variance contribution rate of ≥ 96%. The dimensionality reduction formula is: Z = XW; Where Z is the data matrix after dimensionality reduction, which maps the original high-dimensional feature data to a low-dimensional space through principal component analysis; X is the original voiceprint feature vector matrix; W is the matrix of eigenvectors, which is composed of the eigenvectors of the covariance matrix of the original data and is used to project the original data into the new principal component space; Combine recursive feature elimination with SVM classifier weights to filter the top 20 most discriminative features; We built a Transformer-CNN hybrid deep learning model based on the attention mechanism. Its architecture consists of a 4-layer Transformer encoder, each layer of which includes an 8-head multi-head attention mechanism, layer normalization, and a feed-forward neural network; and 3 CNN convolutional blocks, each of which includes a 3×3 convolutional layer, a 2×2 max pooling layer, and a dropout rate of 0.
3. The AdamW optimizer is used with a learning rate of 0.0008 and a weight decay of 0.001, and the focus loss function is used for training; the focus loss function is: FL(p t )=-a t (1-p t ) γ log(p t ); Where FL is the focus loss; p t The probability that the model predicts that the sample has abnormal aluminum liquid leakage; α t is the category weight, which is used to adjust the importance of samples of different categories in the loss function, and the weight of severe leakage samples is set to 2 to highlight the identification and training of severe abnormal situations; γ is the focusing parameter, which takes a value of 2 to increase the training weight of difficult-to-classify samples.
9. The method for abnormal sound detection and graded response in the foundry industry based on voiceprint recognition according to claim 1, characterized in that: In step S5, it specifically includes: First, a sliding window size of 1.2s and an overlap rate of 60% are used to output a detection result every 0.48s. Then, features are extracted from the processed signal and input into the model of step S4 for detection, and similarity is calculated. Finally, the calculated similarity is compared with a preset graded response threshold, and the detection result is output. The similarity is calculated as: Where x is the feature vector of the real-time monitored sound signal; y is the feature vector of the normal voiceprint; ||x|| and ||y|| are the norms of vectors x and y, respectively, which are used to normalize the vectors and eliminate the influence of vector constants on similarity calculation; (xy) T is the transpose of (xy); ∑ is the normal voiceprint feature covariance matrix, which is obtained by statistically analyzing a large amount of normal sound signal feature data.
10. The method for abnormal sound detection and graded response in the foundry industry based on voiceprint recognition according to claim 1, characterized in that: In step S6, it specifically includes: If the detection result indicates a slight leak in the aluminum casting equipment, that is, the similarity is 0.5-0.75, a local sound and light alarm is triggered, and the abnormal location of the aluminum casting equipment is marked on the inter-factory monitoring screen. The response time is less than 8 seconds. When the detection result is judged to be a moderate leakage of the aluminum casting equipment, that is, the similarity is 0.3-0.5, a remote alarm is issued and the aluminum casting equipment automatically stops feeding, and the response time is less than 15s; When the detection result is judged to be a serious leak in the aluminum casting equipment, that is, when the similarity is less than 0.3, the aluminum casting equipment is shut down urgently, the valve is cut off and the cooling is started, and the response time is less than 5s.
Citation Information
Cited By
Optical fiber sensing voiceprint feature analysis model construction method based on composite neural network
CN120636468A
Power grid digital security risk dynamic early warning system based on deep learning algorithm
CN120932432A
Photovoltaic inverter fault processing method and device based on acoustic characteristics
CN121171259A
High-pressure roller mill sound and vibration abnormity throughout evolution discrimination method
CN121435093A
Evolutionary discrimination method for abnormal precursors of acoustic vibration of high-pressure roller mill
CN121435093B