Device defect diagnosis method and system based on voiceprint features, device and medium
By combining the noise-resistant microphone array and MVDR algorithm with wavelet packet decomposition, multi-scale Mel-frequency cepstral coefficients and a dual-modal reference voiceprint library, the problem of defect diagnosis of industrial equipment in complex noise environments is solved, and efficient and accurate equipment fault detection is achieved.
Patent Information
- Application Number
- CN202510929417.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing industrial equipment defect diagnosis methods and technologies have difficulty in accurately extracting fault features in complex noisy environments and cannot meet the accuracy and reliability requirements for equipment defect diagnosis in industrial production.
An anti-noise microphone array and MVDR algorithm are used for sound data collection and noise reduction processing. Sub-band weighted wavelet packet decomposition and multi-scale Mel-cepstral coefficient construction are combined to establish a dual-modal benchmark voiceprint library. The voiceprint features are optimized through triplet loss, and a temporal convolutional network is used for defect diagnosis decisions.
It effectively suppresses noise interference in complex noise environments, accurately extracts fault features from industrial equipment sounds, improves the accuracy and adaptability of defect diagnosis, and provides high-quality defect diagnosis information.
Smart Images

Figure CN120748441A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial equipment fault diagnosis, and specifically to a method, system, equipment and medium for equipment defect diagnosis based on voiceprint features. Background Art
[0002] In industrial production, the safe and stable operation of equipment is crucial for ensuring production efficiency, product quality, and safety. Traditional quality inspection methods for industrial equipment rely primarily on manual audio inspection or vibration sensors. Manual audio inspection relies on the operator's experience to determine whether defects exist by listening to the operating sounds of the equipment. However, the human ear has a limited hearing range and struggles to distinguish subtle differences between high- and low-frequency sound patterns. This makes it difficult to accurately capture the sound signatures of equipment failures hidden within complex noise, leading to frequent missed and false detections.
[0003] Vibration sensors analyze equipment status by detecting vibration signals. However, vibration sensor installation is complex and requires precise placement at specific locations on the equipment. This not only increases maintenance workload and costs but is also susceptible to interference from the equipment's mechanical structure. For example, the complex vibration transmission path of equipment can cause vibration signals to attenuate and distort during transmission, hindering accurate identification of equipment defects.
[0004] Existing voiceprint recognition technologies, such as Shazam, are primarily optimized for music recognition. These technologies perform well in music scenarios because music signals have relatively stable spectral characteristics and rhythmic patterns. However, in industrial scenarios, equipment operates in complex environments with strong noise interference and severe frequency aliasing of sound signals generated by different devices. Existing voiceprint recognition technologies are not effectively adapted to such environments, making it difficult to accurately extract fault characteristics from the sounds of industrial equipment and thus cannot be directly applied to industrial equipment fault diagnosis.
[0005] Therefore, the existing industrial equipment defect diagnosis methods and technologies have many shortcomings and cannot meet the requirements of accuracy and reliability of equipment defect diagnosis in industrial production. There is an urgent need for a new method and technology that can adapt to complex industrial environments and accurately diagnose equipment defects. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a technical solution: a device defect diagnosis method based on voiceprint features, comprising:
[0007] S1. Obtain device operation sound data;
[0008] S2. Decompose the sound data of the equipment through sub-band weighted wavelet packets and extract industrial voiceprint feature data in layers;
[0009] S3. Based on the industrial voiceprint feature data, multi-scale Mel-frequency cepstral coefficients are constructed on the basis of standard MFCC;
[0010] S4. Establish a dual-modal baseline voiceprint library and reduce the distance between similar defective voiceprints through triplet loss.
[0011] Preferably, S1 collects the device running sound data through an anti-noise microphone array, and the specific steps include:
[0012] S1.1. Microphone selection and deployment: Choose a linear array type suitable for long strip equipment, or a circular array type suitable for central equipment;
[0013] S1.2. Use the MVDR algorithm to implement microphone array noise reduction, ensuring that the target direction signal is distortion-free and minimizing the total power of the output signal.
[0014] Preferably, S1.2 further includes: S1.2.1, setting the target signal direction to θ0, the steering vector to a(θ0), and the received signal vector to x(t)=[x1(t), x2(t), ..., x M (t)] T , the weight vector is w, and the signal model is x(t)=s(t)a(θ0)+n(t), where (s(t) is the target sound source and n(t) is the noise;
[0015] S1.2.2, the optimization weight vector is w, the constraint target direction gain is 1, the constraint formula is: w H a(θ0)=I, and the optimized minimum output power formula is obtained: Among them, R xx =E[x(t)x H (t)] is the covariance matrix of the received signal;
[0016] S1.2.3. The closed-form solution of weights is solved by the Lagrange multiplier method to obtain the optimal weights:
[0017] Preferably, the specific steps of S2 include:
[0018] S2.1, wavelet packet decomposition;
[0019] S2.2. Sub-band feature extraction and weighting, energy ratio of unbalanced working condition Obtain high energy fault data, bearing spalling condition using spectral kurtosis Get the impact signal through envelope entropy Obtain periodic modulation fault data, and the correlation coefficient is obtained by w k =||corr(d k , s ref )||match reference template;
[0020] S2.3. Reconstruct the target signal by weighting: Among them, WPD -1 is the inverse wavelet packet transform, w k For adaptive weights, noise suppression reconstruction:
[0021] Preferably, S2.1 further includes:
[0022] S2.1.1. Select the wavelet basis function. For bearing faults, select db10. For gear impact, select sym8. The number of decomposition levels, j, is determined by the target frequency resolution: f s Indicates the sampling rate;
[0023] S2.1.2, Decomposition process signal x(t) passes through low-pass filter h and high-pass filter g. The coefficient of the k-th node in the j-th layer is:
[0024] Preferably, the specific steps of S3 include:
[0025] S3.1, multi-scale preprocessing;
[0026] S3.1.1. Set scale 1, window length 10-20 ms, frame shift 5-10 ms, and frequency resolution using a wideband 50 filter.
[0027] S3.1.2. Set scale 2, window length 25-40 ms, frame shift 10-20 ms, and frequency resolution using a mid-band 80 filter.
[0028] S3.1.3. Set scale to 3, window length to 50-100 ms, frame shift to 25-50 ms, and frequency resolution to a narrowband 120 filter.
[0029] S3.2. Extract multi-scale MFCCs, retain coefficients of different orders, align time and perform feature fusion.
[0030] Preferably, S4 specifically includes:
[0031] S4.1. Establish a dual-modal baseline voiceprint database;
[0032] S4.1.1. Set up the static library, the voiceprint template of the static library device in normal state, the frequency mean μ f , energy variance
[0033] S4.1.2. Set up a dynamic library to update the voiceprint feature sliding window of qualified products in real time. The window size is adaptive to the production line speed. Use transfer contrast learning and the YAMNet pre-trained model to extract deep features. Use triple loss to narrow the distance between similar defect voiceprints.
[0034] S4.2. Make defect diagnosis decisions using the reference voiceprint library;
[0035] S4.2.1. Real-time calculation of the multi-dimensional deviation between the input voiceprint and the reference library: Among them, MS-MFCC input and MS-MFCC ref denotes the multi-scale MFCC features of input and reference respectively, ||·||2 denotes the L2 norm, R norm is the normalization factor, HDF input and HDF ref represents the harmonic difference characteristics of the input and reference, α and β are weight coefficients;
[0036] S4.2.2. When D>θ1, deep matching is triggered, and a temporal convolutional network is used to detect transient event sequence patterns. The historical defect voiceprint path is matched through dynamic time warping, and the defect type code and confidence are output.
[0037] Preferably, a device defect diagnosis system based on voiceprint features is also provided, including:
[0038] The data acquisition module is responsible for collecting sound data from industrial equipment operation. It uses the MVDR algorithm to achieve microphone array noise reduction. It sets parameters such as the target signal direction, steering vector, received signal vector, and weight vector. By optimizing the weight vector, it minimizes the total power of the output signal while ensuring that the target direction signal is distortion-free, thereby obtaining high-quality sound data.
[0039] The feature extraction module extracts industrial voiceprint feature data from the collected sound data in layers, selects appropriate wavelet basis functions, determines the number of decomposition layers based on the target frequency resolution, and decomposes the sound signal through low-pass and high-pass filters to obtain node coefficients at different levels. For unbalanced working conditions, the energy ratio is used to obtain strong energy fault data; for bearing spalling conditions, the spectral kurtosis is used to obtain the impact signal; and the envelope entropy is used to obtain periodic modulation fault data. The reference template is matched using the correlation coefficient, and adaptive weights are calculated based on the extracted features to perform weighted reconstruction of the target signal. At the same time, noise suppression reconstruction is performed to highlight the fault characteristics and suppress noise interference.
[0040] The voiceprint feature construction module constructs multi-scale Mel-frequency cepstral coefficients based on standard MFCCs to more comprehensively describe the device sound characteristics. It sets the window length, frame shift, and frequency resolution parameters for scales 1, 2, and 3, extracts multi-scale MFCCs, extracts MFCC features at each scale, retains coefficients of different orders, and aligns time for feature fusion to obtain multi-scale MFCC features, thereby more comprehensively reflecting the time-frequency characteristics of the device sound.
[0041] The baseline voiceprint library management module establishes and manages a dual-modal baseline voiceprint library to provide a reference for defect diagnosis. The static library stores voiceprint templates under normal equipment conditions, recording frequency mean and energy variance information as a standard reference for normal equipment operation. The dynamic library updates the voiceprint feature sliding window of qualified products in real time. The window size is adaptively adjusted according to the production line speed. Transfer contrast learning is adopted, and the YAMNet pre-trained model is used to extract deep features. The distance between voiceprints of similar defects is narrowed through triple loss. The voiceprint features in the dynamic library are continuously optimized so that they can adapt to changes in equipment operating status.
[0042] The defect diagnosis decision module makes defect diagnosis decisions on the real-time collected sound data based on the benchmark voiceprint library, outputs the defect type code and confidence, and calculates the multi-dimensional deviation between the input voiceprint and the static and dynamic libraries in the benchmark library in real time, including the L2 norm deviation of the multi-scale MFCC features and the deviation of the harmonic difference features. A comprehensive calculation is performed through normalization factors and weight coefficients to quantify the degree of difference between the input voiceprint and the benchmark voiceprint. When the multi-dimensional deviation exceeds the preset threshold, deep matching is triggered, and a temporal convolutional network is used to detect transient event sequence patterns. The historical defect voiceprint path is matched through dynamic time warping, and the defect type code and confidence are output based on the matching results to provide accurate defect diagnosis information for equipment maintenance personnel.
[0043] Preferably, a device is also provided, comprising one or more processors, one or more memories, and one or more computer program instructions, wherein when the computer program instructions are executed by the processor, any one of the above-mentioned device defect diagnosis methods based on voiceprint features is implemented.
[0044] Preferably, a medium is also provided, storing a computer program, which, when executed by a processor, implements any of the above-mentioned device defect diagnosis methods based on voiceprint features.
[0045] Compared with the existing technology, the advantages of the present invention are: (1) Optimization of industrial strong noise environment: The use of anti-noise microphone array and MVDR algorithm for sound data collection and noise reduction processing can effectively suppress strong noise interference in the industrial environment and ensure the quality of the collected sound data; (2) Accurate feature extraction: Through sub-band weighted wavelet packet decomposition and multiple feature extraction methods, corresponding fault feature data is extracted for different working conditions, and weighted reconstruction and noise suppression reconstruction are performed, which can more accurately extract fault features in the sound of industrial equipment, thereby improving the accuracy and reliability of feature extraction; (3) Multi-scale voiceprint feature construction: Multi-scale Mel-frequency cepstral coefficients are constructed based on the standard MFCC to comprehensively describe the equipment sound characteristics from different scales, which can better reflect the time-frequency characteristics of the equipment sound and make up for the deficiency of existing voiceprint recognition technology in adapting to multi-band aliasing environments in industrial scenarios; (4) Dual-modal reference voiceprint library: A dual-modal reference voiceprint library is established, including a static library and a dynamic library. The static library stores the voiceprint templates of the equipment under normal conditions, and the dynamic library updates the voiceprint feature sliding window of qualified products in real time, and uses transfer contrast learning and triple loss to optimize the voiceprint features, so that the baseline voiceprint library can adapt to changes in the equipment's operating status, thereby improving the accuracy and adaptability of defect diagnosis; (5) Accurate defect diagnosis decision: By calculating the multi-dimensional deviation between the input voiceprint and the baseline library in real time, and triggering deep matching when the deviation exceeds the threshold, using temporal convolutional networks and dynamic time warping technology for defect diagnosis, it can accurately output defect type codes and confidence levels, provide accurate defect diagnosis information for equipment maintenance personnel, and improve the accuracy and reliability of industrial equipment defect diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to make the contents of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings.
[0047] Figure 1 It is a schematic flow diagram of step S1 of the present invention.
[0048] Figure 2 It is a schematic flow diagram of step S2 of the present invention.
[0049] Figure 3 It is a schematic flow diagram of step S3 of the present invention. DETAILED DESCRIPTION
[0050] Example 1
[0051] like Figures 1 to 3As shown, this embodiment provides a specific implementation method of S1. Multiple microphones are arranged in a specific geometric shape, such as linear, annular, spherical, etc. The MVDR algorithm calculates the phase difference of each microphone signal in real time to form a directional "sound beam" pointing to the device, suppressing noise in other directions, such as environmental human voices and fan sounds. If the device is directly in front of the array, the algorithm will enhance the signal in the 0° direction and attenuate the interference in the 90° direction. Assuming that the device sound source is statistically independent of the noise, the microphone can extract the target signal through mathematical decomposition. In addition, the microphone can also use multi-channel noise reduction technology to separate non-correlated sound sources by comparing the signal differences through the correlation of noise on multiple microphones, such as uniform distribution of environmental noise.
[0052] Microphone array types include linear arrays and circular arrays. Linear arrays are suitable for long equipment such as pipelines and conveyor belts, and are also suitable for monitoring motor bearings. Circular arrays use 4-8 microphones to achieve 360-degree coverage of the equipment and are suitable for central equipment such as machine tools and generators. Regardless of the method, the microphones must be deployed separately near the noise source and dedicated to collecting background noise templates. They must have a signal-to-noise ratio of >65dB and a frequency response that covers the equipment's characteristic frequencies, such as 1-10kHz for bearing faults. They must also be waterproof and shockproof with an IP67 rating. The microphones must be within 0.5-2 meters of the equipment to avoid obstructions. For example, when collecting fan noise, the array axis should be aligned with the impeller plane.
[0053] The microphones are integrated with a matching signal acquisition system. Its synchronous acquisition card ensures time alignment of all microphone signals with an error of less than 0.1ms. The sampling rate must be at least twice the target maximum frequency. For example, monitoring 20kHz ultrasonic waves requires a 40kHz sampling rate. Automatic gain prevents clipping, especially for impact sounds such as gear tooth breakage. The system first locates the sound source, determining the device direction based on the arrival time difference. Beamforming then generates a directional beam in real time. Finally, filtering is performed, using a Wiener filter to further suppress residual noise.
[0054] Covariance matrix R xx - Actual estimate: N is the number of time domain sampling points, which must satisfy N>M to avoid ill-conditioned matrices. The steering vector a(θ0) depends on the uniform linear array: in, d is the distance between microphones, c is the speed of sound, and the output signal is: In actual use of the MVDR algorithm, data correction is required to prevent matrix ill-conditioning by diagonal loading: R xx ←R xx +δI, δ is a small constant, δ=0.01·tr(R xx ).
[0055] Process the frequency domain signal in frequency bands: Adaptive update, recursively update the covariance matrix in a time-varying noise environment: λ is the forgetting factor, λ∈(0.95-0.99).
[0056] Example 2
[0057] like Figure 2 As shown, this embodiment optimizes S2, adaptively updates weights, and updates weights with a sliding window to cope with time-varying signals: α smoothing factor, set to 0.9.
[0058] In terms of frequency band fusion strategy, only 3-5 sub-bands with obvious characteristics are weighted to reduce the amount of calculation and merge adjacent related frequency bands, such as gear meshing frequency and its sidebands.
[0059] Example 3
[0060] like Figure 1 、 Figure 2 and Figure 3 As shown, this embodiment provides a specific implementation of the device defect diagnosis method based on voiceprint features:
[0061] S1. Select an appropriate noise-canceling microphone array and deploy it based on the device type. For example, for long, rectangular devices, choose a linear array; for central devices, choose a circular array. Set parameters such as the target signal direction, steering vector, received signal vector, and weight vector. Optimize the weight vector using the MVDR algorithm to minimize the total output signal power while ensuring undistorted signals in the target direction. This allows you to collect device operating sound data.
[0062] S2. Select an appropriate wavelet basis function, such as db10 for bearing faults and sym8 for gear impact. Determine the number of decomposition layers based on the target frequency resolution and decompose the sound signal through low-pass and high-pass filters. For different operating conditions, such as using energy ratio to obtain high-energy fault data for unbalanced conditions and spectral kurtosis to obtain impact signals for bearing spalling conditions, adaptive weights are calculated to perform weighted reconstruction and noise suppression reconstruction of the target signal.
[0063] S3. Set the window length, frame shift, and frequency resolution parameters for scale 1, scale 2, and scale 3. Extract MFCC features at each scale, retain coefficients of different orders, and align them in time for feature fusion to obtain multi-scale MFCC features.
[0064] S4 implementation. Establish a static library to store the voiceprint templates of the device in normal state, and record the frequency mean and energy variance information. Establish a dynamic library to update the voiceprint feature sliding window of qualified products in real time, and the window size is adaptively adjusted according to the production line speed. Adopt transfer contrast learning, use the YAMNet pre-training model to extract deep features, and reduce the distance between similar defective voiceprints through triple loss. Calculate the multi-dimensional deviation between the input voiceprint and the benchmark library in real time. When the deviation exceeds the preset threshold, use the temporal convolutional network to detect the transient event sequence pattern, match the historical defect voiceprint path through dynamic time warping, and output the defect type code and confidence.
[0065] In addition, this embodiment also provides a supporting device defect diagnosis system based on voiceprint features, which includes:
[0066] The data acquisition module collects the device running sound data through the anti-noise microphone array and MVDR algorithm according to the implementation of S1 above, and transmits the data to the subsequent modules.
[0067] The feature extraction module receives the sound data transmitted by the data acquisition module, extracts the industrial voiceprint feature data in layers according to the implementation of S2, and transmits the extracted feature data to the voiceprint feature construction module.
[0068] The voiceprint feature construction module receives the feature data transmitted by the feature extraction module, constructs multi-scale Mel-frequency cepstral coefficients according to the implementation of S3, and transmits the constructed voiceprint feature data to the reference voiceprint library management module.
[0069] The reference voiceprint library management module establishes and manages the dual-modal reference voiceprint library, stores and manages voiceprint templates and feature data according to the implementation of S4, receives the voiceprint feature data transmitted by the voiceprint feature construction module, and updates the voiceprint features in the dynamic library.
[0070] The defect diagnosis decision module receives the sound data collected by the data acquisition module in real time, calculates the multi-dimensional deviation between the input voiceprint and the reference library according to the implementation method of S4, makes a defect diagnosis decision, and outputs the defect type code and confidence level
[0071] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.
Claims
1. The device defect diagnosis method based on voiceprint features is characterized by include: S1. Obtain device operation sound data; S2. Decompose the sound data of the equipment through sub-band weighted wavelet packets and extract industrial voiceprint feature data in layers; S3. Based on the industrial voiceprint feature data, multi-scale Mel-frequency cepstral coefficients are constructed on the basis of standard MFCC; S4. Establish a dual-modal baseline voiceprint library and reduce the distance between similar defective voiceprints through triplet loss.
2. The device defect diagnosis method based on voiceprint features according to claim 1, characterized in that: S1 collects device operation sound data through an anti-noise microphone array. The specific steps include: S1.
1. Microphone selection and deployment: Choose a linear array type suitable for long strip equipment, or a circular array type suitable for central equipment; S1.
2. Use the MVDR algorithm to implement microphone array noise reduction, ensuring that the target direction signal is distortion-free and minimizing the total power of the output signal.
3. The method for diagnosing industrial equipment defects according to claim 2, characterized in that Said S1.2 also includes: S1.2.1, setting the target signal direction to θ0, the steering vector to a(θ0), the received signal vector to x(t)=[x1(t), x2(t), ..., x M (t)] T , the weight vector is w, and the signal model is x(t)=s(t)a(θ0)+n(t), where (s(t) is the target sound source and n(t) is the noise; S1.2.2, the optimization weight vector is w, the constraint target direction gain is 1, the constraint formula is: w H a(θ0)=1, and the optimized minimum output power formula is obtained: Among them, R xx =E[x(t)x H (t)] is the covariance matrix of the received signal; S1.2.
3. The closed-form solution of weights is solved by the Lagrange multiplier method to obtain the optimal weights:
4. The method for diagnosing industrial equipment defects according to claim 1, characterized in that The specific steps of S2 include: S2.1, wavelet packet decomposition; S2.
2. Sub-band feature extraction and weighting, energy ratio of unbalanced working condition Obtain high energy fault data, bearing spalling condition using spectral kurtosis Get the impact signal through envelope entropy Obtain periodic modulation fault data, and the correlation coefficient is obtained by w k =||corr(d k ,s ref )||match reference template; S2.
3. Reconstruct the target signal by weighting: Among them, WPD -1 is the inverse wavelet packet transform, w k For adaptive weights, noise suppression reconstruction:
5. The method for diagnosing defects in industrial equipment according to claim 4, characterized in that Said S2.1 also includes: S2.1.
1. Select the wavelet basis function. For bearing faults, select db10. For gear impact, select sym8. The number of decomposition levels, j, is determined by the target frequency resolution: f s Indicates the sampling rate; S2.1.2, Decomposition process signal x(t) passes through low-pass filter h and high-pass filter g. The coefficient of the k-th node in the j-th layer is:
6. The method for diagnosing defects in industrial equipment according to claim 1, characterized in that The specific steps of S3 include: S3.1, multi-scale preprocessing; S3.1.
1. Set scale 1, window length 10-20 ms, frame shift 5-10 ms, and frequency resolution using a wideband 50 filter. S3.1.
2. Set scale 2, window length 25-40 ms, frame shift 10-20 ms, and frequency resolution using a mid-band 80 filter. S3.1.
3. Set scale to 3, window length to 50-100 ms, frame shift to 25-50 ms, and frequency resolution to a narrowband 120 filter. S3.
2. Extract multi-scale MFCCs, retain coefficients of different orders, align time and perform feature fusion.
7. The method for diagnosing defects in industrial equipment according to claim 1, characterized in that The S4 specifically includes: S4.
1. Establish a dual-modal baseline voiceprint database; S4.1.
1. Set up the static library, the voiceprint template of the static library device in normal state, the frequency mean μ f , energy variance S4.1.
2. Set up a dynamic library to update the voiceprint feature sliding window of qualified products in real time. The window size is adaptive to the production line speed. Use transfer contrast learning and the YAMNet pre-trained model to extract deep features. Use triple loss to narrow the distance between similar defect voiceprints. S4.
2. Make defect diagnosis decisions using the reference voiceprint library; S4.2.
1. Real-time calculation of the multi-dimensional deviation between the input voiceprint and the reference library: Among them, MS-MFCC input and MS-MFCC ref denotes the multi-scale MFCC features of input and reference respectively, ||·||2 denotes the L2 norm, R norm is the normalization factor, HDF input and HDF raf represents the harmonic difference characteristics of the input and reference, α and β are weight coefficients; S4.2.
2. When D>θ1, deep matching is triggered, and a temporal convolutional network is used to detect transient event sequence patterns. The historical defect voiceprint path is matched through dynamic time warping, and the defect type code and confidence are output.
8. The equipment defect diagnosis system based on voiceprint features is characterized by include: The data acquisition module is responsible for collecting sound data from industrial equipment operation. It uses the MVDR algorithm to achieve microphone array noise reduction. It sets parameters such as the target signal direction, steering vector, received signal vector, and weight vector. By optimizing the weight vector, it minimizes the total power of the output signal while ensuring that the target direction signal is distortion-free, thereby obtaining high-quality sound data. The feature extraction module extracts industrial voiceprint feature data from the collected sound data in layers, selects appropriate wavelet basis functions, determines the number of decomposition layers based on the target frequency resolution, and decomposes the sound signal through low-pass and high-pass filters to obtain node coefficients at different levels. For unbalanced working conditions, the energy ratio is used to obtain strong energy fault data; for bearing spalling conditions, the spectral kurtosis is used to obtain the impact signal; and the envelope entropy is used to obtain periodic modulation fault data. The reference template is matched using the correlation coefficient, and adaptive weights are calculated based on the extracted features to perform weighted reconstruction of the target signal. At the same time, noise suppression reconstruction is performed to highlight the fault features and suppress noise interference. The voiceprint feature construction module constructs multi-scale Mel-frequency cepstral coefficients based on standard MFCCs to more comprehensively describe the device sound characteristics. It sets the window length, frame shift, and frequency resolution parameters for scales 1, 2, and 3, extracts multi-scale MFCCs, extracts MFCC features at each scale, retains coefficients of different orders, and aligns time for feature fusion to obtain multi-scale MFCC features, thereby more comprehensively reflecting the time-frequency characteristics of the device sound. The baseline voiceprint library management module establishes and manages a dual-modal baseline voiceprint library to provide a reference for defect diagnosis. The static library stores voiceprint templates under normal equipment conditions, recording frequency mean and energy variance information as a standard reference for normal equipment operation. The dynamic library updates the voiceprint feature sliding window of qualified products in real time. The window size is adaptively adjusted according to the production line speed. Transfer contrast learning is adopted, and the YAMNet pre-trained model is used to extract deep features. The distance between voiceprints of similar defects is narrowed through triple loss. The voiceprint features in the dynamic library are continuously optimized so that they can adapt to changes in equipment operating status. The defect diagnosis decision module makes defect diagnosis decisions on the real-time collected sound data based on the benchmark voiceprint library, outputs the defect type code and confidence, and calculates the multi-dimensional deviation between the input voiceprint and the static and dynamic libraries in the benchmark library in real time, including the L2 norm deviation of the multi-scale MFCC features and the deviation of the harmonic difference features. A comprehensive calculation is performed through normalization factors and weight coefficients to quantify the degree of difference between the input voiceprint and the benchmark voiceprint. When the multi-dimensional deviation exceeds the preset threshold, deep matching is triggered, and a temporal convolutional network is used to detect transient event sequence patterns. The historical defect voiceprint path is matched through dynamic time warping, and the defect type code and confidence are output based on the matching results to provide accurate defect diagnosis information for equipment maintenance personnel.
9. A device, characterized in that The device includes one or more processors, one or more memories, and one or more computer program instructions. When the computer program instructions are executed by the processor, the device defect diagnosis method based on voiceprint features as described in any one of claims 1 to 7 is implemented.
10. A medium storing a computer program, characterized in that: When the computer program is executed by a processor, the device defect diagnosis method based on voiceprint features according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Equipment fault diagnosis method and device based on AI large model, equipment and medium
CN121211288A
Telegraph pole defect detection system and method based on voiceprint characteristics
CN121347674A
Fault diagnosis method and system for control valve
CN121632573A