Motor fault diagnosis system and method based on vibration and voiceprint synchronous analysis

The motor fault diagnosis system, which uses synchronous analysis of vibration and acoustic signature, solves the problems of signal interference and spatiotemporal alignment difficulties in traditional motor fault diagnosis, and achieves high-accuracy and interpretable fault identification.

CN121899641APending Publication Date: 2026-04-21SHENZHEN ZHAOXIN MICROELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN ZHAOXIN MICROELECTRONICS CO LTD
Filing Date
2025-12-30
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In traditional motor fault diagnosis techniques, single-mode signal analysis is susceptible to noise and harmonic interference, and spatiotemporal alignment of multi-source signals is difficult, resulting in low fault identification accuracy, poor robustness, and a lack of physical interpretability in the diagnostic results.

Method used

A motor fault diagnosis system based on synchronous analysis of vibration and acoustic signature is adopted. The system acquires signals synchronously through a data acquisition and processing unit, extracts feature vectors using structural dynamic deformation field method and time domain statistical analysis, performs weighted fusion of features through a feature fusion unit, and classifies and identifies faults using a long short-term memory network.

Benefits of technology

It achieves accurate alignment of multimodal signals in low signal-to-noise ratio and strong interference environments, improves fault type discrimination and spatial positioning capabilities, and enhances fault identification accuracy and physical interpretability of diagnostic results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121899641A_ABST
    Figure CN121899641A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of motor fault diagnosis, in particular to a motor fault diagnosis system and method based on vibration and voiceprint synchronous analysis. The method comprises the following steps: a data acquisition and processing unit synchronously acquires a vibration signal and a voiceprint signal when a motor runs, and preprocesses the vibration signal and the voiceprint signal; the feature extraction unit extracts a vibration feature vector by using a structural dynamic deformation field method based on vibration signal inversion, and extracts a voiceprint feature vector based on time domain statistical analysis; and the feature fusion unit performs feature fusion on the vibration feature vector and the voiceprint feature vector by using a weighted fusion method to generate a fused comprehensive feature. In the feature extraction process, traditional time domain, frequency domain and time-frequency domain vibration features are extracted, and a signal transfer entropy matrix is innovatively introduced to describe the energy flow and dynamic coupling relationship among multiple sensors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of motor fault diagnosis technology, and more specifically, to a motor fault diagnosis system and method based on synchronous analysis of vibration and acoustic signature. Background Technology

[0002] In the field of industrial equipment condition monitoring and fault diagnosis, motors, as core power units, directly affect production safety and efficiency through their operational reliability. Traditional motor fault diagnosis technologies primarily rely on the analysis of single vibration or acoustic signals. However, while vibration signals are sensitive to mechanical faults, they are limited by contact measurement methods, easily affected by sensor installation locations, and insufficiently respond to early, weak faults and electromagnetic anomalies. Acoustic signals, while possessing the advantages of non-contact and wide-area sensing, are highly susceptible to interference from industrial environmental noise, reverberation echoes, and inherent periodic electromagnetic and mechanical harmonics during motor operation, leading to the masking of effective fault characteristics and a low signal-to-noise ratio. Furthermore, existing technologies generally suffer from problems in multi-source signal fusion, such as asynchronous acquisition of vibration and acoustic signals, uncompensated acoustic propagation delays, and lack of spatiotemporal correlation between sensors. They often employ simple feature splicing or average weighting strategies, making it difficult to achieve accurate alignment and effective fusion of these two heterogeneous signals at the temporal and event levels, thus limiting the potential of multimodal information collaborative diagnosis. Furthermore, traditional methods often rely on black-box models for fault classification, lacking physical mechanism modeling of the dynamic propagation and spatial distribution of faults on the motor structure. This results in poor interpretability of diagnostic results, making it difficult to support fault location and tracing. Therefore, this paper proposes a motor fault diagnosis system and method based on simultaneous vibration and acoustic signature analysis. Summary of the Invention

[0003] The purpose of this invention is to provide a motor fault diagnosis system and method based on synchronous analysis of vibration and acoustic signature, so as to solve the problems of low fault identification accuracy, poor robustness and lack of physical interpretability of diagnosis results in traditional motor fault diagnosis, which are caused by the one-sidedness of single modal information, the severe interference of noise and harmonics on acoustic signature signals, the difficulty of spatiotemporal alignment of multi-source signals and the crude fusion mechanism.

[0004] To achieve the above objectives, on the one hand, the present invention aims to provide a motor fault diagnosis system based on synchronous analysis of vibration and acoustic signatures, comprising: The data acquisition and processing unit synchronously acquires vibration signals and acoustic fingerprint signals during motor operation and preprocesses the vibration signals and acoustic fingerprint signals. The feature extraction unit extracts vibration feature vectors using a structural dynamic deformation field method based on vibration signal inversion and extracts acoustic feature vectors based on time-domain statistical analysis. In the process of extracting voiceprint feature vectors, the preprocessed voiceprint signal is subjected to echo and harmonic suppression preprocessing. The feature fusion unit uses a weighted fusion method to fuse the vibration feature vector and the acoustic feature vector to generate a fused comprehensive feature. The fault identification unit inputs the fused comprehensive features into a long short-term memory network model for fault classification and identification, and evaluates the current motor health status level.

[0005] As a further improvement to this technical solution, the data acquisition and processing unit includes a data acquisition module and a data processing module; The data acquisition module synchronously acquires vibration signals and acoustic signals during motor operation; The data processing module preprocesses the vibration signal and acoustic signature signal and establishes a sensor position coordinate system.

[0006] As a further improvement to this technical solution, the feature extraction unit includes a vibration feature extraction module and a voiceprint feature extraction module; The vibration feature extraction module extracts vibration feature vectors based on the preprocessed vibration signal using a structure dynamic deformation field method based on vibration signal inversion, extracting time-domain, frequency-domain, and time-frequency-domain features to generate vibration feature vectors. The voiceprint feature extraction module extracts acoustic parameters based on time-domain statistical analysis and extracts spectral features based on fast Fourier transform to generate a voiceprint feature vector.

[0007] As a further improvement to this technical solution, the vibration feature extraction module extracts vibration feature vectors using a structural dynamic deformation field method based on vibration signal inversion, extracting time-domain, frequency-domain, and time-frequency-domain features to generate vibration feature vectors, including the following steps: S1.1 Perform time series analysis on the preprocessed multi-sensor vibration signal, calculate the statistical parameters characterizing the signal energy and waveform characteristics, generate time domain features, and calculate the signal transfer entropy matrix between sensors based on the multi-sensor vibration signal sequence. S1.2 Input the signal transfer entropy matrix into the pre-trained deformation field inversion model, output the three-dimensional instantaneous deformation velocity field of the predefined grid nodes on the surface of the motor housing, and perform intrinsic orthogonal decomposition on the time series data of the three-dimensional instantaneous deformation velocity field to obtain the time coefficients of the main modes and extract the frequency domain features. S1.3 Perform time-frequency analysis on the main modal time coefficients obtained in step S1.2 to obtain time-frequency domain characteristics; S1.4 Combine the above time-domain features, frequency-domain features and time-frequency-domain features in chronological order to form a multidimensional feature vector, namely the vibration feature vector.

[0008] As a further improvement to this technical solution, in step S1.2, the signal transfer entropy matrix is ​​input into the pre-trained deformation field inversion model, and the instantaneous deformation velocity field of the predefined grid nodes on the surface of the motor housing is output. This involves the following specific steps: The signal transfer entropy matrix is ​​expanded using principal component decomposition, and then encoded using a graph convolutional network to construct a temporal feature tensor. ; Transform the temporal feature tensor The input is fed into a pre-trained deformation field inversion model to extract spatial correlation features corresponding to the equipment structural response, and the response mapping weight matrix is ​​output. The motor geometric constraint matrix and the node adjacency relation matrix are introduced into the response mapping weight matrix to adjust the constraints. Based on the adjusted response mapping weight matrix, the three-dimensional instantaneous deformation velocity vector field of each grid node at the current moment is calculated.

[0009] As a further improvement to this technical solution, the voiceprint feature extraction module extracts acoustic parameters based on time-domain statistical analysis, extracts spectral features based on fast Fourier transform, and generates a voiceprint feature vector, including the following steps: S2.1. Perform echo and harmonic suppression processing on the preprocessed voiceprint signal to generate a purified voiceprint signal. Perform frame segmentation and windowing processing on the purified voiceprint signal to extract time-domain feature parameters that reflect the energy and waveform structure of the sound signal. S2.2 Perform fundamental frequency detection and harmonic analysis on the time-domain signal of the voiceprint signal, calculate the fundamental frequency using the autocorrelation function, and extract the structural characteristic parameters of the fundamental frequency and harmonics; S2.3 Perform a fast Fourier transform on each frame of the audioprint signal to obtain the sound pressure spectrum and power spectral density, and extract the main spectral feature parameters; S2.4 Combine the time-domain characteristic parameters, fundamental frequency and harmonic structure characteristic parameters, and main spectral characteristic parameters according to the time sequence to form the voiceprint feature vector.

[0010] As a further improvement to this technical solution, in step S2.1, the preprocessed voiceprint signal undergoes echo and harmonic suppression preprocessing, including the following steps: S2.11. The input segment of the voiceprint signal is segmented and windowed. The frequency domain spectrum of each frame is obtained by fast Fourier transform, and the echo transfer function is constructed based on the frequency domain modeling method. ; S2.12, Real-time acquisition of motor speed signal According to the speed signal Retrieve the corresponding harmonic template from the pre-trained motor harmonic template library. and harmonic template Superimposed with the echo transfer function, a predicted harmonic echo spectrum is constructed. ; S2.13, Predicting Harmonic Echo Spectrum Perform an inverse Fourier transform to obtain the time-domain predicted echo signal. And predict the echo signal in the time domain. The signal is fed forward and canceled with the input voiceprint signal to generate a cleaned reference signal. ; S2.14, Based on the purification reference signal An adaptive filter is introduced to eliminate residual echoes, using a harmonic template. As input, to purify the reference signal. As the desired signal, output the initial voiceprint signal. ; S2.15, Initial voiceprint signal Transforming to the frequency domain, a secondary suppression of the residual echo component is performed using spectral subtraction, based on the noise power spectrum. Calculate the purified spectrum and transform it to generate the purified voiceprint signal.

[0011] As a further improvement to this technical solution, the feature fusion unit uses a weighted fusion method to perform feature fusion on the vibration feature vector and the acoustic signature feature vector, including the following steps: S3.1. Perform timestamp matching on the sampling data of vibration signal and acoustic signature to establish a unified timeline; S3.2 Calculate the signal delay difference based on the sensor synchronization error and acoustic propagation delay, and perform delay compensation through phase correction; S3.3 Normalize the two types of features after alignment to give them a uniform statistical scale; S3.4. The synchronized vibration features and acoustic signature features are spliced ​​together within the time window D to form a joint feature matrix; S3.5. The joint feature matrix is ​​fused using a weighted fusion method to obtain a comprehensive feature vector. .

[0012] As a further improvement to this technical solution, the fault identification unit includes a fault prediction module and a health management module; The fault prediction module is based on a comprehensive feature vector. The long short-term memory network model is used to analyze the trend of motor operation status; The health management module comprehensively assesses and classifies the health status of the motor based on trend analysis results and real-time operating data.

[0013] On the other hand, the present invention provides a motor fault diagnosis method based on synchronous analysis of vibration and acoustic signature, used in any of the above-mentioned motor fault diagnosis systems based on synchronous analysis of vibration and acoustic signature, comprising the following steps: S4.1 Synchronously acquire vibration and acoustic signals during motor operation, and preprocess the vibration and acoustic signals; S4.2 Extract vibration feature vectors using the structure dynamic deformation field method based on vibration signal inversion, extract acoustic feature vectors based on time-domain statistical analysis, and perform echo and harmonic suppression preprocessing on the preprocessed acoustic signal during the extraction of acoustic feature vectors. S4.3. Use the weighted fusion method to fuse the vibration feature vector and the acoustic feature vector to generate the fused comprehensive feature. S4.4 Input the fused integrated features into the long short-term memory network model for fault classification and identification, and evaluate the current motor health status level.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention relates to a motor fault diagnosis system and method based on synchronous analysis of vibration and acoustic signatures. In the feature extraction process, in addition to extracting traditional time-domain, frequency-domain, and time-frequency-domain vibration features, it innovatively introduces the Signal Transfer Entropy (STE) matrix to characterize the energy flow and dynamic coupling relationship between multiple sensors. This matrix is ​​then input into a pre-trained deformation field inversion model. Combined with motor geometric constraints and node adjacency relationships, the three-dimensional instantaneous deformation velocity field of the shell surface is reconstructed. This process achieves a nonlinear mapping from one-dimensional vibration signals at discrete points to the three-dimensional dynamic response of a continuous structure, overcoming the limitations of traditional methods that rely solely on single-point or statistical features for diagnosis. By performing intrinsic orthogonal decomposition and time-frequency analysis on the deformation field, the spatial distribution and evolution of dominant modes can be identified, effectively distinguishing the differences in structural resonance modes caused by different faults such as bearing loosening, rotor imbalance, and stator eccentricity, thus enhancing the fault type discrimination and spatial location capabilities.

[0015] 2. This invention relates to a motor fault diagnosis system and method based on synchronous analysis of vibration and acoustic signatures. In the acoustic signature preprocessing stage, a feedforward + feedback joint suppression framework is constructed: First, a harmonic template is obtained by looking up a table based on the rotational speed signal, and a predicted spectrum is generated by combining frequency domain echo modeling for feedforward cancellation; then, an NLMS adaptive filter is used with the harmonic template as a reference signal to further eliminate residual components; finally, secondary purification is completed through spectral subtraction. This multi-level purification strategy effectively removes non-fault-related periodic harmonics and reverberation interference, preserving the true acoustic characteristics of the fault. Based on this, the feature fusion unit compensates for acoustic propagation delay through timestamp matching and phase correction, and uses a weighted fusion method to synchronously splice the normalized vibration and acoustic signature features, ensuring accurate alignment of the two heterogeneous signals at the time and event levels. The generated comprehensive feature vector combines the high sensitivity of vibration signals with the non-contact advantage of acoustic signature signals, maintaining stable classification performance even in low signal-to-noise ratio and strong interference environments, significantly improving the accuracy and generalization ability of the Long Short-Term Memory network in fault identification. Attached Figure Description

[0016] Figure 1 This is an overall flowchart of the present invention; The meanings of the labels in the diagram are as follows: 1. Data acquisition and processing unit; 2. Feature extraction unit; 3. Feature fusion unit; 4. Fault identification unit. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0018] Example 1: Please refer to Figure 1 As shown, a motor fault diagnosis system based on synchronous analysis of vibration and acoustic signature is provided, including a data acquisition and processing unit 1 that synchronously acquires vibration signals and acoustic signature signals during motor operation and performs preprocessing on the vibration signals and acoustic signature signals. In this embodiment, the data acquisition and processing unit 1 includes a data acquisition module and a data processing module; The data acquisition module synchronously acquires vibration and acoustic signals during motor operation; at least four single-axis vibration sensors are installed at key locations on the motor housing (such as bearing housing, end cover, and housing center) to form a minimum sensing network. The data processing module preprocesses the vibration and acoustic signature signals and establishes a sensor position coordinate system. Specifically, it denoises, detrends, and normalizes the raw vibration and acoustic signature signals collected by multiple sensors to remove environmental noise, steady-state offsets, and amplitude differences, ensuring signal stability and comparability. Then, it synchronizes the signals based on the sampling timestamps, performs time resampling and interpolation correction on the sampling rates of different types of sensors, aligning the vibration and acoustic signature signals on the same time axis. Next, based on the motor structure model and sensor installation positions, it selects the center of the housing as the origin, defines a three-dimensional rectangular coordinate system (X-axis along the axial direction, Y-axis along the radial direction, and Z-axis along the tangential direction), and calibrates the spatial coordinates of each sensor. Finally, it associates the signal of each sensor with its position parameters in the coordinate system to form a multi-channel time-series data matrix with spatial position information, providing a unified spatial reference framework for subsequent feature extraction and deformation field inversion.

[0019] Feature extraction unit 2 extracts vibration feature vectors based on multi-domain analysis and extracts acoustic feature vectors based on time-domain statistical analysis; In the process of extracting voiceprint feature vectors, the preprocessed voiceprint signal is subjected to echo and harmonic suppression preprocessing. In this embodiment, the feature extraction unit 2 includes a vibration feature extraction module and a voiceprint feature extraction module; The vibration feature extraction module extracts vibration feature vectors based on the preprocessed vibration signal using a structure dynamic deformation field method based on vibration signal inversion, extracting time-domain, frequency-domain, and time-frequency-domain features to generate vibration feature vectors. The voiceprint feature extraction module extracts acoustic parameters based on time-domain statistical analysis and extracts spectral features based on fast Fourier transform to generate a voiceprint feature vector. Furthermore, the vibration feature extraction module utilizes a structural dynamic deformation field method based on vibration signal inversion to extract vibration feature vectors. This is primarily to address the challenge of comprehensively capturing the overall dynamic response and spatial distribution characteristics of the motor structure from local point vibration measurements. Traditional vibration analysis methods typically rely on the time-domain or frequency-domain characteristics of a single or a few sensors, failing to effectively reflect the three-dimensional deformation and dynamic coupling behavior of the motor housing. This method, however, maps one-dimensional vibration signals to a three-dimensional deformation field through a multi-sensor network and deformation field inversion technology, thereby more accurately identifying structural changes and abnormal vibration modes caused by faults. This method can invert the three-dimensional instantaneous deformation velocity field of the motor housing from multi-sensor vibration signals. By fusing time, frequency, and spatial information through the transfer entropy matrix and a pre-trained deformation field inversion model, it provides richer structural dynamic features. Compared to traditional methods, it not only improves the accuracy and robustness of fault feature extraction but also effectively captures the energy flow and coupling relationships between sensors, thereby identifying early faults and complex fault modes earlier and significantly improving the accuracy and reliability of motor fault diagnosis. The vibration feature extraction module uses a structure dynamic deformation field method based on vibration signal inversion to extract vibration feature vectors, extracting time-domain, frequency-domain, and time-frequency-domain features to generate vibration feature vectors, including the following steps: S1.1 Perform time series analysis on the preprocessed multi-sensor vibration signals, calculate statistical parameters characterizing signal energy and waveform properties, and generate time-domain features. Simultaneously, based on the multi-sensor vibration signal sequences, calculate the signal transfer entropy matrix (STE) between sensors to characterize energy flow and dynamic coupling relationships between them. Specifically: First, preprocess the vibration signals of each sensor, including denoising, detrending, and standardization to ensure data synchronization and stability; then, for each pair of sensor signals (e.g., sensor...),... and The information-theoretic method is used to calculate the transfer entropy, where the transfer entropy quantifies the change in the source signal ( ) to target signal ( The intensity of information flow is specifically determined by setting the embedding dimension (e.g., and ) and time delay ( The state space is reconstructed using [a method], and the propagation entropy is estimated based on the conditional probability distribution, as shown in the formula [formula]. (In the formula, For target signal In time The value of is the observed value at the next moment; For target signal In the present and future The vector of values ​​at each time step is used to describe The past state (i.e., the embedding dimension is) (state reconstruction vector); Source signal In the present and future The vector of values ​​at each time step is used to describe Past state (embedding dimension is) ), The joint probability distribution represents the joint probability of the source signal and the target signal occurring at a specified time. (This represents the probability of the target signal taking a value at the next moment, given only its past states). Finally, iterate through all sensor pairs and fill an N×N matrix (N being the number of sensors) with the calculated transfer entropy values, where each element... Indicates from sensor arrive The information transmission intensity, thus forming a complete signal transmission entropy matrix; S1.2 Input the signal transfer entropy matrix into the pre-trained deformation field inversion model, output the three-dimensional instantaneous deformation velocity field of the predefined grid nodes on the surface of the motor housing, realize the mapping from one-dimensional signal to three-dimensional physical field, and perform intrinsic orthogonal decomposition on the time series data of the three-dimensional instantaneous deformation velocity field to obtain the main mode time coefficients and extract frequency domain features. The pre-trained deformation field inversion model employs a multimodal temporal coding-spatial mapping decoding structure, which is essentially a deep neural network architecture that integrates graph convolution and attention mechanisms. Its input is a temporal feature tensor calculated from multi-sensor vibration signals. This tensor integrates the temporal variation of the transfer entropy matrix with the topological coupling features between sensors; the output is the three-dimensional instantaneous deformation velocity field of predefined mesh nodes on the motor housing surface. The intermediate layer of the model consists of three parts: a temporal encoding layer, which uses a one-dimensional convolution and bidirectional LSTM structure to extract time-dependent features; a spatial feature extraction layer, which encodes the sensor topology based on a graph convolutional network (GCN) to learn the coupling relationship between multiple points; and a physical mapping decoding layer, which introduces a multi-head attention mechanism and a fully connected mapping layer to map high-dimensional features into the spatial response weight matrix of the corresponding nodes. The training data comes from a synchronous experimental dataset of finite element simulation and field measurement. The input is multi-point vibration signals (or transfer entropy matrix sequences) under different operating conditions, and the label is the shell deformation velocity field obtained by three-dimensional laser vibration measurement or simulation calculation. The model is pre-trained by minimizing the mean square error between the predicted field and the real field and the structural similarity loss, thereby obtaining the nonlinear mapping capability from signal features to physical response. Furthermore, after preprocessing the vibration and acoustic signature signals and establishing the sensor position coordinate system, the motor housing or structural surface needs to be spatially discretized within this coordinate system to form predefined mesh nodes. Specifically, based on the motor's geometric model, a finite element meshing algorithm is used to generate a predefined set of mesh nodes on the motor housing surface. For regular regions, a uniform mesh (approximately P×Q, where P and Q are the number of nodes in each dimension) can be generated. For complex curved surfaces, unstructured meshing methods such as Delaunay triangulation are used to generate a total of N nodes (N being the total number of discrete nodes). Each node corresponds to a fixed three-dimensional coordinate, and an adjacency matrix is ​​generated based on the spatial distance and geometric adjacency relationship between nodes to represent the topological connections of nodes on the structural surface. This set of mesh nodes and its topological relationships remain unchanged throughout the entire vibration and acoustic signature signal mapping and deformation field inversion process, serving as a unified reference framework for spatial feature mapping and response distribution analysis; this is the predefined mesh node. Furthermore, the signal transfer entropy matrix is ​​input into the pre-trained deformation field inversion model, which outputs the instantaneous deformation velocity field of the predefined mesh nodes on the motor housing surface. This involves the following specific steps: The signal transfer entropy matrix is ​​expanded using principal component decomposition (PCA) and then encoded using a graph convolutional network (GCN). This process involves two aspects: firstly, PCA expands and reduces the dimensionality of the transfer entropy matrix, extracting principal component vectors that represent global information transfer characteristics, thus reducing noise and redundant dimensions; secondly, GCN embeds the structure of the signal transfer entropy matrix to preserve the topological coupling relationships and information flow patterns between sensors. This constructs a temporal feature tensor. ; Transform the temporal feature tensor The input is fed into a pre-trained deformation field inversion model to extract spatial correlation features corresponding to the device structural response, and the output is a response mapping weight matrix (the deformation field inversion model learns the mapping relationship between signal features and physical response through multi-layer convolution and attention mechanisms, and outputs a response mapping weight matrix reflecting the contribution of different sensor inputs to the spatial node response). Specifically, this involves converting the temporal feature tensor... When input into the pre-trained deformation field inversion model, the tensor is first subjected to time-step expansion and normalization to ensure that the features of different sensing nodes are comparable in the time dimension. Subsequently, the model uses temporal convolutional layers or bidirectional LSTM layers to extract the dynamic correlation features of the signal on the time axis, and maps these temporal dynamic features to spatial grid nodes corresponding to the device's geometric topology through spatial attention modules or graph convolutional layers. In this process, the model combines predefined grid node coordinates and node adjacency matrices to learn the mapping rules between temporal signal changes and structural deformation. Finally, the response mapping weight matrix is ​​output through the decoding layer. This matrix reflects the weight distribution of each node's contribution to the overall structural response and is used for subsequent spatial feature fusion and structural state reconstruction. In the response mapping weight matrix, a motor geometric constraint matrix (used to limit the spatial topology between nodes) and a node adjacency relation matrix (used to describe the interaction between mesh nodes on the shell surface) are introduced to perform constraint adjustment (i.e. constraint optimization of weight distribution). Based on the constraint-adjusted response mapping weight matrix, the three-dimensional instantaneous deformation velocity vector field of each mesh node at the current moment is calculated, thereby realizing the accurate inversion mapping from multi-point entropy correlation features to the continuous dynamic response of the shell. Among them, the geometric constraint matrix is ​​used to limit the topological connections and structural boundaries of nodes in three-dimensional space, ensuring that the displacement direction of adjacent nodes is consistent with the actual surface normal of the motor housing; the node adjacency matrix is ​​used to characterize the interaction strength and information propagation path between mesh nodes on the housing surface; based on this, the model performs constraint optimization on the response mapping weight matrix (first, the initial weight matrix is ​​used as input, and then the motor geometric constraint matrix is ​​introduced). Adjacency matrix of nodes Construct a comprehensive constraint loss function ,in, This indicates the deformation field reconstruction error. To measure the consistency between the weight distribution and the geometric structure normal constraints, Used to maintain smoothness and response continuity between adjacent nodes. The weight coefficients are used to measure the consistency between the weight distribution and the geometric normal constraints. To maintain the smoothness and response continuity among adjacent nodes (weight coefficients), the response weights of each node are redistributed through iterative solutions, ensuring that the weight distribution simultaneously satisfies structural continuity and geometric constraints. Finally, the constraint-adjusted weight matrix is ​​jointly calculated with the node coordinates and time derivative terms to obtain the three-dimensional instantaneous deformation velocity vector field of each mesh node at the current moment, i.e., the instantaneous deformation and its rate of change distribution of each node in the x, y, and z directions; S1.3. Perform time-frequency analysis (e.g., STFT) on the main modal time coefficients obtained in step S1.2 to obtain time-frequency domain features (i.e., the energy evolution law of each main spatial mode with time-frequency). Specifically, first, take the time coefficient sequence of each main mode as the input signal and perform local time-domain decomposition on it through short-time Fourier transform (STFT), that is, calculate the instantaneous spectrum distribution of the signal within a sliding time window; within each time window, multiply the signal with a windowing function (e.g., Hamming window or Gaussian window) and then perform fast Fourier transform (FFT) to obtain the time-frequency matrix; then, perform amplitude normalization and energy spectrum calculation on the spectrum results of all time windows to extract the energy distribution features of each mode at different times and frequencies; finally, use the obtained time-frequency envelope, main frequency trajectory, energy center frequency and other parameters as time-frequency domain features for subsequent modal response identification and feature fusion analysis. S1.4 Combine the above time-domain features, frequency-domain features and time-frequency-domain features in chronological order to form a multidimensional feature vector, namely the vibration feature vector.

[0020] Furthermore, the voiceprint feature extraction module extracts acoustic parameters based on time-domain statistical analysis and extracts spectral features based on fast Fourier transform to generate a voiceprint feature vector, including the following steps: S2.1. The preprocessed voiceprint signal undergoes echo and harmonic suppression processing to generate a purified voiceprint signal. This purified signal is then framed and windowed to extract temporal feature parameters reflecting the energy and waveform structure of the sound signal. The framed and windowed processing involves: first, determining the frame length (e.g., 20–40 ms) and frame shift (e.g., 10 ms) based on the sampling rate and analysis precision, dividing the continuous speech signal into a series of short frames to ensure that the signal within each frame can be considered a stationary process; then, multiplying each frame signal by a windowing function (e.g., Hamming or Hanning window) to reduce discontinuities at frame boundaries and spectral leakage effects; next, calculating typical temporal feature parameters for each frame, including at least short-time energy, short-time zero-crossing rate, average amplitude, signal variance, waveform factor, impulse factor, and peak factor, to describe the instantaneous energy changes, waveform shape, and amplitude distribution of the sound signal; finally, concatenating the features of each frame in chronological order to form a temporal feature matrix, laying the foundation for subsequent frequency domain or voiceprint feature extraction. The preprocessing for echo and harmonic suppression primarily targets inherent, fault-independent, strong interference in motor acoustic signature signals. This includes: acoustic echoes: Since motors are typically installed in enclosed or semi-enclosed spaces, the sound they emit is reflected by surrounding objects and mixed with the direct sound, leading to distortion and blurring of the acoustic signature signal; electromagnetic harmonics: Periodic harmonic noise generated by electromagnetic force during motor operation, strictly related to rotational speed, often has a much higher intensity than the weak acoustic features caused by mechanical friction, impact, or other faults. These interferences severely "overwhelm" the true fault characteristics, resulting in inaccurate subsequent feature extraction and even misjudgment. This invention deeply integrates the physical mechanisms of motor operation, constructing an "echo transfer function" and introducing a "motor harmonic template library" driven by real-time rotational speed signals, achieving feedforward, model-driven accurate prediction and cancellation of interference components. Its core function is to selectively and faithfully remove the two main types of interference, echo and electromagnetic harmonics, thereby greatly "purifying" the voiceprint signal. This makes the subsequently extracted acoustic features more realistically reflect the mechanical state, significantly improving the sensitivity and reliability of voiceprint diagnosis and providing a high-quality data foundation for the deep fusion of vibration and voiceprint. The preprocessed speaker signal undergoes echo and harmonic suppression preprocessing, including the following steps: S2.11. Preprocessed voiceprint signal The input segment is segmented and windowed. The frequency domain spectrum of each frame is obtained using Fast Fourier Transform (FFT), and the echo transfer function is constructed based on the frequency domain modeling method. This is used to identify the frequency domain characteristics of the echo path, where, To receive the signal power spectrum, The power spectrum of the direct signal; S2.12, Real-time acquisition of motor speed signal According to the speed signal Retrieve the corresponding harmonic template from the pre-trained motor harmonic template library. and harmonic template Superimposed with the echo transfer function, a predicted harmonic echo spectrum is constructed. (Harmonic template) (Should be used as a reference signal) for feedforward harmonic echo cancellation; The pre-trained motor harmonic template library is a standardized feature library constructed by sampling, spectral analysis, and harmonic modeling of a large number of healthy motor operating signals under various operating conditions (such as load changes, speed fluctuations, and power supply imbalances). During training, the collected vibration and acoustic signature signals are first subjected to FFT or STFT transformation to extract the amplitude, phase, and energy distribution features of the fundamental frequency, even-order, and odd-order harmonics. Then, the feature space is compressed and classified using clustering or principal component analysis (PCA) to form harmonic templates corresponding to different motor types and operating states. Finally, the template features are pre-trained using an autoencoder or convolutional neural network (CNN) to adaptively learn the nonlinear coupling relationships between harmonics, thereby enabling real-time signal matching, deviation detection, and health status comparison in subsequent fault diagnosis. S2.13, Predicting Harmonic Echo Spectrum Perform inverse Fourier transform (IDFT) to obtain the time-domain predicted echo signal. And predict the echo signal in the time domain. With the input voiceprint signal Perform feedforward cancellation to generate a purified reference signal. (At this time, the reference signal is purified) (Error signal of residual echo); S2.14, Based on the purification reference signal An adaptive filter is introduced to eliminate residual echo (the architecture of the adaptive filter mainly consists of an input signal terminal, a weight update unit, an error signal feedback channel, and an output unit. The input signal first passes through a delay line to form an input vector, then enters a weighted summation module, where it is multiplied by the current filter weight vector to generate a filtered output; simultaneously, the system compares this output with the desired signal to calculate the error signal; the error signal is then input to the weight update unit, which dynamically adjusts the filter weights according to the selected adaptive algorithm (NLMS), so that the output error gradually decreases, realizing real-time modeling and self-learning tracking of the target system or noisy environment, thereby possessing capabilities such as signal enhancement, interference suppression, and system identification), using harmonic templates. As input, to purify the reference signal. As the desired signal, output the initial voiceprint signal. The filter output can be used to further eliminate residual echoes and generate a preliminary acoustic signature signal. ; Specifically, the standard normalized least mean square (NLMS) algorithm is used to update the filter weights. : ; ; ; In the formula, For error signals, This is the step size coefficient. To prevent small constants from being divided by zero, This is a transpose operation; Furthermore, the NLMS adaptive update mechanism employed in this invention is not a traditional noise cancellation structure, but rather an adaptive modeling mechanism based on harmonic templates. In this mechanism, the harmonic template... As filter input, it is used to dynamically characterize the periodic features of the voiceprint signal; filter output This represents a real-time estimation of the main harmonic components in the input signal; subsequently, the spectral subtraction module is used to further refine the signal. With the original signal The residuals are suppressed in the frequency domain to further eliminate residual echoes; therefore, in this invention... It is not noise estimation, but harmonic signal estimation. The purification process is completed in the spectral reduction stage. It is functionally equivalent to standard adaptive noise cancellation but structurally more targeted. At each of the above time points, the adaptive filter adjusts according to the current weights. For harmonic templates Weighting is performed to obtain the filter output. This output This is the initial voiceprint signal. That is, the signal after the filter has preliminarily canceled the residual echo, and at the same time, the weights are iteratively updated through the NLMS algorithm. The filter's output becomes increasingly accurate in subsequent time steps, further reducing residual echoes. Ultimately, the iteratively obtained filter output... The signal is input into the spectral subtraction process to further suppress residual echoes and generate the final purified speaker signal. S2.15, Initial voiceprint signal Convert to frequency domain The residual echo component is suppressed twice using spectral subtraction, based on the noise power spectrum. Calculate the cleaned spectrum, transform the cleaned spectrum, and generate the cleaned audioprint signal. ; Specifically: Calculation The amplitude spectrum and phase spectrum are expressed as follows: ; ; According to the noise power spectrum Calculate the amplitude spectrum of the noise. : ; Performing spectral subtraction on the amplitude spectrum yields the purified signal amplitude spectrum: ; The purified amplitude spectrum is combined with the original phase spectrum to reconstruct the purified spectrum: ; right Perform inverse discrete Fourier transform (IFFT) to obtain the cleaned time-domain acoustic signature signal. : ; In the formula, for The phase spectrum reflects the signal at frequency Phase angle at that point, This is the complex exponential form (unit phase factor) corresponding to the phase spectrum, used to preserve the phase information of the original signal. for The real part, for The imaginary part, The purified complex spectrum signal combines the purified amplitude spectrum with the original phase spectrum. This is the lower limit parameter for the spectrum, used to prevent the amplitude after spectral subtraction from becoming negative or too small; the typical value range is 0.001 to 0.01. The imaginary unit, The inverse discrete Fourier transform is used to convert frequency domain signals back to the time domain. S2.2. Fundamental frequency detection and harmonic analysis are performed on the time-domain signal of the voiceprint signal. The fundamental frequency is calculated using the autocorrelation function, and characteristic parameters of the fundamental frequency and harmonic structure are extracted, including the harmonic-to-noise energy ratio (HNR), fundamental frequency stability, and harmonic bandwidth distribution. Specifically, when performing fundamental frequency detection and harmonic analysis on the time-domain signal of the voiceprint signal, the autocorrelation function is first calculated for each windowed frame of the speech signal to measure the similarity of the signal under different delays. The first main peak position (excluding zero delay) is found in the autocorrelation function, and its corresponding delay time is determined. The fundamental period of the signal is represented, from which the fundamental frequency is calculated. Subsequently, the Fast Fourier Transform (FFT) was used to perform spectral analysis on the frame signal, extracting harmonic amplitude and energy distribution at integer multiples of the fundamental frequency to obtain parameters such as the number of harmonics, harmonic intervals, harmonic energy ratios, and the proportion of fundamental frequency energy. Finally, these fundamental frequency and harmonic structural features were combined to form fundamental frequency and harmonic structural feature parameters, which are used to describe the pitch characteristics and harmonic stability of the voiceprint signal, providing a physically interpretable feature basis for subsequent voiceprint recognition and feature fusion. S2.3. Perform Fast Fourier Transform (FFT) on each frame of the voiceprint signal to obtain the sound pressure spectrum and power spectral density (PSD), and extract the main spectral feature parameters: spectral centroid, spectral bandwidth, spectral roll-off point, spectral flatness, band power ratio, and peak distribution characteristics. Specifically, when performing FFT on each frame of the voiceprint signal, first, zero-padding is applied to the windowed frame signal to improve spectral resolution. Then, the time-domain signal is converted to the frequency domain using FFT to obtain the complex form of the spectrum. Next, its amplitude spectrum is calculated to obtain the sound pressure spectrum distribution, and then squared and normalized to obtain the power spectral density (PSD), which is used to characterize the energy intensity of different frequency components. After obtaining the complete spectrum, the main spectral feature parameters are extracted, including at least the dominant frequency, spectral centroid, bandwidth, spectral power ratio, spectral entropy, spectral skewness, and kurtosis, to describe the frequency distribution characteristics and energy concentration of the voiceprint signal. Finally, these features are combined in chronological order to form a frequency domain feature sequence, providing input for subsequent acoustic pattern recognition and voiceprint fusion analysis. S2.4 Combine the time-domain characteristic parameters, fundamental frequency and harmonic structure characteristic parameters, and main spectral characteristic parameters according to the time sequence to form the voiceprint feature vector.

[0021] Feature fusion unit 3 uses a weighted fusion method to fuse the vibration feature vector and the acoustic feature vector to generate a fused comprehensive feature. In this embodiment, the feature fusion unit 3 uses a weighted fusion method to fuse the vibration feature vector and the acoustic signature feature vector, including the following steps: S3.1. Timestamp matching is performed on the sampling data of vibration signal and acoustic pattern signal to establish a unified time axis. Specifically, linear interpolation and time resampling algorithms are used to align the vibration feature vector and acoustic pattern feature vector to the same sampling time window to ensure that the two types of features correspond to the same operating state in the time dimension. S3.2 Calculate the signal delay difference based on the sensor synchronization error and acoustic propagation delay, and compensate for the delay through phase correction so that the characteristic sequence is synchronized at key events (such as speed change, impact vibration, and acoustic burst). In the feature fusion process, to eliminate the influence of sensor synchronization error and acoustic propagation delay on the alignment of vibration features and acoustic signature features, the theoretical delay difference is first calculated based on the sampling timestamp and acoustic propagation path length of each sensor. ,in, , From sound source to sensor , The distance of transmission For the speed of sound, The sampling clock synchronization error is used; then the delay difference is utilized. Phase compensation is performed on the characteristic sequence by multiplying it by a phase factor in the frequency domain. Adjusting the phase of each frequency component aligns the features of different sensors or different modes on the time axis, thereby ensuring that vibration features and acoustic features are synchronized at key events (such as speed changes or impact vibrations), providing a consistent time reference for subsequent weighted fusion. S3.3 Normalize the two types of features after alignment to give them a uniform statistical scale (shift the mean of different feature dimensions to 0 and scale the standard deviation to 1 so that the vibration features and voiceprint features have a uniform scale in terms of numerical range and statistical distribution). S3.4. The synchronized vibration features and acoustic signature features are spliced ​​together within the time window D to form a joint feature matrix; S3.5. The joint feature matrix is ​​fused using a weighted fusion method (i.e., weights are assigned to each type of feature (determined based on feature importance or data quality), and the joint feature matrix is ​​weighted and summed along the feature dimensions) to obtain a comprehensive feature vector. .

[0022] The fault identification unit 4 inputs the fused comprehensive features into the long short-term memory network model for fault classification and identification, and evaluates the current motor health status level; In this embodiment, the fault identification unit 4 includes a fault prediction module and a health management module; The fault prediction module is based on a comprehensive feature vector. The trend analysis of motor operating status is performed using a long short-term memory network model; specifically, the comprehensive feature vector is... A comprehensive feature vector sequence is generated based on the time series. As input to a Long Short-Term Memory (LSTM) network, the temporal dependencies of the feature sequence are modeled using the LSTM's gating mechanism (input gate, forget gate, output gate), extracting long-term and short-term dynamic patterns and generating hidden states. and memory unit Subsequently, the hidden state is mapped to the predicted value or category probability of motor operating status through a fully connected layer. Combined with the historical feature change trend, the network can output the trend analysis results of the motor operating status in the future short time or sliding window, including the development trend of abnormal vibration, the change trend of acoustic feature and potential fault risk, thereby providing the health management module with the basis for time series prediction and realizing dynamic monitoring and early warning of motor status. The health management module comprehensively assesses and classifies the motor's health status based on trend analysis results and real-time operating data. Specifically, based on the trend analysis results output by the fault prediction module and the motor's real-time operating data, the health management module first compares the current comprehensive feature vector with the predicted short-term trend, calculating the deviation of each key indicator (such as vibration amplitude, spectral energy, acoustic signature fundamental frequency, and harmonic stability) from the health baseline. (That is, firstly, a historical health baseline is established for each indicator, including the mean, standard deviation, and normal fluctuation range, which can be determined based on long-term operating data or equipment calibration data.) Then, the current vibration amplitude, spectral energy, acoustic signature fundamental frequency, and harmonic stability are compared with the predicted short-term trend. Real-time characteristics such as harmonic stability are compared with the health baseline of the corresponding indicators. The degree of abnormality of each indicator is quantified by calculating the standardized deviation or normalized residual. For multi-channel or multi-sensor indicators, the results of each sensor can be weighted and averaged or dimensionality reduced and integrated using principal component analysis to obtain a single deviation measure. At the same time, for indicators with obvious trend changes, time weighting or sliding window methods can be introduced to enhance the sensitivity to gradual anomalies. Finally, the deviation value of each key indicator is output, providing a quantitative basis for comprehensive health scoring and fault risk assessment. Then, combined with the statistical distribution and threshold rules of historical operating data, the degree of deviation is risk-scored and weighted to generate a comprehensive health index. Next, based on the preset health level classification standards (such as normal, minor abnormality, warning, and serious fault), the comprehensive health index is mapped to the corresponding level; finally, the health level is fed back to the monitoring system or operation and maintenance platform for visualization, alarm triggering, and maintenance decision support, so as to realize hierarchical management and intelligent maintenance guidance of motor status. Specifically, mapping the comprehensive health index to corresponding levels involves: mapping the comprehensive health index to corresponding levels. Normalized to a scale of 0–100, historical health baselines or expert experience determine the threshold range corresponding to the health level, for example: normal: Minor abnormalities: ;warn: Critical fault: The comprehensive health index calculated in real time Compare with the above intervals: If If so, the mapping is normal; if If so, it will be mapped as a warning.

[0023] Example 2: The difference between Example 2 and Example 1 is that this example introduces a fault diagnosis analysis method used in a motor fault diagnosis system based on synchronous analysis of vibration and acoustic signature.

[0024] A motor fault diagnosis method based on synchronous analysis of vibration and acoustic signature, used in any of the above-mentioned motor fault diagnosis systems based on synchronous analysis of vibration and acoustic signature, includes the following steps: S4.1 Synchronously acquire vibration and acoustic signals during motor operation, and preprocess the vibration and acoustic signals; S4.2 Extract vibration feature vectors using the structure dynamic deformation field method based on vibration signal inversion, extract acoustic feature vectors based on time-domain statistical analysis, and perform echo and harmonic suppression preprocessing on the preprocessed acoustic signal during the extraction of acoustic feature vectors. S4.3. Use the weighted fusion method to fuse the vibration feature vector and the acoustic feature vector to generate the fused comprehensive feature. S4.4 Input the fused integrated features into the long short-term memory network model for fault classification and identification, and evaluate the current motor health status level.

[0025] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.

Claims

1. A motor fault diagnosis system based on synchronous analysis of vibration and acoustic signature, characterized in that, include: The data acquisition and processing unit (1) synchronously acquires the vibration signal and acoustic signal of the motor during operation, and preprocesses the vibration signal and acoustic signal. Feature extraction unit (2), the feature extraction unit (2) extracts vibration feature vectors using the structure dynamic deformation field method based on vibration signal inversion, and extracts acoustic feature vectors based on time domain statistical analysis; In the process of extracting voiceprint feature vectors, the preprocessed voiceprint signal is subjected to echo and harmonic suppression preprocessing. Feature fusion unit (3) uses a weighted fusion method to fuse the vibration feature vector and the acoustic feature vector to generate a fused comprehensive feature; The fault identification unit (4) inputs the fused integrated features into the long short-term memory network model for fault classification and identification, and evaluates the current motor health status level.

2. The motor fault diagnosis system based on synchronous analysis of vibration and acoustic signature as described in claim 1, characterized in that: The data acquisition and processing unit (1) includes a data acquisition module and a data processing module; The data acquisition module synchronously acquires vibration signals and acoustic signals during motor operation; The data processing module preprocesses the vibration signal and acoustic signature signal and establishes a sensor position coordinate system.

3. The motor fault diagnosis system based on synchronous analysis of vibration and acoustic signature as described in claim 1, characterized in that: The feature extraction unit (2) includes a vibration feature extraction module and a voiceprint feature extraction module; The vibration feature extraction module extracts vibration feature vectors based on the preprocessed vibration signal using a structure dynamic deformation field method based on vibration signal inversion, extracting time-domain, frequency-domain, and time-frequency-domain features to generate vibration feature vectors. The voiceprint feature extraction module extracts acoustic parameters based on time-domain statistical analysis and extracts spectral features based on fast Fourier transform to generate a voiceprint feature vector.

4. The motor fault diagnosis system based on synchronous analysis of vibration and acoustic signature as described in claim 3, characterized in that: The vibration feature extraction module extracts vibration feature vectors using a structure dynamic deformation field method based on vibration signal inversion, extracting time-domain, frequency-domain, and time-frequency-domain features to generate vibration feature vectors, including the following steps: S1.1 Perform time series analysis on the preprocessed multi-sensor vibration signal, calculate the statistical parameters characterizing the signal energy and waveform characteristics, generate time domain features, and calculate the signal transfer entropy matrix between sensors based on the multi-sensor vibration signal sequence. S1.2 Input the signal transfer entropy matrix into the pre-trained deformation field inversion model, output the three-dimensional instantaneous deformation velocity field of the predefined grid nodes on the surface of the motor housing, and perform intrinsic orthogonal decomposition on the time series data of the three-dimensional instantaneous deformation velocity field to obtain the time coefficients of the main modes and extract the frequency domain features. S1.3 Perform time-frequency analysis on the main modal time coefficients obtained in step S1.2 to obtain time-frequency domain characteristics; S1.4 Combine the above time-domain features, frequency-domain features and time-frequency-domain features in chronological order to form a multidimensional feature vector, namely the vibration feature vector.

5. The motor fault diagnosis system based on synchronous analysis of vibration and acoustic signature as described in claim 4, characterized in that: In step S1.2, the signal transfer entropy matrix is ​​input into the pre-trained deformation field inversion model, and the instantaneous deformation velocity field of the predefined grid nodes on the surface of the motor housing is output. This involves the following specific steps: The signal transfer entropy matrix is ​​expanded using principal component decomposition, and then encoded using a graph convolutional network to construct a temporal feature tensor. ; Transform the temporal feature tensor The input is fed into a pre-trained deformation field inversion model to extract spatial correlation features corresponding to the equipment structural response, and the response mapping weight matrix is ​​output. The motor geometric constraint matrix and the node adjacency relation matrix are introduced into the response mapping weight matrix to adjust the constraints. Based on the adjusted response mapping weight matrix, the three-dimensional instantaneous deformation velocity vector field of each grid node at the current moment is calculated.

6. The motor fault diagnosis system based on synchronous analysis of vibration and acoustic signature as described in claim 3, characterized in that: The voiceprint feature extraction module extracts acoustic parameters based on time-domain statistical analysis, extracts spectral features based on fast Fourier transform, and generates a voiceprint feature vector, including the following steps: S2.

1. Perform echo and harmonic suppression processing on the preprocessed voiceprint signal to generate a purified voiceprint signal. Perform frame segmentation and windowing processing on the purified voiceprint signal to extract time-domain feature parameters that reflect the energy and waveform structure of the sound signal. S2.2 Perform fundamental frequency detection and harmonic analysis on the time-domain signal of the voiceprint signal, calculate the fundamental frequency using the autocorrelation function, and extract the structural characteristic parameters of the fundamental frequency and harmonics; S2.3 Perform a fast Fourier transform on each frame of the audioprint signal to obtain the sound pressure spectrum and power spectral density, and extract the main spectral feature parameters; S2.4 Combine the time-domain characteristic parameters, fundamental frequency and harmonic structure characteristic parameters, and main spectral characteristic parameters according to the time sequence to form the voiceprint feature vector.

7. The motor fault diagnosis system based on synchronous analysis of vibration and acoustic signature as described in claim 6, characterized in that: In step S2.1, the preprocessed voiceprint signal undergoes echo and harmonic suppression preprocessing, which includes the following steps: S2.

11. The input segment of the voiceprint signal is segmented and windowed. The frequency domain spectrum of each frame is obtained by fast Fourier transform, and the echo transfer function is constructed based on the frequency domain modeling method. ; S2.12, Real-time acquisition of motor speed signal According to the speed signal Retrieve the corresponding harmonic template from the pre-trained motor harmonic template library. and harmonic template Superimposed with the echo transfer function, a predicted harmonic echo spectrum is constructed. ; S2.13, Predicting Harmonic Echo Spectrum Perform an inverse Fourier transform to obtain the time-domain predicted echo signal. And predict the echo signal in the time domain. The signal is fed forward and canceled with the input voiceprint signal to generate a cleaned reference signal. ; S2.14, Based on the purification reference signal An adaptive filter is introduced to eliminate residual echoes, using a harmonic template. As input, to purify the reference signal. As the desired signal, output the initial voiceprint signal. ; S2.15, Initial voiceprint signal Transforming to the frequency domain, a secondary suppression of the residual echo component is performed using spectral subtraction, based on the noise power spectrum. Calculate the purified spectrum and transform it to generate the purified voiceprint signal.

8. The motor fault diagnosis system based on synchronous analysis of vibration and acoustic signature as described in claim 1, characterized in that: The feature fusion unit (3) uses a weighted fusion method to fuse the vibration feature vector and the acoustic feature vector, including the following steps: S3.

1. Perform timestamp matching on the sampling data of vibration signal and acoustic signature to establish a unified timeline; S3.2 Calculate the signal delay difference based on the sensor synchronization error and acoustic propagation delay, and perform delay compensation through phase correction; S3.3 Normalize the two types of features after alignment to give them a uniform statistical scale; S3.

4. The synchronized vibration features and acoustic signature features are spliced ​​together within the time window D to form a joint feature matrix; S3.

5. The joint feature matrix is ​​fused using a weighted fusion method to obtain a comprehensive feature vector. .

9. The motor fault diagnosis system based on synchronous analysis of vibration and acoustic signature as described in claim 8, characterized in that: The fault identification unit (4) includes a fault prediction module and a health management module; The fault prediction module is based on a comprehensive feature vector. The long short-term memory network model is used to analyze the trend of motor operation status; The health management module comprehensively assesses and classifies the health status of the motor based on trend analysis results and real-time operating data.

10. A motor fault diagnosis method based on synchronous analysis of vibration and acoustic signature, used in a motor fault diagnosis system based on synchronous analysis of vibration and acoustic signature as described in any one of claims 1-9, characterized in that: Includes the following steps: S4.1 Synchronously acquire vibration and acoustic signals during motor operation, and preprocess the vibration and acoustic signals; S4.2 Extract vibration feature vectors using the structure dynamic deformation field method based on vibration signal inversion, extract acoustic feature vectors based on time-domain statistical analysis, and perform echo and harmonic suppression preprocessing on the preprocessed acoustic signal during the extraction of acoustic feature vectors. S4.

3. Use the weighted fusion method to fuse the vibration feature vector and the acoustic feature vector to generate the fused comprehensive feature. S4.4 Input the fused integrated features into the long short-term memory network model for fault classification and identification, and evaluate the current motor health status level.