Overturning machine fault detection method and system based on voiceprint detection

By using voiceprint detection technology and an improved RealNVP model, the problems of difficult sensor installation and reliance on manual labor in the fault detection of the flipping machine have been solved, enabling flexible deployment, early fault identification and stable detection, and improving the level of intelligence in the fault detection of the flipping machine.

CN121963783APending Publication Date: 2026-05-01AUTOMOTIVE ENGINEERING CORPORATION +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AUTOMOTIVE ENGINEERING CORPORATION
Filing Date
2026-03-10
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing fault detection technologies for tilting machines suffer from limitations in sensor installation, complex wiring, high maintenance costs, and reliance on human experience, making it difficult to achieve continuous monitoring and early warning.

Method used

A voiceprint-based detection method is adopted, and a voiceprint feature model library is established through acoustic signal processing and an improved RealNVP model to achieve non-contact monitoring and fault identification of the operation status of the tilting machine, including acoustic signal acquisition, preprocessing, frequency domain analysis, voiceprint feature extraction and model matching.

Benefits of technology

It enables flexible deployment of the tilting machine's operating status, has strong adaptability to complex working conditions, and is highly sensitive to early faults, thereby improving the stability and accuracy of fault detection and forming a closed-loop detection alarm and data accumulation mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963783A_ABST
    Figure CN121963783A_ABST
Patent Text Reader

Abstract

The invention discloses a voiceprint detection-based upender fault detection method and system. The method comprises the following steps of: acquiring a sound signal in an operation process of an upender and preprocessing the sound signal; performing framing and fast Fourier transform to obtain frequency domain spectrum data; obtaining a Mel spectrum through a Mel filter bank, performing logarithmic operation on the Mel spectrum to obtain a logarithmic Mel spectrum, and extracting voiceprint features; performing normalization processing and splicing to obtain a voiceprint feature sequence to be detected; establishing a normal voiceprint model and a fault voiceprint model to form a voiceprint model library; using the improved RealNVP model to output a matching result with the voiceprint model library, and obtaining a running state category result; according to the invention, the operation state monitoring and fault identification of the upender are realized, and the stability and accuracy of fault detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial equipment operation status monitoring and intelligent fault diagnosis, and in particular to a fault detection method and system for a flipping machine based on acoustic fingerprint detection. Background Technology

[0002] Tilting machines, commonly found in industrial settings such as automotive production lines, are heavy-duty electromechanical equipment used to flip, invert, and change the posture of workpieces, material frames, or carriers. During long-term operation, key components of the tilting machine, including the drive mechanism, transmission mechanism, bearing assemblies, couplings, reduction mechanisms, and tilting support structure, are subjected to impact loads, alternating loads, and frictional wear, making them prone to abnormal conditions such as loosening, wear, eccentricity, jamming, tooth surface damage, and bearing failure. If these abnormalities persist, they can lead to unstable tilting movements, decreased positioning accuracy, increased energy consumption, and increased noise; in severe cases, they can cause shutdowns, equipment damage, or even safety accidents. Therefore, effectively monitoring the operating status of tilting machines and achieving timely fault identification and early warning is a crucial aspect of industrial equipment health management.

[0003] Existing fault detection methods for tilting machines typically employ contact-based sensing solutions or inspection methods relying on manual experience. The first type of solution, represented by vibration sensors, current sensors, and temperature sensors, monitors the condition by installing sensors on key parts of the equipment. While this method can obtain relatively direct operational status data in some scenarios, it has significant limitations in the actual industrial application of tilting machines: First, tilting machines are mostly heavy-duty equipment with large structures and complex installation environments, limiting sensor placement and resulting in a large workload for installation and maintenance; Second, contact sensors are susceptible to the effects of installation tightness, installation direction, structural resonance, and on-site impact vibrations, leading to insufficient stability and consistency of the acquired signals; Third, the reliability of sensors and their wiring decreases in environments with dust, oil, humidity, or high temperatures, resulting in high maintenance costs; Fourth, for existing equipment already in production, sensor upgrades often require downtime for construction, affecting production cycles. The second type of solution is represented by manual inspection, manual listening, or experience-based judgment. Common practices include on-site personnel determining whether there are abnormalities in the equipment by listening, observing, or conducting periodic inspections. Such solutions rely on personnel experience and are affected by factors such as workload, environmental noise, and subjective differences among personnel, making it difficult to achieve continuous monitoring and early warning, and fault detection is often delayed. Summary of the Invention

[0004] One objective of this invention is to propose a fault detection method and system for a flipping machine based on acoustic print detection. This invention fully utilizes acoustic signal processing technology, acoustic print feature modeling methods, and operational state discrimination technology based on an improved RealNVP model to accurately distinguish between normal operating states and fault states. This invention offers advantages such as no need for contact sensors, flexible deployment, strong adaptability to complex operating conditions, high sensitivity to early faults, and good detection stability.

[0005] The fault detection method for a flipping machine based on voiceprint detection according to an embodiment of the present invention includes the following steps:

[0006] Acoustic acquisition devices are installed at the monitoring points of the tilting machine to collect and preprocess the sound signals during the operation of the tilting machine;

[0007] The preprocessed audio signal is framed and subjected to Fast Fourier Transform to obtain frequency domain spectrum data;

[0008] Frequency domain spectral data is input into a Mel filter bank to obtain a Mel spectrum. Logarithmic operation is performed on the Mel spectrum to obtain a logarithmic Mel spectrum, and voiceprint features are extracted.

[0009] The voiceprint features are normalized and the voiceprint features of adjacent frames are concatenated to obtain the voiceprint feature sequence to be detected.

[0010] A normal voiceprint model is established based on the voiceprint characteristics of the normal operation sound of the flipping machine in the historical voiceprint database, and a fault voiceprint model is established based on the voiceprint characteristics of the fault state sound of the flipping machine in the historical voiceprint database, forming a voiceprint model library.

[0011] The voiceprint feature sequence to be detected is input into the improved RealNVP model, the output is the matching result with the voiceprint model library, and the running status category result is obtained based on the matching result;

[0012] When the running status category result is a fault status, a fault event is generated and an alarm message is output. The audio data, voiceprint feature data and time information corresponding to the fault event are stored in the sample library.

[0013] Optionally, the sound signal refers to the acoustic signal generated during the operation of the tilting machine. The acoustic signal is formed by the change in sound pressure in the air medium over time caused by the operation of the mechanical parts of the tilting machine. The sound signal is an analog sound signal that can be received by the acoustic acquisition device. The preprocessing includes DC removal and detrending, bandpass filtering, and pre-emphasis.

[0014] Optionally, obtaining the frequency domain spectral data specifically includes:

[0015] The preprocessed audio signal is divided into frame-level audio sequences according to the preset frame length and preset frame shift.

[0016] Perform a discrete Fourier transform on the frame-level sound sequence for each frame to obtain the complex frequency domain transform result corresponding to the frame;

[0017] The amplitude spectrum and power spectrum of the frame are generated based on the results of the complex frequency domain transform.

[0018] The amplitude and power spectra of each frame are arranged in chronological order according to the frame number to form frequency domain spectrum data.

[0019] Optionally, the extraction of the voiceprint features specifically includes:

[0020] The frequency domain spectrum data is grouped according to the frame number to obtain the frame-level frequency domain spectrum data corresponding to each frame.

[0021] The frame-level frequency domain spectrum data of each frame is input into the Mel filter bank. For each frame and each filter in the Mel filter bank, the data corresponding to each frequency index in the frame-level frequency domain spectrum data of each frame is multiplied point by point with the filter response corresponding to the same frequency index and summed to obtain the Mel band energy of the frame on the filter. The Mel band energy corresponding to each filter in the Mel filter bank is used to form the Mel spectrum.

[0022] Logarithmic Mel spectrum is obtained by performing a logarithmic operation on the Mel spectrum;

[0023] Voiceprint features are extracted based on log-Mel spectrum.

[0024] Optionally, obtaining the voiceprint feature sequence to be detected specifically includes:

[0025] The voiceprint features are segmented into frames to obtain a sequence of voiceprint feature vectors that correspond one-to-one with each frame.

[0026] Normalization is performed on the voiceprint feature vector of each frame in the voiceprint feature vector sequence to obtain a normalized voiceprint feature vector sequence.

[0027] The normalized voiceprint feature vector sequence is subjected to adjacent frame splicing processing to obtain the splicing vector sequence corresponding to each frame. The adjacent frame splicing processing is to select the voiceprint feature vectors of multiple consecutive frames including the frame and adjacent frames as the center of each frame, and connect them in order according to the frame number to form the splicing vector corresponding to the frame.

[0028] The spliced ​​vector sequence is connected in chronological order according to the corresponding audio segments to form the voiceprint feature sequence to be detected.

[0029] Optionally, the formation of the voiceprint model library specifically includes:

[0030] The voiceprint feature sample sets labeled as the normal operation state of the tilting machine and the voiceprint feature sample sets labeled as the fault state of the tilting machine are obtained from the historical voiceprint database to obtain the normal sample set and the fault sample set.

[0031] Each voiceprint feature sample in the normal sample set is vectorized to obtain a set of normal voiceprint feature vectors.

[0032] Each voiceprint feature sample in the fault sample set is vectorized to obtain a set of fault voiceprint feature vectors.

[0033] A normal voiceprint model is established based on the set of normal voiceprint feature vectors;

[0034] A fault soundprint model is established based on the set of fault soundprint feature vectors.

[0035] Normal voiceprint models and faulty voiceprint models are stored and combined to form a voiceprint model library.

[0036] Optionally, obtaining the running status category result specifically includes:

[0037] The voiceprint feature sequence to be detected is input into an improved RealNVP model. The improved RealNVP model includes a feature embedding and conditional modeling module, a coupling transformation module, a multi-scale latent variable modeling module, and a probability evaluation and state discrimination module. The feature embedding and conditional modeling module performs dimensional alignment and concatenation on the embedding representation to form an input feature vector sequence and obtain a conditional vector. The coupling transformation module performs multi-level reversible coupling transformation on the input feature vector sequence to obtain intermediate transformation results. The multi-scale latent variable modeling module uses a reversible pyramid decomposition mechanism to perform multi-scale decomposition on the intermediate transformation results to form a latent variable vector. The probability evaluation and state discrimination module calculates the log-likelihood value of the voiceprint feature sequence to be detected.

[0038] In the feature embedding and conditional modeling module, the voiceprint feature sequence to be detected is input into the embedding network segment by segment in chronological order to obtain the embedding representation. The embedding representation is then dimensionally aligned and spliced ​​to form an input feature vector sequence. An aggregation operation is then performed on the input feature vector sequence to obtain the conditional vector.

[0039] In the coupled transformation module, under the constraint of the condition vector, a multi-level reversible coupled transformation is performed on the input feature vector sequence to obtain intermediate transformation results. Each level of reversible coupled transformation refers to dividing the input feature vector sequence into a first sub-vector and a second sub-vector, calculating the scaling transformation result and translation transformation result based on the first sub-vector and the condition vector, and performing the scaling transformation and translation transformation on the second sub-vector.

[0040] In the multi-scale latent variable modeling module, a reversible pyramid decomposition mechanism is used to perform multi-scale decomposition processing on the intermediate transformation results. The reversible pyramid decomposition mechanism includes dividing the intermediate transformation results into high-scale components and low-scale components level by level, and continuing to perform the next level of decomposition on each level of low-scale components to form latent variable representations corresponding to different scale levels. The latent variable representations are combined according to the scale level order to form a latent variable vector.

[0041] In the probability assessment and state discrimination module, the log-likelihood value of the voiceprint feature sequence to be detected is calculated based on the latent variable vector and the volume change information corresponding to the multi-level reversible coupling transformation.

[0042] The log-likelihood value is matched with the model in the voiceprint model library to obtain the operating status category result of the flipping machine.

[0043] Optionally, the generation and storage of the fault event specifically includes:

[0044] When the operating status category result is a fault status, a fault event is generated. The fault event includes a fault event identifier, a device identifier, a monitoring point identifier, and the fault occurrence time.

[0045] Alarm information is generated based on fault events. The alarm information includes alarm type, alarm time, device identifier, monitoring point identifier, and operating status category result.

[0046] The audio data, voiceprint feature data and time information corresponding to the fault event are associated to form a sample record, and the sample record is stored in the sample library.

[0047] A fault detection system for a flipping machine based on voiceprint detection according to an embodiment of the present invention includes:

[0048] The acoustic acquisition and preprocessing module is used to acquire sound signals during the operation of the tilting machine at the monitoring points of the tilting machine, and to preprocess the sound signals.

[0049] The frequency domain spectrum data generation module is used to perform frame segmentation and fast Fourier transform on the preprocessed audio signal to obtain frequency domain spectrum data;

[0050] The voiceprint feature extraction module is used to input frequency domain spectrum data into a Mel filter bank to obtain a Mel spectrum, perform a logarithmic operation on the Mel spectrum to obtain a logarithmic Mel spectrum, and extract voiceprint features based on the logarithmic Mel spectrum;

[0051] The voiceprint feature sequence construction module is used to normalize the voiceprint features and concatenate the voiceprint features of adjacent frames to obtain the voiceprint feature sequence to be detected.

[0052] The voiceprint model library construction module is used to build a normal voiceprint model based on the voiceprint features of the normal operation sound of the flipping machine in the historical voiceprint database, and to build a fault voiceprint model based on the voiceprint features of the fault state sound of the flipping machine in the historical voiceprint database, thus forming a voiceprint model library.

[0053] The operation status discrimination module is used to input the voiceprint feature sequence to be detected into the improved RealNVP model, output the matching result with the voiceprint model library, and obtain the operation status category result of the flipping machine based on the matching result;

[0054] The fault handling and sample storage module is used to generate a fault event and output alarm information when the running status category result is a fault state, and to store the audio data, voiceprint feature data and time information corresponding to the fault event into the sample library.

[0055] The beneficial effects of this invention are:

[0056] This invention, by introducing a fault detection method and system for tilting machines based on acoustic fingerprint detection, effectively overcomes several shortcomings of existing tilting machine fault detection technologies in terms of overall technical approach and implementation, achieving significant beneficial effects. Firstly, this invention uses the naturally generated sound signals during the tilting machine's operation as the primary detection object. Operational status monitoring can be achieved simply by deploying acoustic acquisition devices at monitoring points, eliminating the need for additional contact sensors such as vibration, displacement, or current sensors on key parts of the tilting machine. This avoids problems in existing technologies, such as limited sensor installation, complex wiring, high maintenance costs, and difficulties in modifying existing equipment. It significantly improves the flexibility and engineering adaptability of the system deployment, making it particularly suitable for widespread application in complex industrial environments and existing production lines.

[0057] Secondly, this invention constructs a voiceprint feature extraction process from the time domain to the frequency domain and then to the perceptual domain by sequentially performing preprocessing, framing and fast Fourier transform, Mel filter bank processing, and logarithmic operations on the acquired sound signals. This allows the obtained voiceprint features to better match the human ear's perception of sound changes, while effectively suppressing interference from environmental noise and changes in operating conditions. Further normalization of the voiceprint features and splicing adjacent frame voiceprint features form a temporally continuous sequence of detectable voiceprint features. This allows potential subtle acoustic anomalies during the operation of the flipping machine to be continuously accumulated and amplified at the feature level, thereby improving the detection capability for early and progressive faults and avoiding the insensitivity of existing methods based on single-frame features or simple spectral thresholding to instantaneous anomalies.

[0058] Furthermore, this invention establishes acoustic modeling for the normal and fault states of the flipping machine based on a historical acoustic model database. By constructing an acoustic model library, it provides a reliable reference for subsequent state discrimination, eliminating reliance on manual experience or fixed thresholds and instead modeling based on historical data distribution, thus enhancing the objectivity and consistency of fault detection results. Building upon this foundation, this invention introduces an improved RealNVP model for probabilistic modeling and matching discrimination of the acoustic feature sequences to be detected. It utilizes a reversible deep generation model to accurately characterize the acoustic feature distribution and distinguishes between normal and fault states through log-likelihood values. This effectively solves the problem of insufficient generalization ability of traditional classification models under conditions of scarce fault samples and imbalanced sample distribution, improving the stability and reliability of operational state discrimination.

[0059] Finally, when the invention detects a fault in the tilting machine, it can automatically generate a fault event and output alarm information. Simultaneously, it stores the corresponding audio data, voiceprint feature data, and time information in a sample library, providing data support for subsequent model updates, fault analysis, and maintenance decisions, forming a closed-loop mechanism of "detection—alarm—data accumulation." Through this closed-loop mechanism, the invention not only achieves real-time monitoring and fault early warning of the tilting machine's operating status but also lays the foundation for full lifecycle health management and continuous performance optimization of the equipment, comprehensively improving the intelligence level and practical application value of tilting machine fault detection. Attached Figure Description

[0060] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0061] Figure 1 This is an overall flowchart of the fault detection method and system for a flipping machine based on voiceprint detection proposed in this invention;

[0062] Figure 2 This is a schematic diagram of the acoustic signature model library of the fault detection method and system for flipping machines based on acoustic signature detection proposed in this invention.

[0063] Figure 3 This is a schematic diagram of the improved RealNVP model of the fault detection method and system for flipping machines based on voiceprint detection proposed in this invention; Detailed Implementation

[0064] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0065] refer to Figures 1-3The fault detection method for a flipping machine based on voiceprint detection includes the following steps:

[0066] Acoustic acquisition devices are installed at the monitoring points of the tilting machine to collect and preprocess the sound signals during the operation of the tilting machine;

[0067] The preprocessed audio signal is framed and subjected to Fast Fourier Transform to obtain frequency domain spectrum data;

[0068] Frequency domain spectral data is input into a Mel filter bank to obtain a Mel spectrum. Logarithmic operation is performed on the Mel spectrum to obtain a logarithmic Mel spectrum, and voiceprint features are extracted.

[0069] The voiceprint features are normalized and the voiceprint features of adjacent frames are concatenated to obtain the voiceprint feature sequence to be detected.

[0070] A normal voiceprint model is established based on the voiceprint characteristics of the normal operation sound of the flipping machine in the historical voiceprint database, and a fault voiceprint model is established based on the voiceprint characteristics of the fault state sound of the flipping machine in the historical voiceprint database, forming a voiceprint model library.

[0071] The voiceprint feature sequence to be detected is input into the improved RealNVP model, the output is the matching result with the voiceprint model library, and the running status category result is obtained based on the matching result;

[0072] When the running status category result is a fault status, a fault event is generated and an alarm message is output. The audio data, voiceprint feature data and time information corresponding to the fault event are stored in the sample library.

[0073] In this embodiment, the sound signal refers to the acoustic signal generated during the operation of the tilting machine. The acoustic signal is formed by the change of sound pressure in the air medium over time caused by the operation of the mechanical parts of the tilting machine. The sound signal is an analog sound signal that can be received by the acoustic acquisition device. The preprocessing includes DC removal and detrending, bandpass filtering, and pre-emphasis.

[0074] In this embodiment, obtaining the frequency domain spectral data specifically includes:

[0075] The preprocessed audio signal is divided into frame-level audio sequences according to the preset frame length and preset frame shift.

[0076] Perform a discrete Fourier transform on the frame-level sound sequence for each frame to obtain the complex frequency domain transform result corresponding to the frame;

[0077] The complex frequency domain transformation result is specifically as follows: for each frame-level sound sequence, according to the discrete Fourier transform rules, each discrete sampling point in the frame is multiplied point by point with the complex exponential basis function under the corresponding frequency index, and the product results of all sampling points are accumulated to obtain the complex numerical frequency domain representation of the frame at each frequency index; the complex numerical frequency domain representation contains both real and imaginary part information, constituting the complex frequency domain transformation result corresponding to the frame;

[0078] The amplitude spectrum and power spectrum of the frame are generated based on the results of the complex frequency domain transform.

[0079] The amplitude spectrum and power spectrum are specifically obtained as follows: For the complex frequency domain transformation result of each frame, the modulus value corresponding to each frequency index is calculated. The modulus value is determined by the real part and the imaginary part of the complex value to obtain the amplitude spectrum corresponding to the frame. Based on the obtained amplitude spectrum, the modulus value at each frequency index is squared and normalized according to the number of discrete sampling points contained in the frame to obtain the power spectrum corresponding to the frame. The above processing is performed on all frequency indices in the complex frequency domain transformation result in sequence to form a complete frame-level amplitude spectrum and frame-level power spectrum.

[0080] The amplitude and power spectra of each frame are arranged in chronological order according to the frame number to form frequency domain spectrum data.

[0081] In this embodiment, the extraction of voiceprint features specifically includes:

[0082] The frequency domain spectrum data is grouped according to the frame number to obtain the frame-level frequency domain spectrum data corresponding to each frame.

[0083] The frame-level frequency domain spectrum data of each frame is input into the Mel filter bank. For each frame and each filter in the Mel filter bank, the data corresponding to each frequency index in the frame-level frequency domain spectrum data of each frame is multiplied point by point with the filter response corresponding to the same frequency index and summed to obtain the Mel band energy of the frame on the filter. The Mel band energy corresponding to each filter in the Mel filter bank is used to form the Mel spectrum.

[0084] Logarithmic Mel spectrum is obtained by performing a logarithmic operation on the Mel spectrum;

[0085] The logarithmic operation is specifically as follows: for the energy value corresponding to each Mel frequency band in the Mel spectrum, the energy value is added to a preset non-zero constant, and the result of the addition is subjected to a logarithmic operation with the natural logarithm as the base to obtain the corresponding logarithmic Mel spectrum value; the above logarithmic operation is performed on the energy values ​​of all Mel frequency bands in sequence to form a complete logarithmic Mel spectrum;

[0086] Voiceprint features are extracted based on log-Mel spectrum;

[0087] The extraction process is as follows: for the log-Mel spectrum of each frame, a discrete cosine transform is performed on the log-Mel spectrum according to a predetermined cepstral analysis rule to obtain the corresponding cepstral coefficients; the cepstral coefficients are combined with the log-Mel spectrum to form the voiceprint features corresponding to the frame; the above processing is performed sequentially on the log-Mel spectrum of all frames to obtain voiceprint features used to characterize the operating state of the flipping machine.

[0088] In this embodiment, obtaining the voiceprint feature sequence to be detected specifically includes:

[0089] The voiceprint features are segmented into frames to obtain a sequence of voiceprint feature vectors that correspond one-to-one with each frame.

[0090] Normalization is performed on the voiceprint feature vector of each frame in the voiceprint feature vector sequence to obtain a normalized voiceprint feature vector sequence.

[0091] The normalization process specifically involves: performing a normalization operation on each frame of the voiceprint feature vector sequence according to the feature dimension. The operation process is as follows: for each voiceprint feature component, first subtract the mean of the corresponding dimension in the historical voiceprint database, and then divide by the standard deviation of the corresponding dimension in the historical voiceprint database to obtain the normalized voiceprint feature vector corresponding to the frame; perform the above normalization process on all frames in the voiceprint feature vector sequence in sequence to form a normalized voiceprint feature vector sequence.

[0092] The normalized voiceprint feature vector sequence is subjected to adjacent frame splicing processing to obtain the splicing vector sequence corresponding to each frame. The adjacent frame splicing processing is to select the voiceprint feature vectors of multiple consecutive frames including the frame and adjacent frames as the center of each frame, and connect them in order according to the frame number to form the splicing vector corresponding to the frame.

[0093] The spliced ​​vector sequence is connected in chronological order according to the corresponding audio segments to form the voiceprint feature sequence to be detected.

[0094] In this embodiment, the formation of the voiceprint model library specifically includes:

[0095] The voiceprint feature sample sets labeled as the normal operation state of the tilting machine and the voiceprint feature sample sets labeled as the fault state of the tilting machine are obtained from the historical voiceprint database to obtain the normal sample set and the fault sample set.

[0096] Each voiceprint feature sample in the normal sample set is vectorized to obtain a set of normal voiceprint feature vectors.

[0097] Each voiceprint feature sample in the fault sample set is vectorized to obtain a set of fault voiceprint feature vectors.

[0098] A normal voiceprint model is established based on the set of normal voiceprint feature vectors;

[0099] The establishment process is as follows: sum all normal voiceprint feature vectors in the normal voiceprint feature vector set and normalize them according to the number of normal voiceprint feature vectors to obtain a normal mean vector. Calculate the normal covariance matrix using the normal mean vector. The calculation involves performing an outer product operation on the difference vector between each normal voiceprint feature vector and the normal mean vector to obtain an outer product matrix. Sum the outer product matrices and normalize them by subtracting one from the number of normal voiceprint feature vectors to obtain the normal covariance matrix. The normal mean vector and the normal covariance matrix together constitute a normal voiceprint model.

[0100] A fault soundprint model is established based on the set of fault soundprint feature vectors.

[0101] The establishment process is as follows: sum all fault soundprint feature vectors in the fault soundprint feature vector set and normalize them according to the number of fault soundprint feature vectors to obtain the fault mean vector. Calculate the fault covariance matrix using the fault mean vector. The calculation involves performing an outer product operation on the difference vector between each fault soundprint feature vector and the fault mean vector to obtain an outer product matrix. Sum the outer product matrices and normalize them by subtracting one from the number of fault soundprint feature vectors to obtain the fault covariance matrix. The fault soundprint model is constructed using the fault mean vector and the fault covariance matrix.

[0102] Normal voiceprint models and faulty voiceprint models are stored and combined to form a voiceprint model library.

[0103] In this embodiment, obtaining the running status category result specifically includes:

[0104] The voiceprint feature sequence to be detected is input into an improved RealNVP model. The improved RealNVP model includes a feature embedding and conditional modeling module, a coupling transformation module, a multi-scale latent variable modeling module, and a probability evaluation and state discrimination module. The feature embedding and conditional modeling module performs dimensional alignment and concatenation on the embedding representation to form an input feature vector sequence and obtain a conditional vector. The coupling transformation module performs multi-level reversible coupling transformation on the input feature vector sequence to obtain intermediate transformation results. The multi-scale latent variable modeling module uses a reversible pyramid decomposition mechanism to perform multi-scale decomposition on the intermediate transformation results to form a latent variable vector. The probability evaluation and state discrimination module calculates the log-likelihood value of the voiceprint feature sequence to be detected.

[0105] In the feature embedding and conditional modeling module, the voiceprint feature sequence to be detected is input into the embedding network segment by segment in chronological order to obtain the embedding representation. The embedding representation is then dimensionally aligned and spliced ​​to form an input feature vector sequence. An aggregation operation is then performed on the input feature vector sequence to obtain the conditional vector.

[0106] The embedding representation specifically involves dividing the voiceprint feature sequence to be detected into consecutive voiceprint feature segments in chronological order, and inputting each voiceprint feature segment into a pre-constructed embedding network for feature mapping. The embedding network performs multi-layer nonlinear transformations on the input voiceprint feature segments and outputs an embedding representation that is different in dimension from the original voiceprint features but semantically consistent.

[0107] In the coupled transformation module, under the constraint of the condition vector, a multi-level reversible coupled transformation is performed on the input feature vector sequence to obtain intermediate transformation results. Each level of reversible coupled transformation refers to dividing the input feature vector sequence into a first sub-vector and a second sub-vector, calculating the scaling transformation result and translation transformation result based on the first sub-vector and the condition vector, and performing the scaling transformation and translation transformation on the second sub-vector.

[0108] In the multi-scale latent variable modeling module, a reversible pyramid decomposition mechanism is used to perform multi-scale decomposition processing on the intermediate transformation results. The reversible pyramid decomposition mechanism includes dividing the intermediate transformation results into high-scale components and low-scale components level by level, and continuing to perform the next level of decomposition on each level of low-scale components to form latent variable representations corresponding to different scale levels. The latent variable representations are combined according to the scale level order to form a latent variable vector.

[0109] In the probability assessment and state discrimination module, the log-likelihood value of the voiceprint feature sequence to be detected is calculated based on the latent variable vector and the volume change information corresponding to the multi-level reversible coupling transformation.

[0110] The calculation process is as follows: based on the prior distribution of the latent variables, calculate the logarithmic probability value of the latent variable vector under the prior distribution; during the multi-level reversible coupling transformation process, obtain the volume change information corresponding to each level of reversible coupling transformation, and accumulate the volume change information after converting it into logarithmic form; add the logarithmic probability value of the latent variable vector to the logarithmic value corresponding to the volume change information at each level to obtain the logarithmic likelihood value of the voiceprint feature sequence to be detected;

[0111] The log-likelihood value is matched with the model in the voiceprint model library to obtain the operating status category result of the flipping machine;

[0112] The matching process specifically involves: comparing the log-likelihood value corresponding to the voiceprint feature sequence to be detected with the log-likelihood reference ranges corresponding to the normal voiceprint model and the fault voiceprint model in the voiceprint model library, respectively; when the log-likelihood value falls within the log-likelihood reference range corresponding to the normal voiceprint model, the operating state category of the flipping machine is determined to be normal; when the log-likelihood value falls within the log-likelihood reference range corresponding to the fault voiceprint model, the operating state category of the flipping machine is determined to be faulty; when the log-likelihood value does not satisfy the log-likelihood reference ranges corresponding to both the normal and fault voiceprint models, the operating state category of the flipping machine is determined based on the distance relationship with each of the reference ranges.

[0113] In this embodiment, the generation and storage of the fault event specifically includes:

[0114] When the operating status category result is a fault status, a fault event is generated. The fault event includes a fault event identifier, a device identifier, a monitoring point identifier, and the fault occurrence time.

[0115] The specific fault event is as follows: after the system determines that the operating status of the tilting machine is a fault state, a fault event generation process is triggered. The fault event generation process includes assigning a unique fault event identifier to the current fault state, determining the corresponding equipment identifier based on the currently operating tilting machine, and determining the corresponding monitoring point identifier based on the deployment location of the acoustic acquisition device. While generating the fault event identifier, equipment identifier, and monitoring point identifier, the current system time is obtained as the fault occurrence time, and the fault event identifier, equipment identifier, monitoring point identifier, and fault occurrence time are combined to form a complete fault event.

[0116] Alarm information is generated based on fault events. The alarm information includes alarm type, alarm time, device identifier, monitoring point identifier, and operating status category result.

[0117] The audio data, voiceprint feature data and time information corresponding to the fault event are associated to form a sample record, and the sample record is stored in the sample library.

[0118] A fault detection system for a tilting machine based on voiceprint detection includes:

[0119] The acoustic acquisition and preprocessing module is used to acquire sound signals during the operation of the tilting machine at the monitoring points of the tilting machine, and to preprocess the sound signals.

[0120] The frequency domain spectrum data generation module is used to perform frame segmentation and fast Fourier transform on the preprocessed audio signal to obtain frequency domain spectrum data;

[0121] The voiceprint feature extraction module is used to input frequency domain spectrum data into a Mel filter bank to obtain a Mel spectrum, perform a logarithmic operation on the Mel spectrum to obtain a logarithmic Mel spectrum, and extract voiceprint features based on the logarithmic Mel spectrum;

[0122] The voiceprint feature sequence construction module is used to normalize the voiceprint features and concatenate the voiceprint features of adjacent frames to obtain the voiceprint feature sequence to be detected.

[0123] The voiceprint model library construction module is used to build a normal voiceprint model based on the voiceprint features of the normal operation sound of the flipping machine in the historical voiceprint database, and to build a fault voiceprint model based on the voiceprint features of the fault state sound of the flipping machine in the historical voiceprint database, thus forming a voiceprint model library.

[0124] The operation status discrimination module is used to input the voiceprint feature sequence to be detected into the improved RealNVP model, output the matching result with the voiceprint model library, and obtain the operation status category result of the flipping machine based on the matching result;

[0125] The fault handling and sample storage module is used to generate a fault event and output alarm information when the running status category result is a fault state, and to store the audio data, voiceprint feature data and time information corresponding to the fault event into the sample library.

[0126] Example 1:

[0127] This embodiment is applied to the pretreatment electrophoresis turning machine production line in the painting workshop of an automobile manufacturing company. The production line is equipped with several turning machines to perform pretreatment and electrophoresis-related processes in the automotive painting process. These turning machines operate under long-term high-load and frequent start-stop conditions. Their key components include the turning bearing assembly, gear transmission mechanism, and support structure. In actual production, it was found that after a period of operation, these turning machines are prone to problems such as bearing wear and abnormal gear meshing. In the early stages, these problems often only manifest as subtle changes in operating sound, making them difficult to detect promptly through manual inspection or simple parameter monitoring.

[0128] In this production line, an acoustic acquisition device is arranged near the main rotating shaft of the rotating machine to collect sound signals generated during the machine's operation. The acoustic acquisition device is an industrial-grade condenser microphone, fixedly installed on the top of the equipment casing, without direct contact with the rotating machine body, thus avoiding modifications to the equipment structure.

[0129] During normal production, the sound signals from the start-up, acceleration, uniform rotation, and stopping phases of the tilting machine are continuously collected. These signals are first processed by DC removal, bandpass filtering, and pre-emphasis to eliminate low-frequency interference and background noise. The processed sound signals are then divided into short time frames and subjected to a Fast Fourier Transform (FFT) to obtain the corresponding frequency domain spectrum data. Subsequently, the frequency domain spectrum data is input into a Mel filter bank to form a Mel spectrum, and a logarithmic operation is performed on the Mel spectrum to obtain the logarithmic Mel spectrum.

[0130] Based on this, a discrete cosine transform is performed on the log-Mel spectrum of each frame to obtain cepstral coefficients, which are then combined with the log-Mel spectrum to form the voiceprint feature. To reduce the influence of differences in voiceprint feature distribution under different acquisition times and operating conditions, mean-variance normalization is performed on the voiceprint features, and the voiceprint features of adjacent frames are spliced ​​together to form a voiceprint feature sequence to be detected that contains temporally continuous information.

[0131] In the initial stage of system deployment, a historical acoustic signature database was constructed based on three consecutive months of historical operating data from the tilting machine on the production line. The database includes acoustic signature samples labeled with characteristics under normal tilting machine operation and those with confirmed fault conditions such as bearing wear and abnormal gear meshing. Based on these samples, normal acoustic signature models and fault acoustic signature models were established, forming an acoustic signature model library.

[0132] During actual operation, the real-time generated voiceprint feature sequence to be detected is input into the improved RealNVP model. This model performs probabilistic modeling on the voiceprint feature sequence and calculates its matching results with normal and faulty voiceprint models in the voiceprint model library, thereby outputting the operating status category of the flipper. When the system determines that the flipper is in a faulty state, it automatically generates a fault event and outputs alarm information, while storing the audio data, voiceprint feature data, and time information for the corresponding time period into the sample library for subsequent analysis and model updates.

[0133] To verify the practical effectiveness of the method of this invention, the system was continuously run on the production line for 90 days, and the detection results were compared and analyzed with manual inspection and post-maintenance records. The results show that this invention can issue early warnings through changes in acoustic signature characteristics before obvious mechanical failures occur in the turning machine, especially demonstrating a high recognition rate for early bearing wear and minor gear meshing abnormalities. Furthermore, due to the use of a non-contact acoustic acquisition method, the system deployment did not affect the production cycle time.

[0134] Table 1. Comparison of Fault Detection Results of Tilting Machines

[0135] Indicator Name Manual inspection method Vibration sensor monitoring method The present invention voiceprint detection method Average fault early warning time (days) 1.5 4.2 9.8 Bearing fault identification accuracy (%) 68.4 82.7 94.6 Accuracy rate of gear meshing abnormality identification (%) 61.9 79.3 92.1 False alarm rate (%) 12.6 8.4 3.1 Number of unplanned outages (90 days) 6 3 1 Cost per deployment (relative value) Low higher medium Equipment structural modification requirements none have none

[0136] As shown in Table 1, the present invention significantly improves the early warning time for faults. The average early warning time for manual inspection is 1.5 days, and for vibration sensor monitoring it is 4.2 days, while the present invention reaches 9.8 days, which is about 6.5 times faster than manual inspection and about 133% faster than vibration monitoring, allowing for a more sufficient window for handling during operation and maintenance.

[0137] In terms of recognition accuracy, this invention also outperforms the comparative scheme. Regarding bearing fault recognition accuracy, manual inspection achieves 68.4%, vibration sensor monitoring achieves 82.7%, and this invention achieves 94.6%, representing improvements of 26.2 and 11.9 percentage points respectively. For gear meshing anomaly recognition accuracy, manual inspection achieves 61.9%, vibration sensor monitoring achieves 79.3%, and this invention achieves 92.1%, representing improvements of 30.2 and 12.8 percentage points respectively. This demonstrates that this invention has a more stable recognition capability for different types of mechanical anomalies.

[0138] Regarding false alarms and production impact, the false alarm rate of this invention is 3.1%, lower than the 12.6% of manual inspection and the 8.4% of vibration sensor monitoring. In 90 days of operation, the number of unplanned downtimes with this invention was 1, fewer than the 6 times for manual inspection and the 3 times for vibration sensor monitoring. Combined with the "Requirements for Equipment Structure Modification" in the table, it can be seen that this invention can be deployed without modifying the equipment structure, ensuring high detection performance while also considering ease of engineering implementation and overall cost control.

[0139] In summary, the fault detection method for a tilting machine based on voiceprint detection of this invention is superior to existing comparative solutions in terms of early fault warning capability, identification accuracy, false alarm control, and actual production stability. It can effectively solve the technical problems of difficulty in timely detection of early faults in tilting machines, insufficient detection stability, and high operation and maintenance costs, and has good industrial application value.

Claims

1. A fault detection method for a flipping machine based on voiceprint detection, characterized in that, Includes the following steps: Acoustic acquisition devices are installed at the monitoring points of the tilting machine to collect and preprocess the sound signals during the operation of the tilting machine; The preprocessed audio signal is framed and subjected to Fast Fourier Transform to obtain frequency domain spectrum data; Frequency domain spectral data is input into a Mel filter bank to obtain a Mel spectrum. Logarithmic operation is performed on the Mel spectrum to obtain a logarithmic Mel spectrum, and voiceprint features are extracted. The voiceprint features are normalized and the voiceprint features of adjacent frames are concatenated to obtain the voiceprint feature sequence to be detected. A normal voiceprint model is established based on the voiceprint characteristics of the normal operation sound of the flipping machine in the historical voiceprint database, and a fault voiceprint model is established based on the voiceprint characteristics of the fault state sound of the flipping machine in the historical voiceprint database, forming a voiceprint model library. The voiceprint feature sequence to be detected is input into the improved RealNVP model, the output is the matching result with the voiceprint model library, and the running status category result is obtained based on the matching result; When the running status category result is a fault status, a fault event is generated and an alarm message is output. The audio data, voiceprint feature data and time information corresponding to the fault event are stored in the sample library.

2. The fault detection method for a flipping machine based on voiceprint detection according to claim 1, characterized in that, The sound signal refers to the acoustic signal generated during the operation of the tilting machine. The acoustic signal is formed by the change of sound pressure in the air medium over time caused by the operation of the mechanical parts of the tilting machine. The sound signal is an analog sound signal that can be received by the acoustic acquisition device. The preprocessing includes DC removal and detrending, bandpass filtering, and pre-emphasis.

3. The fault detection method for a flipping machine based on voiceprint detection according to claim 1, characterized in that, The acquisition of the frequency domain spectral data specifically includes: The preprocessed audio signal is divided into frame-level audio sequences according to the preset frame length and preset frame shift. Perform a discrete Fourier transform on the frame-level sound sequence for each frame to obtain the complex frequency domain transform result corresponding to the frame; The amplitude spectrum and power spectrum of the frame are generated based on the results of the complex frequency domain transform. The amplitude and power spectra of each frame are arranged in chronological order according to the frame number to form frequency domain spectrum data.

4. The fault detection method for a flipping machine based on voiceprint detection according to claim 1, characterized in that, The extraction of voiceprint features specifically includes: The frequency domain spectrum data is grouped according to the frame number to obtain the frame-level frequency domain spectrum data corresponding to each frame. The frame-level frequency domain spectrum data of each frame is input into the Mel filter bank. For each frame and each filter in the Mel filter bank, the data corresponding to each frequency index in the frame-level frequency domain spectrum data of each frame is multiplied point by point with the filter response corresponding to the same frequency index and summed to obtain the Mel band energy of the frame on the filter. The Mel band energy corresponding to each filter in the Mel filter bank is used to form the Mel spectrum. Logarithmic Mel spectrum is obtained by performing a logarithmic operation on the Mel spectrum; Voiceprint features are extracted based on log-Mel spectrum.

5. The fault detection method for a flipping machine based on voiceprint detection according to claim 1, characterized in that, The specific steps to obtain the voiceprint feature sequence to be detected include: The voiceprint features are segmented into frames to obtain a sequence of voiceprint feature vectors that correspond one-to-one with each frame. Normalization is performed on the voiceprint feature vector of each frame in the voiceprint feature vector sequence to obtain a normalized voiceprint feature vector sequence. The normalized voiceprint feature vector sequence is subjected to adjacent frame splicing processing to obtain the splicing vector sequence corresponding to each frame. The adjacent frame splicing processing is to select the voiceprint feature vectors of multiple consecutive frames including the frame and adjacent frames as the center of each frame, and connect them in order according to the frame number to form the splicing vector corresponding to the frame. The spliced ​​vector sequence is connected in chronological order according to the corresponding audio segments to form the voiceprint feature sequence to be detected.

6. The fault detection method for a flipping machine based on voiceprint detection according to claim 1, characterized in that, The formation of the voiceprint model library specifically includes: The voiceprint feature sample sets labeled as the normal operation state of the tilting machine and the voiceprint feature sample sets labeled as the fault state of the tilting machine are obtained from the historical voiceprint database to obtain the normal sample set and the fault sample set. Each voiceprint feature sample in the normal sample set is vectorized to obtain a set of normal voiceprint feature vectors. Each voiceprint feature sample in the fault sample set is vectorized to obtain a set of fault voiceprint feature vectors. A normal voiceprint model is established based on the set of normal voiceprint feature vectors; A fault soundprint model is established based on the set of fault soundprint feature vectors. Normal voiceprint models and faulty voiceprint models are stored and combined to form a voiceprint model library.

7. The fault detection method for a flipping machine based on voiceprint detection according to claim 1, characterized in that, The specific steps for obtaining the operational status category result include: The voiceprint feature sequence to be detected is input into an improved RealNVP model. The improved RealNVP model includes a feature embedding and conditional modeling module, a coupling transformation module, a multi-scale latent variable modeling module, and a probability evaluation and state discrimination module. The feature embedding and conditional modeling module performs dimensional alignment and concatenation on the embedding representation to form an input feature vector sequence and obtain a conditional vector. The coupling transformation module performs multi-level reversible coupling transformation on the input feature vector sequence to obtain intermediate transformation results. The multi-scale latent variable modeling module uses a reversible pyramid decomposition mechanism to perform multi-scale decomposition on the intermediate transformation results to form a latent variable vector. The probability evaluation and state discrimination module calculates the log-likelihood value of the voiceprint feature sequence to be detected. In the feature embedding and conditional modeling module, the voiceprint feature sequence to be detected is input into the embedding network segment by segment in chronological order to obtain the embedding representation. The embedding representation is then dimensionally aligned and spliced ​​to form an input feature vector sequence. An aggregation operation is then performed on the input feature vector sequence to obtain the conditional vector. In the coupled transformation module, under the constraint of the condition vector, a multi-level reversible coupled transformation is performed on the input feature vector sequence to obtain intermediate transformation results. Each level of reversible coupled transformation refers to dividing the input feature vector sequence into a first sub-vector and a second sub-vector, calculating the scaling transformation result and translation transformation result based on the first sub-vector and the condition vector, and performing the scaling transformation and translation transformation on the second sub-vector. In the multi-scale latent variable modeling module, a reversible pyramid decomposition mechanism is used to perform multi-scale decomposition processing on the intermediate transformation results. The reversible pyramid decomposition mechanism includes dividing the intermediate transformation results into high-scale components and low-scale components level by level, and continuing to perform the next level of decomposition on each level of low-scale components to form latent variable representations corresponding to different scale levels. The latent variable representations are combined according to the scale level order to form a latent variable vector. In the probability assessment and state discrimination module, the log-likelihood value of the voiceprint feature sequence to be detected is calculated based on the latent variable vector and the volume change information corresponding to the multi-level reversible coupling transformation. The log-likelihood value is matched with the model in the voiceprint model library to obtain the operating status category result of the flipping machine.

8. The fault detection method for a flipping machine based on voiceprint detection according to claim 1, characterized in that, The generation and storage of the fault events specifically include: When the operating status category result is a fault status, a fault event is generated. The fault event includes a fault event identifier, a device identifier, a monitoring point identifier, and the fault occurrence time. Alarm information is generated based on fault events. The alarm information includes alarm type, alarm time, device identifier, monitoring point identifier, and operating status category result. The audio data, voiceprint feature data and time information corresponding to the fault event are associated to form a sample record, and the sample record is stored in the sample library.

9. A fault detection system for a tilting machine based on voiceprint detection, comprising the fault detection method for a tilting machine based on voiceprint detection as described in any one of claims 1 to 8, characterized in that, include: The acoustic acquisition and preprocessing module is used to acquire sound signals during the operation of the tilting machine at the monitoring points of the tilting machine, and to preprocess the sound signals. The frequency domain spectrum data generation module is used to perform frame segmentation and fast Fourier transform on the preprocessed audio signal to obtain frequency domain spectrum data; The voiceprint feature extraction module is used to input frequency domain spectrum data into a Mel filter bank to obtain a Mel spectrum, perform a logarithmic operation on the Mel spectrum to obtain a logarithmic Mel spectrum, and extract voiceprint features based on the logarithmic Mel spectrum; The voiceprint feature sequence construction module is used to normalize the voiceprint features and concatenate the voiceprint features of adjacent frames to obtain the voiceprint feature sequence to be detected. The voiceprint model library construction module is used to build a normal voiceprint model based on the voiceprint features of the normal operation sound of the flipping machine in the historical voiceprint database, and to build a fault voiceprint model based on the voiceprint features of the fault state sound of the flipping machine in the historical voiceprint database, thus forming a voiceprint model library. The operation status discrimination module is used to input the voiceprint feature sequence to be detected into the improved RealNVP model, output the matching result with the voiceprint model library, and obtain the operation status category result of the flipping machine based on the matching result; The fault handling and sample storage module is used to generate a fault event and output alarm information when the running status category result is a fault state, and to store the audio data, voiceprint feature data and time information corresponding to the fault event into the sample library.