Cableway wheel set anomaly detection method and system, terminal and storage medium

By deploying environmental isolated microphone arrays and convolutional autoencoder models on the brackets on both sides of the cable carriages, audio data is collected in real time and feature reconstruction error calculations are performed, the high cost and missed detection problems of existing cable carriage detection methods are solved, real-time early warning and high accuracy detection of early faults are achieved.

CN120544609AInactive Publication Date: 2025-08-26SHANDONG LANGCHAO SMART CULTURAL TOURISM IND DEV CO LTD

Patent Information

Application Number
CN202511028561.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-08-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing cable carriage wheel set abnormality detection methods have problems such as high installation and maintenance costs, inability to realize online real-time monitoring, and relying on a large number of labeled fault samples, resulting in insufficient generalization capabilities of model, and traditional methods are prone to missed early failure detection.

Method used

By deploying an environmentally isolated microphone array on the brackets on both sides of the cable carriage, audio data is collected in real time, and feature reconstruction error calculation is performed using a convolutional autoencoder model, combining Log-Mel spectrum feature extraction and 5-frame context splicing technology to achieve real-time early warning of early failures.

Benefits of technology

Real-time online monitoring and timely early warning of cable carriage sets are realized, deployment and maintenance costs are reduced, early identification accuracy is improved, and anti-interference ability and model generalization ability in complex environments are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544609A_ABST
    Figure CN120544609A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of equipment detection, and particularly relates to a cableway wheel set anomaly detection method and system, a terminal and a storage medium, and the method comprises the steps: deploying acoustic sensors on supports at two sides of a cableway wheel set, collecting audio data of the cableway wheel set in a working state in real time, and generating corresponding audio features; inputting the generated audio features into a pre-trained convolutional auto-encoder model, and outputting a reconstructed feature vector; calculating a mean square error of the generated audio feature and the reconstructed feature vector, and taking the calculated mean square error as an abnormal score; and comparing the abnormal score with a preset abnormal score threshold, and judging that the cableway wheel set is abnormal when the abnormal score exceeds the preset abnormal score threshold. According to the method, real-time reconstruction error calculation is carried out on the audio features by adopting the convolutional auto-encoder model, early warning can be triggered at the early failure stage of the wheel set component, and the problem of missing detection delay caused by insufficient sensitivity of a traditional vibration or thermal imaging method is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of equipment detection, and in particular relates to a method, system, terminal and storage medium for detecting abnormality of a cableway wheel group. Background Art

[0002] As a vital facility for mountain transportation and tourism, the safe operation of cableway systems is directly linked to the equipment's service life and the safety of life and property. During routine operations and maintenance, cableway sheaves (including key components like pulleys, bearings, and brackets) are subject to prolonged, high-load, continuous operation. These components are prone to abnormal vibration and noise due to issues such as poor lubrication, component cracks, bearing wear, or loose installation. Failure to promptly detect and address these anomalies not only accelerates equipment aging and shortens service life, but can also lead to serious safety incidents.

[0003] With the development of acoustic detection technology, fault diagnosis methods based on acoustic analysis have gradually become a research hotspot. However, existing acoustic detection methods, such as fault diagnosis based on vibration sensing, detection devices based on ultrasonic waves, and inspection systems based on infrared thermal imaging, all have certain limitations. Specifically, vibration sensors and ultrasonic devices need to be fixedly installed at each wheel assembly location and calibrated regularly, resulting in high installation and maintenance costs and easy damage in harsh environments. Infrared thermal imaging and ultrasonic detection often require downtime or manual inspection of equipment, which cannot achieve online real-time monitoring and has the probability of detection blind spots and human omissions. In addition, most of these methods rely on a large number of labeled fault samples for training and are highly dependent on fault samples. In actual operation, there are few abnormal instances, many types, and data collection is difficult, resulting in insufficient model generalization capabilities. Summary of the Invention

[0004] In response to the above-mentioned deficiencies in the prior art, the present invention provides a method, system, terminal and storage medium for detecting abnormalities in a cableway wheel group. By deploying an environmentally isolated microphone array on the upper part of the brackets on both sides of the cableway wheel group, real-time collection and online analysis of audio data in the working state of the wheel group are realized, and a convolutional autoencoder model is used to perform real-time reconstruction error calculation of audio features. This can trigger an early warning at the early stage of failure such as tiny cracks in the wheel group components and poor initial lubrication, avoiding the problem of missed detection and delay caused by insufficient sensitivity of traditional vibration or thermal imaging methods.

[0005] In a first aspect, the present invention provides a method for detecting abnormality of a cableway wheel assembly, comprising: S1. Deploy acoustic sensors on the brackets on both sides of the cableway sheave to collect audio data in real time when the cableway sheave is in operation. S2. Preprocessing the collected audio data of the cableway wheel assembly in the working state to generate a feature vector; S3, input the generated feature vector into the pre-trained convolutional autoencoder model and output the reconstructed feature vector; S4. Calculate the mean square error between the feature vector of the input convolutional autoencoder model and the reconstructed feature vector, and use the calculated mean square error as the anomaly score; S5. Compare the abnormality score with a preset abnormality score threshold, and when the abnormality score exceeds the preset abnormality score threshold, determine that the cableway wheel assembly has an abnormality.

[0006] A further improvement of this technical solution is that step S2 includes: S2.1. Perform short-time Fourier transform (STFT) on the collected audio data with a window size of 1024 points and a frame shift of 512 points to generate a spectrogram with a dimension of 513 × 312. S2.2, converting the generated spectrogram into a Mel spectrogram through a 128-order Mel filter bank; S2.3, performing logarithmic transformation on the Mel spectrum to generate a Log-Mel spectrum; S2.4, perform 5-frame context splicing on the Log-Mel spectrum to generate a feature matrix with a dimension of 128×5; S2.5. Flatten the feature matrix with the generated dimension of 128×5 into a 640-dimensional feature vector.

[0007] A further improvement of the present technical solution is that the encoder in the convolutional autoencoder model in step S3 includes: S3.1a. Set up the first convolutional layer, use 64 filters of size 5 to perform one-dimensional convolution on the 640-dimensional feature vector, after batch normalization and ReLU activation, the output dimension is 636×64 feature matrix; S3.1b. Set the first pooling layer to perform max pooling with a stride of 2 on the feature matrix of dimension 636×64, and output a feature matrix of dimension 318×64. S3.1c. Set up the second convolutional layer, use 128 filters of size 5 to perform one-dimensional convolution on the feature matrix of dimension 318 × 64, and output a feature matrix of dimension 314 × 128; S3.1d. Set the second pooling layer to perform max pooling with a stride of 2 on the feature matrix of dimension 314 × 128, and output a feature matrix of dimension 157 × 128. S3.1e. Set the third convolutional layer to perform one-dimensional convolution on the feature matrix of dimension 157 × 128 using 256 filters of size 5, and output a feature matrix of dimension 153 × 256. S3.1f. Set the third pooling layer to perform max pooling with a stride of 2 on the feature matrix of dimension 153×256, and output a feature matrix of dimension 76×256. S3.1g. Set up a flattening layer and a fully connected layer to flatten the 76×256 feature matrix into a one-dimensional vector. Then, pass it through a fully connected layer with 128 neurons and a ReLU activation to output a feature vector with a dimension of 128. S3.1h. Set up a latent space layer and linearly map the feature vector of dimension 128 to 16 dimensions through a fully connected layer containing 16 neurons to generate a latent vector z of dimension 16.

[0008] A further improvement of the technical solution is that the decoder in the convolutional autoencoder model in step S3 includes: S3.2a. Set up a fully connected expansion layer to linearly expand the latent vector z of dimension 16 through a fully connected expansion layer containing 80 × 256 neurons to generate a feature vector of dimension 20480, and reshape it into a feature matrix of dimension 80 × 256. S3.2b. Set up the first deconvolution layer, use 256 filters of size 5 to perform one-dimensional deconvolution on the feature matrix of dimension 80×256, and generate a feature matrix of dimension 76×256 after batch normalization and ReLU activation; S3.2c. Set the first upsampling layer to upsample the feature matrix of dimension 76×256 by a factor of 2 using a pre-stored bilinear interpolation algorithm to generate a feature matrix of dimension 152×256. S3.2d. Set up the second deconvolution layer, use 128 filters of size 5 to perform one-dimensional deconvolution on the feature matrix of dimension 152×256, and generate a feature matrix of dimension 148×128 after batch normalization and ReLU activation. S3.2e. Set a second upsampling layer to upsample the feature matrix of dimension 148 × 128 by a factor of 2 using a pre-stored bilinear interpolation algorithm to generate a feature matrix of dimension 296 × 128. S3.2f. Set up the third deconvolution layer, use 64 filters of size 5 to perform one-dimensional deconvolution on the feature matrix of dimension 296×128, and generate a feature matrix of dimension 292×64 after batch normalization and ReLU activation; S3.2g. Set a third upsampling layer to upsample the feature matrix of dimension 292 × 64 by a factor of 2 using a pre-stored bilinear interpolation algorithm to generate a feature matrix of dimension 584 × 64. S3.2h. Set the output convolution layer and use a filter of size 5 to perform one-dimensional convolution on the feature matrix of dimension 584×64 to generate a reconstructed feature vector of dimension 640×1.

[0009] A further improvement of this technical solution is that step S4 includes: The 640-dimensional feature vector of the input convolutional autoencoder model is divided into frame vector , each frame contains F frequency band features; the reconstructed feature vector with a dimension of 640×1 output by the convolutional autoencoder model is divided into frame vector , each frame contains F frequency band features; Calculate the frame-level mean square error between the feature vector and the reconstructed feature vector : ; The anomaly score S is defined as: .

[0010] In a second aspect, the present invention provides a cableway wheel assembly abnormality detection system, comprising: The acoustic signal acquisition module is used to collect audio data of the cableway wheel group in real time under the working state through acoustic sensors deployed on the brackets on both sides of the cableway wheel group; A data preprocessing module is used to preprocess the collected audio data of the cableway wheel group in the working state to generate a feature vector; The feature vector reconstruction module is used to input the feature vector generated by the data preprocessing module into the pre-trained convolutional autoencoder model and output the reconstructed feature vector; Anomaly score calculation module, used to calculate the mean square error between the feature vector of the input convolutional autoencoder model module and the reconstructed feature vector, and use the calculated mean square error as the anomaly score; The abnormality determination module is used to compare the abnormality score calculated by the abnormality score calculation module with a preset abnormality score threshold, and determine that an abnormality exists in the cableway wheel assembly when the abnormality score exceeds the preset abnormality score threshold.

[0011] A further improvement of this technical solution is that the acoustic sensor adopts an environmental isolation microphone, with a pair installed on each side of the bracket.

[0012] A further improvement of this technical solution is that the data preprocessing module is specifically used to: Perform short-time Fourier transform (STFT) on the collected audio data with a window size of 1024 points and a frame shift of 512 points to generate a spectrum with a dimension of 513×312; The generated spectrogram is converted into a Mel spectrogram through a 128-order Mel filter bank; Perform logarithmic transformation on the Mel spectrum to generate a Log-Mel spectrum; Perform 5-frame context splicing on the Log-Mel spectrum to generate a feature matrix with a dimension of 128×5; The generated feature matrix with a dimension of 128×5 is flattened into a 640-dimensional feature vector.

[0013] In a third aspect, the present invention provides a terminal, comprising: processor, memory, wherein The memory is used to store computer programs, The processor is used to call and run the computer program from the memory, so that the terminal executes the above-mentioned terminal method.

[0014] In a fourth aspect, the present invention provides a computer storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the methods described in the above aspects.

[0015] The beneficial effects of the present invention are: Real-time online monitoring and timely warning capabilities: This invention deploys environmentally isolated microphone arrays on the upper portions of brackets on both sides of the cableway sheave assembly, enabling real-time collection and online analysis of audio data during the sheave assembly's operating state. Compared to traditional manual visual inspections and periodic shutdown inspections, this solution continuously monitors equipment status without interrupting cableway operation. In particular, the use of a convolutional autoencoder (CNN-AE) model to perform real-time reconstruction error calculations on audio features can trigger early warnings for early faults, such as tiny cracks in wheel assembly components or initial poor lubrication. This accelerates fault detection to the incipient stage and avoids the missed detection delays often associated with traditional vibration or thermal imaging methods due to insufficient sensitivity.

[0016] Significantly Reduced Deployment and Maintenance Costs: Existing acoustic detection technologies, such as vibration sensors and ultrasonic devices, require fixed hardware installation at each wheel location and periodic calibration and maintenance. This new technology, however, utilizes a contactless microphone array, requiring only a pair of environmentally isolated sensors on a single bracket, thus reducing costs. Furthermore, model training and inference are performed on general-purpose edge computing devices or cloud servers, eliminating the need for frequent sensor calibration and reducing operational and maintenance costs.

[0017] Strong environmental adaptability and anti-interference capabilities: To address the complex environmental noise (such as wind noise and passenger noise) and equipment structure resonance interference in mountain cableway scenarios, this invention innovatively introduces Log-Mel spectrum feature extraction and 5-frame context splicing technology. A 128-order Mel filter group is used to simulate the nonlinear frequency perception characteristics of the human ear, and a logarithmic transformation is combined to eliminate the influence of differences in recording conditions and improve feature discrimination. At the model design level, the 95% normal sample reconstruction error quantile is used as the dynamic threshold, which reduces the false alarm rate compared to the traditional fixed threshold method. In actual measurements, it can effectively distinguish between environmental noise and real fault signals.

[0018] High Sensitivity and Generalization: Traditional detection methods based on spectral analysis or machine learning rely on a large number of labeled fault samples. This new method, however, employs an unsupervised learning framework, requiring only normal audio data for model training. By extracting hierarchical time-frequency features through a convolutional neural network, the model improves its accuracy for identifying six typical fault types, including bearing wear and loose mounting. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 A schematic flow chart of a method according to an embodiment of the present invention.

[0021] Figure 2 A schematic block diagram of a system according to an embodiment of the present invention.

[0022] Figure 3 A schematic diagram of the structure of a terminal provided by an embodiment of the present invention.

[0023] 210 is an acoustic signal acquisition module, 220 is a data preprocessing module, 230 is a feature vector reconstruction module, 240 is an anomaly score calculation module, and 250 is an anomaly determination module. DETAILED DESCRIPTION

[0024] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions of the present invention will be clearly and completely described below in conjunction with the drawings in the specific embodiments. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0026] Figure 1 The present invention provides a schematic flow chart of a method for detecting abnormalities in a cableway wheel assembly. Figure 1 The execution subject can be a cableway wheel group abnormality detection system. According to different requirements, the order of the steps in the flow chart can be changed, and some can be omitted.

[0027] like Figure 1 As shown, the method includes: S1. Deploy acoustic sensors on the brackets on both sides of the cableway sheave to collect audio data in real time when the cableway sheave is in operation. S2. Preprocessing the collected audio data of the cableway wheel assembly in the working state to generate a feature vector; S3, input the generated feature vector into the pre-trained convolutional autoencoder model and output the reconstructed feature vector; S4. Calculate the mean square error between the feature vector of the input convolutional autoencoder model and the reconstructed feature vector, and use the calculated mean square error as the anomaly score; S5. Compare the abnormality score with a preset abnormality score threshold, and when the abnormality score exceeds the preset abnormality score threshold, determine that the cableway wheel assembly has an abnormality.

[0028] To facilitate understanding of the present invention, the cableway wheel group abnormality detection method provided by the present invention is further described below based on the principle of the cableway wheel group abnormality detection method of the present invention and the process of abnormality detection of the cableway wheel group in the embodiment.

[0029] Wherein, step S2 includes: S2.1. Perform short-time Fourier transform (STFT) on the collected audio data with a window size of 1024 points and a frame shift of 512 points to generate a spectrogram with a dimension of 513 × 312. S2.2, converting the generated spectrogram into a Mel spectrogram through a 128-order Mel filter bank; S2.3, performing logarithmic transformation on the Mel spectrum to generate a Log-Mel spectrum; S2.4, perform 5-frame context splicing on the Log-Mel spectrum to generate a feature matrix with a dimension of 128×5; S2.5. Flatten the feature matrix with the generated dimension of 128×5 into a 640-dimensional feature vector.

[0030] Specifically, a short-time Fourier transform (STFT) is performed on the collected audio data, with a window size of 1024 points and a frame shift of 512 points. The STFT transforms the time-domain signal into a spectrogram, capturing the frequency distribution of the audio signal at different time points. The resulting spectrogram has dimensions of 513 × 312, where 513 represents the number of points on the frequency axis and 312 represents the number of frames on the time axis.

[0031] The generated spectrogram is processed through a 128-order Mel filter bank to convert it into a Mel-spectrogram. The Mel filter bank simulates the nonlinear frequency perception of human hearing, mapping the frequency axis to the Mel scale, thereby extracting features relevant to human auditory perception. The resulting Mel-spectrogram has a dimension of 128 × 312, where 128 represents the number of Mel filters and 312 represents the number of frames on the time axis.

[0032] Perform a logarithmic transformation on the Mel-spectrogram to generate a Log-Mel spectrogram. This transformation converts the spectrogram's amplitude values ​​to a logarithmic scale, eliminating the effects of recording conditions or ambient noise while enhancing the distinguishing power of features. The resulting Log-Mel spectrogram still has a dimension of 128 × 312.

[0033] The Log-Mel spectrogram is concatenated for five frames to generate a 128×5 feature matrix. Specifically, for each time instant t, five frames (t−2:t+2) are concatenated along the frequency band to form a feature matrix containing all five frames of context. This introduces temporal context and enhances the model's ability to perceive short-term dynamics.

[0034] The resulting 128×5 feature matrix is ​​flattened into a 640-dimensional feature vector. Flattening converts the two-dimensional feature matrix into a one-dimensional vector for input into the subsequent convolutional autoencoder model. The flattened feature vector has a length of 640 and serves as the input feature of the model.

[0035] In addition, the 640-dimensional feature vector generated after preprocessing is input into the pre-trained convolutional autoencoder model. The dimension of the feature vector is 640, represented as a one-dimensional array, where each element corresponds to a feature value.

[0036] Specifically, the encoder in the convolutional autoencoder model in step S3 includes: S3.1a. Set up the first convolutional layer, use 64 filters of size 5 to perform one-dimensional convolution on the 640-dimensional feature vector, after batch normalization and ReLU activation, the output dimension is 636×64 feature matrix; S3.1b. Set the first pooling layer to perform max pooling with a stride of 2 on the feature matrix of dimension 636×64, and output a feature matrix of dimension 318×64. S3.1c. Set up the second convolutional layer, use 128 filters of size 5 to perform one-dimensional convolution on the feature matrix of dimension 318 × 64, and output a feature matrix of dimension 314 × 128; S3.1d. Set the second pooling layer to perform max pooling with a stride of 2 on the feature matrix of dimension 314 × 128, and output a feature matrix of dimension 157 × 128. S3.1e. Set the third convolutional layer to perform one-dimensional convolution on the feature matrix of dimension 157 × 128 using 256 filters of size 5, and output a feature matrix of dimension 153 × 256. S3.1f. Set the third pooling layer to perform max pooling with a stride of 2 on the feature matrix of dimension 153×256, and output a feature matrix of dimension 76×256. S3.1g. Set up a flattening layer and a fully connected layer to flatten the 76×256 feature matrix into a one-dimensional vector. Then, pass it through a fully connected layer with 128 neurons and a ReLU activation to output a feature vector with a dimension of 128. S3.1h. Set up a latent space layer and linearly map the feature vector of dimension 128 to 16 dimensions through a fully connected layer containing 16 neurons to generate a latent vector z of dimension 16.

[0037] Furthermore, the decoder in the convolutional autoencoder model in step S3 includes: S3.2a. Set up a fully connected expansion layer to linearly expand the latent vector z of dimension 16 through a fully connected expansion layer containing 80 × 256 neurons to generate a feature vector of dimension 20480, and reshape it into a feature matrix of dimension 80 × 256. S3.2b. Set up the first deconvolution layer, use 256 filters of size 5 to perform one-dimensional deconvolution on the feature matrix of dimension 80×256, and generate a feature matrix of dimension 76×256 after batch normalization and ReLU activation; S3.2c. Set the first upsampling layer to upsample the feature matrix of dimension 76×256 by a factor of 2 using a pre-stored bilinear interpolation algorithm to generate a feature matrix of dimension 152×256. S3.2d. Set up the second deconvolution layer, use 128 filters of size 5 to perform one-dimensional deconvolution on the feature matrix of dimension 152×256, and generate a feature matrix of dimension 148×128 after batch normalization and ReLU activation. S3.2e. Set a second upsampling layer to upsample the feature matrix of dimension 148 × 128 by a factor of 2 using a pre-stored bilinear interpolation algorithm to generate a feature matrix of dimension 296 × 128. S3.2f. Set up the third deconvolution layer, use 64 filters of size 5 to perform one-dimensional deconvolution on the feature matrix of dimension 296×128, and generate a feature matrix of dimension 292×64 after batch normalization and ReLU activation; S3.2g. Set a third upsampling layer to upsample the feature matrix of dimension 292 × 64 by a factor of 2 using a pre-stored bilinear interpolation algorithm to generate a feature matrix of dimension 584 × 64. S3.2h. Set the output convolution layer and use a filter of size 5 to perform one-dimensional convolution on the feature matrix of dimension 584×64 to generate a reconstructed feature vector of dimension 640×1.

[0038] The convolutional autoencoder model uses unsupervised learning, trained only on normal cableway sheave audio data. The mean squared error (MSE) is used as the loss function during training, and the Adam optimization algorithm is used as the optimizer. The initial learning rate is 0.001, and the training cycle is 100 training cycles (or iterations), or stopped early when the reconstruction error on the validation set no longer decreases. The formula for calculating the mean squared error (MSE) is: ;in, is the input feature vector element; To reconstruct vector elements; is the feature dimension, .

[0039] During the training process of the convolutional autoencoder model, the mean square error (MSE) is used as the loss function to measure the reconstruction ability of the model. Specifically: Input: The input of the model is the preprocessed audio feature vector in normal state; Output: The goal of the model is to reconstruct the input feature vector as accurately as possible; Loss function: This function measures the model's reconstruction error by calculating the mean squared error (MSE) between the input feature vector and the reconstructed feature vector output by the model. The smaller the loss function value, the better the model's reconstruction ability. Optimization goal: During training, this loss function is minimized through an optimization algorithm (such as the Adam optimizer) so that the model can learn the characteristic distribution of normal audio data and effectively identify anomalies in subsequent anomaly detection.

[0040] A further improvement of this technical solution is that step S4 includes: The 640-dimensional feature vector of the input convolutional autoencoder model is divided into frame vector , each frame contains F frequency band features; the reconstructed feature vector with a dimension of 640×1 output by the convolutional autoencoder model is divided into frame vector , each frame contains F frequency band features; Calculate the frame-level mean square error between the feature vector and the reconstructed feature vector : ; The anomaly score S is defined as: .

[0041] The 640-dimensional feature vector of the input convolutional autoencoder model is divided into multiple frame vectors, each frame contains F frequency band features. Assume that each frame contains frequency band features, the 640-dimensional feature vector can be divided into frame vectors.

[0042] The reconstructed feature vector of 640×1 output by the convolutional autoencoder model is also divided into multiple frame vectors, each frame contains In this way, the reconstructed feature vector can also be divided into 5 frame vectors.

[0043] In addition, step S5 includes: Based on the audio data of the cableway sheave under normal operating conditions, a convolutional autoencoder model was trained to calculate the reconstruction error distribution of the normal audio data. The 95th percentile of this distribution was selected as the anomaly score threshold Z, which is used to distinguish between normal and abnormal conditions.

[0044] During real-time monitoring, the anomaly score S of each audio clip is calculated as follows: Preprocess each audio clip to generate a 640-dimensional feature vector; Input the feature vector into the convolutional autoencoder model to obtain the reconstructed feature vector; Calculate the frame-level mean squared error between the feature vector and the reconstructed feature vector ; The average or maximum of the mean squared errors of all frames is taken as the anomaly score S.

[0045] Compare the calculated anomaly score S with the preset anomaly score threshold Z: If yes, it is determined that there is an abnormality in the cableway wheel assembly; If yes, it is determined that the cableway wheel assembly is operating normally.

[0046] Specifically, the calculation formula of the anomaly score threshold Z is: ;in, represents the 95th percentile; represents the anomaly score of the i-th training sample; Z is the total number of training samples. This threshold is used to detect anomalies in real-time audio data. This method can detect potential cableway sheave failures at an early stage, improving the operational safety and reliability of the cableway system.

[0047] In some embodiments, the cableway sheave abnormality detection system 200 may include multiple functional modules composed of computer program segments. The computer program of each program segment in the cableway sheave abnormality detection system 200 may be stored in a memory of a computer device and executed by at least one processor to perform (see Figure 1 Description) Function of abnormality detection of cableway sheaves.

[0048] In this embodiment, the cableway wheel assembly abnormality detection system 200 can be divided into multiple functional modules according to the functions it performs, such as Figure 2 As shown. The functional modules may include: an acoustic signal acquisition module 210, a data preprocessing module 220, a feature vector reconstruction module 230, an anomaly score calculation module 240, and an anomaly determination module 250. A module as referred to in the present invention refers to a series of computer program segments that can be executed by at least one processor and can perform fixed functions, and is stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0049] Specifically, the acoustic signal acquisition module is used to collect audio data of the cableway wheel group in the working state in real time through acoustic sensors deployed on the brackets on both sides of the cableway wheel group; the data preprocessing module is used to preprocess the collected audio data of the cableway wheel group in the working state to generate a feature vector; the feature vector reconstruction module is used to input the feature vector generated by the data preprocessing module into the pre-trained convolutional autoencoder model and output the reconstructed feature vector; the anomaly score calculation module is used to calculate the mean square error between the feature vector input to the convolutional autoencoder model module and the reconstructed feature vector, and use the calculated mean square error as the anomaly score; the anomaly judgment module is used to compare the anomaly score calculated by the anomaly score calculation module with a preset anomaly score threshold, and when the anomaly score exceeds the preset anomaly score threshold, judge that there is an anomaly in the cableway wheel group.

[0050] The acoustic sensors utilize environmentally isolated microphones (EnviroMic-200), with a pair mounted on each side bracket. These microphones can capture subtle sound changes, making them suitable for detecting abnormal noise from cableway sheaves. They effectively isolate ambient noise, such as wind noise and passenger conversations, ensuring that the collected audio data primarily reflects the operating status of the cableway sheaves. They are also capable of long-term, stable operation in harsh outdoor environments, making them suitable for mountain cableway operations.

[0051] Furthermore, the data preprocessing module is specifically used to: Perform short-time Fourier transform (STFT) on the collected audio data with a window size of 1024 points and a frame shift of 512 points to generate a spectrum with a dimension of 513×312; The generated spectrogram is converted into a Mel spectrogram through a 128-order Mel filter bank; Perform logarithmic transformation on the Mel spectrum to generate a Log-Mel spectrum; Perform 5-frame context splicing on the Log-Mel spectrum to generate a feature matrix with a dimension of 128×5; The generated feature matrix with a dimension of 128×5 is flattened into a 640-dimensional feature vector.

[0052] The Short-Time Fourier Transform (STFT) transforms audio signals from the time domain to the time-frequency domain, preserving information in both dimensions. With a window size of 1024 points and a frame shift of 512 points, the transformed spectrogram achieves a good balance between time and frequency resolution. This not only captures the frequency variations of the audio signal over a relatively short period of time, but also reflects, to a certain extent, the distribution of different frequency components, providing a rich information foundation for subsequent feature extraction. The resulting spectrogram has a dimension of 513×312, where 513 represents the number of points on the frequency axis and 312 represents the number of frames on the time axis. This dimensional spectrogram is well-suited for subsequent processing, providing a suitable input format for subsequent filter bank processing and feature extraction, ensuring the consistency and efficiency of the entire data preprocessing process.

[0053] Converting the spectrogram to a Mel spectrogram using a 128-order Mel filter bank maps the frequency axis to the Mel scale. The Mel scale is a nonlinear frequency scale closely related to human auditory perception. It better simulates the human ear's perception of sounds of different frequencies, making the extracted features more consistent with human auditory perception, thereby improving the model's ability to analyze and understand audio signals. The use of the Mel filter bank reduces the dimensionality of the spectrogram to a certain extent, reducing the original 513 frequency points to 128 Mel frequency bands. It also highlights the frequency components in the audio signal that are more critical to auditory perception and removes some redundant information. This helps improve the efficiency and effectiveness of subsequent feature extraction, allowing the model to focus more on features that are important for anomaly detection.

[0054] Performing a logarithmic transformation on the Mel spectrogram to generate a Log-Mel spectrogram converts the amplitude values ​​of the spectrogram to a logarithmic scale. This transformation can effectively compress spectrograms with a wide amplitude range, making the differences between frequency components of different amplitudes more obvious, enhancing the ability to distinguish features, and helping the model better identify and distinguish the subtle differences between normal and abnormal audio signals. Logarithmic transformation can also eliminate amplitude differences caused by factors such as recording equipment and recording environment, making the feature representation of audio data collected under different conditions more consistent, improving the model's adaptability and robustness to audio data under different recording conditions, and reducing misjudgments caused by different recording conditions.

[0055] The Log-Mel spectrogram is concatenated with five frames of context to generate a feature matrix with a dimension of 128×5. This allows the model to incorporate information from the preceding and following frames into each time frame, forming a feature matrix containing five frames of contextual information. The introduction of this contextual information allows the model to not only focus on the features of the current frame, but also consider changes in the features of the preceding and following frames, thereby better capturing short-term dynamic information in the audio signal. This improves the model's perception of the temporal continuity and changing trends of the audio signal, and facilitates more accurate detection of abnormal signals. The feature matrix generated through contextual concatenation expands the original 128-dimensional features of a single frame to a 640-dimensional feature matrix, enhancing the expressive power of the features and providing richer feature information for subsequent model training. This helps improve the model's classification and recognition performance of audio signals, enabling the model to more effectively learn the characteristic patterns of audio signals, thereby improving the accuracy of anomaly detection.

[0056] Flattening the generated 128×5 feature matrix into a 640-dimensional feature vector converts the two-dimensional feature matrix into a one-dimensional feature vector, making it compatible with the input format of the convolutional autoencoder model. This flattening operation ensures that the feature vector can be smoothly input into the model for subsequent processing and analysis. It is a key step in connecting the data preprocessing module and the model training module, ensuring the normal operation of the entire anomaly detection system. The flattened feature vector simplifies the data structure, facilitating subsequent calculations and processing. During model training, the one-dimensional feature vector enables more efficient matrix operations and parameter updates, improving the efficiency and speed of model training while also reducing the complexity of model training, making the entire anomaly detection system more efficient and stable.

[0057] Figure 3 This is a structural diagram of a terminal 300 provided in an embodiment of the present invention. The terminal 300 can be used to execute the cableway wheel group abnormality detection method provided in an embodiment of the present invention.

[0058] The terminal 300 may include a processor 310, a memory 320, and a communication module 330. These components communicate via one or more buses. Those skilled in the art will appreciate that the server structure shown in the figure does not limit the present invention. The server structure may be a bus structure or a star structure, and may include more or fewer components than shown, or may combine certain components or arrange the components differently.

[0059] Memory 320 can be used to store execution instructions of processor 310. Memory 320 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk. When the execution instructions in memory 320 are executed by processor 310, terminal 300 can perform some or all of the steps in the above-described method embodiments.

[0060] The processor 310 is the control center of the storage terminal. It uses various interfaces and lines to connect various parts of the entire electronic terminal. It executes various functions of the electronic terminal and / or processes data by running or executing software programs and / or modules stored in the memory 320, and calling data stored in the memory. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions. For example, the processor 310 can only include a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single computing core or multiple computing cores.

[0061] The communication module 330 is used to establish a communication channel so that the storage terminal can communicate with other terminals, receive user data sent by other terminals, or send user data to other terminals.

[0062] The present invention also provides a computer storage medium, wherein the computer storage medium may store a program that, when executed, may include some or all of the steps of each embodiment provided herein. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0063] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solutions in the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code, and includes instructions for causing a computer terminal (which can be a personal computer, a server, or a second terminal, a network terminal, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0064] In this specification, the same or similar parts between the various embodiments can be referred to each other. In particular, for the terminal embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiment.

[0065] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or modules, and can be electrical, mechanical or other forms.

[0066] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of the present embodiment according to actual needs.

[0067] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0068] Although the present invention has been described in detail with reference to the accompanying drawings and in conjunction with preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, persons of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and such modifications or substitutions shall be within the scope of the present invention. Any changes or substitutions that can be easily conceived by persons skilled in the art within the technical scope disclosed in the present invention shall be within the scope of protection of the present invention.

Claims

1. A method for detecting abnormality of a cableway wheel group, characterized in that: include: S1. Deploy acoustic sensors on the brackets on both sides of the cableway sheave to collect audio data in real time when the cableway sheave is in operation. S2. Preprocessing the collected audio data of the cableway wheel assembly in the working state to generate a feature vector; S3, input the generated feature vector into the pre-trained convolutional autoencoder model and output the reconstructed feature vector; S4. Calculate the mean square error between the feature vector of the input convolutional autoencoder model and the reconstructed feature vector, and use the calculated mean square error as the anomaly score; S5. Compare the abnormality score with a preset abnormality score threshold, and when the abnormality score exceeds the preset abnormality score threshold, determine that the cableway wheel assembly has an abnormality.

2. The cableway wheel assembly abnormality detection method according to claim 1, characterized in that: Step S2 includes: S2.

1. Perform short-time Fourier transform (STFT) on the collected audio data with a window size of 1024 points and a frame shift of 512 points to generate a spectrogram with a dimension of 513 × 312. S2.2, converting the generated spectrogram into a Mel spectrogram through a 128-order Mel filter bank; S2.3, performing logarithmic transformation on the Mel spectrum to generate a Log-Mel spectrum; S2.4, perform 5-frame context splicing on the Log-Mel spectrum to generate a feature matrix with a dimension of 128×5; S2.

5. Flatten the feature matrix with the generated dimension of 128×5 into a 640-dimensional feature vector.

3. The cableway wheel assembly abnormality detection method according to claim 2, characterized in that: The encoder in the convolutional autoencoder model in step S3 includes: S3.1a. Set up the first convolutional layer, use 64 filters of size 5 to perform one-dimensional convolution on the 640-dimensional feature vector, after batch normalization and ReLU activation, the output dimension is 636×64 feature matrix; S3.1b. Set the first pooling layer to perform max pooling with a stride of 2 on the feature matrix of dimension 636×64, and output a feature matrix of dimension 318×64. S3.1c. Set up the second convolutional layer, use 128 filters of size 5 to perform one-dimensional convolution on the feature matrix of dimension 318 × 64, and output a feature matrix of dimension 314 × 128; S3.1d. Set the second pooling layer to perform max pooling with a stride of 2 on the feature matrix of dimension 314 × 128, and output a feature matrix of dimension 157 × 128. S3.1e. Set the third convolutional layer to perform one-dimensional convolution on the feature matrix of dimension 157 × 128 using 256 filters of size 5, and output a feature matrix of dimension 153 × 256. S3.1f. Set the third pooling layer to perform max pooling with a stride of 2 on the feature matrix of dimension 153×256, and output a feature matrix of dimension 76×256. S3.1g. Set up a flattening layer and a fully connected layer to flatten the 76×256 feature matrix into a one-dimensional vector. Then, pass it through a fully connected layer with 128 neurons and a ReLU activation to output a feature vector with a dimension of 128. S3.1h. Set up a latent space layer and linearly map the feature vector of dimension 128 to 16 dimensions through a fully connected layer containing 16 neurons to generate a latent vector z of dimension 16.

4. The cableway wheel assembly abnormality detection method according to claim 3, characterized in that: The decoder in the convolutional autoencoder model in step S3 includes: S3.2a. Set up a fully connected expansion layer to linearly expand the latent vector z of dimension 16 through a fully connected expansion layer containing 80 × 256 neurons to generate a feature vector of dimension 20480, and reshape it into a feature matrix of dimension 80 × 256. S3.2b. Set up the first deconvolution layer, use 256 filters of size 5 to perform one-dimensional deconvolution on the feature matrix of dimension 80×256, and generate a feature matrix of dimension 76×256 after batch normalization and ReLU activation; S3.2c. Set the first upsampling layer to upsample the feature matrix of dimension 76×256 by a factor of 2 using a pre-stored bilinear interpolation algorithm to generate a feature matrix of dimension 152×256. S3.2d. Set up the second deconvolution layer, use 128 filters of size 5 to perform one-dimensional deconvolution on the feature matrix of dimension 152×256, and generate a feature matrix of dimension 148×128 after batch normalization and ReLU activation. S3.2e. Set a second upsampling layer to upsample the feature matrix of dimension 148 × 128 by a factor of 2 using a pre-stored bilinear interpolation algorithm to generate a feature matrix of dimension 296 × 128. S3.2f. Set up the third deconvolution layer, use 64 filters of size 5 to perform one-dimensional deconvolution on the feature matrix of dimension 296×128, and generate a feature matrix of dimension 292×64 after batch normalization and ReLU activation; S3.2g. Set a third upsampling layer to upsample the feature matrix of dimension 292 × 64 by a factor of 2 using a pre-stored bilinear interpolation algorithm to generate a feature matrix of dimension 584 × 64. S3.2h. Set the output convolution layer and use a filter of size 5 to perform one-dimensional convolution on the feature matrix of dimension 584×64 to generate a reconstructed feature vector of dimension 640×1.

5. The cableway wheel assembly abnormality detection method according to claim 4, characterized in that: Step S4 includes: The 640-dimensional feature vector of the input convolutional autoencoder model is divided into frame vector , each frame contains F frequency band features; the reconstructed feature vector with a dimension of 640×1 output by the convolutional autoencoder model is divided into frame vector , each frame contains F frequency band features; Calculate the frame-level mean square error between the feature vector and the reconstructed feature vector : ; The anomaly score S is defined as: .

6. A cableway wheel abnormality detection system, characterized in that: include: The acoustic signal acquisition module is used to collect audio data of the cableway wheel group in real time under the working state through acoustic sensors deployed on the brackets on both sides of the cableway wheel group; A data preprocessing module is used to preprocess the collected audio data of the cableway wheel group in the working state to generate a feature vector; The feature vector reconstruction module is used to input the feature vector generated by the data preprocessing module into the pre-trained convolutional autoencoder model and output the reconstructed feature vector; Anomaly score calculation module, used to calculate the mean square error between the feature vector of the input convolutional autoencoder model module and the reconstructed feature vector, and use the calculated mean square error as the anomaly score; The abnormality determination module is used to compare the abnormality score calculated by the abnormality score calculation module with a preset abnormality score threshold, and determine that an abnormality exists in the cableway wheel assembly when the abnormality score exceeds the preset abnormality score threshold.

7. The cableway wheel assembly abnormality detection system according to claim 6, characterized in that: The acoustic sensors use environmentally isolated microphones, with a pair mounted on each side of the bracket.

8. The cableway wheel assembly abnormality detection system according to claim 6, characterized in that: The data preprocessing module is specifically used for: Perform short-time Fourier transform (STFT) on the collected audio data with a window size of 1024 points and a frame shift of 512 points to generate a spectrum with a dimension of 513×312; The generated spectrogram is converted into a Mel spectrogram through a 128-order Mel filter bank; Perform logarithmic transformation on the Mel spectrum to generate a Log-Mel spectrum; Perform 5-frame context splicing on the Log-Mel spectrum to generate a feature matrix with a dimension of 128×5; The generated feature matrix with a dimension of 128×5 is flattened into a 640-dimensional feature vector.

9. A terminal, characterized in that: include: processor; a memory for storing execution instructions of the processor; The processor is configured to execute the method according to any one of claims 1 to 5.

10. A computer-readable storage medium storing a computer program, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Device signal anomaly detection method and device based on auto-encoder, and medium

    CN119004085A

  • Method for detecting abnormal sound of track inspection robot based on convolutional auto-encoder

    CN119296580A

  • Sound anomaly detection method and device based on Transform model, equipment and medium

    CN120340527A

  • Fault signal locating and identifying method of industrial equipment based on microphone array

    US20230152187A1

  • Power plant equipment state auditory monitoring method merging frequency band top-down attention mechanism

    WO2023245991A1

Cited By

  • Power equipment fault detection method, device, equipment and medium

    CN122050436A