Audio data-based device anomaly detection method, apparatus, device, medium, and product

By analyzing the time-frequency features and calculating the distribution offset based on audio data, the accuracy problem of device anomaly detection in existing technologies is solved, and accurate detection of device anomalies is achieved without the need for negative sample features.

CN122024766BActive Publication Date: 2026-08-04中国电气装备集团科学技术研究院有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
中国电气装备集团科学技术研究院有限公司
Filing Date
2026-04-10
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing methods for detecting device anomalies based on audio data have poor accuracy, mainly because the audio data during device operation is non-stationary and the time-frequency coupling characteristics are complex, making it difficult to stably characterize the difference between normal and abnormal states of the device. Furthermore, the low probability of anomalies leads to a scarcity of audio sample data.

Method used

By extracting the log-Mel spectrum of the target audio data, time-frequency features are obtained, distribution parameters and offsets at each time-frequency position are calculated, and the distribution offset is measured using the reference distribution parameters of the positive samples to determine the device anomaly detection results.

Benefits of technology

Accurate detection of equipment anomalies was achieved without the need for negative sample features, improving detection precision and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024766B_ABST
    Figure CN122024766B_ABST
Patent Text Reader

Abstract

The application discloses a device anomaly detection method and device based on audio data, equipment, medium and product. The method comprises the following steps: extracting time-frequency features of target audio data according to the log-mel spectrum of the target audio data; extracting channel feature vectors matched with each time-frequency position from the time-frequency features of the target audio data, and determining distribution parameters of each time-frequency position according to the channel feature vectors matched with each time-frequency position; calculating distribution offsets of each time-frequency position according to the distribution parameters of each time-frequency position and reference distribution parameters; aggregating the distribution offsets of each time-frequency position according to the frequency dimension to obtain time-dimension distribution offsets, and determining a device anomaly detection result. The present scheme solves the problem of poor accuracy of the existing device anomaly detection method based on audio data, and can realize accurate detection of device anomalies without reference to negative sample features by measuring the distribution offset of the audio signal according to the reference distribution parameters of the positive samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, device, medium, and product for detecting device anomalies based on audio data. Background Technology

[0002] Equipment malfunctions can be detected through audio data during operation, such as bearing wear, structural loosening, and electrical faults. By deploying acoustic sensors near the equipment to collect audio data during operation and analyzing this data, it is possible to monitor the equipment's operating status and provide early warnings of potential faults without physical contact with the equipment.

[0003] Currently, existing methods for detecting device anomalies based on audio data mainly include: (1) judging device anomalies by comparing the characteristics of audio signals such as short-time energy and zero-crossing rate with corresponding thresholds; (2) training a deep learning model to learn features from a large number of audio samples and using the trained deep learning model to detect device anomalies.

[0004] For method (1), the audio data during equipment operation has obvious non-stationarity and time-frequency coupling characteristics. The spectral structure and energy distribution of sound will change under different operating conditions. Artificially designed features such as short-time energy and zero-crossing rate are insufficient to express the complex time-frequency structure changes in audio signals, making it difficult to stably characterize the difference between normal and abnormal states of the equipment, thus affecting the accuracy of equipment anomaly detection. For method (2), the probability of anomalies occurring throughout the entire operating life of the equipment is low and the types are complex. It is difficult to collect audio data during abnormal periods of the equipment. Therefore, there is a lack of audio sample data of equipment anomalies, which limits the application of supervised learning models in equipment anomaly detection. Summary of the Invention

[0005] This invention provides a device anomaly detection method, apparatus, equipment, medium, and product based on audio data, to solve the problem of poor accuracy in existing audio data-based device anomaly detection methods. By measuring the distribution offset of the audio signal based on the reference distribution parameters of positive samples, accurate detection of device anomalies can be achieved without referencing the features of negative samples.

[0006] According to one aspect of the present invention, a device anomaly detection method based on audio data is provided, the method comprising: Acquire target audio data, the target audio data including at least one audio segment to be detected, the audio segment being a time-domain discrete audio signal; Based on the log-Mel spectrum of the target audio data, the time-frequency features of the target audio data are extracted; the time-frequency features include three feature dimensions: time, frequency, and channel. Extract the channel feature vector matching each time-frequency position from the time-frequency features of the target audio data, and determine the distribution parameters of each time-frequency position based on the channel feature vector matching each time-frequency position; the distribution parameters include the mean vector and the covariance matrix. The distribution offset of each time-frequency position is calculated based on the distribution parameters of each time-frequency position and the reference distribution parameters of each time-frequency position obtained in advance; the reference distribution parameters include the reference mean vector and reference covariance matrix of each time-frequency position determined based on the audio data of the normal operation period of the device. The distribution offsets of each time and frequency position are aggregated according to the frequency dimension to obtain the distribution offset in the time dimension. Based on the distribution offset in the time dimension, the equipment anomaly detection result is determined.

[0007] According to another aspect of the present invention, a device anomaly detection apparatus based on audio data is provided, the apparatus comprising: An audio data acquisition module is used to acquire target audio data, the target audio data including at least one audio segment to be detected, the audio segment being a time-domain discrete audio signal; The time-frequency feature extraction module is used to extract the time-frequency features of the target audio data based on the log-Mel spectrum of the target audio data; the time-frequency features include three feature dimensions: time, frequency, and channel. The distribution parameter determination module is used to extract the channel feature vectors matching each time-frequency position from the time-frequency features of the target audio data, and determine the distribution parameters of each time-frequency position based on the channel feature vectors matching each time-frequency position; the distribution parameters include the mean vector and the covariance matrix; The offset determination module is used to calculate the distribution offset of each time-frequency position based on the distribution parameters of each time-frequency position and the reference distribution parameters of each time-frequency position obtained in advance; the reference distribution parameters include the reference mean vector and reference covariance matrix of each time-frequency position determined based on the audio data of the normal operation period of the device. The detection result determination module is used to aggregate the distribution offset of each time and frequency position according to the frequency dimension to obtain the distribution offset in the time dimension, and determine the equipment anomaly detection result based on the distribution offset in the time dimension.

[0008] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the device anomaly detection method based on audio data according to any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the device anomaly detection method based on audio data according to any embodiment of the present invention.

[0010] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the device anomaly detection method based on audio data as described in any embodiment of the present invention.

[0011] The technical solution of this invention involves acquiring target audio data, which includes at least one audio segment to be detected, wherein the audio segment is a time-domain discrete audio signal; extracting time-frequency features of the target audio data based on the log-Mel spectrum of the target audio data; the time-frequency features include three feature dimensions: time, frequency, and channel; extracting channel feature vectors matching each time-frequency position from the time-frequency features of the target audio data; determining the distribution parameters of each time-frequency position based on the matching channel feature vectors; the distribution parameters include a mean vector and a covariance matrix; calculating the distribution offset of each time-frequency position based on the distribution parameters of each time-frequency position and pre-acquired reference distribution parameters of each time-frequency position; the reference distribution parameters include a reference mean vector and a reference covariance matrix of each time-frequency position determined based on audio data during normal operation of the device; aggregating the distribution offsets of each time-frequency position according to the frequency dimension to obtain the distribution offset in the time dimension; and determining the device anomaly detection result based on the distribution offset in the time dimension. This technical solution solves the problem of poor accuracy in existing audio data-based device anomaly detection methods. By measuring the distribution offset of the audio signal based on the reference distribution parameters of positive samples, accurate device anomaly detection can be achieved without the need to refer to negative sample features.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart of a device anomaly detection method based on audio data according to Embodiment 1 of the present invention; Figure 2 This is a flowchart of a device anomaly detection method based on audio data according to Embodiment 2 of the present invention; Figure 3 This is a flowchart of a device anomaly detection method based on audio data according to Embodiment 3 of the present invention; Figure 4 This is a schematic diagram of a device anomaly detection device based on audio data according to Embodiment 4 of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device that implements the device anomaly detection method based on audio data according to an embodiment of the present invention. Detailed Implementation

[0015] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be used interchangeably where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices. The acquisition, storage, use, and processing of data in the technical solutions of this application all comply with the relevant provisions of national laws and regulations.

[0017] Example 1 Figure 1 This is a flowchart illustrating a device anomaly detection method based on audio data, as provided in Embodiment 1 of the present invention. This embodiment is applicable to anomaly detection scenarios in industrial equipment, particularly for situations where audio signals are used for device anomaly detection. This method can be executed by an audio data-based device anomaly detection device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes: S110. Acquire target audio data, wherein the target audio data includes at least one audio segment to be detected, and the audio segment is a time-domain discrete audio signal.

[0018] This solution can be executed by electronic devices such as computers and servers. These devices can communicate with acoustic sensors deployed near the device to acquire target audio data collected by the sensors during device operation. The target audio data can be a time-domain discrete audio signal that has not undergone device anomaly detection, and can be represented as... , It is an integer representing the sampling number.

[0019] To achieve continuous detection of long-duration audio data, the electronic device can segment the target audio data into time windows according to a preset sliding window length, obtaining the audio segments corresponding to each time window, which can be represented as follows: , Indicates the time window index. Indicates from audio data The first one extracted from the middle Audio segments within a time window.

[0020] The target audio data may include one or more audio segments. If the target audio data includes multiple audio segments, each audio segment may be an audio segment of a continuous time window or an audio segment of a non-continuous time window.

[0021] S120. Extract the time-frequency features of the target audio data based on the log-Mel spectrum of the target audio data; the time-frequency features include three feature dimensions: time, frequency, and channel.

[0022] Understandably, discrete-time audio signals are one-dimensional signals. Electronic devices can convert target audio data into a log-Mel spectrum, that is, map the one-dimensional signal into a two-dimensional "time-frequency band" matrix to explicitly characterize abnormal time-frequency features such as impulses, harmonics, and friction. Electronic devices can determine the log-Mel spectrum of the target audio data based on short-time Fourier transform and Mel filter banks.

[0023] Before determining the log-Mel spectrum of the target audio data, electronic devices can perform preprocessing operations such as resampling, amplitude normalization, DC component removal, and abnormal amplitude processing on the target audio data to reduce signal errors caused by differences in the sampling environment and noise disturbances caused by the sampling environment.

[0024] Specifically, the log-Mel spectrum can be expressed as: ;in, Indicates the audio segment index. Indicates the time frame index. This indicates the Mel band index or Mel filter index. Indicates frequency index, Indicates the first The audio clip in the first The complex spectrum after performing a short-time Fourier transform on each time frame. Indicates the first The Mel frequency band at frequency The weight of the position, Indicates the first The audio clip in the first The first time frame Logarithmic energy in a Mel band This indicates a preset positive number, usually an extremely small positive number, such as 0.00001, used to avoid... The problem of the logarithm not having a value when it is 0.

[0025] After obtaining the log-Mel spectrum of the target audio data, the electronic device can input the log-Mel spectrum into an audio pattern recognition model to extract the time-frequency features output from the intermediate layers of the audio pattern recognition model. The audio pattern recognition model can be a pre-trained deep learning model for audio pattern recognition, such as PANNs (Pre-trained Audio Neural Networks). The electronic device can determine the target layer from the intermediate layers of the audio pattern recognition model based on the dimensional requirements of the time-frequency features and the network structure of the audio pattern recognition model, and use the features output by the target layer as the time-frequency features of the target audio data. It should be noted that the time-frequency features of the target audio data are three-dimensional features, including time, frequency, and channel dimensions.

[0026] Specifically, the time-frequency characteristics of audio segments in the target audio data can be represented as: ,in, Indicates the audio segment index. This represents a time-frequency feature extraction network. Indicates the first Log-Mel spectrum of an audio segment The time dimension representing time-frequency features The frequency dimension representing the time-frequency characteristics, Channel dimension representing time-frequency characteristics.

[0027] Optionally, after extracting the time-frequency features of the target audio data, the method further includes: If the channel dimension of the time-frequency features of the target audio segment is greater than a preset dimension threshold, then the channel dimension of the time-frequency features of the target audio segment is reduced based on a preset dimensionality reduction method; the preset dimensionality reduction method is one of random channel sampling, linear projection, or principal component analysis.

[0028] To improve the stability of statistical distribution results while reducing the computational overhead of distribution metrics, electronic devices can limit the channel dimension of time-frequency features by setting a dimension threshold. For target audio segments where the channel dimension of the time-frequency features is greater than the preset dimension threshold, the electronic device can use a preset dimensionality reduction method to reduce the channel dimension of the time-frequency features of the target audio data to below the preset dimension threshold.

[0029] Specifically, the preset dimensionality reduction method can be any of the following: random channel sampling, linear projection, and principal component analysis (PCA). Random channel sampling reduces the channel dimensionality by randomly selecting a subset of channels from the time-frequency features, independent of the data structure or distribution of the time-frequency features; dimensionality reduction is achieved through simple random selection. Linear projection maps high-dimensional time-frequency features to a low-dimensional space, typically using a projection matrix. Principal component analysis reduces dimensionality by finding the direction of maximum variance in the time-frequency features (the principal component) and projecting the features onto that direction. Both linear projection and PCA map the original data to a new space, preserving the main directions of change in the original data and offering strong interpretability.

[0030] When the channel dimension of the time-frequency features of an audio segment exceeds a preset dimension threshold, this scheme performs targeted dimensionality reduction on the time-frequency features of the audio segment. This helps to save hardware resources such as computing and storage, avoids the impact of excessively high-dimensional channel features on the statistical distribution results, and ensures the stability of the statistical distribution results.

[0031] S130. Extract the channel feature vector matching each time-frequency position from the time-frequency features of the target audio data, and determine the distribution parameters of each time-frequency position based on the channel feature vector matching each time-frequency position; the distribution parameters include the mean vector and the covariance matrix.

[0032] As is easily understood, the time-frequency feature of the target audio data is a three-dimensional feature, with each "time-frequency" position in the time-frequency feature corresponding to a channel-dimensional feature vector. Electronic devices can extract the channel feature vectors matching each time-frequency position from the time-frequency feature of the target audio data, construct a statistical sample set, and statistically analyze the channel feature vectors matching each time-frequency position to determine the mean vector and covariance matrix of each time-frequency position.

[0033] Specifically, if the target audio data includes multiple audio segments, the electronic device constructs a statistical sample set from these audio segments. It then statistically analyzes the channel feature vectors matching the same time-frequency position for each audio segment in the statistical sample set to determine the mean vector and covariance matrix for that time-frequency position. If the target audio data includes only one audio segment, the electronic device can construct a statistical sample set from multiple time-frequency positions within a preset neighborhood of each time-frequency position. It then statistically analyzes the channel feature vectors corresponding to each time-frequency position in the statistical sample set to determine the mean vector and covariance matrix for that time-frequency position.

[0034] S140. Calculate the distribution offset of each time-frequency position based on the distribution parameters of each time-frequency position and the reference distribution parameters of each time-frequency position obtained in advance; the reference distribution parameters include the reference mean vector and reference covariance matrix of each time-frequency position determined based on the audio data of the normal operation period of the device.

[0035] Electronic devices can pre-determine the reference mean vector and reference covariance matrix for each time-frequency position based on audio data during normal operation. The audio data during normal operation can include multiple audio segments. The electronic device can extract the time-frequency features of each audio segment based on its log-Mel spectrum, treat each audio segment as a positive sample, extract the matching channel feature vector for each time-frequency position from the time-frequency features of each audio segment, and determine the reference distribution parameters for each time-frequency position.

[0036] Specifically, the reference mean vector at each time-frequency location can be represented as: The reference covariance matrix for each time-frequency location can be expressed as: ,in, This represents the positive sample index, which is the index of each audio segment in the audio data during the normal operation of the device. This represents the number of positive samples, that is, the number of audio segments in the audio data during the normal operation of the device. Indicates a time index. Indicates frequency index, Indicates positive samples In time and frequency position The channel feature vector.

[0037] Understandably, the reference mean vector and reference covariance matrix determined based on positive samples can characterize the Gaussian distribution at each time-frequency position under normal operating conditions. Similarly, the mean vector and covariance matrix determined based on the samples to be tested can characterize the Gaussian distribution at each time-frequency position under the testing condition. After obtaining the distribution parameters at each time-frequency position, the electronic device can use the reference distribution parameters at each time-frequency position as a standard, and calculate the distribution offset of the distribution parameters at each time-frequency position based on metrics such as KL divergence, JS divergence, and Hellinger distance. In other words, it can calculate the degree of shift of the Gaussian distribution to be tested relative to the normal Gaussian distribution.

[0038] S150. Aggregate the distribution offsets of each time and frequency position according to the frequency dimension to obtain the distribution offset in the time dimension. Based on the distribution offset in the time dimension, determine the equipment anomaly detection result.

[0039] As is easily understood, the distribution offset at each time-frequency position represents the distribution offset in both time and frequency dimensions. After obtaining the distribution offset at each time-frequency position, the electronic device can aggregate the distribution offsets along the frequency dimension to obtain the distribution offset in the time dimension. Based on the distribution offset in the time dimension, the electronic device can determine whether a device malfunction occurs within the time window corresponding to each audio segment, the start and end times of the device malfunction, and other device malfunction detection results.

[0040] The technical solution of this invention involves acquiring target audio data, which includes at least one audio segment to be detected, wherein the audio segment is a time-domain discrete audio signal; extracting time-frequency features of the target audio data based on the log-Mel spectrum of the target audio data; the time-frequency features include three feature dimensions: time, frequency, and channel; extracting channel feature vectors matching each time-frequency position from the time-frequency features of the target audio data; determining the distribution parameters of each time-frequency position based on the matching channel feature vectors; the distribution parameters include a mean vector and a covariance matrix; calculating the distribution offset of each time-frequency position based on the distribution parameters of each time-frequency position and pre-acquired reference distribution parameters of each time-frequency position; the reference distribution parameters include a reference mean vector and a reference covariance matrix of each time-frequency position determined based on audio data during normal operation of the device; aggregating the distribution offsets of each time-frequency position according to the frequency dimension to obtain the distribution offset in the time dimension; and determining the device anomaly detection result based on the distribution offset in the time dimension. This technical solution solves the problem of poor accuracy in existing audio data-based device anomaly detection methods. By measuring the distribution offset of the audio signal based on the reference distribution parameters of positive samples, accurate device anomaly detection can be achieved without the need to refer to negative sample features.

[0041] Example 2 Figure 2This is a flowchart of a device anomaly detection method based on audio data provided in Embodiment 2 of the present invention. This embodiment is a refinement based on the above embodiment. Figure 2 As shown, the method includes: S210. Acquire target audio data, wherein the target audio data is an audio segment to be detected, and the audio segment is a time-domain discrete audio signal.

[0042] In this scheme, the target audio data is an audio segment to be detected, such as an audio segment in the current time window.

[0043] S220. Extract the time-frequency features of the audio segment based on the log-Mel spectrum of the audio segment; the time-frequency features include three feature dimensions: time, frequency, and channel.

[0044] S230. Extract the channel feature vectors matching each time-frequency position from the time-frequency features of the audio segment.

[0045] S240. For each time-frequency location, a statistical sample set is formed from the time-frequency locations within a preset time neighborhood of that time-frequency location. Based on the channel feature vectors matched by each time-frequency location in the statistical sample set, the distribution parameters of that time-frequency location are determined. The distribution parameters include the mean vector and the covariance matrix.

[0046] Since the target audio data only includes one audio segment, a single audio segment cannot constitute a statistical sample set. Therefore, for each time-frequency position, the electronic device can construct a statistical sample set from each time-frequency position within a preset time neighborhood of that time-frequency position.

[0047] The preset time neighborhood can be determined based on the time neighborhood radius, and can be the left neighbor, right neighbor, or eccentric neighbor of the time-frequency location. In a specific example, the time-frequency location... The statistical sample set can be represented as: , Represents the radius of the time neighborhood, and is a positive number.

[0048] Electronic devices can statistically analyze the channel feature vectors matched at each time-frequency position in a statistical sample set to obtain the mean vector and covariance matrix at that time-frequency position. Specifically, the mean vector at that time-frequency position can be represented as: The covariance matrix at this time-frequency location can be expressed as: ,in, This indicates the number of time-frequency locations in the statistical sample set. Indicates a time interval index. Represents the time neighborhood radius, taken as a positive number. Indicates time-frequency position The channel feature vector.

[0049] After determining the distribution parameters at each time-frequency location, the method further includes: For each time-frequency location, the covariance matrix of that time-frequency location is shrunk based on a preset shrinkage coefficient and a target dimension identity matrix; the target dimension is the same as the dimension of the covariance matrix of that time-frequency location.

[0050] To ensure real-time anomaly detection, the number of audio segments to be detected as statistical samples is usually small. With a limited number of statistical samples, noise amplification and ill-conditioned problems in the covariance matrix are easily caused. To improve the stability and accuracy of statistical estimation, electronic devices can perform shrinkage processing on the covariance matrix at each time-frequency position based on a preset shrinkage coefficient and a target dimension identity matrix. This pulls the unstable covariance matrix towards a more robust identity matrix, thereby reducing the overall estimation error.

[0051] Specifically, the process of shrinking the covariance matrix can be represented as follows: , Indicates time-frequency position The covariance matrix before shrinkage processing Indicates the shrinkage coefficient. Indicates and An identity matrix with consistent dimensions.

[0052] S250. Calculate the distribution offset of each time-frequency position based on the distribution parameters of each time-frequency position and the reference distribution parameters of each time-frequency position obtained in advance; the reference distribution parameters include the reference mean vector and reference covariance matrix of each time-frequency position determined based on the audio data of the normal operation period of the device.

[0053] In one feasible solution, calculating the distribution offset of each time-frequency location based on the distribution parameters of each time-frequency location and the pre-acquired reference distribution parameters of each time-frequency location includes: For each time-frequency location, the Wasserstein distance of that time-frequency location is calculated based on its distribution parameters and reference distribution parameters, and this Wasserstein distance is used as the distribution offset of that time-frequency location.

[0054] Electronic devices can utilize the Wasserstein distance to measure the degree of deviation of a Gaussian distribution under test relative to a normal Gaussian distribution. When measuring the deviation between two Gaussian distributions, the Wasserstein distance can simultaneously capture the geometric difference between the mean and covariance, providing a smooth and continuous distance metric that avoids the gradient vanishing problem of traditional divergence when the distributions do not overlap.

[0055] Specifically, the formula for calculating the Wasserstein distance can be expressed as: , Indicates time-frequency position The reference mean vector, Indicates time-frequency position The reference covariance matrix, Indicates time-frequency position The mean vector, Indicates time-frequency position The covariance matrix, This represents the trace function for a matrix.

[0056] S260. Aggregate the distribution offsets of each time and frequency position according to the frequency dimension to obtain the distribution offset in the time dimension. Based on the distribution offset in the time dimension, determine the equipment anomaly detection result.

[0057] Based on the above scheme, determining the equipment anomaly detection result according to the distribution offset in the time dimension includes: For each moment in the time dimension, if the distribution offset at that moment is greater than or equal to a preset offset threshold, the device is determined to be abnormal at that moment; otherwise, the device is determined to be normal at that moment.

[0058] Specifically, the equipment anomaly detection result at each moment can be represented as: ;in, Indicates time The distribution offset, This indicates the preset offset threshold. Indicates time Equipment malfunction Indicates time The equipment is working properly.

[0059] Electronic devices can determine the start and end times of device anomalies based on the device anomaly detection results at each moment in the time dimension covered by the target audio data.

[0060] This solution addresses the case where the target audio data is a single audio segment to be detected. It extends the "single-point feature" of a single time-frequency location to the "local feature" of the time-frequency neighborhood of that location, constructing a Gaussian distribution that can characterize the features of that time-frequency location. This solves the problem of poor accuracy in existing audio data-based device anomaly detection methods. By measuring the distribution offset of the audio signal based on the reference distribution parameters of positive samples, accurate detection of device anomalies can be achieved without referencing negative sample features.

[0061] Example 3 Figure 3 This is a flowchart of a device anomaly detection method based on audio data provided in Embodiment 3 of the present invention. This embodiment is a refinement based on the above embodiment. Figure 3 As shown, the method includes: S310. Acquire target audio data, wherein the target audio data consists of at least two audio segments to be detected, and the audio segments are time-domain discrete audio signals.

[0062] In this approach, the target audio data consists of multiple audio segments to be detected, such as audio segments from multiple time windows within a day.

[0063] S320. Extract the time-frequency features of each audio segment based on the log-Mel spectrum of each audio segment; the time-frequency features include three feature dimensions: time, frequency, and channel.

[0064] S330. Extract the channel feature vector matching each time-frequency position from the time-frequency features of each audio segment.

[0065] S340. For each time-frequency position, construct a statistical sample set from each audio segment, and determine the distribution parameters of that time-frequency position based on the channel feature vector matched in each audio segment; the distribution parameters include the mean vector and the covariance matrix.

[0066] For each time-frequency position, the electronic device can use each audio segment as a statistical sample, statistically analyze the channel feature vectors matching that time-frequency position in each audio segment, and calculate the mean vector and covariance matrix for that time-frequency position. Specifically, the mean vector for that time-frequency position can be represented as: The covariance matrix at this time-frequency location can be expressed as: ,in, This indicates the number of audio segments in the statistical sample set. Indicates the audio segment index. Indicates the first digit in the statistical sample set. Time-frequency position of each audio segment The channel feature vector.

[0067] S350. Calculate the distribution offset of each time-frequency position based on the distribution parameters of each time-frequency position and the reference distribution parameters of each time-frequency position obtained in advance; the reference distribution parameters include the reference mean vector and reference covariance matrix of each time-frequency position determined based on the audio data of the normal operation period of the device.

[0068] S360. Aggregate the distribution offsets of each time and frequency position according to the frequency dimension to obtain the distribution offset in the time dimension. Based on the distribution offset in the time dimension, determine the equipment anomaly detection result.

[0069] This solution addresses the scenario where the target audio data consists of multiple audio segments to be detected. Based on the channel feature vectors of multiple audio segments at the same time-frequency position, a Gaussian distribution is constructed to characterize the features of that time-frequency position. This solves the problem of poor accuracy in existing audio data-based device anomaly detection methods. By measuring the distribution offset of the audio signal based on the reference distribution parameters of positive samples, accurate detection of device anomalies can be achieved without referencing the features of negative samples.

[0070] Example 4 Figure 4 This is a schematic diagram of a device anomaly detection device based on audio data provided in Embodiment 4 of the present invention. Figure 4 As shown, the device includes: The audio data acquisition module 410 is used to acquire target audio data, the target audio data including at least one audio segment to be detected, the audio segment being a time-domain discrete audio signal; The time-frequency feature extraction module 420 is used to extract the time-frequency features of the target audio data based on the log-Mel spectrum of the target audio data; the time-frequency features include three feature dimensions: time, frequency, and channel. The distribution parameter determination module 430 is used to extract the channel feature vector matching each time-frequency position from the time-frequency features of the target audio data, and determine the distribution parameters of each time-frequency position based on the channel feature vector matching each time-frequency position; the distribution parameters include the mean vector and the covariance matrix; The offset determination module 440 is used to calculate the distribution offset of each time-frequency position based on the distribution parameters of each time-frequency position and the reference distribution parameters of each time-frequency position obtained in advance; the reference distribution parameters include the reference mean vector and reference covariance matrix of each time-frequency position determined based on the audio data of the normal operation period of the device. The detection result determination module 450 is used to aggregate the distribution offset of each time and frequency position according to the frequency dimension to obtain the distribution offset in the time dimension, and determine the equipment anomaly detection result based on the distribution offset in the time dimension.

[0071] In one feasible approach, the target audio data contains one audio segment. The distribution parameter determination module 430 is specifically used for: Extract the channel feature vector matching each time-frequency position from the time-frequency features of the audio segment; For each time-frequency location, a statistical sample set is constructed from all time-frequency locations within a preset time neighborhood of that time-frequency location. Based on the channel feature vectors matched by each time-frequency location in the statistical sample set, the distribution parameters of that time-frequency location are determined.

[0072] In another feasible approach, the target audio data contains at least two audio segments; The distribution parameter determination module 430 is also used for: Extract the channel feature vector matching each time-frequency position from the time-frequency features of each audio segment; For each time-frequency position, a statistical sample set is constructed from each audio segment. Based on the channel feature vector matched at that time-frequency position in each audio segment, the distribution parameters of that time-frequency position are determined.

[0073] Optionally, the device further includes: The dimensionality reduction module is used to reduce the channel dimension of the time-frequency features of the target audio data by a preset dimensionality reduction method if the channel dimension of the time-frequency features of the target audio data is greater than a preset dimensionality threshold after the time-frequency features of the target audio data are extracted. The preset dimensionality reduction method is one of random channel sampling, linear projection or principal component analysis.

[0074] Preferably, the device further includes: The shrinkage processing module is used to shrink the covariance matrix of each time-frequency location after determining the distribution parameters of each time-frequency location, based on a preset shrinkage coefficient and a target dimension identity matrix; the target dimension is the same as the dimension of the covariance matrix of the time-frequency location.

[0075] In a preferred embodiment, the offset determination module 440 is specifically used for: For each time-frequency location, the Wasserstein distance of that time-frequency location is calculated based on its distribution parameters and reference distribution parameters, and this Wasserstein distance is used as the distribution offset of that time-frequency location.

[0076] The device anomaly detection device based on audio data provided in the embodiments of the present invention can execute the device anomaly detection method based on audio data provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0077] Example 5 Figure 5A schematic diagram of an electronic device 510 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0078] like Figure 5 As shown, the electronic device 510 includes at least one processor 511 and a memory communicatively connected to the at least one processor 511. The memory may be a read-only memory (ROM) 512, a random access memory (RAM) 513, etc. The memory stores computer programs executable by the at least one processor. The processor 511 can perform various appropriate actions and processes based on the computer program stored in the ROM 512 or loaded from storage unit 518 into the RAM 513. The RAM 513 may also store various programs and data required for the operation of the electronic device 510. The processor 511, ROM 512, and RAM 513 are interconnected via a bus 514. An input / output (I / O) interface 515 is also connected to the bus 514.

[0079] Multiple components in electronic device 510 are connected to I / O interface 515, including: input unit 516, such as keyboard, mouse, etc.; output unit 517, such as various types of displays, speakers, etc.; storage unit 518, such as disk, optical disk, etc.; and communication unit 519, such as network card, modem, wireless transceiver, etc. Communication unit 519 allows electronic device 510 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0080] Processor 511 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 511 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 511 performs the various methods and processes described above, such as device anomaly detection methods based on audio data.

[0081] In some embodiments, the audio data-based device anomaly detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 518. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 510 via ROM 512 and / or communication unit 519. When the computer program is loaded into RAM 513 and executed by processor 511, one or more steps of the audio data-based device anomaly detection method described above may be performed. Alternatively, in other embodiments, processor 511 may be configured to perform the audio data-based device anomaly detection method by any other suitable means (e.g., by means of firmware).

[0082] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0083] Computer programs used to implement the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable audio data-based device anomaly detection device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0084] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0085] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0086] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0087] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0088] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0089] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A device anomaly detection method based on audio data, characterized by, The method includes: Acquire target audio data, the target audio data including at least one audio segment to be detected, the audio segment being a time-domain discrete audio signal; Based on the log-Mel spectrum of the target audio data, the time-frequency features of the target audio data are extracted; the time-frequency features include three feature dimensions: time, frequency, and channel. Extract the channel feature vector matching each time-frequency position from the time-frequency features of the target audio data, and determine the distribution parameters of each time-frequency position based on the channel feature vector matching each time-frequency position; the distribution parameters include the mean vector and the covariance matrix. For each time-frequency location, the Wasserstein distance of that time-frequency location is calculated based on its distribution parameters and reference distribution parameters, and this Wasserstein distance is used as the distribution offset of that time-frequency location. The reference distribution parameters include the reference mean vector and reference covariance matrix of each time-frequency location determined based on the audio data during normal operation of the device. The distribution offsets at each time and frequency position are aggregated according to the frequency dimension to obtain the distribution offsets in the time dimension. For each moment in the time dimension, if the distribution offset at that moment is greater than or equal to a preset offset threshold, the device is determined to be abnormal at that moment; otherwise, the device is determined to be normal at that moment.

2. The method of claim 1, wherein, The target audio data contains one audio segment. The step of extracting channel feature vectors matching each time-frequency position from the time-frequency features of the target audio data, and determining the distribution parameters of each time-frequency position based on the channel feature vectors matching each time-frequency position, includes: Extract the channel feature vector matching each time-frequency position from the time-frequency features of the audio segment; For each time-frequency location, a statistical sample set is constructed from all time-frequency locations within a preset time neighborhood of that time-frequency location. Based on the channel feature vectors matched by each time-frequency location in the statistical sample set, the distribution parameters of that time-frequency location are determined.

3. The method of claim 1, wherein, The target audio data contains at least two audio segments; The step of extracting channel feature vectors matching each time-frequency position from the time-frequency features of the target audio data, and determining the distribution parameters of each time-frequency position based on the channel feature vectors matching each time-frequency position, includes: Extract the channel feature vector matching each time-frequency position from the time-frequency features of each audio segment; For each time-frequency position, a statistical sample set is constructed from each audio segment. Based on the channel feature vector matched at that time-frequency position in each audio segment, the distribution parameters of that time-frequency position are determined.

4. The method of claim 1, wherein, After extracting the time-frequency features of the target audio data, the method further includes: If the channel dimension of the time-frequency features of the target audio data is greater than a preset dimension threshold, then the channel dimension of the time-frequency features of the target audio data is reduced based on a preset dimensionality reduction method; the preset dimensionality reduction method is one of random channel sampling, linear projection, or principal component analysis.

5. The method of claim 1, wherein, After determining the distribution parameters at each time-frequency location, the method further includes: For each time-frequency location, the covariance matrix of that time-frequency location is shrunk based on a preset shrinkage coefficient and a target dimension identity matrix; the target dimension is the same as the dimension of the covariance matrix of that time-frequency location.

6. An apparatus for device anomaly detection based on audio data, the apparatus comprising: The device includes: An audio data acquisition module is used to acquire target audio data, the target audio data including at least one audio segment to be detected, the audio segment being a time-domain discrete audio signal; The time-frequency feature extraction module is used to extract the time-frequency features of the target audio data based on the log-Mel spectrum of the target audio data; the time-frequency features include three feature dimensions: time, frequency, and channel. The distribution parameter determination module is used to extract the channel feature vectors matching each time-frequency position from the time-frequency features of the target audio data, and determine the distribution parameters of each time-frequency position based on the channel feature vectors matching each time-frequency position; the distribution parameters include the mean vector and the covariance matrix; The offset determination module is used to calculate the Wasserstein distance of each time-frequency position based on the distribution parameters of that time-frequency position and the reference distribution parameters of that time-frequency position, and use the Wasserstein distance of that time-frequency position as the distribution offset of that time-frequency position; the reference distribution parameters include the reference mean vector and reference covariance matrix of each time-frequency position determined based on the audio data of the normal operation period of the device. The detection result determination module is used to aggregate the distribution offset of each time and frequency position according to the frequency dimension to obtain the distribution offset in the time dimension. For each moment in the time dimension, if the distribution offset at that moment is greater than or equal to a preset offset threshold, the device is determined to be abnormal at that moment; otherwise, the device is determined to be normal at that moment.

7. An electronic device, comprising: The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the device anomaly detection method based on audio data according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the device anomaly detection method based on audio data as described in any one of claims 1-5.

9. A computer program product, characterised in that, The method includes a computer program that, when executed by a processor, implements the device anomaly detection method based on audio data according to any one of claims 1-5.