Abnormal Sound Detection via Statistical Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In environments with limited power supply, existing systems cannot transmit raw or compressed sound data due to restricted bandwidth, making it impossible to report abnormal sounds effectively from facilities to remote servers.

Innovation Solution

An abnormal sound detection system that uses a terminal to compute and transmit statistical data from sound recordings, allowing a server to learn a normal sound model and recreate artificial sounds for confirmation, even with limited bandwidth, by employing a logarithmic mel spectrogram and pseudo-spectrogram reconstruction techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If raw sound data or typical compressed sound data is transmitted, then the quality of sound transmission is improved, but the transmission traffic exceeds the limited battery-driven bandwidth

Engineering Contradiction:
Improvesound data qualityVSAvoidtransmission traffic
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential statistical features (spectral centroid, spectral rolloff, spectral flux, zero-crossing rate, and chroma features) from the complete sound signal, transmitting only these extracted parameters instead of the full sound data. This extraction approach maintains sufficient information for abnormal sound detection while dramatically reducing transmission traffic to fit within battery-driven bandwidth limits.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a simplified representation (copy) of the sound data in the form of statistical feature vectors that capture the essential characteristics needed for abnormal sound detection. These feature vectors serve as a compressed copy that preserves diagnostic information while occupying minimal transmission bandwidth.

Inventive Principle:
Principle #26Copying

2Quantity of substance

If audio encoder is used for compression, then the sound data is compressed, but the encoding process consumes excessive power for long-term battery operation

Engineering Contradiction:
Improvesound data sizeVSAvoidencoding power consumption
Core Design Contradiction:
Quantity of substanceVSUse of energy by moving object

Solution Approach 1:

The patent replaces the complex mechanical/audio encoding process (FFT, DCT, quantization) with a simpler statistical feature extraction approach. Instead of using traditional audio encoders that require significant computational resources, the system calculates basic statistical parameters directly from the sound signal, dramatically reducing power consumption while achieving effective data compression.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Quantity of substance

If only abnormality presence/absence is reported, then the transmission traffic is reduced, but the user cannot hear and confirm the actual abnormal sound

Engineering Contradiction:
Improvetransmission trafficVSAvoidsound information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent introduces statistical feature vectors as an intermediary representation that bridges the gap between complete sound data and simple abnormality flags. These feature vectors serve as a middle ground, providing sufficient information for both automated detection and human auditory confirmation without requiring transmission of full-resolution sound data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11164594B2Abnormal sound detection system, artificial sound creation system, and artificial sound creating method
Publication Date: 2021.11.02 HITACHI LTD
  • US11164594B2 patent drawing
  • US11164594B2 patent drawing
  • US11164594B2 patent drawing

AI summary

Confirmation can be made what sound has been made under a restriction in which transmittable traffic is small. An abnormal sound detection system including an artificial sound creating function is configured, the abnormal sound detection system including a statistic calculation unit configured to calculate a statistic set expressing sizes of a direct current component, an alternating current component, and a noise component in an amplitude time series at each of frequencies of a sound inputted at a terminal, a statistic transmitting unit configured to transmit the statistic set from the terminal to a server, a statistic receiving unit configured to receive the statistic set in the server, and an artificial sound reproducing unit configured to reproduce a cyclostationary artificial sound based on the statistic set received in the server.