Audio Classification Using Spectrogram Segmentation for Hazard Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for recognizing special acoustic signals in road traffic lack sensitivity and accuracy, leading to potential false alarms and incorrect reactions, especially in diverse international settings where different signal types are used.

Innovation Solution

A method for classifying temporally sequential digital audio data using frequency spectrograms, which enhances the identification of hazardous situations by analyzing overtones and fundamental tones, employing techniques like short-time Fourier transform and wavelet transform, and utilizing classifiers such as artificial neural networks to achieve high sensitivity and low error rates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If high sensitivity is used to ensure accurate classification of special signals, then the error rate decreases, but false alarms may increase

Engineering Contradiction:
Improveclassification accuracyVSAvoidfalse alarm rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The audio signal is divided into multiple time intervals, and frequency spectrograms are calculated for each interval. This segmentation allows the system to analyze the temporal structure of special signals and distinguish them from continuous noise, improving classification accuracy while maintaining reliability through pattern recognition across multiple segments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the audio signal from the time domain to the frequency domain using Fourier transform, creating frequency spectrograms. This dimensional transformation enables the system to identify characteristic frequency patterns of special signals (such as the specific frequency ranges of Martin's horn, wail, yelp, and rumbler) that are not apparent in the time domain, thereby improving measurement precision without increasing false alarms.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If the system distinguishes between various special signals used around the world, then adaptability improves, but device complexity increases

Engineering Contradiction:
Improvesignal type recognitionVSAvoidclassification system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent employs a unified classification approach using frequency spectrograms and pattern recognition that can identify multiple types of special signals (Martin's horn, wail, yelp, rumbler, and other hazardous acoustic signals) through a single system. The method universally applies Fourier transform and temporal pattern analysis across different signal types, enabling the system to distinguish between various international special signals without requiring separate dedicated systems for each signal type, thus improving adaptability while controlling complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If frequency spectrograms are calculated for multiple time intervals, then measurement precision improves, but loss of time increases

Engineering Contradiction:
Improvesignal identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent calculates frequency spectrograms for multiple stepwise progressive time intervals in advance, creating a temporal profile of the audio signal. This preliminary analysis of frequency patterns across time intervals enables the classification system to quickly identify characteristic patterns of special signals without requiring real-time exhaustive analysis, thereby improving measurement precision while minimizing time loss through proactive signal characterization.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11404074B2Method for the classification of temporally sequential digital audio data
Publication Date: 2022.08.02 ROBERT BOSCH GMBH
  • US11404074B2 patent drawing
  • US11404074B2 patent drawing

AI summary

A method is for the classification of temporally sequential digital audio data that describe acoustic signals that identify hazardous situations. The method includes calculating a large number of frequency spectrograms for stepwise progressive time intervals of the temporally sequential audio data, and forming a specific number of frequency segments for each octave of each individual frequency spectrogram. The frequency segments include a subset of the individual frequency spectrograms. The method further includes adding corresponding frequency segments of the octaves of each individual frequency spectrogram, calculating frequency components through the formation of mean values for the individual, added frequency segments in each individual frequency spectrogram, and generating a classification vector using a classifier and the number of frequency components of the large number of frequency spectrograms. The classifier is configured to classify signals, described by the associated temporally sequential digital audio data, that identify hazardous situations.