Notch Filter Data Augmentation for Sound Event Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound event detection and classification systems using machine learning face recognition rate issues due to insufficient and varied training data, especially when applied to equipment with different frequency responses.

Innovation Solution

The method involves augmenting training data by transforming its frequency or time components using a notch filter with randomly set center frequencies and bandwidths, and selectively deleting data intervals based on uniform distribution probability variables to create diverse and robust training datasets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If existing data augmentation schemes such as block mixing are used to artificially add noise or trim and mix audio files, then the number of training samples is increased, but the recognition rate decreases when the classifier is applied to equipment having a frequency response different from the training samples

Engineering Contradiction:
Improvenumber of training samplesVSAvoidrecognition rate
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies parameter changes by using a notch filter with randomly selected center frequencies and bandwidths to transform the frequency components of training data. This simulates different frequency responses of recording equipment, enabling the classifier to maintain high recognition rates across various devices while still increasing the effective number of training samples through data transformation rather than simple duplication

Inventive Principle:
Principle #35Parameter changes

2Reliability

If training data is collected from multiple sources to increase diversity, then the recognition rate improves, but the cost and complexity of data collection increases

Engineering Contradiction:
Improverecognition rateVSAvoiddata collection complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates transformed copies of existing training data by applying notch filters with random parameters. Instead of collecting new data from multiple sources, the system generates synthetic variations of original recordings that simulate different frequency responses, thereby maintaining high recognition rates without the complexity of multi-source data collection

Inventive Principle:
Principle #26Copying

Solution Approach 2:

By changing the frequency domain parameters of existing training data through random notch filtering, the system generates diverse training samples that mimic data from different recording environments and equipment, achieving data diversity through parameter transformation rather than physical data collection

Inventive Principle:
Principle #35Parameter changes

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enhances the recognition rate by increasing the diversity and number of training data without the need for additional data collection, thereby reducing costs and improving performance across different equipment frequency responses.

Implementation Method 1

obtaining training data having a transformed frequency component from the original data by filtering the original data using a filter configured to remove a component of a predetermined frequency band

Methodology Applied
Scientific EffectNotch filter: Filter (electronic)

Data Source

PatentUS11657325B2Apparatus and method for augmenting training data using notch filter
Publication Date: 2023.05.23 ELECTRONICS & TELECOMM RES INST
  • US11657325B2 patent drawing
  • US11657325B2 patent drawing
  • US11657325B2 patent drawing

AI summary

Disclosed is an apparatus and method for augmenting training data using a notch filter. The method may include obtaining original data, and obtaining training data having a modified frequency component from the original data by filtering the original data using a filter configured to remove a component of a predetermined frequency band.