Notch Filter Data Augmentation for Sound Event Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound event detection and classification systems using machine learning face recognition rate issues due to insufficient and varied training data, especially when applied to equipment with different frequency responses.
Innovation Solution
The method involves augmenting training data by transforming its frequency or time components using a notch filter with randomly set center frequencies and bandwidths, and selectively deleting data intervals based on uniform distribution probability variables to create diverse and robust training datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If existing data augmentation schemes such as block mixing are used to artificially add noise or trim and mix audio files, then the number of training samples is increased, but the recognition rate decreases when the classifier is applied to equipment having a frequency response different from the training samples
Solution Approach 1:
The patent applies parameter changes by using a notch filter with randomly selected center frequencies and bandwidths to transform the frequency components of training data. This simulates different frequency responses of recording equipment, enabling the classifier to maintain high recognition rates across various devices while still increasing the effective number of training samples through data transformation rather than simple duplication
2Reliability
If training data is collected from multiple sources to increase diversity, then the recognition rate improves, but the cost and complexity of data collection increases
Solution Approach 1:
The patent creates transformed copies of existing training data by applying notch filters with random parameters. Instead of collecting new data from multiple sources, the system generates synthetic variations of original recordings that simulate different frequency responses, thereby maintaining high recognition rates without the complexity of multi-source data collection
Solution Approach 2:
By changing the frequency domain parameters of existing training data through random notch filtering, the system generates diverse training samples that mimic data from different recording environments and equipment, achieving data diversity through parameter transformation rather than physical data collection
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enhances the recognition rate by increasing the diversity and number of training data without the need for additional data collection, thereby reducing costs and improving performance across different equipment frequency responses.
Implementation Method 1
obtaining training data having a transformed frequency component from the original data by filtering the original data using a filter configured to remove a component of a predetermined frequency band
Data Source
AI summary
Disclosed is an apparatus and method for augmenting training data using a notch filter. The method may include obtaining original data, and obtaining training data having a modified frequency component from the original data by filtering the original data using a filter configured to remove a component of a predetermined frequency band.


