Sound Event Detection Module for Repeating Audio Signals

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems are hindered by the need for extensive training corpora, which limits their effectiveness when recognizing sounds from users not included in the training data, leading to reduced recognition rates due to speaker-specific parameter model differences.

Innovation Solution

A sound event detecting module and method that identifies repeating sound events by recognizing sound sections, generating feature vectors, comparing them using similarity score matrices, and determining high correlation counts to detect recurring sounds without relying on extensive training corpora, utilizing components like sound end recognizing units, similarity comparing units, and correlation arbitrating units.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speech recognition systems use training corpora to create parameter models, then recognition accuracy for trained speakers is improved, but recognition accuracy for untrained speakers deteriorates significantly

Engineering Contradiction:
Improverecognition accuracyVSAvoidadaptability to untrained speakers
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system enables speakers to train the model themselves by providing a small set of representative sound examples (e.g., 5-10 samples per speaker). This self-training mechanism allows each speaker to personalize the model without requiring extensive pre-collected training data, thereby improving recognition accuracy for untrained speakers while maintaining system versatility

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the approach from using fixed, pre-trained parameter models to dynamically adapting model parameters based on individual speaker's sound characteristics. By allowing speakers to provide their own training samples and adjusting the model parameters accordingly, the system achieves high recognition accuracy for each speaker without requiring extensive universal training corpora

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If extensive training corpora are collected to improve recognition rate, then recognition rate is improved, but system complexity and data collection requirements increase

Engineering Contradiction:
Improverecognition rateVSAvoiddata collection requirement
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the essential training data requirement from the complex process of collecting extensive training corpora. Instead of requiring large-scale pre-collected datasets, the system extracts only the necessary training samples directly from the speaker themselves during actual use, significantly reducing data collection complexity while maintaining high recognition rates

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary model creation using a small set of representative sound samples provided by the speaker before actual recognition tasks. This preliminary action of model training using minimal data (5-10 samples per speaker) establishes a personalized model that achieves high recognition accuracy without requiring extensive pre-collected training data

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If speaker-specific parameter models are created from different speakers, then recognition accuracy for each speaker is improved, but the influence of training corpora on detection technology increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidindependence from training corpora
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

Each speaker trains their own parameter model using their voice samples, making the system independent from external training corpora. This self-service approach allows the model to capture individual speaker characteristics without being influenced by generic training data, thereby improving recognition accuracy while reducing dependence on training corpora

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies local quality by creating speaker-specific parameter models tailored to each individual's voice characteristics rather than using a universal model. Each speaker's model is optimized for their specific acoustic properties, improving recognition accuracy for that speaker while reducing the influence of generic training corpora on the detection technology

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS8655655B2Sound event detecting module for a sound event recognition system and method thereof
Publication Date: 2014.02.18 IND TECH RES INST
  • US8655655B2 patent drawing
  • US8655655B2 patent drawing
  • US8655655B2 patent drawing

AI summary

A sound event detecting module for detecting whether a sound event with characteristic of repeating is generated. A sound end recognizing unit recognizes ends of sounds according to a sound signal to generate sound sections and multiple sets of feature vectors of the sound sections correspondingly. A storage unit stores at least M sets of feature vectors. A similarity comparing unit compares the at least M sets of feature vectors with each other, and correspondingly generates a similarity score matrix, which stores similarity scores of any two of the sound sections of the at least M of the sound sections. A correlation arbitrating unit determines the number of sound sections with high correlations to each other according to the similarity score matrix. When the number is greater than one threshold value, the correlation arbitrating unit indicates that the sound event with the characteristic of repeating is generated.