Sound Recognition Apparatus Using Segmented Sound Unit Groups
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound recognition systems face difficulties in accurately identifying various usual sounds due to their reliance on acoustic features for phonemes, which are insufficient in describing the diverse characteristics of objects, events, and operation states.
Innovation Solution
A sound recognition apparatus that calculates sound feature values, converts them into labels using correlated data, and identifies sound events by segmenting sound unit groups based on probability models, allowing for the recognition of usual sounds with varying acoustic features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If acoustic features for phonemes are used for sound recognition, then speech recognition is effective, but recognition of usual sounds with varying acoustic features is insufficient
Solution Approach 1:
The patent segments sounds into multiple sound units with different acoustic features (e.g., frequency, temporal variations) and processes each segment independently. This allows the system to capture the diverse characteristics of usual sounds while maintaining structured analysis, resolving the contradiction between precise measurement and adaptability to varying sound types.
Solution Approach 2:
The patent extracts multiple acoustic feature parameters (frequency characteristics, temporal variations, spectral features) from sound signals and uses these varied parameters to represent different sound types. By changing and utilizing multiple parameters simultaneously, the system achieves both precise recognition and adaptability to diverse usual sounds.
2Productivity
If phoneme-based acoustic models are used, then speech processing is efficient, but usual sound recognition capability is limited
Solution Approach 1:
The patent creates a universal sound recognition framework that processes both speech and usual sounds using the same acoustic feature extraction and segmentation methodology. This multi-functional approach maintains processing efficiency while improving usual sound recognition accuracy by treating all sound types uniformly with appropriate feature analysis.
Solution Approach 2:
The patent dynamically adjusts the acoustic feature extraction and segmentation process based on the characteristics of the input sound. For usual sounds with varying features, the system adapts its analysis parameters and segmentation granularity, maintaining efficiency while improving recognition accuracy through dynamic parameter adjustment.
3Measurement precision
If detailed acoustic feature analysis is performed on all sounds, then recognition accuracy improves, but processing load increases
Solution Approach 1:
The patent divides sound signals into discrete sound units and segments, analyzing acoustic features at the segment level rather than processing entire sounds uniformly. This segmentation reduces computational load by enabling localized feature extraction and processing, while maintaining high recognition accuracy through detailed analysis of relevant segments.
Solution Approach 2:
The patent applies detailed acoustic feature analysis selectively to portions of sounds that contain discriminative information for recognition. By performing partial analysis on critical segments rather than exhaustive analysis of all sound data, the system achieves high recognition accuracy with reduced processing load.
Data Source
AI summary
A sound recognition apparatus can include a sound feature value calculating unit configured to calculate a sound feature value based on a sound signal, and a label converting unit configured to convert the sound feature value into a corresponding label with reference to label data in which sound feature values and labels indicating sound units are correlated. A sound identifying unit is configured to calculate a probability of each sound unit group sequence that a label sequence is segmented for each sound unit group with reference to segmentation data. The segmentated data indicates a probability that a sound unit sequence will be segmented into at least one sound unit group. The sound identity unit can also identify a sound event corresponding to the sound unit group sequence selected based on the calculated probability.


