Sound Recognition Confidence Adjustment via Learned Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current sound recognition systems face challenges in accurately identifying non-verbal sound events and scenes from audio signals, particularly in dynamic environments where prior knowledge and device-specific properties are not adequately considered, leading to inefficiencies and inaccuracies.
Innovation Solution
A method that adjusts sound class scores in audio signals by using a learned model to process confidence values associated with properties like energy and environment, allowing for dynamic adaptation without retraining the machine learning model, and incorporating external knowledge to improve recognition accuracy across various devices and conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If sound class scores are adjusted using a learned model that incorporates device-specific properties and environmental factors, then recognition accuracy is improved, but system complexity increases
Solution Approach 1:
The system performs preliminary actions by pre-processing audio frames to extract device-specific properties and environmental factors before sound class classification. This preliminary processing enables the learned model to receive enriched input data that captures contextual information, thereby improving recognition accuracy without requiring complex modifications to the core classification algorithm.
Solution Approach 2:
The patent introduces an intermediary learned model that acts as a mediator between raw audio data and sound class scores. This intermediary component processes device-specific properties and environmental factors, transforming them into adjusted sound class scores that reflect contextual conditions. The learned model serves as a bridge that reconciles raw observations with contextual knowledge, improving accuracy while maintaining manageable system complexity through modular architecture.
2Adaptability or versatility
If the system dynamically adapts to changing environments and devices, then recognition robustness is improved, but computational resources increase
Solution Approach 1:
The system applies partial action by selectively processing only the most relevant device-specific properties and environmental factors for each audio frame. Rather than comprehensively analyzing all possible parameters, the learned model focuses on key features that have the greatest impact on sound class classification. This selective processing enables dynamic adaptation to changing environments while constraining computational resource consumption to essential operations only.
3Stability of the object's composition
If sound class scores are adjusted based on device-specific properties, then consistency across devices is improved, but processing time increases
Solution Approach 1:
The patent implements parameter changes by modifying sound class scores based on device-specific properties extracted from audio frames. The learned model adjusts classification parameters (sound class scores) according to device characteristics such as microphone sensitivity, frequency response, and environmental conditions. This parameter adjustment mechanism ensures that recognition results remain consistent across different devices by compensating for device-specific variations in the scoring parameters.
Data Source
AI summary
A method for recognising at least one of a non-verbal sound event and a scene in an audio signal comprising a sequence of frames of audio data, the method comprising: for each frame of the sequence: receiving at least one sound class score, wherein each sound class score is representative of a degree of affiliation of the frame with a sound class of a plurality of sound classes; for a sound class score of the at least one sound class scores: determining a confidence that the sound class score is representative of a degree of affiliation of the frame with the sound class by processing a value for a property associated with the frame, wherein the value is processed using a learned model for the property; adjusting the sound class score for the frame based at least on the determined confidence.


