Hierarchical Audio Session Classification Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Electronic devices often erroneously upmix audio content, losing artistic intent and introducing spatial and timbre artifacts, particularly when converting stereo music to surround sound or vice versa, and struggle to accurately classify media as movie, music, or voice, which affects audio reproduction settings and multimedia indexing.
Innovation Solution
A hierarchical classification model that uses predetermined criteria to classify audio sessions based on source and employs machine learning analysis of metadata to accurately determine the type of content, reducing processing time and resource usage while increasing accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If audio content is upmixed from stereo to surround sound, then spatial coverage is improved, but artistic intent is lost and spatial/timbre artifacts are introduced
Solution Approach 1:
The system performs preliminary classification of the audio session type (movie, music, voice, gaming) before applying any audio processing. This preliminary identification allows the system to select appropriate processing paths that preserve artistic intent for music while still providing spatial enhancement for movie content, thus resolving the contradiction between spatial coverage and artistic intent preservation
2Measurement precision
If machine learning analysis is used to classify audio sessions, then classification accuracy is improved, but processing time and resource usage increase
Solution Approach 1:
The classification system is segmented into multiple stages: first using lightweight predetermined criteria (source identification, metadata analysis) for quick initial classification, then applying machine learning analysis only to sessions that require more sophisticated classification. This segmentation reduces overall processing time while maintaining high accuracy for complex cases
Solution Approach 2:
The system applies partial machine learning analysis by using it only when predetermined criteria are insufficient, rather than applying full machine learning to all audio sessions. This selective approach maintains high classification accuracy while significantly reducing processing time and resource consumption for straightforward cases
Data Source
AI summary
Examples of methods for audio session classification are described herein. In some examples, a method may include determining, at a first classification stage, whether an audio session is classifiable with predetermined criteria. In some examples, the method may include classifying, at a second classification stage, the audio session based on a machine learning analysis of metadata in a case that the audio session is not classifiable at the first classification stage.


