Audio Feature Scene Type Identification via Machine Learning Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in accurately identifying scene types in video content due to varying relationships between scene types and video tendencies.
Innovation Solution
An apparatus and method that utilize a processor and memory to generate relations between audio feature amounts and scene types through machine learning, creating an identification model that classifies audio features into clusters and determines scene types based on these relations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If scene types are identified from video content, then video analysis capability is utilized, but identification accuracy deteriorates when relations between scene types and video tendencies vary greatly
Solution Approach 1:
The patent segments the identification task by separating video analysis from audio analysis. Instead of relying solely on video content, the system extracts audio feature amounts and performs cluster analysis to identify scene types independently through multiple modalities, thereby improving accuracy when video-based identification fails
Solution Approach 2:
The patent introduces audio feature amounts as an intermediary element between the content and scene type identification. By analyzing audio characteristics and their cluster distributions, the system mediates the identification process to achieve accurate scene type recognition even when direct video-based identification is unreliable
2Measurement precision
If machine learning is used to generate identification models, then scene type identification accuracy is improved, but processing complexity increases
Solution Approach 1:
The patent performs preliminary cluster analysis on audio feature amounts before final scene type identification. By pre-processing the audio data into clusters and establishing cluster-to-scene-type mappings in advance, the system reduces the complexity of real-time identification while maintaining high accuracy through the pre-built identification model
Data Source
AI summary
An apparatus for generating relations between feature amounts of audio and scene type includes at least one processor and a memory. The memory is operatively coupled to the at least one processor. The processor is configured to set one of the scene types to each of clusters classifying the feature amounts of audio in one or more pieces of content. The processor is also configured to generate a plurality of pieces of learning data, each representative of a feature amount, from among the feature amounts of the audio, that belongs to each cluster and the scene type set for each cluster. The processor is also configured to generate an identification model representative of relations between the feature amounts of audio and the scene types by performing machine learning using the plurality of pieces of learning data.


