Audio Feature Extraction for ACR Database Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing number of content pieces handled by Automatic Content Recognition (ACR) services using only audio information leads to a significant increase in database capacity and operational burdens, including increased processing time and maintenance costs, as the data amount of feature point information grows.
Innovation Solution
An information processing device and method that extracts feature point information only from the main sound of audio content, reducing the need for extensive database capacity and operational resources by focusing on the main sound, even when sub sounds are listened to, thus minimizing database preparation and maintenance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If feature point information is extracted from all audio channels including sub sounds, then content identification accuracy is improved, but database capacity and operational burdens increase significantly
Solution Approach 1:
The patent extracts only the necessary component (main sound feature point information) from the complete audio content, excluding sub sounds and other non-essential audio channels. This selective extraction reduces the quantity of data stored in the database while maintaining sufficient accuracy for content identification, directly resolving the contradiction between identification reliability and database capacity.
Solution Approach 2:
The patent segments audio content into distinct components (main sound and sub sounds) and processes only the main sound for feature point extraction. By dividing the audio content and selecting only the essential segment for identification purposes, the system reduces database burden while preserving identification accuracy.
2Reliability
If feature point information is extracted from all audio channels including sub sounds, then content identification accuracy is improved, but processing time and operational costs increase
Solution Approach 1:
The patent extracts only the essential main sound component for feature point extraction, eliminating the need to process sub sounds and other redundant audio channels. This extraction approach significantly reduces processing time while maintaining identification accuracy, as the main sound contains sufficient information for reliable content recognition.
Solution Approach 2:
The patent applies partial action by processing only the necessary portion (main sound) of the complete audio content rather than all audio channels. This partial processing approach is sufficient for achieving accurate content identification while dramatically reducing the time and computational resources required.
3Quantity of substance
If feature point information is extracted only from main sound, then database capacity is reduced, but content identification may fail when sub sounds are listened to
Solution Approach 1:
The patent introduces a switching mechanism that acts as an intermediary between the user's audio channel selection and the feature point extraction process. When sub sounds are selected for listening, the switch automatically extracts feature points from the main sound instead, ensuring identification functionality is maintained across different audio channel selections while keeping the database compact.
Solution Approach 2:
The patent creates a universal identification system that works regardless of which audio channel the user is listening to. By designing the system to always reference main sound feature points even when sub sounds are played, the system achieves multi-functional adaptability - supporting different listening preferences while maintaining a single, reduced-capacity database.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An ACR service is realized by an information provision server which has not prepared feature point information of a sub sound because only the feature point information is extracted from a main sound and an inquiry is made to the information provision server even when the sub sound is listened to in a client device. In addition, even when content that has a main sound and a plurality of pieces of audio information is distributed, it is not necessary for the information provision server to prepare feature point information of sub sounds, and the capacity of the database may not increase.