Interesting Section Extracting Device Using Audio Likelihood Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for extracting interesting sections from audio-visual content require significant user effort and time, as they necessitate setting offset times or audio feature conditions specific to each type of content, making the process labor-intensive and inefficient.
Innovation Solution
An interesting section extracting device that uses anchor models to generate likelihood vectors from audio signals, allowing for automatic identification and extraction of interesting sections by specifying a single time, reducing the need for manual intervention and content-specific settings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If offset time is set for each type of AV content to extract interesting sections, then the extraction accuracy is improved, but the user workload and time required for setting increases significantly
Solution Approach 1:
The system automatically determines offset times by analyzing audio features (speech presence, music presence, silence periods) without requiring user intervention. The controller autonomously identifies appropriate in-point and out-point offsets based on content analysis, eliminating the manual setting process while maintaining extraction accuracy
Solution Approach 2:
The system dynamically adjusts offset time parameters based on the analyzed audio characteristics of each content type. By changing the offset parameters automatically according to detected speech/music patterns and silence periods, the system adapts to different content types without requiring manual reconfiguration
2Ease of operation
If manual operation of controller is required to determine start and end times, then the user has control over extraction, but the operation complexity and skill requirement increases
Solution Approach 1:
The controller automatically performs the functions of determining in-points and out-points by analyzing audio features, eliminating the need for manual user operation. The system serves itself by autonomously identifying interesting sections based on speech/music detection and silence period analysis
Solution Approach 2:
The manual mechanical operation of pressing buttons to set timestamps is replaced by an automated electronic system that analyzes audio signals and automatically determines extraction points, substituting user action with electronic signal processing
3Measurement precision
If audio feature conditions are set for each type of AV content, then the extraction precision is improved, but the device complexity and setup effort increases
Solution Approach 1:
The system uses a universal audio analysis framework that automatically adapts to different content types by detecting their characteristics. The same controller and analysis algorithms work across news programs, documentaries, and other formats without requiring separate configuration, achieving both precision and simplicity
Data Source
AI summary
An interesting section extracting device extracts an interesting section of interest to a user from a video file with reference to an audio signal included in the video file such that a specified time is included in the interesting section. The interesting section extracting device includes an interface device that obtains the specified time; and a likelihood vector generating unit that calculates, in one-to-one correspondence with first unit sections of the audio signal, likelihoods for anchor models that respectively represent features of a plurality of types of sound pieces and generates likelihood vectors having the calculated likelihoods as components thereof. An interesting section extracting unit calculates a first feature section as candidate section, which is candidate for the interesting section to be extracted, by using likelihood vectors and extract, as the interesting section, part of the first feature section including the specified time.


