Electronic Apparatus with Live-View Sound Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face limitations in selecting appropriate background sounds for videos as they need to manually search and download sound sources, which restricts the variety of options available.
Innovation Solution
An electronic device equipped with a camera, microphone, display, and processor that automatically identifies suitable sound sources based on live view image features, such as object information, location, and metadata, and outputs the selected sound during video recording.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If users manually search and download sound sources, then they can insert background sounds in videos, but the variety of available sound sources is limited and the process is time-consuming
Solution Approach 1:
The system automatically analyzes the video content and selects appropriate sound sources without requiring user intervention for searching or downloading. The processor extracts features from the video, matches them with sound source metadata, and automatically inserts the selected sound, making the system serve itself rather than requiring manual user actions.
Solution Approach 2:
Sound sources are pre-tagged with metadata describing their characteristics, themes, and适用 scenarios. This preliminary organization of sound sources with descriptive metadata enables the system to quickly match and select appropriate sounds based on video content features, eliminating the need for users to search through large libraries manually.
2Ease of operation
If users manually search for sound sources, then they can insert background sounds, but the process requires using separate applications and multiple steps
Solution Approach 1:
The system merges the video editing function with the sound source selection and insertion function into a single integrated application. The processor within the electronic device handles both video analysis and sound source management, eliminating the need for users to switch between separate video editing applications and sound libraries, thereby simplifying the overall operation process.
Solution Approach 2:
The electronic device's processor is designed to perform multiple functions: video capture, video analysis, sound source metadata matching, and sound insertion. This multi-functional capability allows the device to handle the entire sound insertion process within a single application, reducing the complexity of using multiple separate applications.
3Adaptability or versatility
If the system automatically identifies sound sources based on live view image features, then diverse and relevant audio options are provided, but the system complexity increases
Solution Approach 1:
The system uses metadata as an intermediary layer between video content features and sound sources. Instead of directly analyzing and matching complex video features with sound sources, the processor extracts features from the video, compares them with pre-stored metadata tags associated with sound sources, and selects matches based on this intermediate comparison layer, simplifying the overall matching complexity.
Solution Approach 2:
The system transforms the complex task of sound source selection into a parameter-matching problem. By representing both video content and sound sources as sets of extractable features and metadata parameters (such as theme, mood, objects present), the system simplifies the selection process to comparing and matching these parameters, reducing the effective complexity of the identification system.
Data Source
AI summary
An electronic apparatus includes: a microphone; a camera; a display; a speaker; and a processor operatively connected to the camera, the display, and the speaker, where the processor displays, on the display, a live view image obtained through the camera, when a camera application is executed; obtain a characteristic text for the live view image based on ae piece of object information included in the live view image, displays the characteristic text together with the live view image, identifies a sound source including metadata which matches the characteristic text, from among sound sources; displays information about the sound source together with the live view image, and when a user manipulation for capturing a moving image is input, outputs a first sound source from among the sound source through the speaker and capture a moving image through the microphone and the camera.


