Audio Object Editing for Intuitive Video Sound Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio technologies lack intuitive and efficient methods for personalized audio editing in audio and video recording scenarios, limiting user experience and flexibility in adjusting sound features and environments.
Innovation Solution
An electronic device provides controls for adjusting audio object, sound field, and audio mixing parameters, allowing users to edit audio features such as volume, timbre, and spatial position directly on a video image, with capabilities for audio object separation, sound field editing, and real-time audio mixing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If traditional audio editing methods are used, then audio editing functionality is provided, but the editing process is complex and not intuitive for users
Solution Approach 1:
The patent segments the audio stream into multiple independent audio objects (e.g., speaker audio, background audio, music audio) that can be individually controlled. This segmentation allows users to edit specific audio components without affecting others, making the editing process more intuitive and manageable while reducing the perceived complexity of the system.
Solution Approach 2:
The patent introduces an audio object as an intermediary between the audio stream and the editing controls. This audio object serves as a mediator that links specific audio content to editing parameters, enabling intuitive user interaction through visual representations (such as bounding boxes or icons) while managing the complexity of audio processing internally.
2Adaptability or versatility
If audio editing controls are provided, then personalized audio editing is enabled, but the user interface becomes more complex
Solution Approach 1:
The patent applies local quality by providing different types of editing controls tailored to specific audio objects and their characteristics. For example, different controls are provided for speaker audio versus background audio, and different spatial adjustment options are provided based on the audio object type. This targeted approach enables personalized editing without overwhelming the user with all possible controls simultaneously.
Solution Approach 2:
The patent implements dynamic control interfaces that adapt based on the selected audio object, current editing mode, and user preferences. The user interface dynamically displays only the relevant controls for the currently selected audio object, allowing for high adaptability and customization capability while maintaining a clean, simple interface by hiding irrelevant controls.
3Measurement precision
If audio object separation is implemented, then precise audio control is achieved, but processing complexity increases
Solution Approach 1:
The patent creates simplified representations or copies of audio objects (such as visual indicators, metadata tags, or control proxies) that correspond to the separated audio components. These copies enable precise identification and control of audio objects through user-friendly interfaces without requiring the user to directly interact with the complex audio processing algorithms, thus achieving high precision while managing processing complexity.
Data Source
AI summary
An electronic device including a camera configured to shoot an image, a microphone array configured to capture audio, wherein the electronic device is configured for displaying a first video image in a first video, where the first video image comprises a first audio object, and wherein the first video comprises first audio, displaying a first audio object control on a screen, where the first audio object control is associated with adjusting a sound feature of the first audio object, and where adjusting the sound feature of the first audio object includes one or more of adjusting volume of the first audio object, adjusting a timbre of the first audio object, or adjusting a spatial position of the first audio object, and adjusting, in response to a first operation on the first audio object control, the sound feature of the first audio object in the first audio.


