Video-Linked Audio Object Editing for Intuitive Sound Field Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio technologies lack intuitive and efficient methods for personalized audio editing, particularly in video recording and video call scenarios, limiting user experience and flexibility in adjusting audio features such as volume, timbre, and spatial position.

Innovation Solution

An electronic device provides intuitive audio editing controls on a video image, allowing users to adjust audio objects, sound fields, and mixing parameters directly, enabling features like volume adjustment, timbre change, spatial positioning, and audio object separation, with real-time editing capabilities during recording and playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If traditional audio editing methods are used, then audio processing can be performed, but the operation becomes complex and non-intuitive, reducing ease of operation

Engineering Contradiction:
Improveaudio editing operationVSAvoidaudio editing control
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the audio stream into multiple independent audio objects (e.g., speaker audio, background audio, music audio) that can be individually controlled. Each audio object is associated with its corresponding video object, allowing users to edit specific audio components without affecting others, thereby simplifying the editing operation while maintaining system complexity for advanced functionality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an audio object as an intermediary between the video content and the audio processing controls. This audio object serves as a mediator that links specific video objects (speakers, instruments) with their corresponding audio streams, enabling intuitive control through visual association while managing the complexity of multi-track audio editing through a unified interface.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If audio editing controls are provided for each audio object, then personalized audio editing is enabled, but the control interface becomes more complex

Engineering Contradiction:
Improveaudio editing capabilityVSAvoidcontrol interface
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal audio object control mechanism that can handle multiple audio objects simultaneously through a standardized interface. The same control elements (volume sliders, mute buttons, spatial position controls) can be applied to any audio object, providing versatile editing capability while avoiding the need for different control schemes for different audio types, thus managing interface complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent adds a spatial dimension to audio control by associating audio objects with their visual positions in the video frame. Audio controls are displayed in proximity to the corresponding video objects, creating a spatial mapping between visual and audio editing dimensions. This allows users to intuitively understand which audio control corresponds to which sound source without complex menus or labels.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Manufacturing precision

If audio objects are separated and individually controlled, then audio quality and flexibility are improved, but processing complexity increases

Engineering Contradiction:
Improveaudio editing precisionVSAvoidaudio processing system
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent performs audio object separation and association with video objects during the video recording or playback preparation phase, before the actual editing operation. This preliminary processing organizes the audio stream into structured audio objects with identified sources and spatial positions, enabling precise individual control later without requiring complex real-time processing during editing, thus achieving high precision with manageable system complexity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4708923A1Audio processing method and related apparatus
Publication Date: 2026.03.11 HUAWEI TECH CO LTD
  • EP4708923A1 patent drawingFigure 1
  • EP4708923A1 patent drawingFigure 2
  • EP4708923A1 patent drawingFigure 3

AI summary

This application provides an audio processing method and a related apparatus. An electronic device may determine audio object information and sound field environment information based on a video image and an audio in a video. Based on the foregoing information, the electronic device may provide an audio editing control in a user interface, so that a user performs audio editing such as audio object editing, sound field editing, and audio mixing editing. The electronic device may process the audio based on an audio editing operation of the user, so that a sound effect that can be heard by the user after a processed audio is rendered can achieve an effect of adjustment by the user. The foregoing audio editing operations are simple and intuitive, so that the user can conveniently perform personalized editing on the audio in an audio and video recording scenario like a video recording scenario or a video call scenario, to improve user experience of audio and video recording.