Management Server Viewpoint-Specific Audio Video Synchronization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems fail to provide users with an immersive audio experience that aligns with the camera position during live events, as they receive a single mixed sound regardless of camera switching, lacking realism.

Innovation Solution

A management server connected to user and distributor terminals via a network, which extracts and processes audio data based on camera position, calculating audio coefficients to simulate the sound heard at the desired camera position, and synchronizes this with video data to deliver an immersive audio-visual experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single mixed audio is transmitted regardless of camera switching, then the system complexity is reduced, but the audio realism and immersion at different viewpoints deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidaudio realism
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent segments the audio signal into multiple independent channels, each corresponding to a specific camera viewpoint. Instead of transmitting a single mixed audio stream, the system transmits separate audio data for each camera position, allowing the terminal to select and play the audio that matches the currently displayed video camera angle, thereby achieving viewpoint-specific audio realism.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by providing different audio characteristics for different spatial locations (camera positions). Each camera viewpoint has its own optimized audio mix that reflects what would be heard from that specific position in the venue, rather than using a uniform mixed audio for all viewpoints.

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If audio is processed and transmitted for each camera position, then the audio realism at different viewpoints is improved, but the network bandwidth and processing requirements increase

Engineering Contradiction:
Improveaudio realismVSAvoidnetwork bandwidth
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The audio processing is performed in advance during the broadcasting phase. Audio is recorded and processed for each camera position beforehand, and the pre-processed audio data is packaged together with the corresponding video data. This eliminates the need for real-time audio processing and mixing at the terminal, reducing computational requirements and network bandwidth needs during playback.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If audio and video are synchronized according to camera position, then the immersion and presence experience is enhanced, but the system complexity and coordination requirements increase

Engineering Contradiction:
Improveaudio-video synchronizationVSAvoidsystem coordination
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent merges the audio and video data streams at the packaging level. Each camera viewpoint's audio and video data are bundled together as a unified data package, ensuring that when the terminal selects a specific camera angle for video display, the corresponding pre-synchronized audio is automatically played, maintaining perfect audio-video synchronization without requiring complex real-time coordination.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12132940B2Management server
Publication Date: 2024.10.29 ALIEN MUSIC ENTERPRISE INC
  • US12132940B2 patent drawing
  • US12132940B2 patent drawing
  • US12132940B2 patent drawing

AI summary

[Summary][Problem to be Solved]To construct a management server that can provide video and audio according to the position, where a camera films, to user terminals.[Solution]It is configured that a management server comprises a camera extraction section receiving a desired camera from a user terminal, an audio extraction section receiving audio generation data from a distributor terminal, a video extraction section receiving video generation data corresponding to a filming subject from the distributor terminal, a distribution audio generation section generating distribution audio to be distributed to the user terminal, a video generation section generating distribution video to be distributed to the user terminal based on the video of the filming subject and the distribution audio.