Media Signal Synchronization via Characteristic Feature Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for synchronizing audio and video signals, such as side information embedding, test signal generation, and end of path checking, are limited by issues like data corruption, intrusive procedures, and compatibility with only certain types of content, necessitating a method that can correct 'lip sync' errors in-service and across various media signals and transmission paths.
Innovation Solution
A system and method that involves receiving synchronized input media signals, extracting characteristic features, calculating signal delays through a network, and outputting a synchronization signal to determine the extent of desynchronization between media signals, allowing for real-time correction without intrusive methods or content-specific limitations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If side information embedding method is used to measure and correct lip sync errors, then synchronization accuracy can be improved, but the side information may be dropped or corrupted in the transmission path
Solution Approach 1:
The patent uses audio and video characteristic features (intermediaries) extracted from the media signals themselves to determine delay, rather than relying on separate side information channels. The audio characteristic extraction module and video characteristic extraction module derive timing information directly from the content of the signals, making the measurement process more reliable against data loss
Solution Approach 2:
The system performs self-measurement by extracting features from its own input and output signals. The delay determination module compares characteristic features from input signals with those from output signals, allowing the system to autonomously measure its own delay without external reference signals or vulnerable side channels
2Measurement precision
If test signal generation and detection method is used, then synchronization measurement can be performed, but the system must be taken out of service and cannot detect delay changes during operation
Solution Approach 1:
The patent enables continuous synchronization measurement during normal system operation. The characteristic extraction modules continuously process input and output signals, and the delay determination module continuously calculates delay values, allowing the system to detect and correct lip sync errors without interruption or taking the system out of service
Solution Approach 2:
The system implements continuous feedback by comparing characteristic features of input signals with output signals in real-time. The delay determination module uses this feedback to continuously determine delay values, enabling dynamic adjustment and correction of synchronization errors during operation rather than relying on periodic out-of-service checks
3Ease of operation
If end of path checking method is used to correct lip sync errors, then synchronization can be adjusted, but the method is limited to certain types of content such as video containing facial movements and audio containing corresponding speech
Solution Approach 1:
The patent employs multiple characteristic extraction modules that can handle different types of media signals. The audio characteristic extraction module and video characteristic extraction module are designed to extract relevant features from various content types, not limited to specific scenarios like facial movements and speech, thereby providing universal applicability across different content genres
Solution Approach 2:
The system adapts to different content types by changing the parameters and methods of characteristic extraction. Depending on the input signal characteristics, the extraction modules can adjust which features to prioritize, enabling the delay determination module to effectively measure delay across diverse content types rather than being restricted to predetermined scenarios
Data Source
AI summary
The embodiments described herein provide a method and system for determining the extent to which a plurality of media signals are out of sync with each other. The method includes: receiving a first input media signal and a second input media signal wherein the first and second input media signals are in sync with each other; extracting at least one first characteristic feature from the first input media signal; extracting at least one second characteristic feature from the second input media signal; receiving a first output media signal and a second output media signal wherein the first output signal corresponds to the first input media signal after being transmitted through a network, and the second output signal corresponds to the second input media signal after being transmitted through the network; extracting the at least one first characteristic feature from the first output media signal; extracting the at least one second characteristic feature from the second output media signal; calculating a first signal delay based on the at least one first characteristic feature extracted from the first input and output media signals; calculating a second signal delay based on the at least one second characteristic feature extracted from the second input and output media signals; and outputting a synchronization signal based on the difference between the first and second delay signals wherein the synchronization signal represents the extent to which the first and second output media signals are out of sync with each other.


