Videotelephony Image Compression via Segmented Region Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current videotelephony systems face challenges in achieving a balance between image and sound synchronization and quality, particularly in maintaining fluidity and reducing desynchronization between image and sound, due to limitations in passband and existing compression methods.
Innovation Solution
A method involving principal component analysis (PCA) and independent component analysis (ICA) for image compression, focusing on key areas like the mouth and eyes, with adaptive learning bases and dynamic updating to enhance compression efficiency and synchronization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional compression methods are used to reduce image data rate, then transmission bandwidth is reduced, but image quality and synchronization with sound deteriorate
Solution Approach 1:
The image is divided into multiple zones with different compression priorities. The mouth zone (and other active regions) is segmented as a separate region of interest that receives higher compression priority and lower compression ratio, ensuring better quality and synchronization. Less important zones use higher compression ratios, reducing overall data rate while maintaining critical synchronization quality.
Solution Approach 2:
Different compression qualities are applied to different regions of the image. The mouth zone and other active regions are identified and assigned local quality parameters that ensure high fidelity and temporal synchronization. Background or less important areas use lower quality settings, optimizing the trade-off between overall data rate and critical synchronization requirements.
2Speed
If high compression ratios are applied to refresh the entire image rapidly, then fluidity is improved, but data transmission requirements increase and desynchronization occurs
Solution Approach 1:
Instead of refreshing the entire image at high rate, only the mouth zone and other active regions are refreshed rapidly. The rest of the image uses lower refresh rates, reducing overall data transmission requirements while maintaining the fluidity and synchronization critical for speech perception in the important regions.
Solution Approach 2:
High refresh rate and low compression are applied partially only to the mouth zone and active regions, rather than excessively to the entire image. This partial application of high-quality compression maintains critical fluidity and synchronization where needed, while reducing overall data rate by using lower quality settings in less important areas.
3Device complexity
If the entire image is compressed uniformly, then processing is simplified, but critical regions like the mouth do not achieve sufficient compression for good synchronization
Solution Approach 1:
The image is segmented into multiple zones with different compression parameters. The mouth zone is identified and separated from the rest of the image, allowing dedicated compression processing optimized for synchronization. This segmentation increases processing complexity slightly but dramatically improves synchronization quality in critical regions compared to uniform compression.
Solution Approach 2:
Different compression algorithms and parameters are applied locally to different regions. The mouth zone uses compression settings optimized for temporal fidelity and synchronization, while other regions use different settings. This local quality approach improves synchronization reliability despite increased processing complexity.
Data Source
AI summary
A method of compression of videotelephony images characterized by: creating (10) a learning base containing images; centering the learning base about zero; determining component images by principal component analysis (12); and keeping a number of significant principal components (14).


