Display Apparatus Voice Conversion Entity Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current IPTV systems lack the capability to customize the voice of entities in video content, limiting user preferences for voice changes in multimedia services.
Innovation Solution
A display apparatus and method that detects entities in video frames, allows user selection of voice samples, and replaces the original voice with the selected sample based on lip movement detection, using a voice conversion method and apparatus with modules for face detection, UI, storage, and audio processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If face replacement method is used in IPTV system, then entity face can be replaced with other faces, but voice customization capability is still lacking
Solution Approach 1:
The system segments the voice processing task by separating face detection, voice sample selection, and voice conversion into distinct functional modules. The face detection module identifies entities independently, the UI module handles voice sample selection separately, and the voice conversion module processes audio independently, allowing voice customization to be added without complicating the existing face replacement infrastructure.
Solution Approach 2:
The display apparatus is enhanced with multi-functionality by integrating both face replacement and voice customization capabilities within the same system. The voice conversion module can operate independently or in conjunction with face replacement, making the system versatile enough to handle multiple multimedia processing tasks without requiring separate dedicated systems.
2Adaptability or versatility
If voice conversion is added to customize entity voice, then user preference for voice changes is enabled, but system complexity increases
Solution Approach 1:
Voice samples are pre-collected and stored in the system's memory before actual voice conversion is needed. The UI module allows users to select from pre-prepared voice samples, eliminating the need for real-time voice processing and analysis. This preliminary preparation significantly reduces the computational complexity during actual video playback and voice conversion operations.
Solution Approach 2:
The UI module serves as an intermediary between the user and the voice conversion process. It provides a simplified interface for selecting voice samples and manages the mapping between selected voices and detected entities, abstracting away the complex audio processing details from the user while enabling versatile voice customization options.
3Measurement precision
If lip movement detection is used to replace voice, then voice synchronization is improved, but detection complexity increases
Solution Approach 1:
The system focuses lip movement detection specifically on the region of interest (the entity's face and lips) rather than analyzing the entire video frame. By localizing the detection area and using face detection results to guide lip movement analysis, the system achieves accurate voice-lip synchronization while reducing the computational complexity of detecting and measuring lip movements across the whole scene.
Data Source
AI summary
The voice conversion method of a display apparatus includes: in response to the receipt of a first video frame, detecting one or more entities from the first video frame; in response to the selection of one of the detected entities, storing the selected entity; in response to the selection of one of a plurality of previously-stored voice samples, storing the selected voice sample in connection with the selected entity; and in response to the receipt of a second video frame including the selected entity, changing a voice of the selected entity based on the selected voice sample and outputting the changed voice.


