Audio-Driven Virtual Image Generation via TRIZ Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtual image generation methods require expensive hardware and significant computational resources, especially in VR, AR, and MR scenarios, reducing efficiency and increasing costs due to the need for video processing.
Innovation Solution
A method that extracts audio features to acquire expression and pose parameters, generating auxiliary information for texture and geometric shapes, allowing for virtual image creation without video processing, thus reducing hardware and computational requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If video-based methods are used for virtual image generation, then generation accuracy is improved, but hardware cost and computational resource consumption increase significantly
Solution Approach 1:
The patent extracts and utilizes only the audio component from multimedia input, separating it from video processing. By focusing solely on audio features for driving virtual image generation, the system eliminates the need for expensive image capture devices and video processing hardware while maintaining generation accuracy through audio-driven expression and pose parameters.
Solution Approach 2:
The patent replaces the mechanical/video-based processing system with an audio-based system. Instead of processing video streams that require complex hardware, the system processes audio signals to extract features and generate virtual images, substituting a simpler audio processing mechanism for the complex video processing mechanism.
2Measurement precision
If video processing is used for virtual image generation, then generation accuracy is improved, but computational resource consumption and processing time increase
Solution Approach 1:
The patent extracts only the necessary audio features from the input signal, avoiding the computational overhead of processing entire video streams. By focusing extraction on audio components that directly influence expression and pose, the system reduces computational resource consumption while maintaining the accuracy needed for realistic virtual image generation.
Solution Approach 2:
Instead of the conventional approach of generating virtual images from video input, the patent inverts the process by generating virtual images from audio input. This inversion allows the system to bypass computationally intensive video processing while achieving similar or better generation accuracy through audio-driven parameters.
3Device complexity
If audio-based methods are used for virtual image generation, then hardware cost is reduced, but generation accuracy may deteriorate
Solution Approach 1:
The patent changes the input parameter from video data to audio data, and accordingly changes the processing parameters to audio features, expression parameters, and pose parameters. This parameter transformation allows the system to use simpler, cheaper hardware while maintaining generation accuracy by focusing computational resources on the most relevant audio-driven parameters for realistic virtual image synthesis.
Data Source
AI summary
Embodiments of the present disclosure relate to a method, an electronic device, and a computer program product for generating a virtual image. The method includes extracting an audio feature of an audio input of a target object; and acquiring an expression parameter and a pose parameter associated with the target object based on the audio feature. The method further includes generating, based on the audio feature, auxiliary information related to a texture for at least a portion of the target object and a geometric shape of at least a portion of the target object. The method further includes generating a virtual image of the target object based on the expression parameter, the pose parameter, and the auxiliary information.


