Audio-to-Visual Image Generation with VQ-VAE Manifold Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computer vision algorithms require direct line-of-sight for effective operation, which is challenging in indoor environments due to occluding objects and limited scene visibility, and audio-based methods for visual information extraction are complex and limited to detecting sounding objects or require complex instrumentation.

Innovation Solution

A two-stage method using a vector-quantized variational auto-encoder (VQ-VAE) to learn a data manifold for visual modalities and an audio transformation network (AT-net) to map audio data to visual representations, enabling the generation of visual images from audio data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If computer vision algorithms are used for visual tasks, then visual information can be processed, but direct line-of-sight is required which is challenging in indoor environments due to occluding objects

Engineering Contradiction:
Improveeffectiveness of visual algorithmsVSAvoidoccluding objects blocking vision
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent introduces audio data as an intermediary medium to bridge the gap between sounding objects and visual representation. Instead of requiring direct visual line-of-sight, the system uses audio signals as a mediator that can penetrate occluding objects and reach the microphone array, thereby resolving the contradiction between visual algorithm effectiveness and occlusion problems

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical/optical vision system with an acoustic system. By substituting the requirement for direct visual contact with audio-based sensing, the system eliminates the line-of-sight constraint while maintaining the ability to detect and represent objects in the environment

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If audio-based methods are used for visual information extraction, then line-of-sight is not required, but the methods are complex and limited to detecting sounding objects

Engineering Contradiction:
Improveability to work without line-of-sightVSAvoidcomplexity of audio-based methods
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal audio-to-visual transformation system that can handle multiple types of visual information (depth maps, semantic segmentation, RGB images) through a single unified framework. The VQ-VAE model serves as a multi-functional translator that converts audio representations into various visual modalities, eliminating the need for separate complex instrumentation for each visual task

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent transforms the problem from detecting specific acoustic properties of sounding objects to learning a general transformation mapping between audio and visual parameter spaces. By changing the approach from object-specific detection to parameter-space transformation through manifold learning, the system achieves versatility without requiring complex specialized instrumentation

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If end-to-end models are used for audio-to-visual transformation, then the system is simpler, but the accuracy of visual reconstruction is limited

Engineering Contradiction:
Improvesimplicity of end-to-end modelVSAvoidaccuracy of visual reconstruction
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the audio-to-visual transformation process into distinct stages: audio encoding, manifold mapping, and visual decoding. By dividing the transformation into separate functional components (audio encoder, VQ-VAE manifold, visual decoder), the system achieves both computational tractability and high reconstruction accuracy, avoiding the limitations of both simple end-to-end models and overly complex multi-stage systems

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12394105B2Systems and methods for generating a visual image from audio data, and systems and methods for training the same
Publication Date: 2025.08.19 HUAWEI TECH CANADA CO LTD
  • US12394105B2 patent drawing
  • US12394105B2 patent drawing
  • US12394105B2 patent drawing

AI summary

Systems and methods for generating a visual image from audio data and for training the same. The method may include: mapping audio data registered with a microphone array onto closest visual representations in a data manifold for latent representation of images of a visual modality; and generating a visual image of the visual modality from the closest visual representations.