Reality Content Generation With 3D Audio and Depth-Aligned Scenes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The high costs and labor-intensive nature of manually creating virtual reality content, such as virtual worlds and environments, pose challenges in achieving satisfactory visual quality.

Innovation Solution

A method utilizing generative adversarial networks (GANs) to generate virtual scenes, audio content, and depth maps from text, adjusting these based on sound attributes, and combining them to create 3D audio content for improved visual quality in reality services.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual creation of virtual reality content is used, then visual quality can be controlled, but production cost and time increase significantly

Engineering Contradiction:
Improvevisual qualityVSAvoidproduction time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent uses generative adversarial networks (GANs) to create virtual scenes, audio content, and depth maps by copying and generating content from text descriptions rather than manual creation. The GAN model learns from existing data to synthesize new virtual reality content, replacing the time-consuming manual modeling and rendering processes while maintaining high visual quality through the network's generated outputs.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces manual mechanical creation processes with automated AI-based systems. Instead of manually modeling 3D scenes and generating audio content, the system uses GANs to automatically generate virtual scenes, audio content, and depth maps from text inputs, substituting human creativity with machine learning algorithms that can produce similar or superior results more efficiently.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If manual creation of virtual reality content is used, then visual quality can be controlled, but production cost increases

Engineering Contradiction:
Improvevisual qualityVSAvoidproduction cost
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent uses generative adversarial networks (GANs) to create virtual scenes, audio content, and depth maps by copying and generating content from text descriptions rather than manual creation. The GAN model learns from existing data to synthesize new virtual reality content, replacing the time-consuming manual modeling and rendering processes while maintaining high visual quality through the network's generated outputs.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces manual mechanical creation processes with automated AI-based systems. Instead of manually modeling 3D scenes and generating audio content, the system uses GANs to automatically generate virtual scenes, audio content, and depth maps from text inputs, substituting human creativity with machine learning algorithms that can produce similar or superior results more efficiently.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If text-based generation is used, then production effort is reduced, but alignment precision of sound attributes may deteriorate

Engineering Contradiction:
Improveproduction efficiencyVSAvoidsound attribute alignment
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent incorporates feedback mechanisms where the GAN model iteratively refines generated content based on evaluations of sound attribute alignment. The system assesses the coherence between generated audio content and visual scenes, making adjustments to ensure proper spatial and temporal alignment of sound attributes with corresponding visual elements, thereby maintaining precision despite automated generation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent uses generative adversarial networks (GANs) to create virtual scenes, audio content, and depth maps by copying and generating content from text descriptions rather than manual creation. The GAN model learns from existing data to synthesize new virtual reality content, replacing the time-consuming manual modeling and rendering processes while maintaining high visual quality through the network's generated outputs.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12511849B2Method for improving visual quality of reality service content, host, and computer readable storage medium
Publication Date: 2025.12.30 HTC CORP
  • US12511849B2 patent drawing
  • US12511849B2 patent drawing
  • US12511849B2 patent drawing

AI summary

The embodiments of the disclosure provide a method for improving a visual quality of a reality service content, a host, and a computer readable storage medium. The method includes: generating a first virtual scene, an audio content, and a depth map based on a text, wherein the audio content includes an audio component and the depth map includes depth information; determining a sound attribute corresponding to an audio source based on the audio component and the depth information; adjusting the first virtual scene as a second virtual scene at least based on the sound attribute corresponding to the audio source; determining a 3D audio content at least based on the sound attribute and the audio content; and combining the 3D audio content with the second virtual scene into the reality service content.