Multimodal Machine Learning for 3D Audio Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio technologies primarily utilize stereo sound, which limits the ability to create immersive three-dimensional audio experiences, as most media creators lack the tools and knowledge to design and implement 3D audio environments.

Innovation Solution

The use of multimodal machine learning methods and neural networks to automatically generate three-dimensional audio by segmenting and associating image and audio elements from multimedia content, such as videos and games, to create spatially accurate soundtracks that mimic real-world environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional stereo sound is used, then device complexity is reduced and ease of operation is improved, but audio immersion and spatial accuracy are limited

Engineering Contradiction:
Improveease of audio implementationVSAvoidaudio immersion capability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system automatically generates 3D audio from existing stereo media without requiring manual intervention or specialized knowledge from content creators. The machine learning model processes input audio and images to autonomously produce spatially accurate soundtracks, making 3D audio implementation self-service rather than requiring expert expertise

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms audio from two-dimensional stereo channels to three-dimensional spatial distribution by changing the parameter representation. The machine learning model maps audio parameters to spatial coordinates, converting traditional stereo audio into immersive 3D audio experiences while maintaining compatibility with existing media

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If manual 3D audio design is performed, then audio quality and immersion are improved, but device complexity and time consumption increase

Engineering Contradiction:
Improveaudio immersion capabilityVSAvoidtime for audio production
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The machine learning model is pre-trained on extensive datasets of audio-visual pairs, performing the complex analysis and generation work in advance. During actual use, the model quickly processes new input media by applying learned patterns, eliminating the need for time-consuming manual 3D audio design while maintaining high quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual audio engineering processes with an automated machine learning system. Instead of requiring audio professionals to manually create 3D audio experiences, the system uses neural networks to automatically process media and generate immersive soundtracks, substituting mechanical manual work with intelligent automation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If machine learning-based 3D audio generation is implemented, then audio immersion and spatial accuracy are improved, but computational resources and processing time increase

Engineering Contradiction:
Improveaudio immersion capabilityVSAvoidcomputational energy consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The system segments the audio processing task into distinct stages: input media analysis, spatial mapping, and audio generation. The machine learning model processes audio and visual inputs separately before integrating them to produce final 3D audio, reducing computational load compared to processing everything simultaneously while maintaining immersion quality

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240220866A1Multimodal machine learning for generating three-dimensional audio
Publication Date: 2024.07.04 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20240220866A1 patent drawing
  • US20240220866A1 patent drawing

AI summary

Methods and systems use one or more machine learning models to automatically generate three-dimensional sound. A multimodal content item is accessed by a computing device. Three-dimensional sound is automatically generated by the computing device using the one or more machine learning models based on the multimodal content item.