AI Audio Score Composition for Visual Content

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies for generating audio scores for visual content, such as films and live events, are limited in their ability to create dynamic and adaptive soundscapes, often relying on pre-recorded and scripted audio that cannot respond effectively to real-time or improvised situations.

Innovation Solution

Training Artificial Intelligence (AI) models to analyze visual and textual features from multimedia datasets, enabling the generation of audio scores that synchronize with visual content in real-time, using genre, instrument, sound effect, and anticipation models to create dynamic and contextually relevant audio.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If pre-recorded and scripted audio is used for visual content, then production process is simplified and predictable, but the audio cannot respond effectively to real-time or improvised situations

Engineering Contradiction:
Improveadaptability to real-time eventsVSAvoidcomplexity of audio generation system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The AI model analyzes visual features from the video content itself and autonomously generates corresponding audio scores without requiring external human intervention or pre-scripted audio files. The system serves itself by extracting features directly from the visual input and composing appropriate audio responses in real-time.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The audio generation system transitions from static pre-recorded audio to dynamic real-time composition. The AI model continuously analyzes changing visual features and generates audio scores that adapt dynamically to the current state of the video content, enabling responsiveness to improvised situations.

Inventive Principle:
Principle #15Dynamics

2Productivity

If AI models generate audio scores in real-time, then adaptability to visual content is improved, but computational resources and processing time increase

Engineering Contradiction:
Improvereal-time audio generation speedVSAvoidcomputational energy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The audio generation process is divided into distinct stages: visual feature extraction, feature analysis, and audio score composition. Each stage processes only the necessary information for that specific task, reducing overall computational burden while maintaining real-time performance capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system extracts only the most relevant visual features needed for audio composition rather than processing all visual data. This partial action approach reduces computational energy consumption while still generating sufficiently accurate audio scores in real-time.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If pre-planned audio is used for films, then copyright issues are avoided, but emotional impact and immersion are reduced

Engineering Contradiction:
Improvecopyright complianceVSAvoidemotional resonance with visual content
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

Instead of using pre-existing copyrighted audio works, the system generates original audio scores by copying and transforming visual features into corresponding audio representations. This creates unique, copyright-free audio content that is specifically tailored to match the emotional and narrative elements of each visual scene.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20240420670A1Artificial intelligence models for composing audio scores
Publication Date: 2024.12.19 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20240420670A1 patent drawing
  • US20240420670A1 patent drawing
  • US20240420670A1 patent drawing

AI summary

A computing system, and corresponding method, are disclosed that extracts visual features from a visual dataset, including features related to facial recognition or human expressions. The computing system further identifies audio features corresponding to the visual features. It utilizes an audio-scoring artificial intelligence (AI) engine to compose an audio score to present the visual features. The AI engine is trained to generate audio scores based on the audio features corresponding to the visual features, enhancing the overall presentation of the visual dataset.