Physics-Guided Impact Sound Synthesis From Silent Video

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current impact sound synthesis methods require sophisticated environments for physics simulation and time-consuming parameter selection, leading to impracticality in capturing complex object interactions and often generate unfaithful sounds due to the lack of essential physics knowledge.

Innovation Solution

Reconstructing physics priors from audio and video training data to train a generative model, integrating physics knowledge into the synthesis process using a diffusion model, and processing visual latent vectors and Gaussian noise to generate realistic impact sounds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If physics-based synthesis models are used to simulate impact sounds, then the accuracy of sound generation can be improved, but the device complexity and time required for parameter selection increase significantly

Engineering Contradiction:
Improveaccuracy of sound generationVSAvoidcomplexity of simulation environment
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces complex physics-based simulation systems with a data-driven deep learning model. Instead of using sophisticated physics engines and manual parameter selection, the system trains a neural network on recorded impact sounds to learn the mapping from visual inputs to audio outputs, thereby substituting mechanical/physical simulation with computational learning

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates a virtual copy of physical impact sounds through data collection and training. By recording actual impact sounds and using them to train the generative model, the system creates a digital replica that can synthesize impact sounds without requiring physical simulation environments, effectively copying the essence of physics-based sound generation

Inventive Principle:
Principle #26Copying

2Ease of operation

If end-to-end black box model training is used, then the ease of operation is improved, but the reliability of sound generation deteriorates due to lack of physics knowledge

Engineering Contradiction:
Improveease of model trainingVSAvoidfidelity of generated sound
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent segments the sound generation process into distinct components: a visual encoder that processes video inputs, a physics-informed generative model that creates audio from visual latent representations, and a decoder that outputs the final sound. This segmentation allows each component to be optimized independently while maintaining overall system reliability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces visual latent representations as an intermediary between the video input and audio generation. Instead of directly mapping pixels to sound waves, the system uses latent vectors as a bridge that captures essential visual features, allowing the generative model to focus on creating physically plausible sounds without being constrained by pixel-level details

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250356834A1Impact sound synthesis using physics-driven diffusion model
Publication Date: 2025.11.20 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20250356834A1 patent drawing
  • US20250356834A1 patent drawing
  • US20250356834A1 patent drawing

AI summary

According to one embodiment, a method, computer system, and computer program product for predicting and synthesizing audio of an impact depicted in a video is provided. The present invention may include reconstructing physics priors from received audio and video training data; training a generative model for impact sound synthesis using the reconstructed physics priors to guide the generative model in learning a correspondence between video inputs and impact sounds; receiving silent video input to produce a visual latent vector representation, wherein the video input depicts an impact between two or more physical objects; and processing the visual latent vector representation, the reconstructed physics priors, and Gaussian noise through the trained generative model to perform the impact sound synthesis.