Audiovisual Data Search Using Generative Latent Codes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing lossy compression algorithms for audio and visual data fail to leverage conditional distributions and require manual feature selection, leading to limited readability and malleability, and are agnostic to content-specific statistical regularities.

Innovation Solution

A system utilizing generative models with latent codes to process, store, and transmit audiovisual data by generating content on the fly from partial instructions, employing trained generator mappings and compressor mappings to convert between uncompressed data and latent codes, enabling efficient compression and security through specialized hardware.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional lossy compression algorithms are used, then compression is achieved, but the algorithms are agnostic to content-specific statistical regularities and require manual feature selection

Engineering Contradiction:
Improvecompression effectivenessVSAvoidcontent adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system uses generative models that automatically learn and adapt to content-specific statistical regularities through training on representative datasets. The model performs self-optimization by adjusting its latent space mapping and generator functions to capture the statistical structure of the specific content type, eliminating the need for manual feature selection and achieving both reliable compression and content adaptability

Inventive Principle:
Principle #25Self-service

2Reliability

If generative models with latent codes are used, then compression and security are improved, but computing power requirements increase

Engineering Contradiction:
Improvecompression qualityVSAvoidcomputing power consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The generative model is pre-trained on representative datasets before actual compression tasks. This preliminary training phase allows the model to learn optimal latent space representations and statistical regularities in advance, so that during actual compression operations, the model can efficiently process data without requiring excessive computing power for learning or feature extraction

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If latent codes are used for compression, then storage and transmission efficiency are improved, but the system complexity increases

Engineering Contradiction:
Improvedata sizeVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system segments the compression process into distinct functional components: a compressor that maps input data to latent codes, a generative model that defines the latent space structure, and a generator that reconstructs data from latent codes. This segmentation allows each component to be optimized independently and simplifies the overall system architecture while achieving efficient data representation in the latent space

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12411886B2Systems and methods for searching audiovisual data using latent codes from generative networks and models
Publication Date: 2025.09.09 UNKNOT INC
  • US12411886B2 patent drawing
  • US12411886B2 patent drawing
  • US12411886B2 patent drawing

AI summary

Systems and methods for viewing, storing, transmitting, searching, and editing application-specific audiovisual content (or other unstructured data) are disclosed in which edge devices generate content on the fly from a partial set of instructions rather than merely accessing the content in its final or near-final form. An image processing architecture may include a generative model that may be a deep learning model. The generative model may include a latent space comprising a plurality of latent codes and a trained generator mapping. The trained generator mapping may convert points in the latent space to uncompressed data points, which in the case of audiovisual content may be generated image frames. The generative model may be capable of closely approximating (up to noise or perceptual error) most or all potential data points in the relevant compression application, which in the case of audiovisual content may be source images.