Image Compression Using Machine Learning Caption Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image compression techniques face a trade-off between compression ratio and image quality, limiting the achievable compression ratio without significant loss of image quality.

Innovation Solution

The use of machine learning techniques, specifically an encoder apparatus that generates encoded data representative of an image along with caption data, and a decoder apparatus that uses trained machine learning models to reconstruct images from the encoded data, assisted by the caption data to improve image quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If lossy compression algorithms are used to achieve larger compression ratios, then storage and transmission requirements are reduced, but image quality deteriorates

Engineering Contradiction:
Improvecompression ratioVSAvoidimage quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent introduces an intermediary representation layer between the original image and the compressed data. This intermediary uses machine learning models to create a latent space representation that captures essential image features, allowing for high compression ratios while maintaining acceptable image quality through intelligent feature extraction and reconstruction

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the fundamental parameters of image representation by transitioning from traditional pixel-based compression to machine learning-based latent space compression. This parameter change enables the system to achieve much higher compression ratios (e.g., 10:1 or greater) while maintaining image quality, as the compression operates on semantic features rather than raw pixel data

Inventive Principle:
Principle #35Parameter changes

2Productivity

If traditional image compression techniques are used, then processing requirements are reduced, but the achievable compression ratio is limited

Engineering Contradiction:
Improvecompression ratioVSAvoidprocessing requirements
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical compression algorithms (like JPEG, PNG, ZIP) with machine learning-based compression systems. This substitution uses neural networks to learn optimal compression parameters and representations, achieving superior compression ratios while the computational complexity is managed through efficient model architectures and hardware acceleration

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250173911A1Apparatus, systems and methods for image processing
Publication Date: 2025.05.29 SONY INTERACTIVE ENTERTAINMENT LLC
  • US20250173911A1 patent drawing
  • US20250173911A1 patent drawing
  • US20250173911A1 patent drawing

AI summary

A decoder apparatus comprises receiving circuitry to receive caption data indicative of a language-based description for a first image and encoded data representative of the first image; and decoder circuitry comprising one or more trained machine learning models operable to generate a reconstructed image in dependence on the caption data and the encoded data, the reconstructed image having a higher image quality than an image quality associated with the encoded data representative of the first image.