Media Encoding Pixel Classification Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video processing and image processing technologies face challenges in efficiently encoding and decoding media content, particularly due to the large search space of encoding options and the lack of reliable perceptual metrics for lossy compression, which can result in suboptimal visual quality and compression efficiency.

Innovation Solution

The proposed solution involves classifying each pixel in a media item as either Photographic or Non-Photographic content, generating a pixel-clusters map, and applying different encoding techniques (lossy for Photographic content and lossless for Non-Photographic content) to optimize visual quality and compression efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a single encoding technique is applied to the entire media item, then the encoding process is simple and fast, but the visual quality and compression efficiency are suboptimal

Engineering Contradiction:
Improvecompression efficiencyVSAvoidencoding process complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides the media item into multiple regions based on content type (photographic vs. non-photographic). Each region is then encoded using the most suitable encoding technique for its specific content characteristics. This segmentation allows optimization of compression efficiency for each region while maintaining overall manageable complexity through automated classification.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different encoding techniques are applied to different regions of the media item based on their specific content characteristics. Photographic regions use one encoding approach optimized for their properties, while non-photographic regions use another approach. This local quality approach ensures optimal compression and visual quality for each region rather than using a uniform approach throughout.

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If lossy encoding is applied to maximize compression efficiency, then the file size is reduced, but the visual quality deteriorates

Engineering Contradiction:
Improvefile sizeVSAvoidvisual quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent applies different encoding strategies to different regions based on their content type and importance. Critical regions with photographic content may use lossy encoding to reduce file size, while regions with non-photographic content or high visual importance use lossless encoding to preserve quality. This localized approach optimizes the balance between file size and visual quality for each region.

Inventive Principle:
Principle #3Local quality

3Manufacturing precision

If the search space of encoding options is exhausted to find the optimal encoding, then the encoding quality is maximized, but the processing time increases significantly

Engineering Contradiction:
Improveencoding qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary classification of media regions into photographic and non-photographic content types before encoding. This preliminary action guides the selection of appropriate encoding techniques in advance, eliminating the need to exhaustively search through all possible encoding options. The classification results are used to directly select optimal encoding parameters, significantly reducing processing time while maintaining high encoding quality.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12243274B2System, device, and method for improved encoding and enhanced compression of images, videos, and media content
Publication Date: 2025.03.04 CLOUDINARY LTD
  • US12243274B2 patent drawing
  • US12243274B2 patent drawing
  • US12243274B2 patent drawing

AI summary

System, device, and method for improved encoding and enhanced compression of images, videos, and media content. A method includes: (a) receiving a source image, and analyzing its content on a pixel-by-pixel basis, and classifying each pixel as either (I) a pixel associated with Photographic content, or (II) a pixel associated with Non-Photographic content; (c) generating a pixel-clusters map that indicates (i) clusters of pixels that were classified as Photographic content, and (ii) clusters of pixels that were classified as Non-Photographic content; (d) generating a composed image, by: (d1) applying a first encoding technique, particularly lossy encoding, to encode pixel-clusters that were classified as Photographic content; (d2) applying a second, different, encoding technique, particularly lossless encoding, to encode pixel-clusters that were classified as Non-Photographic content.