Media Organization System Using Hash-Based Duplicate Elimination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The proliferation of digital media content leads to inefficient organization, resulting in significant redundant storage and processing of duplicate images, videos, and music, which consumes excessive memory, time, and computing resources, increasing energy usage and carbon footprint.

Innovation Solution

A media organization system that utilizes computing resources to identify and eliminate duplicates by analyzing media content through techniques such as cross-referencing, attribute parsing, hash generation, content analysis, and machine learning to determine repetitive features and attributes, thereby optimizing storage and processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If traditional media organization methods are used, then media content can be stored and accessed, but significant redundant storage and processing of duplicate content occurs, consuming excessive memory, time, and computing resources

Engineering Contradiction:
Improvestorage capacity utilizationVSAvoidenergy consumption
Core Design Contradiction:
Quantity of substanceVSLoss of energy

Solution Approach 1:

The system performs preliminary analysis of media content by generating hash values and extracting attributes before full storage occurs. This preliminary action identifies potential duplicates early in the ingestion process, preventing redundant storage from occurring in the first place, thereby reducing both storage waste and the energy required for subsequent processing

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces hash values and attribute data as intermediary representations of media content. These intermediaries serve as proxies that enable rapid comparison and duplicate detection without requiring full content analysis, significantly reducing the computational energy needed while maintaining accurate duplicate identification

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If traditional indexing and search operations are performed on large media content, then content retrieval is enabled, but energy requirements increase proportionally with the amount of total media content

Engineering Contradiction:
Improvemedia retrieval efficiencyVSAvoidenergy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system segments media content into distinct searchable attributes (hash values, metadata, content characteristics) that can be independently indexed and queried. This segmentation allows the search operation to work with smaller, more efficient data structures rather than scanning entire media files, reducing energy consumption while maintaining retrieval capability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces traditional mechanical full-content scanning with a computational substitution approach using hash-based lookup and attribute matching. This substitution enables constant-time or logarithmic-time retrieval operations instead of linear-time scanning, dramatically reducing the energy required for search operations as media collections grow

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If duplicate content is not identified and eliminated, then all media files are retained for user access, but redundant usage of memory, time, and computing resources increases significantly

Engineering Contradiction:
Improvemedia access availabilityVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system creates simplified copies of media content in the form of hash values and attribute metadata that capture essential identification features. These copies enable efficient duplicate detection and comparison operations without requiring complex analysis of the original media files, reducing processing complexity while preserving the ability to identify and manage duplicates

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20240112079A1Machine-learning techniques for carbon footprint optimization from improved organization of media
Publication Date: 2024.04.04 KUMAR RATIN
  • US20240112079A1 patent drawing
  • US20240112079A1 patent drawing
  • US20240112079A1 patent drawing

AI summary

Apparatuses, systems, and techniques to enable optimizations in storage and processing of media based on identification of repititions between two or more media content. In at least one embodiment, one or more repition of content is identified based on properties of media included.