Hash-Based File Management for Duplicate Content Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing size of content files due to high visual detail requirements leads to significant storage demands and prolonged download/access times, while data compression and optimization methods impose processing burdens or are time-consuming for content creators.
Innovation Solution
A system and method that uses unique file identifiers, such as hash keys, to identify and eliminate redundancy by storing pointers to existing content, reducing duplicate files through automated content management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data compression is applied to reduce file sizes, then storage requirements and download times are improved, but processing burden on the device increases due to decompression requirements
Solution Approach 1:
The system performs preliminary actions by generating unique identifiers for files before storage and maintaining an index of these identifiers. When files need to be accessed or compared, the pre-generated identifiers enable rapid matching without requiring decompression or complex analysis of the actual file content, thus reducing processing burden while maintaining storage efficiency.
Solution Approach 2:
The patent introduces unique identifiers (hash keys) as an intermediary between files and the storage system. Instead of directly managing or comparing large file contents, the system uses these compact identifiers as mediators to track, compare, and manage files, significantly reducing the processing burden while maintaining accurate file management capabilities.
2Quantity of substance
If data optimization is performed to reduce redundant data, then storage efficiency is improved, but time investment for content creators increases
Solution Approach 1:
The system implements self-service by automatically generating unique identifiers for files and maintaining the index without requiring manual intervention from content creators. The automated identification and comparison process eliminates the need for creators to manually optimize or track redundant data, improving storage efficiency while minimizing time investment.
Solution Approach 2:
The patent replaces manual mechanical processes (content creators manually identifying and removing redundant data) with an automated computational system that generates unique identifiers and automatically tracks file redundancy. This substitution eliminates the time-consuming manual optimization process while achieving superior storage efficiency.
3Quantity of substance
If high-capacity storage devices are purchased to accommodate large content files, then local storage capability is improved, but cost and device complexity increase
Solution Approach 1:
The system performs preliminary identification and tracking of files using unique identifiers before storage allocation. By maintaining an index of identifiers and their storage locations, the system can efficiently manage existing storage capacity without requiring additional high-capacity devices, thus avoiding increased device complexity while maintaining adequate local storage capability.
4Quantity of substance
If file identifiers and comparison systems are implemented to eliminate redundancy, then storage efficiency is improved, but system complexity increases
Solution Approach 1:
The patent transforms the file management approach by changing the parameter used for identification from complex file content analysis to simple unique identifiers (hash keys). This parameter change simplifies the comparison and redundancy detection process, improving storage efficiency while minimizing the increase in system complexity through the use of straightforward identifier matching.
Data Source
AI summary
A content obtaining system for obtaining content comprising a plurality of files, the system comprising, a database obtaining unit configured to obtain a database of unique identifiers corresponding to respective files stored by the content obtaining system, an identifier receiving unit configured to receive unique identifiers for one or more of the plurality of files in the content to be obtained, an identifier comparison unit configured to compare the received unique identifiers with the database, and to identify whether any of the identifiers match, and a content management unit configured to, in the case that there is no match for a given identifier, obtain and store the corresponding file for that identifier, and, in the case that there is a match for a given identifier, generate and store a reference to the location of the file stored by the content obtaining system for which the match was identified.


