Cloud Data Placement Optimization via Content Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The growing volume and complexity of data generated by individuals, enterprises, and organizations pose challenges in efficiently managing and storing data, particularly with the rise of 'big data,' as conventional database management tools struggle to determine which data is duplicate or obsolete, leading to increased storage costs in cloud IT services.
Innovation Solution
A system that utilizes a processor to extract content and context from data, classify it based on its characteristics, and determine the optimal storage location among various cloud-based IT services, dynamically monitoring and modifying data placement to optimize costs, security, and availability by employing analytic and semantic techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If entities purchase more local storage to accommodate growing data volume, then storage capacity increases, but storage costs increase proportionally
Solution Approach 1:
The patent extracts data from the original storage environment and places it in a cloud-based storage system. The system identifies and extracts duplicate and obsolete data elements, separating them from the active data set, and stores them in a cloud repository, thereby reducing the amount of data that needs to be maintained in primary storage and reducing overall storage costs.
Solution Approach 2:
The patent introduces a cloud-based storage system as an intermediary between the entity's data management needs and physical storage infrastructure. This intermediary cloud system provides automated data management services including duplicate detection, data classification, and intelligent storage placement, eliminating the need for the entity to directly manage storage infrastructure and reducing storage costs.
2Ease of operation
If entities rely on cloud IT services for data storage, then local storage maintenance is eliminated, but storage costs become perpetual and accumulate over time
Solution Approach 1:
The patent performs preliminary actions by automatically identifying, classifying, and archiving duplicate and obsolete data before it consumes excessive cloud storage resources. The system proactively manages data lifecycle stages, moving data to appropriate storage tiers based on its value and accessibility requirements, thereby preventing unnecessary accumulation of storage costs.
Solution Approach 2:
The patent changes the storage parameters of data by dynamically adjusting where data is stored based on its characteristics. The system modifies storage location, access frequency, and retention period parameters for different data elements, transitioning data between cloud storage tiers and local storage as appropriate, thereby optimizing the balance between operational ease and cost accumulation.
3Productivity
If conventional database management tools are used to process big data, then data processing is performed, but duplicate and obsolete data cannot be effectively identified
Solution Approach 1:
The patent implements a feedback mechanism that continuously analyzes stored data to identify duplicates and obsolete elements. The system provides feedback loops that monitor data access patterns, storage utilization, and data relationships, automatically adjusting data management strategies and triggering actions such as data archiving or deletion based on the analyzed information, thereby maintaining high data quality while preserving processing productivity.
Solution Approach 2:
The patent replaces conventional mechanical database management approaches with an intelligent system that uses automated analysis and classification algorithms. Instead of relying on traditional database tools that lack sophisticated duplicate detection capabilities, the system substitutes a smart platform that automatically identifies duplicate and obsolete data through content analysis, metadata examination, and pattern recognition, thereby improving data quality without compromising processing productivity.
4Reliability
If all data is retained in cloud storage to ensure availability, then data accessibility is maintained, but storage costs increase perpetually
Solution Approach 1:
The patent segments data into different categories based on its value, accessibility requirements, and retention needs. The system divides the data set into active data, duplicate data, and obsolete data segments, placing each segment in the most appropriate storage location. This segmentation allows the organization to maintain high availability for critical data while reducing costs by archiving or eliminating redundant data segments.
Solution Approach 2:
The patent applies local quality by providing different storage characteristics to different data elements based on their specific requirements. Critical, frequently accessed data receives high-availability cloud storage with rapid access, while duplicate and obsolete data are placed in lower-cost archival storage or eliminated entirely. This differentiated approach ensures data availability where needed while minimizing overall storage costs.
Data Source
AI summary
The disclosed embodiments included a system, apparatus, method, and computer program product for optimizing the placement of data utilizing cloud-based IT services. The apparatus comprises a processor that executes computer-readable program code embodied on a computer program product. By executing that computer-readable program code, the processor extracts content from data and determines the context in which that data was generated, modified, and/or accessed. The processor also classifies the data based on its content and context, determines the cost of storing the data at each a plurality of locations, and specifies which of those locations the data is to be stored based on the classification of that data and the cost of storing that data at each of the plurality of locations.


