Cloud Storage Repository Automating Data Ingestion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current cloud storage solutions lack efficient tools for managing large-scale data ingestion and providing human interfaces for knowledge workers, particularly in enterprises, where millions or billions of items require efficient querying and maintenance, and there is a need for automated tools to handle access control and data synchronization across multiple users.

Innovation Solution

A method and system that link cloud storage repositories to user directory services, providing a graphical user interface for knowledge workers to manage access control, assemble data maps, synchronize changes, and centrally manage delta comparisons, using data source connectors to interface with cloud storage networks and handle policy-based ingestion logic, while ensuring secure data transfer without exposing cloud storage API security.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If cloud storage repositories are linked to user directory services and data source connectors are used to traverse and capture data from third-party systems, then access control management and data synchronization are automated, but the system complexity and implementation difficulty increase significantly

Engineering Contradiction:
Improveautomated access control managementVSAvoidsystem implementation complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent introduces a cloud storage repository as an intermediary layer between user directory services and third-party data sources. This repository acts as a central hub that manages access control information and coordinates data synchronization, automating the management process while abstracting the complexity from individual components.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The cloud storage repository is designed to perform multiple functions: storing access control information, resolving user identities, synchronizing data changes, and interfacing with both directory services and third-party systems. This multi-functionality consolidates automation capabilities into a single system that handles diverse tasks.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If automated tools are implemented for managing access control and data synchronization across multiple users, then management efficiency improves, but the computational resources and processing time required increase

Engineering Contradiction:
Improvemanagement efficiencyVSAvoidcomputational resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary actions by pre-resolving user identities and pre-establishing access control mappings in the cloud storage repository before data synchronization operations begin. This preparation work is done once and reused across multiple synchronization cycles, reducing computational overhead during actual data management operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying mechanisms where access control information and user identity mappings are replicated and cached in the cloud storage repository. Instead of querying directory services repeatedly during synchronization, the system uses pre-copied information, significantly reducing computational resource consumption during data management operations.

Inventive Principle:
Principle #26Copying

3Stability of the object's composition

If data source connectors traverse and capture data from third-party systems with centralized access control resolution, then data consistency across users is improved, but the time required for identity resolution and data submission increases

Engineering Contradiction:
Improvedata consistencyVSAvoididentity resolution time
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

User identities and access control information are resolved in advance and stored in the cloud storage repository before data capture operations begin. This preliminary resolution ensures data consistency across users while minimizing the time required during actual synchronization, as identity lookup is replaced by reference to pre-resolved information.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11575674B2Methods and systems relating to network based storage
Publication Date: 2023.02.07 COHESITY INC
  • US11575674B2 patent drawing
  • US11575674B2 patent drawing
  • US11575674B2 patent drawing

AI summary

Cloud storage provides for accessible interfaces, near-instant elasticity and scalability, multi-tenancy, and metered resources within a framework of distributed resources acting to provide highly fault tolerant solutions with high data durability. However, cloud storage also has drawbacks and limitations with information uploading and how information is subsequently accessed. To date the lack of automated tools for managing tens, hundreds and thousands of users and/or documents within enterprises and organizations means that for most migrating is a massive undertaking. Accordingly, knowledge workers require a human interface to the data ingested from third-party systems that manages the data in original folder contexts/locations for each knowledge worker within the interfaces. It would be further beneficial for knowledge workers to have tools for incremental ingestion of changes from data sources to a cloud storage repository as well as determining/centralizing cloud storage repository ingestion from the data sources.