Distributed Job Manager Service for Scalable Data Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional data storage management systems face performance limitations in large or active deployments due to centralized architecture, leading to bottlenecks that slow down core data protection operations.
Innovation Solution
Implementing a distributed architecture where multiple components act as local job managers, each responsible for managing storage management jobs independently, reducing reliance on a centralized storage manager and enabling concurrent job execution and metadata management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a centralized storage manager is used to manage all storage management jobs, then system control and coordination are simplified, but system performance and scalability deteriorate due to bottlenecks in large or active deployments
Solution Approach 1:
The patent divides the centralized job manager into multiple distributed job managers deployed across different nodes in the storage system. Each distributed job manager handles a subset of storage management jobs independently, eliminating the single-point bottleneck and enabling parallel processing of backup, archive, and restore operations across the distributed storage network.
2Productivity
If multiple distributed components act as local job managers, then system scalability and concurrent job execution improve, but system complexity and coordination overhead increase
Solution Approach 1:
The patent combines the job manager functionality with existing distributed storage nodes, so that storage nodes simultaneously perform data storage and job management functions. This integration eliminates the need for separate centralized job management infrastructure and leverages the existing distributed architecture to handle both data and control plane workloads efficiently.
Solution Approach 2:
Each distributed job manager operates autonomously, making local decisions about job execution, resource allocation, and metadata management without requiring constant coordination with a central controller. This self-service capability reduces coordination overhead and enables faster local decision-making for time-sensitive storage operations.
3Reliability
If a centralized storage manager manages all jobs, then job coordination and metadata management are centralized, but the system experiences performance limits and bottlenecks in large deployments
Solution Approach 1:
The patent segments the metadata management responsibility across multiple distributed job managers, each maintaining local metadata about the jobs they manage. This distributed metadata approach parallelizes metadata operations and eliminates the single-point performance limit while maintaining consistency through coordinated metadata updates across the distributed system.
Data Source
AI summary
An improved system architecture for scaling deployments of data management-as-a-service (DMaaS) distributes “job manager” functionality so that a large-scale data storage management system may be deployed in a cloud computing environment as DMaaS without experiencing performance bottlenecks that slow down core data protection operations at scale. The disclosed solution creates an infrastructure where job manager features run on any number of distinct machines or compute resources that act as local job managers. In the illustrative DMaaS, any number of components that scale horizontally (“the distributed components”) carry job management responsibility for storage management operations such as backup jobs, auxiliary copy jobs, archive jobs, etc. These distributed components are configured to perform locally as job managers and to see each storage management job through from beginning to end without centralized control. These distributed components or local job managers also maintain, collect, and locally store job metadata.


