Distributed Worker Pool for Cloud File Scanning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional Cloud Access Security Broker (CASB) systems face challenges in efficiently scanning large volumes of files in cloud applications, leading to latency and poor user experience due to increased loads from cloud deployments.
Innovation Solution
The proposed CASB system employs a distributed worker pool approach with a message broker and various types of workers operating in parallel to efficiently process content through queues, allowing for distributed file crawling, policy enforcement, and integration with cloud-based security systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional CASB systems scan large volumes of files in cloud applications, then security monitoring and policy enforcement are improved, but system latency increases and user experience deteriorates
Solution Approach 1:
The patent divides the CASB system into a distributed worker pool where multiple workers independently scan files across cloud applications. Each worker operates as an independent unit that can process files in parallel, segmenting the overall scanning task to reduce latency while maintaining comprehensive security monitoring.
Solution Approach 2:
The patent transitions from a single-point scanning architecture to a distributed multi-dimensional worker pool architecture. Workers are assigned to different cloud applications and file sets across multiple dimensions (different cloud providers, different file types, different user accounts), enabling parallel processing that reduces scan latency while maintaining thorough security inspection.
2Adaptability or versatility
If cloud deployments increase the load on traditional CASB systems, then coverage of cloud applications is improved, but system performance deteriorates
Solution Approach 1:
The patent implements a universal worker pool architecture where workers can be dynamically assigned to scan different types of cloud applications (Office 365, Dropbox, Box, Google Drive, Salesforce, etc.). Each worker is designed with multi-functionality to handle various cloud service types, allowing the system to expand coverage without proportionally increasing performance degradation.
Solution Approach 2:
The patent employs dynamic worker assignment and scaling mechanisms where the controller can dynamically allocate workers to different cloud applications based on current load and requirements. The system can adapt its worker distribution in real-time to maintain performance while expanding cloud application coverage.
3Productivity
If distributed worker pool processes files in parallel through queues, then scanning efficiency is improved, but system complexity increases
Solution Approach 1:
The patent introduces a message broker as an intermediary component that mediates between the controller and workers, and between workers and cloud applications. The message broker handles job assignment, result collection, and coordination, simplifying the overall system architecture by centralizing communication logic while enabling parallel processing efficiency.
Solution Approach 2:
The patent implements self-service mechanisms where workers autonomously retrieve files from cloud applications, process them according to policies, and report results without constant intervention from the controller. Workers self-manage their execution and the system automatically handles task distribution and result aggregation, reducing the operational complexity of managing parallel processing.
Data Source
AI summary
Systems and methods for operating a scanning system, implemented either on-premises or in a cloud-based service, for crawling and analyzing files stored in one or more data repositories. The scanning system includes a controller, a message broker, and a distributed pool of workers, and, in one embodiment, a method includes receiving, by the controller, policy and configuration data associated with at least one organization; generating, by the controller, job assignments corresponding to files to be analyzed according to the received policy and configuration data; publishing the job assignments to the message broker for parallel distribution among the distributed pool of workers; retrieving and scanning, by at least one worker, the files from the one or more data repositories in accordance with the assigned job; and executing, where required by the policy and configuration data, at least one policy-based action on the files within the data repositories.


