File Scanning Remediation via Owner Data Retrieval
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Scanning large data systems for target files meeting specific criteria is cumbersome and time-consuming, often resulting in inefficient and inaccurate results, as conventional methods struggle to process millions or billions of files efficiently and fail to accurately identify and communicate with the appropriate file owners for remediation actions.
Innovation Solution
A computer-implemented method and apparatus for improved target file scanning and processing, which includes scanning file repositories based on scan criteria, retrieving file owner data from a personnel database using file owner identifiers, and providing scan alerts to the appropriate users for remediation actions, enabling efficient and accurate identification and processing of target files, even in systems with vast numbers of files.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional file scanning methods are used on large data systems, then the scan can be performed, but the scanning process becomes prohibitively time-consuming and computationally impossible for systems with hundreds of millions or billions of files
Solution Approach 1:
The patent segments the large-scale file scanning task into multiple phases: first identifying candidate files using efficient indexing and metadata filtering, then performing detailed classification only on those candidates. This segmentation allows the system to process billions of files by avoiding exhaustive analysis of every file, thereby dramatically improving scanning productivity while reducing time loss.
Solution Approach 2:
The patent performs preliminary actions by pre-processing and indexing file metadata before the actual scanning operation. Candidate files are pre-identified based on basic criteria, and only these pre-selected files undergo detailed classification analysis. This preliminary action reduces the computational burden during the main scanning process, enabling faster processing of large file systems.
2Measurement precision
If conventional scanning methods are used, then all files can be processed, but the accuracy of identifying target files and their owners deteriorates due to computational limitations
Solution Approach 1:
The patent extracts and prioritizes key identifying features from files during the candidate identification phase, such as file type, location, and basic metadata. By extracting these critical attributes early, the system maintains high accuracy in identifying target files without needing to analyze every file in complete detail, thus preserving measurement precision while managing computational resources effectively.
Solution Approach 2:
The patent implements feedback mechanisms where scan results and classification outcomes are used to refine and update the scanning process. The system learns from identified target files and adjusts its candidate selection criteria, improving the accuracy and reliability of both target file identification and owner attribution over time through iterative optimization.
3Measurement precision
If manual file scanning and owner identification is performed, then accurate results can be achieved, but the process becomes cumbersome and requires significant manual effort
Solution Approach 1:
The patent implements self-service automation where the system automatically performs file scanning, candidate identification, classification, and owner retrieval without requiring manual intervention. The automated workflow maintains high accuracy through structured processes while eliminating the cumbersomeness of manual operation, allowing the system to serve itself in completing the entire scanning and remediation workflow.
Solution Approach 2:
The patent creates a universal scanning system that handles multiple file types, classification criteria, and remediation actions through a single integrated platform. This multi-functional approach maintains accuracy by applying consistent processes across diverse scenarios while simplifying operation, as users interact with a unified system rather than multiple separate tools or manual procedures.
4Reliability
If exhaustive file scanning is performed, then no target files are missed, but computational resources are wasted processing files that do not require remediation
Solution Approach 1:
The patent segments the file population into distinct groups: candidate files that may require remediation and non-candidate files that do not. By using metadata filtering and indexing to create these segments, the system achieves reliable detection of all potential target files while avoiding wasteful processing of files that clearly do not meet remediation criteria, thus optimizing computational resource efficiency.
Solution Approach 2:
The patent applies partial action by performing detailed classification analysis only on candidate files rather than all files in the system. This approach ensures that no target files are missed among the candidate set while avoiding excessive computational expenditure on files that are clearly not candidates, achieving a balance between detection completeness and resource efficiency.
Data Source
AI summary
Embodiments of the present disclosure enable improved methodologies of scanning large file repositories and managing target files identified from such scanning efficiently and effectively. Embodiments of the present disclosure scan any number of file repositories of a data system to identify particular target files that satisfy scan criteria, and process the target files identified therefrom. The target files may be processed to identify file owner data and utilize the file owner data for any of a myriad of purposes, for example to provide scan alert(s) corresponding to the target files to such users. Any of a number of file remediation actions may be performed based on the scan results, for example by the users receiving scan alert(s) and/or automatically in the embodiments described.


