Remote Metadata Scanning for Restricted Content Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer forensic methods for identifying restricted content on a network are time-consuming, impractical, and inefficient, as they require physical access to hard drives and often produce numerous false positives, making it difficult to scan large networks effectively.
Innovation Solution
A method and system for remotely accessing computer systems to retrieve metadata, comparing it to a database of known restricted content, and flagging matches, which allows for efficient identification of restricted content without physically handling hard drives or scanning entire file systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If full forensic scanning of hard drives is performed, then restricted content can be identified, but the process becomes extremely time-consuming and impractical for large networks
Solution Approach 1:
The patent extracts only the essential identifying features (metadata, hash values, file names, file types) from files instead of scanning entire file contents. This extraction approach maintains detection accuracy for restricted content while dramatically reducing scanning time and resource requirements, making network-wide scanning practical.
Solution Approach 2:
The patent segments the file system scanning process into discrete manageable units by using file metadata and hash values as identification markers. Instead of processing continuous large file contents, the system scans segmented metadata structures, enabling efficient parallel processing and reducing overall scanning time while maintaining detection capability.
2Measurement precision
If physical access to hard drives is required for scanning, then comprehensive content analysis can be performed, but it becomes difficult to scan remote systems in networks that cannot be easily accessed
Solution Approach 1:
The patent introduces an intermediary approach by scanning file system metadata and using hash value comparison instead of requiring direct physical access to file contents. The metadata acts as an intermediary that can be accessed remotely through standard file system interfaces, enabling forensic analysis without physical drive removal while maintaining detection accuracy through hash matching.
3Measurement precision
If algorithms analyze stored data for restricted content, then content can be discovered, but thousands of false positives are produced requiring extensive verification
Solution Approach 1:
The patent uses hash values (digital fingerprints) as copies or representations of file contents for comparison purposes. Instead of analyzing actual file contents which lead to false positives, the system compares hash value copies against known restricted content hashes, maintaining detection capability while eliminating false positives caused by misinterpretation of actual file data.
Data Source
AI summary
Embodiments of the present invention provide a method and apparatus for remotely accessing a computer system or network to identify storage devices and to retrieve metadata from the storage devices that are respectively unique to files stored in the storage devices. An agent is executed by the computer system and receives scanning instructions from a server. A scanning tool compares the metadata retrieved from the computer system or network to a database or list of known metadata of known restricted content. Metadata retrieved from the computer system or network that matches metadata from the database or list of known restricted content is flagged and the file associated with the matching metadata is flagged and reported as potentially storing restricted content. During the scanning, restricted content itself is not scanned, not copied, not transferred and not stored.


