Flexible Data Validation for Security Analytics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Computer security systems face challenges in integrating and correlating data from diverse and arbitrarily structured sources, as existing solutions are ill-suited to handle emerging data sets and require significant intermediate processing, leading to scalability issues and data consistency problems.
Innovation Solution
A data fusion environment with a scalable architecture that enables arbitrary structuring of data, using a system that ingests data from various sources, applies transformations, and extracts features without imposing structure requirements, supporting Turing complete analytics and extensible computational logic, allowing for simultaneous processing of multiple analytics instances.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If distributed computation is employed to manage large data sets, then processing capability is improved, but system complexity and specialization requirements increase
Solution Approach 1:
The patent implements a universal data processing framework that can handle multiple types of security data (network traffic, logs, machine scanning results) through a single distributed computation system. The framework uses standardized data structures and processing pipelines that work across different data sources, eliminating the need for separate specialized systems for each data type while maintaining high processing capability.
2Adaptability or versatility
If generic frameworks are used for distributed computation, then flexibility is improved, but intermediate processing requirements increase
Solution Approach 1:
The patent applies preliminary action by pre-defining data schemas, validation rules, and processing pipelines before data ingestion. The system establishes data contracts and processing templates in advance, which automatically guide the distributed computation process. This eliminates the need for extensive intermediate processing and adaptation during runtime, as the framework is already configured to handle the specific security data types.
3Adaptability or versatility
If data is received from various sources with different structuring, then data diversity is improved, but correlation difficulty increases
Solution Approach 1:
The patent implements homogeneity by transforming diverse data from different sources into a unified internal representation. The system uses standardized data schemas, common field names, and consistent data types across all input sources. This homogeneous internal structure enables efficient correlation and analysis while still accepting heterogeneous external data inputs, resolving the contradiction between data diversity and correlation difficulty.
Data Source
AI summary
A method and apparatus for extracting and displaying a feature data set is provided. A method comprises: retrieving a digitally stored first data set from a first digital data storage source; selecting a first data set type of a plurality of data set types for the first data set based at least in part on the first source, and creating and storing an association of the first data set type to the first data set; selecting a first validation process from among a plurality of validation processes based at least in part on the first data set type; executing program instructions corresponding to the first validation process using at least a portion of the first data set to determine if the first data set is valid; in response to determining that the first data set is valid, assigning a validator instruction set to the first data set; assigning at least a portion of the first data set to a first analytics instruction set of a plurality of analytics instruction sets based on the first type and the validator instruction set; causing execution of the first analytics instruction set using at least a portion of the first data set to extract and store a feature data set representing features of the first data set; in response to a query, causing a feature represented in the feature data set to be displayed on a computer display device using a graphical user interface.


