Automated Data Mining for Threat Assertion Corroboration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current cybersecurity systems, such as SIEM solutions, are inefficient in identifying the root cause of unknown threats and require significant manual analysis, leading to increased workload and processing loads on data archival systems, which hampers the ability to quickly address cybersecurity threats.

Innovation Solution

An automated data mining technique that uses a confidence schema to rank-order hypotheses based on the occurrence of indicators in historical data, reducing the number of data queries and enabling faster decision-making for security analysts by corroborating threat assertions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If automated techniques exhaustively test all observables for all hypotheses against historical data, then hypothesis validation completeness is improved, but processing load and computational expense increase

Engineering Contradiction:
Improvehypothesis validation completenessVSAvoidprocessing load
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the hypothesis validation process into multiple stages: initial filtering of hypotheses based on confidence scores, selective testing of observables for high-priority hypotheses, and progressive deepening of analysis only where needed. This segmentation allows the system to achieve sufficient validation completeness without exhaustively processing all hypotheses to the same depth, thereby reducing overall processing load.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements partial action by testing only a subset of observables for each hypothesis rather than all possible observables. The system determines the appropriate level of testing based on factors such as hypothesis priority, available evidence, and resource constraints, performing just enough validation to achieve confident decision-making without the excessive processing of complete exhaustive analysis.

Inventive Principle:
Principle #16Partial or excessive action

2Use of energy by moving object

If manual approaches are used to validate hypotheses by searching for indicators, then processing load on data archival systems is reduced, but analyst time and productivity decrease

Engineering Contradiction:
Improveprocessing load on data archival systemsVSAvoidanalyst time
Core Design Contradiction:
Use of energy by moving objectVSLoss of time

Solution Approach 1:

The patent implements self-service by enabling the system to automatically perform hypothesis validation and indicator searching without requiring manual analyst intervention for each step. The automated techniques include querying data archival systems, evaluating observables, and determining hypothesis validity independently, which reduces both processing load on archival systems and the time analysts spend on repetitive validation tasks.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent incorporates feedback mechanisms where the system learns from validation results and adjusts its hypothesis testing strategy. The system uses feedback from initial validation attempts to refine subsequent queries, focus on the most promising hypotheses, and optimize the balance between automated processing and manual review, thereby reducing both archival system load and analyst time requirements.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If the number of data queries against historical data is increased to improve hypothesis support, then measurement precision is improved, but processing time and system expense increase

Engineering Contradiction:
Improvehypothesis support accuracyVSAvoidinvestigation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-processing and pre-ranking hypotheses based on available confidence scores and initial evidence before conducting detailed validation queries. This preliminary sorting allows the system to identify high-priority hypotheses that warrant deeper investigation with additional data queries, while lower-priority hypotheses can be resolved with minimal querying, thereby optimizing the balance between measurement precision and processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent dynamically adjusts the parameters of data queries based on the validation stage and hypothesis priority. The system modifies query depth, scope, and specificity according to the confidence level already established and the remaining uncertainty to be resolved, ensuring that additional queries provide maximum information gain per unit of processing time spent.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10686830B2Corroborating threat assertions by consolidating security and threat intelligence with kinetics data
Publication Date: 2020.06.16 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10686830B2 patent drawing
  • US10686830B2 patent drawing
  • US10686830B2 patent drawing

AI summary

A cognitive security analytics platform is enhanced by providing a computationally- and storage-efficient data mining technique to improve the confidence and support for one or more hypotheses presented to a security analyst. The approach herein enables the security analyst to more readily validate a hypothesis and thereby corroborate threat assertions to identify the true causes of a security offense or alert. The data mining technique is entirely automated but involves an efficient search strategy that significantly reduces the number of data queries to be made against a data store of historical data. To this end, the algorithm makes use of maliciousness information attached to each hypothesis, and it uses a confidence schema to sequentially test indicators of a given hypothesis to generate a rank-ordered (by confidence) list of hypotheses to be presented for analysis and response by the security analyst.