Data Clustering System for Risky Trading Investigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Analysts face challenges in efficiently selecting and prioritizing relevant data items within large electronic collections for investigations, such as risky trading, due to the difficulty in processing and analyzing vast amounts of data, and existing systems require manual repetition of searches, leading to time-consuming and resource-intensive investigations.

Innovation Solution

A data analysis system that automatically generates memory-efficient clustered data structures, analyzes them, and provides an interactive user interface for efficient evaluation, allowing analysts to dynamically re-group and filter data clusters based on automated scoring and prioritization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If analysts manually search and process large electronic collections of data items, then they can identify relevant data for investigations, but the processing time and resource consumption increase significantly

Engineering Contradiction:
Improverelevance identification accuracyVSAvoidinvestigation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the large electronic collection of data items into multiple clusters, where each cluster contains a subset of related data items. This segmentation allows analysts to navigate and evaluate data in smaller, more manageable groups rather than processing the entire collection at once, thereby reducing the time and resources required while maintaining the ability to identify relevant data through systematic cluster evaluation.

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If all data items are stored in electronic memory for analysis, then complete data availability is achieved, but memory consumption becomes excessive

Engineering Contradiction:
Improvedata availabilityVSAvoidmemory consumption
Core Design Contradiction:
Quantity of substanceVSWeight of stationary object

Solution Approach 1:

The patent divides the large set of data items into multiple smaller clusters that are stored and processed separately in electronic memory. Each cluster contains a manageable subset of data items that can be loaded into memory for analysis. This segmentation enables the system to work with complete data availability on a cluster-by-cluster basis rather than requiring all data to be simultaneously present in memory, thereby significantly reducing peak memory consumption while maintaining data accessibility.

Inventive Principle:
Principle #1Segmentation

3Ease of operation

If clusters of data items are created for organization, then data navigation becomes easier, but the number of clusters to evaluate increases the complexity of analysis

Engineering Contradiction:
Improvedata navigationVSAvoidcluster evaluation complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent extracts and applies multiple evaluation criteria to each data cluster to generate scores that indicate the likelihood of risky activity. By taking out the complex evaluation task and automating it through systematic scoring based on predefined criteria (such as trading patterns, data item relationships, and risk indicators), the system reduces the manual complexity of evaluating numerous clusters. Analysts can then prioritize clusters based on these automated scores rather than manually assessing each cluster, thereby maintaining ease of navigation while reducing evaluation complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

4Adaptability or versatility

If manual searching and analysis methods are used, then flexibility in investigation approach is maintained, but productivity and efficiency decrease

Engineering Contradiction:
Improveinvestigation flexibilityVSAvoiddata processing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent implements a dynamic system where data clusters are automatically evaluated using multiple criteria that can be adjusted and refined. The scoring mechanism allows for dynamic prioritization of clusters based on risk indicators and evaluation results. This dynamic approach maintains flexibility in investigation strategies while significantly improving productivity through automated cluster evaluation and prioritization, enabling analysts to quickly identify high-risk clusters without sacrificing investigative adaptability.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11928733B2Systems and user interfaces for holistic, data-driven investigation of bad actor behavior based on clustering and scoring of related data
Publication Date: 2024.03.12 PALANTIR TECHNOLOGIES INC
  • US11928733B2 patent drawing
  • US11928733B2 patent drawing
  • US11928733B2 patent drawing

AI summary

Embodiments of the present disclosure relate to a data analysis system that may automatically generate memory-efficient clustered data structures, automatically analyze those clustered data structures, automatically tag and group those clustered data structures, and provide results of the automated analysis and grouping in an optimized way to an analyst. The automated analysis of the clustered data structures (also referred to herein as data clusters) may include an automated application of various criteria, rules, indicators, or scenarios so as to generate scores, reports, alerts, or conclusions that the analyst may quickly and efficiently use to evaluate the groups of data clusters. In particular, the groups of data clusters may be dynamically re-grouped and/or filtered in an interactive user interface so as to enable an analyst to quickly navigate among information associated with various groups of data clusters and efficiently evaluate those data clusters in the context of, for example, a risky trading investigation.