LLM Actor Attribution Cybersecurity Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Attributing cybersecurity events to specific actors is challenging due to the vast amount of data involved and the difficulty in mapping patterns across multiple analysts, especially when attackers operate on the Dark Web, making it hard to identify relationships and connections between different incidents.

Innovation Solution

A large language model (LLM) is trained on text data attributed with high confidence to specific actors, using a multiclass classification approach, pre-trained on data from analyst notes and Dark Web conversations, and fine-tuned for actor attribution tasks to predict actor involvement in cybersecurity incidents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional manual analysis methods are used to attribute cybersecurity events to actors, then analysts can identify patterns and connections, but the process is extremely time-consuming and resource-intensive due to the vast amount of data involved

Engineering Contradiction:
Improveactor attribution accuracyVSAvoidanalysis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces a large language model as an intermediary between the vast cybersecurity data and human analysts. The LLM is trained on attributed cybersecurity events and serves as a mediator that processes data, identifies patterns, and generates actor attributions, thereby reducing the time burden on analysts while maintaining attribution accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary action by pre-training the LLM on labeled cybersecurity data with known actor attributions. This preliminary training enables the model to learn patterns and relationships in advance, so that when actual attribution tasks are performed, the heavy analytical work has already been prepared, significantly reducing real-time analysis time

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If analysts manually map patterns across multiple cybersecurity incidents to identify actor relationships, then connections can be discovered, but the complexity of coordinating multiple analysts and data sources increases significantly

Engineering Contradiction:
Improvepattern connection completenessVSAvoidanalysis system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent merges multiple data sources, analyst notes, and Dark Web conversations into a unified training dataset for the LLM. By combining these diverse information sources into a single model training process, the system achieves comprehensive pattern connection without the organizational complexity of coordinating multiple analysts and separate analysis tools

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The LLM serves as a universal system that handles multiple functions: processing analyst notes, analyzing Dark Web conversations, identifying patterns across incidents, and generating attributions. This single multi-functional model replaces the need for multiple specialized tools and coordinated analyst efforts, reducing system complexity while maintaining information completeness

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If more resources are allocated to cybersecurity event analysis, then more thorough investigation is possible, but resource consumption increases significantly

Engineering Contradiction:
Improveinvestigation reliabilityVSAvoidcomputational resource usage
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent creates a computational copy of analyst knowledge and reasoning processes within the LLM. By training the model on existing attributed cases, it replicates expert analysis capabilities in software form, enabling thorough investigations without proportionally increasing human analyst resources or computational overhead

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250007926A1Large language models for actor attributions
Publication Date: 2025.01.02 CROWDSTRIKE
  • US20250007926A1 patent drawing
  • US20250007926A1 patent drawing
  • US20250007926A1 patent drawing

AI summary

Systems and methods of actor attribution utilizing a machine learning (ML) model, such as a large language model (LLM), are provided. The method includes generating a first ML model based on first data associated with a first cybersecurity incident of a plurality of cybersecurity incidents. The method includes training the first ML model based on actor attribution associated with the first cybersecurity incident to generate a second ML model. The method includes receiving second data that is associated with a second cybersecurity incident of the plurality of cybersecurity incidents. The method includes producing, by a processing device for the second ML model using the second data, an attribution of the second cybersecurity incident to an actor.