Distributed Representations for ML Feature Engineering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning (ML) approaches for computer security and IT infrastructure require manual feature engineering, large amounts of labeled training data, and train models from scratch, leading to long training times and reduced performance.
Innovation Solution
The techniques employ self-supervised learning to generate distributed representations of computing processes and events using a neural network, eliminating the need for manual feature engineering and enabling effective training with limited labeled data, while allowing information transfer across ML models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual feature engineering is used to identify features for ML models, then model performance can be optimized for specific tasks, but domain expertise is required and the process becomes time-consuming and complex
Solution Approach 1:
The system performs self-service by automatically generating distributed representations of processes and events without requiring manual feature engineering. The neural network learns features autonomously from raw data, eliminating the need for domain expert intervention while maintaining high performance.
Solution Approach 2:
The patent applies preliminary action by pre-training a neural network to generate distributed representations that capture essential features of processes and events. These pre-learned representations can be directly used by downstream ML models, eliminating the need for subsequent manual feature engineering.
2Measurement precision
If large amounts of labeled training data are used to train ML models, then model performance improves, but data acquisition becomes difficult and time-consuming
Solution Approach 1:
The system employs self-supervised learning where the neural network learns from unlabeled process and event data autonomously. The model generates its own training signals by predicting future events or reconstructing input sequences, eliminating the need for manual labeling while achieving high performance.
Solution Approach 2:
The patent performs preliminary pre-training on large amounts of unlabeled data to learn distributed representations. This preliminary action allows the model to capture underlying patterns before fine-tuning on smaller labeled datasets, significantly reducing data acquisition time.
3Measurement precision
If ML models are trained from scratch for each problem/task, then models can be optimized for specific tasks, but training time increases and training data demands increase
Solution Approach 1:
The patent applies preliminary action by pre-training a neural network on general process and event data to learn transferable distributed representations. These pre-learned features can be reused across multiple tasks, eliminating the need to train from scratch and significantly reducing training time and data requirements for each new task.
Solution Approach 2:
The system creates universal distributed representations that can be applied across multiple different ML tasks in computer security and IT infrastructure. The same pre-trained neural network serves as a feature extractor for various problems, achieving task-specific performance without retraining from scratch.
Data Source
AI summary
Techniques for generating distributed representations of computing processes and events are provided. According to one set of embodiments, a computer system can receive occurrence data pertaining to a plurality of computing processes and a plurality of events associated with the plurality of computing processes. The computer system can then generate, based on the occurrence data, (1) a set of distributed process representations that includes, for each computing process, a representation that encodes a sequence of events associated with the computing process in the occurrence data, and (2) a set of distributed event representations that includes, for each event, a representation that encodes one or more event properties associated with the event and one or more events that occur within a window of the event in the occurrence data.


