Distributed Representations for ML Feature Engineering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning (ML) approaches for computer security and IT infrastructure require manual feature engineering, large amounts of labeled training data, and train models from scratch, leading to long training times and reduced performance.

Innovation Solution

The techniques employ self-supervised learning to generate distributed representations of computing processes and events using a neural network, eliminating the need for manual feature engineering and enabling effective training with limited labeled data, while allowing information transfer across ML models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual feature engineering is used to identify features for ML models, then model performance can be optimized for specific tasks, but domain expertise is required and the process becomes time-consuming and complex

Engineering Contradiction:
Improvefeature selection accuracyVSAvoidfeature engineering complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically generating distributed representations of processes and events without requiring manual feature engineering. The neural network learns features autonomously from raw data, eliminating the need for domain expert intervention while maintaining high performance.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies preliminary action by pre-training a neural network to generate distributed representations that capture essential features of processes and events. These pre-learned representations can be directly used by downstream ML models, eliminating the need for subsequent manual feature engineering.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If large amounts of labeled training data are used to train ML models, then model performance improves, but data acquisition becomes difficult and time-consuming

Engineering Contradiction:
Improvemodel performanceVSAvoiddata acquisition time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system employs self-supervised learning where the neural network learns from unlabeled process and event data autonomously. The model generates its own training signals by predicting future events or reconstructing input sequences, eliminating the need for manual labeling while achieving high performance.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary pre-training on large amounts of unlabeled data to learn distributed representations. This preliminary action allows the model to capture underlying patterns before fine-tuning on smaller labeled datasets, significantly reducing data acquisition time.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If ML models are trained from scratch for each problem/task, then models can be optimized for specific tasks, but training time increases and training data demands increase

Engineering Contradiction:
Improvetask-specific model performanceVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training a neural network on general process and event data to learn transferable distributed representations. These pre-learned features can be reused across multiple tasks, eliminating the need to train from scratch and significantly reducing training time and data requirements for each new task.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates universal distributed representations that can be applied across multiple different ML tasks in computer security and IT infrastructure. The same pre-trained neural network serves as a feature extractor for various problems, achieving task-specific performance without retraining from scratch.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11928466B2Distributed representations of computing processes and events
Publication Date: 2024.03.12 VMWARE INC
  • US11928466B2 patent drawing
  • US11928466B2 patent drawing
  • US11928466B2 patent drawing

AI summary

Techniques for generating distributed representations of computing processes and events are provided. According to one set of embodiments, a computer system can receive occurrence data pertaining to a plurality of computing processes and a plurality of events associated with the plurality of computing processes. The computer system can then generate, based on the occurrence data, (1) a set of distributed process representations that includes, for each computing process, a representation that encodes a sequence of events associated with the computing process in the occurrence data, and (2) a set of distributed event representations that includes, for each event, a representation that encodes one or more event properties associated with the event and one or more events that occur within a window of the event in the occurrence data.