Application Discovery via Log Clustering in Hybrid Clouds

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing application discovery services struggle to accurately capture application components across complex and evolving cloud environments, particularly in hybrid multi-cloud setups where applications are spread across data centers, multiple clouds, and the edge.

Innovation Solution

An automated computer-implemented method for application discovery using log messages generated by event sources in a cloud infrastructure. The method constructs a data frame of probability distributions of event types, applies clustering techniques to form clusters corresponding to applications, and displays an interactive GUI to visualize and select clusters for further analysis and performance optimization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional application discovery services use workload naming conventions, tags, and agent-based methodologies to capture system configuration and performance data, then they can provide basic application identification, but they fail to accurately capture application components in isolated network environments and across hybrid multi-cloud setups

Engineering Contradiction:
Improveapplication discovery accuracyVSAvoidapplicability across cloud environments
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces flow log data as an intermediary medium to bridge the gap between isolated application components and the discovery service. Instead of relying on agents within isolated environments or direct access to application components, the system uses network flow logs as a mediator that captures communication patterns between components, enabling accurate discovery across hybrid multi-cloud environments without requiring direct access to isolated resources.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a universal application discovery approach that works across diverse cloud environments (on-premises, public cloud, private cloud, edge) by using a common methodology based on flow log analysis. This universal approach replaces environment-specific agent-based methods with a single technique that can discover applications regardless of their deployment location or network isolation status.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Quantity of substance

If application discovery services gather detailed information from multiple sources including VM inventory, configuration, and performance history, then they can provide comprehensive data, but they become limited and not applicable across the variety of different and complex cloud environments

Engineering Contradiction:
Improveinformation completenessVSAvoidcloud environment compatibility
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent extracts the essential discovery mechanism from the cumbersome multi-source data collection approach. Instead of gathering comprehensive data from VM inventory, configuration, and performance history across multiple sources, the system extracts only the necessary flow log data that contains the critical information about application component communications, eliminating the need for environment-specific data collection methodologies.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent inverts the traditional approach by not trying to access application components directly through agents and configuration data, but rather observing their behavior indirectly through network flow logs. This inversion allows the system to discover applications by analyzing their communication patterns rather than by directly querying their configuration, making it applicable across all cloud environments.

Inventive Principle:
Principle #13The other way round (Inversion)

3Reliability

If existing application discovery services use agent-based methodologies to capture running processes and network connections, then they can identify applications on accessible systems, but they cannot capture application components that are isolated in a network

Engineering Contradiction:
Improveapplication identification reliabilityVSAvoiddeployment complexity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent enables the application components to effectively self-identify through their own network communications. Instead of requiring external agents to probe and identify applications, the system analyzes the self-generated network flow logs that automatically record communication patterns, allowing isolated components to be discovered through their own operational behavior without requiring deployment of additional software on those components.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250130871A1Methods and systems for application discovery from log messages
Publication Date: 2025.04.24 VMWARE INC
  • US20250130871A1 patent drawing
  • US20250130871A1 patent drawing
  • US20250130871A1 patent drawing

AI summary

This disclosure is directed to automated computer-implemented methods for application discovery from log messages generated by event sources of applications executing in a cloud infrastructure. The methods are executed by an operations manager that constructs a data frame of probability distributions of event types of the log messages generated by the event sources in a time period. The operations manager executes clustering techniques that are used to form clusters of the probability distributions in the data frame, where each of the clusters corresponds to one of the applications. The operations manager displays the clusters of the probability distributions in a two-dimensional map of applications in a graphical user interface that enables a user to select one of the clusters in the map of applications that corresponds to one of the applications and launch clustering of probability distributions of the user-selected cluster to discover two or more instances of the application.