Distributed Graph Query Processing for Scalable Business Process Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Business Process Mining systems are limited by their inflexibility in data analysis, scalability issues, and reliance on predefined models, which leads to distorted performance monitoring and incorrect actions due to uncontrolled variables and unpredictable events, and struggle with real-time data updates and large datasets.

Innovation Solution

A method and system for processing queries on a graph dataset stored across multiple nodes, dividing queries into atoms, calculating execution costs, determining optimal execution paths, and executing these atoms in parallel across nodes to produce a result set, allowing for flexible and scalable analysis without relying on predefined models or data partitioning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If data is stored in a centralized database, then data can be easily accessed and queried, but the system becomes less scalable and real-time updates become difficult

Engineering Contradiction:
Improvedata accessVSAvoidscalability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent divides the centralized database into multiple distributed nodes, each storing a portion of the graph data. This segmentation allows the system to scale by adding more nodes while maintaining data accessibility through distributed query execution across the network.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from a single centralized database to a multi-dimensional distributed architecture where data is stored across multiple nodes. This dimensional expansion enables both scalability and real-time updates by allowing parallel data access and modification operations across different nodes.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of manufacture

If predefined data models are used, then data structure is standardized and easy to manage, but the system loses flexibility in analyzing new or uncontrolled variables

Engineering Contradiction:
Improvedata model standardizationVSAvoiddata analysis flexibility
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic graph data model where the schema is not fixed but can evolve to accommodate new variables and relationships. The graph structure allows flexible addition of nodes and edges representing new data types without requiring predefined models, enabling analysis of uncontrolled variables while maintaining ease of data entry through standardized graph operations.

Inventive Principle:
Principle #15Dynamics

3Device complexity

If queries are executed sequentially on large datasets, then query processing is simple, but analysis time increases significantly

Engineering Contradiction:
Improvequery processing simplicityVSAvoidanalysis time
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The patent divides complex queries into smaller sub-queries that can be executed in parallel across different distributed nodes. Each node processes a portion of the query independently, and results are aggregated to produce the final answer, significantly reducing analysis time while maintaining manageable query processing complexity through standardized query decomposition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system enables continuous query execution across multiple nodes simultaneously rather than sequential processing. Parallel query execution allows multiple analysis operations to proceed concurrently, maintaining continuous progress on large dataset analysis without the time loss associated with sequential processing.

Inventive Principle:
Principle #20Continuity of useful action

4Speed

If real-time data updates are implemented, then system responsiveness improves, but data consistency and integrity become more difficult to maintain

Engineering Contradiction:
Improveupdate speedVSAvoiddata consistency
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The patent implements preliminary indexing and caching mechanisms that prepare data structures in advance for rapid updates. When updates occur, the system uses pre-established indexes to quickly locate and modify data without requiring full dataset scans, maintaining both real-time update speed and data consistency through optimized update protocols.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP2831767B1Method and system for processing data queries
Publication Date: 2019.12.25 BRITISH TELECOM PLC
  • EP2831767B1 patent drawingFigure 1
  • EP2831767B1 patent drawingFigure 2
  • EP2831767B1 patent drawingFigure 3

AI summary

The invention relates to a method and system that provide a high performance and extremely scalable triple store within the Resource Description Framework (or alternative data models), with optimized query execution. An embodiment of the invention provides a data storage and analysis system to support scalable monitoring and analysis of business processes along multiple configurable perspectives and levels of granularity. This embodiment analyses data from processes that have been already executed and from ongoing processes, as a continuous flow of information. From the point of view of the data analysis, this embodiment provides to define and monitor processes which is based on no initial domain knowledge about the process and such that the process will be built only from the incoming flow of information. This approach allows data to dynamically define the process model. Another embodiment of the invention provides a grid infrastructure that allows storage of data across many grid nodes and distribution of the workload, avoiding the bottleneck represented by constantly querying a database. This embodiment permits efficient analysis of a large amount of data.