Decoupled Query Processing for Multi-Node Analytics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing query processing and data analytics systems face inefficiencies when analyzing large datasets across different architectures, requiring multiple technologies for online and offline data analysis, which leads to computational expense and time-consuming processes.

Innovation Solution

The system decouples data movement and parallel computation, allowing for high-performance query processing and data analytics across diverse scales by re-keying and redistributing organized data, optimizing searches in various computing environments, and enabling complex parallel computations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple technologies are used to analyze data across different architectures (single-core, multi-core, multi-node), then the system can achieve broader data analysis coverage, but the computational expense and time consumption increase significantly

Engineering Contradiction:
Improvedata analysis coverageVSAvoidquery processing speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent creates a universal query processing framework that can handle single-core, multi-core, and multi-node architectures through a single technology. The system uses a unified data model and execution engine that automatically adapts to different architectural scales, eliminating the need for multiple specialized technologies while maintaining broad data analysis coverage across all architectures.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent segments the query processing system into modular components (data model layer, execution engine layer, and optimization layer) that can independently operate across different architectural scales. This segmentation allows the system to maintain versatility across architectures while improving productivity by enabling parallel processing at each segment without requiring multiple complete technology stacks.

Inventive Principle:
Principle #1Segmentation

2Productivity

If data is redistributed across multiple cores or nodes, then parallel computation capability improves, but network traffic and data movement overhead increase

Engineering Contradiction:
Improveparallel computation capabilityVSAvoidnetwork traffic overhead
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent applies preliminary action by performing data re-keying and local aggregation operations before redistribution across cores or nodes. This preprocessing step reduces the volume of data that needs to be moved across the network, thereby improving parallel computation capability while minimizing network traffic overhead. The system prepares data in an optimized format that facilitates efficient parallel processing without excessive data movement.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements local quality by allowing each core or node to perform computations on locally-resident data partitions without requiring all data to be redistributed. The system optimizes queries by keeping frequently accessed data locally and only redistributing necessary portions, thereby maintaining high parallel computation capability while reducing network traffic and energy loss from data movement.

Inventive Principle:
Principle #3Local quality

3Ease of operation

If raw data is redistributed between cores, then data availability for computation improves, but the complexity of data organization and management increases

Engineering Contradiction:
Improvedata accessibilityVSAvoiddata organization complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary layer (the unified data model and execution engine) that manages data organization and redistribution automatically. This intermediary handles the complexity of data organization by providing standardized interfaces and automated data management routines, thereby improving data accessibility across cores while hiding the underlying organizational complexity from users and applications.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the organizational parameters of data by transforming raw data into a standardized format with consistent schemas and metadata structures. This parameter change simplifies data organization by imposing uniform structures on distributed data, making it easier to access and manage across multiple cores while reducing the complexity associated with heterogeneous data formats and organization schemes.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10896178B2High performance query processing and data analytics
Publication Date: 2021.01.19 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10896178B2 patent drawing
  • US10896178B2 patent drawing
  • US10896178B2 patent drawing

AI summary

High performance query processing and data analytics can be performed across architecturally diverse scales, such as single core, multi-core and/or multi-nodes. The high performance query processing and data analytics can include a separation of query computation, keying data, and data movement and parallel computation, thereby enhancing the capabilities of the query processing and data analytics, while allowing the specification of complex forms of data parallel computation that may execute across real-time and offline. The decoupling of data movement and parallel computation, as described herein can improve query processing and data analytics speed, can provide for the optimization of searches in a plurality of computing environments, and can provide the ability to search through a larger space of execution plans.