Global Compiled Code Store for Query Engine Cold Start

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data processing systems face performance bottlenecks due to high compilation times for queries, especially during the 'cold start' scenario where query engines have not seen many queries, leading to increased query performance times.

Innovation Solution

Implementing a global compiled code store that shares compiled code across query engines, allowing reuse of generated code and minimizing the 'cold start' effect by pre-populating local caches with likely-to-be-used code from the global store.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If code generation is performed at run-time to optimize query execution, then query execution performance is improved with smaller instruction footprint and fewer branches, but compilation time increases leading to slower query performance during cold start

Engineering Contradiction:
Improvequery execution speedVSAvoidcompilation time
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The system performs code generation and compilation in advance by maintaining a code cache that stores pre-compiled query execution code. When a query is received, the system checks the code cache first and retrieves pre-compiled code if available, avoiding run-time compilation overhead. This preliminary action resolves the contradiction by separating the compilation phase from the execution phase, allowing fast execution without cold start penalties.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If multiple query engines independently generate and compile code, then each engine can optimize its own queries, but redundant compilation occurs across engines leading to wasted resources and time

Engineering Contradiction:
Improvequery processing throughputVSAvoidcompilation resource consumption
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The system merges the code caches of multiple query engines into a shared code cache that is accessible by all engines. When one query engine generates and compiles code, it stores the compiled result in the shared cache where other engines can retrieve and reuse it. This merging eliminates redundant compilation across engines, reducing resource consumption while maintaining high query processing throughput.

Inventive Principle:
Principle #5Merging (Combining)

3Loss of time

If a shared code cache is implemented across query engines, then compilation time is reduced through code reuse, but system complexity increases due to cache management requirements

Engineering Contradiction:
Improvecompilation timeVSAvoidcache management complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The shared code cache implements self-service mechanisms where query engines automatically check the cache for existing compiled code, retrieve relevant code segments, and store newly compiled code without requiring external coordination. The cache uses automatic invalidation and versioning to handle updates. This self-service approach reduces cache management complexity while still achieving fast code retrieval and compilation time reduction.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11853301B1Sharing compiled code for executing queries across query engines
Publication Date: 2023.12.26 AMAZON TECH INC
  • US11853301B1 patent drawing
  • US11853301B1 patent drawing
  • US11853301B1 patent drawing

AI summary

Compiled portions of code generated to perform a query plan at a query engine may be shared with other query engines. A data store, separate from the query engines, may store compiled portions of query code generated for different queries. If a query engine does not have a locally stored compiled portion of query code, then the separate data store may be accessed in order to obtain a compiled portion of query code, allowing reuse of compiled query code across different queries engines for queries directed to different databases.