Database Defragmentation Service Micro-Partition Consolidation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional relational database management systems require extensive computing and storage resources, are costly, and have limited scalability, making them inefficient for large data storage and management, especially in disaster-prone environments.

Innovation Solution

A network-based database system with compute service managers and execution platforms that utilize micro-partitioning and metadata management to optimize query processing, allowing for efficient data storage and retrieval across multiple geographic regions, and includes a defragmentation service to consolidate and manage data effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If traditional relational database management systems are used for data storage and management, then data can be stored and accessed, but extensive computing and storage resources are required, making the system costly and inefficient for large data storage

Engineering Contradiction:
Improvedata storage capacityVSAvoidcomputing resources
Core Design Contradiction:
Quantity of substanceVSUse of energy by moving object

Solution Approach 1:

The patent divides database files into multiple micro-partitions distributed across different storage devices. Each micro-partition contains a portion of the database data, allowing parallel processing and distributed storage. This segmentation enables the system to handle large volumes of data without requiring centralized computing resources, thus reducing the computing resource overhead while maintaining large storage capacity.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If traditional database systems are deployed with on-premises infrastructure, then data can be managed in-house, but significant capital investment in hardware and infrastructure is required

Engineering Contradiction:
Improvein-house database managementVSAvoidhardware infrastructure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements automated metadata management and self-organizing micro-partition structures that enable the database system to manage itself without requiring complex infrastructure. The metadata automatically tracks the location and contents of micro-partitions, allowing the system to adapt to distributed storage environments without manual configuration or complex hardware management, thus providing in-house database management capabilities with simplified infrastructure.

Inventive Principle:
Principle #25Self-service

3Productivity

If database files are continuously modified with insertions and deletions, then data can be updated dynamically, but file fragmentation occurs reducing query performance

Engineering Contradiction:
Improvedata update speedVSAvoidquery response time
Core Design Contradiction:
ProductivityVSSpeed

Solution Approach 1:

The patent performs metadata updates and micro-partition reorganization in advance before they are strictly necessary. When files become fragmented, the system proactively reorganizes micro-partitions and updates metadata structures before query performance degrades significantly. This preliminary action maintains optimal query performance while allowing continuous data modifications, as the system prepares the storage structure in advance rather than reacting to performance degradation.

Inventive Principle:
Principle #10Preliminary action

4Quantity of substance

If data is stored across multiple computing devices for scalability, then storage capacity increases, but the system becomes highly susceptible to data loss during power outages or disasters

Engineering Contradiction:
Improvestorage capacityVSAvoiddata safety
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent merges data redundancy and metadata management at the micro-partition level across distributed storage devices. Each micro-partition is tracked by metadata that records its location and contents, allowing the system to reconstruct data across multiple devices. This merging of redundancy mechanisms with distributed storage provides both the storage capacity needed for scalability and the data safety required to prevent loss during disasters, as the system can recover by reconstructing micro-partitions from available data across the distributed network.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11593306B1File defragmentation service
Publication Date: 2023.02.28 SNOWFLAKE INC
  • US11593306B1 patent drawing
  • US11593306B1 patent drawing
  • US11593306B1 patent drawing

AI summary

The subject technology selects a most recently created file from a set of files stored in a source table. The subject technology iterates, in the source table, starting from the most recently created file up to an age threshold to select a first set of files for performing a first defragmentation process. The subject technology sets an indication corresponding to a particular file that is a last file, from the first set of files, that meets the age threshold. The subject technology performs the first defragmentation process on the selected first set of files. The subject technology determines that the first defragmentation process was successful.