Unified OS for Petabyte Data Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional computer architectures are inadequate for handling data-intensive computing tasks, particularly with petabyte-scale datasets, due to limitations in CPU performance, I/O bandwidth, and storage capacity, leading to inefficiencies in data processing and analysis.

Innovation Solution

A data-intensive computer system comprising multiple server systems forming a unified processing and storage infrastructure, with a unifying operating system environment that coordinates distributed processes across processing and storage sub-systems, enabling efficient data storage, retrieval, and analysis by treating the database as a layer in the memory hierarchy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional computer architecture is used, then hardware components are simple and well-understood, but the system cannot handle petabyte-scale data processing requirements

Engineering Contradiction:
Improvedata processing capabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the monolithic computer architecture into distributed server systems, each handling specific data processing tasks. The database is divided into multiple partitions distributed across different servers, enabling parallel processing of large datasets while maintaining manageable complexity at each node.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new architectural dimension by integrating the database directly into the memory hierarchy of the distributed system. This creates a multi-layered architecture where data can be accessed at different speeds and locations, enabling petabyte-scale processing without proportionally increasing system complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Quantity of substance

If data is stored in traditional storage systems, then storage capacity is limited, but moving data to processing units creates excessive I/O bandwidth requirements

Engineering Contradiction:
Improvestorage capacityVSAvoidI/O bandwidth consumption
Core Design Contradiction:
Quantity of substanceVSLoss of energy

Solution Approach 1:

The patent merges the storage and processing functions by integrating the database directly into the memory hierarchy of the distributed server system. This eliminates the traditional separation between storage and processing, allowing data to be processed in-place without excessive I/O operations.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The unified operating system environment acts as an intermediary that manages data placement and access across the distributed system. It optimizes data location and retrieval, reducing unnecessary I/O operations while maintaining access to petabyte-scale storage capacity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If distributed processing is implemented, then data processing throughput increases, but coordinating processes across multiple systems becomes complex

Engineering Contradiction:
Improvestream processing throughputVSAvoidprocess coordination complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The unified operating system environment provides universal process coordination capabilities across all distributed server systems. It implements a standardized interface and management layer that handles process scheduling, data distribution, and resource allocation uniformly, simplifying coordination complexity while maintaining high throughput.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Quantity of substance

If petabyte-scale databases are implemented, then data storage capacity is sufficient, but data retrieval and access time increase

Engineering Contradiction:
Improvedatabase storage capacityVSAvoiddata access time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The database is segmented into distributed partitions across multiple server systems, each accessible through the unified operating system. This segmentation allows parallel data retrieval operations and reduces access time by eliminating single-point bottlenecks while maintaining petabyte-scale storage capacity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9462048B2System and method for program and resource allocation within a data-intensive computer
Publication Date: 2016.10.04 JOHNS HOPKINS UNIVERSITY
  • US9462048B2 patent drawing
  • US9462048B2 patent drawing
  • US9462048B2 patent drawing

AI summary

A system and method for operating a data-intensive computer is provided. The data-intensive computer includes a processing sub-system formed by a plurality of processing node servers and a database sub-system formed by a plurality of database servers configured to form a collective database in excess of a petabyte of storage. The data-intensive computer also includes an operating system sub-system formed by a plurality of operating system servers that extend a unifying operating system environment across the processing sub-system, the database sub-system, and the operating system sub-system to act as components in a single data-intensive computer. The operating system sub-system is configured to coordinate execution of a single application as distributed processes having at least one of the distributed processes executed on the processing sub-system and at least one of the distributed processes executed on the database sub-system.