Distributed Database Task Distribution via Node Heterogeneity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing distributed database systems face inefficiencies in data processing time due to variations in node performance, as they do not effectively account for differences in computing resources and task execution performance across nodes, leading to suboptimal task distribution and prolonged processing times.

Innovation Solution

A distributed database system with a computing power determination unit, device selection unit, and task distribution control unit that identifies and utilizes optimal computing devices based on their performance differences to distribute data accordingly, ensuring tasks are executed efficiently across nodes with varying capabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If accelerators are mounted on nodes to improve processing speed, then data processing performance is improved, but variation in processing performance between nodes increases due to system heterogeneity

Engineering Contradiction:
Improvedata processing speedVSAvoidperformance uniformity between nodes
Core Design Contradiction:
ProductivityVSStability of the object's composition

Solution Approach 1:

The patent assigns different data amounts to different nodes based on their specific computing capabilities. High-performance nodes with accelerators are assigned larger data amounts, while standard nodes are assigned smaller data amounts. This local differentiation in data distribution compensates for hardware heterogeneity and achieves uniform processing completion times across all nodes.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts the data processing parameter (data amount) based on the computing power of each node. By changing the data amount parameter in proportion to each node's computing capability, the patent ensures that all nodes complete their tasks simultaneously despite having different hardware configurations.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If more nodes are added to process large data amounts, then processing capacity increases, but system scale and cost increase

Engineering Contradiction:
Improvedata processing capacityVSAvoidnumber of nodes
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent optimizes the data amount parameter assigned to each node based on its computing power. By maximizing the utilization of each node's processing capability through proportional data assignment, the system achieves higher overall processing capacity with fewer nodes, thereby reducing system scale and costs.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If data amounts are evenly distributed to all nodes, then task distribution is simple, but processing time increases due to performance variations among nodes

Engineering Contradiction:
Improvetask distribution simplicityVSAvoiddata processing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

Instead of uniform data distribution, the patent implements local quality differentiation by assigning data amounts proportional to each node's computing power. This approach optimizes processing time by ensuring that high-performance nodes process more data while standard nodes process less, with all nodes completing their tasks simultaneously.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10936377B2Distributed database system and resource management method for distributed database system
Publication Date: 2021.03.02 HITACHI VANTARA LTD
  • US10936377B2 patent drawing
  • US10936377B2 patent drawing
  • US10936377B2 patent drawing

AI summary

The data processing times of data processing nodes are heterogeneous, and hence the execution time of a whole system is not optimized. A task is executed using a plurality of optimal computing devices by distributing a data amount of data to be processed with a processing command of the task for the plurality of optimal computing devices depending on a difference in computing power between the plurality of optimal computing devices, to thereby execute the task in a distributed manner using the plurality of optimal computing devices.