Distributed Data Processing System Column Integration Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In distributed data processing systems handling IoT or big data, the repetition of similar data types across bases leads to increased communication and storage costs, as conventional methods of integrating columns result in repeated data joins, failing to effectively reduce data amounts.

Innovation Solution

A data processing method where a central server uses node-cut processing to collect attribute information and relationship data from base servers, performing replacement calculations to identify data combinations that can be reduced, and notifies base servers to optimize column integration, thereby reducing data amounts transferred to the central server.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is collected from base servers without alteration, then data completeness is maintained, but communication and storage costs increase

Engineering Contradiction:
Improvedata completenessVSAvoidcommunication and storage costs
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent extracts redundant data from base server datasets by identifying and removing duplicate data across different bases. The extraction process focuses on removing unnecessary data elements while preserving unique and valuable information, thereby reducing communication and storage costs without compromising data completeness.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent merges data from multiple base servers by identifying common data elements that can be consolidated. Through integration and deduplication processes, the system combines datasets efficiently, eliminating redundant transmissions and storage while maintaining the完整性 of the overall data set.

Inventive Principle:
Principle #5Merging (Combining)

2Quantity of substance

If columns are integrated at each base using conventional statistical processing, then data amount is reduced, but repeated data joins occur at the central server

Engineering Contradiction:
Improvedata amountVSAvoidprocessing efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent applies preliminary action by performing data integration and deduplication at base servers before data transmission to the central server. By pre-processing data to eliminate redundancies in advance, the system prevents repeated data joins at the central server, thereby improving overall processing efficiency while reducing data transmission volumes.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements local quality by allowing each base server to perform customized data integration based on its specific data characteristics and local requirements. This localized processing approach enables efficient reduction of data amounts at each base while ensuring that the integrated data is optimized for subsequent central server processing.

Inventive Principle:
Principle #3Local quality

3Ease of manufacture

If different columns are excluded at each base for integration, then local data optimization is achieved, but data exclusion patterns vary across bases causing repeated joins

Engineering Contradiction:
Improvelocal data optimizationVSAvoiddata integration complexity
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The patent implements universality by establishing a standardized data integration framework that can be applied across all base servers. The system identifies universal data patterns and integration rules that work consistently across different bases, enabling local optimization while maintaining compatibility and preventing repeated joins through uniform processing approaches.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11132235B2Data processing method, distributed data processing system and storage medium
Publication Date: 2021.09.28 HITACHI LTD
  • US11132235B2 patent drawing
  • US11132235B2 patent drawing
  • US11132235B2 patent drawing

AI summary

A data processing method of a distributed data processing system, in which the base server collects and standardizes data and generates base data by node cut processing, the central server collects the attribute information of the column of the base data from a plurality of base servers and the relationship between the integration source and the integration destination of the base data by the node cut processing of the base server as base column integrated information, a combination of an integration source and an integration destination capable of reducing the data amount as a result of an replacement for calculating a combination of an integration source and an integration destination capable of reducing the data amount by exchanging the integration source and the integration destination when data is combined, the combination is notified to the base server as an exchange instruction.