Semantic Sub-File Storage for Parallel Computing Bandwidth

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing parallel storage systems lack efficient techniques for storing sub-files with semantically meaningful boundaries and metadata, leading to increased data processing and transfer bandwidth costs and reduced disk space efficiency.

Innovation Solution

The method involves generating files as a plurality of semantically meaningful sub-files with associated metadata, allowing for user-specified semantic information to be stored with sub-files in a parallel computing system, enabling replication based on meaningful boundaries and resolutions, and using storage tiering to optimize data storage and retrieval.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is stored as linear arrays of bytes with traditional metadata, then storage capacity is satisfied, but data processing costs and transfer bandwidth increase

Engineering Contradiction:
Improvestorage capacityVSAvoiddata processing costs and transfer bandwidth
Core Design Contradiction:
Quantity of substanceVSLoss of energy

Solution Approach 1:

The patent segments data into sub-files with semantically meaningful boundaries rather than storing as continuous linear arrays. Each sub-file contains data with specific semantic properties (e.g., electron density, temperature) that can be independently processed, reducing unnecessary data transfer and processing overhead while maintaining full storage capacity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces semantic information as an intermediary layer between raw data and application processing. This semantic metadata acts as a mediator that enables intelligent data selection, filtering, and optimization, allowing systems to process only relevant data portions rather than transferring entire datasets, thus reducing bandwidth costs.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If traditional file storage is used, then storage capacity is maintained, but disk space efficiency decreases

Engineering Contradiction:
Improvestorage capacityVSAvoiddisk space efficiency
Core Design Contradiction:
Quantity of substanceVSLoss of substance

Solution Approach 1:

The patent applies local quality by storing different semantic information at different levels of granularity. Critical semantic metadata is stored with sub-files for immediate access, while summary-level semantic information is stored separately. This allows the system to maintain full disk space capacity while improving space efficiency through intelligent data organization and selective storage.

Inventive Principle:
Principle #3Local quality

3Device complexity

If sub-files are created without semantic boundaries, then storage is simplified, but data analysis and query processing become more difficult

Engineering Contradiction:
Improvestorage structureVSAvoiddata analysis and query processing
Core Design Contradiction:
Device complexityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent performs preliminary action by pre-computing and storing semantic information alongside sub-files during the data collection phase. This advance preparation of semantic metadata enables complex data analysis and query processing to be performed efficiently during execution, as the semantic structure is already in place rather than requiring complex post-processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8949255B1Methods and apparatus for capture and storage of semantic information with sub-files in a parallel computing system
Publication Date: 2015.02.03 TRIAD NATIONAL SECURITY LLC
  • US8949255B1 patent drawing
  • US8949255B1 patent drawing
  • US8949255B1 patent drawing

AI summary

Techniques are provided for storing files in a parallel computing system using sub-files with semantically meaningful boundaries. A method is provided for storing at least one file generated by a distributed application in a parallel computing system. The file comprises one or more of a complete file and a plurality of sub-files. The method comprises the steps of obtaining a user specification of semantic information related to the file; providing the semantic information as a data structure description to a data formatting library write function; and storing the semantic information related to the file with one or more of the sub-files in one or more storage nodes of the parallel computing system. The semantic information provides a description of data in the file. The sub-files can be replicated based on semantically meaningful boundaries.