Semantic Sub-File Storage for Parallel Computing Bandwidth
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing parallel storage systems lack efficient techniques for storing sub-files with semantically meaningful boundaries and metadata, leading to increased data processing and transfer bandwidth costs and reduced disk space efficiency.
Innovation Solution
The method involves generating files as a plurality of semantically meaningful sub-files with associated metadata, allowing for user-specified semantic information to be stored with sub-files in a parallel computing system, enabling replication based on meaningful boundaries and resolutions, and using storage tiering to optimize data storage and retrieval.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is stored as linear arrays of bytes with traditional metadata, then storage capacity is satisfied, but data processing costs and transfer bandwidth increase
Solution Approach 1:
The patent segments data into sub-files with semantically meaningful boundaries rather than storing as continuous linear arrays. Each sub-file contains data with specific semantic properties (e.g., electron density, temperature) that can be independently processed, reducing unnecessary data transfer and processing overhead while maintaining full storage capacity.
Solution Approach 2:
The patent introduces semantic information as an intermediary layer between raw data and application processing. This semantic metadata acts as a mediator that enables intelligent data selection, filtering, and optimization, allowing systems to process only relevant data portions rather than transferring entire datasets, thus reducing bandwidth costs.
2Quantity of substance
If traditional file storage is used, then storage capacity is maintained, but disk space efficiency decreases
Solution Approach 1:
The patent applies local quality by storing different semantic information at different levels of granularity. Critical semantic metadata is stored with sub-files for immediate access, while summary-level semantic information is stored separately. This allows the system to maintain full disk space capacity while improving space efficiency through intelligent data organization and selective storage.
3Device complexity
If sub-files are created without semantic boundaries, then storage is simplified, but data analysis and query processing become more difficult
Solution Approach 1:
The patent performs preliminary action by pre-computing and storing semantic information alongside sub-files during the data collection phase. This advance preparation of semantic metadata enables complex data analysis and query processing to be performed efficiently during execution, as the semantic structure is already in place rather than requiring complex post-processing.
Data Source
AI summary
Techniques are provided for storing files in a parallel computing system using sub-files with semantically meaningful boundaries. A method is provided for storing at least one file generated by a distributed application in a parallel computing system. The file comprises one or more of a complete file and a plurality of sub-files. The method comprises the steps of obtaining a user specification of semantic information related to the file; providing the semantic information as a data structure description to a data formatting library write function; and storing the semantic information related to the file with one or more of the sub-files in one or more storage nodes of the parallel computing system. The semantic information provides a description of data in the file. The sub-files can be replicated based on semantically meaningful boundaries.


