Seismic Data Cataloging Through De-Duplication and Format Standardization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The rapid growth of seismic data volumes, often in isolated and vendor-specific proprietary formats, poses a challenge for effective management and cataloging in exploration and production (E&P) companies, with 85% to 90% of data being seismic data.
Innovation Solution
A method involving de-duplication, conversion, and metadata extraction of seismic data files across different formats, followed by ingestion into a cloud storage platform, utilizing machine-learning models to standardize byte locations and checksum validation for accurate cataloging.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If seismic data files are stored in multiple vendor-specific proprietary formats, then data compatibility and accessibility are improved, but data management complexity and storage redundancy increase
Solution Approach 1:
The patent introduces a standardized intermediate format (SEGY) as a mediator between various vendor-specific proprietary formats and the central data management system. The system automatically converts data from multiple source formats into this unified intermediate format, which then serves as the standard for storage and management, thereby reducing complexity while maintaining compatibility.
Solution Approach 2:
The system changes the parameter of data format standardization by enforcing a unified SEGY format as the target format for all seismic data. This parameter change transforms the data management approach from maintaining multiple formats to consolidating into a single standard, reducing redundancy and management complexity.
2Adaptability or versatility
If multiple copies of seismic data are generated for different workflows, then data availability for various purposes is improved, but data volume and storage requirements increase
Solution Approach 1:
The patent extracts only the essential data and metadata information needed for different workflows from the full seismic data sets. By separating and extracting relevant parameters (such as amplitude, depth, location) into metadata structures, the system enables multiple workflows to access specific data subsets without duplicating the entire large seismic data files.
Solution Approach 2:
The system segments seismic data management into separate components: the master seismic data stored in standardized format, and multiple metadata catalogs that index and describe different aspects of the data for various workflows. This segmentation allows different workflows to query and access specific metadata without requiring copies of the full seismic data.
3Reliability
If seismic data is cataloged and standardized, then data integrity and manageability are improved, but processing time and computational resources increase
Solution Approach 1:
The system performs preliminary actions by automatically converting all incoming seismic data to the standardized SEGY format and extracting metadata during the data ingestion process. This preliminary standardization and metadata extraction occurs before the data is stored in the database, so that when data is later queried or accessed, it is already in the required format, reducing processing time during subsequent operations.
Solution Approach 2:
The patent implements self-service through automated processes where the system independently performs format conversion, metadata extraction, and validation without requiring manual intervention. The automated cataloging process continuously maintains data integrity by checking and correcting metadata fields, reducing both time and resource requirements compared to manual cataloging.
Data Source
AI summary
Systems and methods for seismic data cataloging are provided. A method includes: receiving first seismic data files (SDFs) in a first file format (FFF), each including a seismic display pattern, de-duplicating the first SDFs to generate second SDFs in the FFF, identifying seismic three-dimensional (3D) files in the FFF and seismic two-dimensional (2D) files in the FFF from among the second SDFs, extracting header information from each seismic 3D and 2D file, converting each seismic 3D file to a corresponding plurality of seismic files in a second file format (SFF), each including a respective seismic display pattern, generating a corresponding histogram for each respective seismic display pattern for each seismic 3D file and the plurality of seismic files in the SFF, comparing each corresponding histogram for respective corresponding pairs of files to determine whether both have a same amplitude, if not, repeating the converting, generating, and comparing.


