Multi-source marine environment data storage method and system based on HDF5

By building a storage framework and spherical irregular grid model that is dynamically decoupled by domain-type-source, the data standard splitting and redundancy problems in marine environmental data storage are solved, and efficient integration and rapid retrieval of multi-source data is achieved, which is suitable for national-level marine data centers and scientific research institutions.

CN120336278AActive Publication Date: 2025-07-18QINGDAO INNOVATION & DEV CENT OF HARBIN ENG UNIV +1

Patent Information

Application Number
CN202510804363.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-18
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The prior art has problems such as data standard fragmentation, multi-source heterogeneous data storage island effect, earth spherical data storage redundancy and topological conflict, and traditional storage technology in the marine environment data storage, resulting in low data sharing efficiency and spatial expression redundancy.

Method used

Using a multi-source marine environment data storage method based on HDF5, a storage framework with dynamic decoupling of domain-type-source is built, an ontology expression model of spherical irregular grid is designed, a spatiotemporal composite coding mechanism and a unified interface of hybrid grids is introduced, and a cross-scale and multi-modal data is supported to automatically align and efficient retrieval of cross-scale and multi-modal data.

Benefits of technology

It realizes unified management and rapid retrieval of interdisciplinary and multimodal data, reduces redundant storage in high-latitude areas, improves the integration efficiency and retrieval performance of multi-source data, and provides an efficient data sharing base.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336278A_ABST
    Figure CN120336278A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data storage, and discloses a multi-source marine environment data storage method and system based on HDF5. According to the method, a hierarchical storage framework is constructed, a domain-type-source three-level directory architecture is constructed, a space-time dimension dynamic extension mechanism is executed, and space-time composite coding is performed; designing an HDF5 nested packet storage architecture, wherein grid topological structure definition, dynamic resolution self-adaptive storage, multi-source data processing and metadata management are carried out; constructing distributed data association and indexes, including data standardization preprocessing and cross-source data association; the storage performance is optimized and expanded, wherein density sensitive blocking and layered compression are carried out. According to the method, redundant storage of a high-latitude region is greatly reduced, and dynamic adaptation of a local encryption grid is supported. According to the invention, open circulation and collaborative utilization of interdisciplinary data resources are facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data storage, and particularly relates to a multi-source marine environment data storage method and system based on HDF5. Background Art

[0002] With the rapid development of marine science and environmental monitoring technologies, global research institutions have established a massive data acquisition network covering multiple disciplines such as atmosphere, ocean, geology, and ecology. The data types span multi-modal forms such as numerical simulation, in-situ observation, satellite remote sensing, AI inversion, and video surveillance. However, the integration and management of multi-source data face three core contradictions: First, the data standards among research institutions are fragmented. For example, the NetCDF-CF convention is commonly used in the atmosphere field, while ocean buoy data mostly follows the WMO binary format, and environmental monitoring data relies on the ISO 19115 meta-model. The differences in the standard systems result in a large amount of time being consumed for format conversion and semantic alignment when sharing cross-institutional data. Second, the heterogeneity of data types is significant. The storage requirements for scientific data (such as temperature-salinity profiles), images (such as satellite cloud images), videos (such as deep-sea exploration images), and AI-generated products (such as wave prediction tensors) are diverse. Traditional solutions need to rely on dispersed storage systems (such as databases for storing numerical values and object storage for storing images), exacerbating the data silo effect. Third, the spatial particularity of spherical data on the Earth (such as the convergence of meridians and the latitude dependence of resolution) conflicts with the rigid structure of traditional storage formats. For example, when storing data with a resolution of 0.1° at the equator in a regular grid, a large amount of redundant storage is generated in high-latitude regions due to the sudden reduction in the area of grid cells, and it is impossible to be compatible with the locally encrypted grids required for vortex tracking.

[0003] The current mainstream technical solutions expose significant limitations in addressing the above challenges: (1) Although relational databases support structured metadata management, it is difficult to support the storage requirements of irregular grids and high-dimensional time-series data, and the scalability is limited by the fixed table structure. (2) Although scientific data formats such as NetCDF have the ability to store multi-dimensional arrays, their preset grid models (such as regular longitude-latitude grids) cannot dynamically adapt to spherical irregular topologies (such as adaptive encrypted grids or unstructured grids generated by AI), and lack native support for multi-modal data such as images and videos. (3) Emerging cloud-native storage optimizes query performance through columnar storage, but lacks a dedicated indexing mechanism for spherical data on the Earth, resulting in low spatio-temporal query efficiency and prone to spatial distortion when retrieving cross-latitude data. The contradictions among the data organization granularity, spatial expression ability, and multi-source compatibility of the above methods have become the key obstacles restricting the sharing of atmospheric and oceanic environment data in China.

[0004] In the prior art, a storage and management method for a multi-resolution block-stacked grid based on HDF5 proposed a hierarchical storage architecture that combines regular grids and adaptive encrypted grids, and realized the multi-resolution management of data by constructing a multi-level K-d tree index. Although this method optimized the spatial query efficiency, its storage model still relied on fixed topological rules, could not be compatible with the spatial irregularity of satellite remote sensing and buoy measured data, and did not define the ontology semantics of spherical grids.

[0005] Through the above analysis, the problems and defects existing in the prior art are as follows: Problem 1: The fragmentation of data standards leads to low sharing efficiency. There are significant differences in data standard systems among scientific research institutions (such as the NetCDF-CF convention, the WMO binary format, and the ISO 19115 meta-model), and the data standards in the atmospheric field and ocean data are incompatible, resulting in the complication of the data integration process and hindering multi-disciplinary collaborative research.

[0006] Problem 2: The storage island effect of multi-source heterogeneous data is serious. Scientific data, images, videos, and AI-generated products need to rely on dispersed storage systems due to different storage requirements. The traditional solution lacks a unified storage framework, exacerbating data islands. For example, numerical simulation data and remote sensing images need to be managed separately, making it difficult to realize cross-modal data correlation analysis.

[0007] Problem 3: Redundancy in the storage of spherical data on the Earth and topological conflicts. The regular grid storage method generates up to 45% redundant storage in high-latitude regions due to the sudden reduction of the grid cell area, and cannot adapt to local encrypted grids. Fixed longitude and latitude grids cannot dynamically respond to unstructured grids or adaptive encryption requirements, resulting in an imbalance between spatial expression accuracy and storage efficiency.

[0008] Problem 4: Insufficient adaptability of traditional storage technologies. Relational databases are limited by fixed table structures and are difficult to support the storage of irregular grids and high-dimensional time-series data; although scientific data formats such as NetCDF support multi-dimensional arrays, their preset grid models cannot dynamically adapt to unstructured grids generated by AI and lack multi-modal data compatibility. Summary of the Invention

[0009] To overcome the problems existing in the related technologies, the disclosed embodiments of the present invention provide a multi-source ocean environment data storage method and system based on HDF5. The purpose of the present invention is to propose an adaptive storage method for multi-source ocean environment data based on HDF5. By constructing a storage framework with dynamic decoupling of domain-type-source, and designing an ontology expression model for spherical irregular grids, the technical barriers to multi-source data integration and precise representation of the Earth's space are broken through. This method innovatively introduces a spatio-temporal composite coding mechanism and a unified interface for hybrid grids, supports automatic alignment and efficient retrieval of cross-scale and multi-modal data, and provides a highly compatible and low-redundancy integrated data foundation for ocean digital twins and data sharing.

[0010] The technical solution is as follows: A multi-source marine environment data storage method based on HDF5, comprising the following steps: S1: Construct a hierarchical logical directory and coding system; taking the domain - type - source as the main line, construct a three - level directory structure, combined with a spatio - temporal dimension dynamic expansion mechanism and a spatio - temporal composite coding method, to form a unified data logical organization system, providing a directory index basis and refined time - space expression for subsequent data storage and retrieval; S2: Design a storage structure and multi - source adaptation mechanism; based on the constructed directory and spatio - temporal coding system, design an HDF5 nested grouping storage structure, supporting the topological definition storage of regular and irregular grids, dynamic resolution adaptive storage, and multi - modal data storage of images and videos; construct a metadata management mechanism to ensure data description consistency and cross - source compatibility; S3: Construct a data index and association mechanism for multi - source collaboration; aiming at the integration requirements of multi - source heterogeneous data, design a standardized pre - processing process, unify the time reference and physical unit system, construct a distributed spatio - temporal index mechanism, and achieve accurate matching and efficient linkage query of cross - source data; S4: Carry out performance optimization and storage compression strategies; after completing logical organization, physical storage, and data indexing, introduce density - sensitive block division and multi - level compression strategies, dynamically adjust the block division method according to grid density and access patterns, and adopt compression algorithms of ZFP and JPEG2000 to achieve high - performance and high - compression - ratio optimization of multi - source data storage.

[0011] In step S1, taking the domain - type - source as the main line, construct a three - level directory structure, including: creating multiple main groups under the root directory to represent the disciplinary domain attributes; establishing subgroups of different types of products under each domain group; grouping by data source within the type subgroup to determine the data collection device type.

[0012] In step S1, the spatio - temporal dimension dynamic expansion mechanism includes: Time - axis organization, creating a time - series directory nested by / Year / Month / Day within the data source group; storing the timestamp array in / Index / TimeIndex, and establishing a mapping relationship between the time directory and the data set; Spherical spatio - temporal hybrid index, introducing a latitude - dependent resolution function to dynamically adjust the grid cell size, and generating a unique grid code through the H3 geospatial grid algorithm; embedding a spatio - temporal R - tree index structure in the HDF5 file, where the index node contains the time range, spatial grid code, and physical address of the data block; H3 hexagonal grid coding, using the H3 geospatial grid to divide the sphere into multi - resolution hexagonal cells; Hierarchical spatio-temporal coding dynamically selects the coding granularity and dimension according to the spatio-temporal characteristics of the data. At the same time, for the high-dimensional data of the profile data of the ocean and the atmosphere, a vertical dimension identifier is embedded in the coding, which is associated with the depth of the ocean-atmosphere profile data.

[0013] Furthermore, the latitude-dependent resolution function is: ; where is the resolution at latitude , is the equatorial resolution, is the polar resolution, and is the resolution switching threshold; Generate a unique grid code through the H3 geographic grid algorithm, including: (1) Convert longitude and latitude to three-dimensional coordinates; ; ; ; where is the longitude; (2) Determine the resolution level. Determine the resolution level according to the latitude-dependent resolution function, including levels 0-15; (3) Calculate the center of the hexagon. Calculate the center position of the hexagon formed by multiple nodes; Let the longitude and latitude coordinates of the vertices of the hexagon be: ; Then the geometric center is: ; ; where and are the longitude and latitude of the th vertex of the hexagon, respectively; (4) Generate the H3 code. Generate a 64-bit H3 index according to the center position of the hexagon.

[0014] In step S2, the topological definition includes: The regular grid stores the coordinate axes as independent datasets and is linked to the main dataset through attributes; The irregular grid stores the three-dimensional coordinate matrix of each node and is linked to the main dataset through attributes; the irregular grid creates an `Adjacency_List` dataset, which records the adjacent cell indices of each grid cell in a sparse matrix format; a grid_type field is defined in the dataset attributes, and a grid type description is attached; an independent Group is created for each time step to store the grid topology and geometric data of that time step.

[0015] In step S2, dynamic resolution adaptive storage includes: storing the basic coordinates of the grid nodes in the form of a three-dimensional floating-point array, and declaring the coordinate system standard in the attributes; creating an auxiliary dataset under the same grouping to store the planar projection coordinates and annotating the projection method; regular grids record a fixed resolution, irregular grids divide the resolution identification by region, and record the differences between the equator and the poles in the attributes. Multi-source data processing includes: directly storing regular grids as multi-dimensional arrays; linking irregular grids through independent coordinate datasets and attributes; decoding image data into pixel arrays; and treating video data as multi-dimensional arrays in the time dimension. The metadata management includes designing the basic metadata of the dataset for different types of data. The basic metadata of the dataset includes scientific data metadata, image data metadata, and video data metadata.

[0016] In step S3, for the integration requirements of multi-source heterogeneous data, a standardized preprocessing process is designed to unify the time reference and physical unit system, including: Designing a standardized preprocessing process to uniformly convert the raster data of formatted satellite images, and resampling the time series data of the buoy to the unified time reference unit system for normalization; defining a standard unit conversion rule library, which is declared through the `unit` field of the attribute. Achieving precise matching and efficient linkage query of cross-source data, including spatio-temporal alignment of cross-source data and spatio-temporal matching of multi-source data.

[0017] In step S4, the density-sensitive block division includes: dynamically adjusting the HDF5 block size according to the grid node density, dynamically adjusting the block dimension according to the data access mode, dividing time series data by the time axis; dividing spatial raster data by longitude and latitude. Achieving high-performance and high-compression ratio multi-source data storage optimization, including: using the ZFP floating-point compression algorithm for lossless compression of floating-point scientific data and coordinate data; applying JPEG2000 lossy compression for image / video data.

[0018] Furthermore, the node density is defined as: ; wherein, is the total number of nodes in the current area, is the area of the region; The block size is defined as: ; Another object of the present invention is to provide a multi-source marine environmental data storage system based on HDF5. This system implements the multi-source marine environmental data storage method based on HDF5. This system includes: A hierarchical storage framework construction module, which is used to construct a hierarchical logical directory and coding system; taking domain - type - source as the main line, constructing a three-level directory structure, and combining a spatio-temporal dimension dynamic expansion mechanism and a spatio-temporal composite coding method to form a unified data logical organization system, providing a directory index basis and a refined time - space expression for subsequent data storage and retrieval; An HDF5 nested grouping storage architecture design module, which is used to design a storage structure and a multi-source adaptation mechanism; based on the constructed directory and spatio-temporal coding system, designing an HDF5 nested grouping storage structure, supporting the topological definition storage of regular and irregular grids, dynamic resolution adaptive storage, and multi-modal data storage of images and videos; constructing a metadata management mechanism to ensure data description consistency and cross-source compatibility; A distributed data association and index construction module, which is used to construct a data index and association mechanism for multi-source collaboration; aiming at the integration requirements of multi-source heterogeneous data, designing a standardized preprocessing process, unifying the time reference and physical unit system, constructing a distributed spatio-temporal index mechanism, and realizing accurate matching and efficient linkage query of cross-source data; An optimization and expansion module, which is used to carry out performance optimization and storage compression strategies; after completing logical organization, physical storage, and data indexing, introducing density-sensitive block division and multi-level compression strategies, dynamically adjusting the block division method according to grid density and access patterns, and adopting compression algorithms of ZFP and JPEG2000 to realize the optimization of multi-source data storage with high performance and high compression ratio.

[0019] Combining all the above technical solutions, the beneficial effects of the present invention are as follows: First, the present invention solves the bottleneck problems of existing marine environmental data storage technologies in aspects such as the fragmentation of multi-source data standards, redundant spatial expression, insufficient adaptability of storage architectures, and low spatio-temporal retrieval efficiency, and proposes a multi-source marine environmental data adaptive storage method based on HDF5. This method breaks through the contradiction among data organization granularity, spatial topology adaptability, and multi-source compatibility of traditional storage solutions by constructing a dynamically decoupled storage framework of domain - type - source, designing an ontological expression model of spherical irregular grids, and integrating a spatio-temporal composite coding mechanism and a unified interface for hybrid grids.

[0020] Second, compared with the prior art, the present invention improves the integration efficiency and spatial representation accuracy of multi-source marine environmental data by constructing a dynamically decoupled hierarchical storage framework and a spherical grid ontology expression model. Through the four-level directory architecture of domain-type-source-modal, the unified management and rapid retrieval of interdisciplinary and multi-modal data are realized, effectively solving the data island problem caused by standard fragmentation and storage dispersion in traditional solutions. At the same time, through the spatio-temporal composite coding mechanism and the latitude adaptive resolution design, the redundant storage in high-latitude regions is greatly reduced, and the dynamic adaptation of local encrypted grids is supported. The present invention provides an efficient data sharing foundation for global marine research institutions and industrial applications, contributing to the open circulation and collaborative utilization of interdisciplinary data resources.

[0021] Third, the multi-source marine environmental data adaptive storage method proposed by the present invention has good engineering feasibility and interdisciplinary adaptation ability. Its transformation can greatly reduce the redundant storage in high-latitude regions, significantly improve the integration efficiency and retrieval performance of multi-source data, and is applicable to scenarios such as national-level marine data centers, research institutions, and ship meteorological systems. It has broad commercial application prospects and promotion value in key fields such as smart ocean, ocean power strategy, climate change monitoring, and offshore resource development.

[0022] Fourth, the current international mainstream marine environmental data management solutions still mostly adopt fixed-structure NetCDF or traditional databases, lacking unified support for irregular grids, heterogeneous modalities (images, videos, AI prediction tensors, etc.), and spherical coordinate systems, and not forming an index structure deeply coupled with the earth space characteristics. The present invention first introduces a spherical hybrid grid expression mechanism based on H3 hexagonal coding, and realizes a dynamic decoupled directory design and index compression mechanism across modalities and sources, achieving a principle-level breakthrough in supporting complex marine data storage and query, and filling the technical gap in the data organization and coding level in this field internationally. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure; Figure 1 is a flowchart of a multi-source marine environmental data storage method based on HDF5 provided by an embodiment of the present invention; Figure 2 is a three-level directory architecture diagram of domain-type-source provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings. Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0025] The innovation of the present invention lies in: by constructing a multi-level directory system with domain-type-source dynamic decoupling, combining a spherical adaptive resolution function and H3 hexagonal grid encoding, a spatio-temporal composite index and a multi-modal unified storage mechanism are proposed; the system supports the coexistence of regular and irregular grids, the adaptive access and compression optimization of multi-source data, realizes high-dimensional redundancy reduction, efficient cross-source retrieval, and the integrated management of complex environment data, providing technical support for the sharing and intelligent application of marine environmental information.

[0026] Dynamic decoupling hierarchical storage framework: Construct a four-level directory architecture of domain-type-source-modal, and through the spatio-temporal dimension dynamic expansion mechanism and spatio-temporal composite encoding, realize the unified management and efficient retrieval of interdisciplinary and multi-modal data; Spherical grid ontology expression model: Design a unified interface for hybrid grids and a latitude adaptive resolution mechanism, and fuse the topological structure definitions of regular / irregular grids to solve the conflict between the spatial particularity of spherical data on the earth and the rigid structure of traditional storage formats; Distributed data association and performance optimization: By combining density-sensitive block partitioning and hierarchical compression techniques with a cross-source spatio-temporal alignment algorithm, improve storage efficiency and reduce high-dimensional redundancy, and support the dynamic adaptation and fast access of unstructured data and high-dimensional time-series data.

[0027] Example 1, as Figure 1 shown, the multi-source marine environmental data storage method based on HDF5 provided by the embodiment of the present invention: S1: Construct a hierarchical logical directory and encoding system; taking domain-type-source as the main line, construct a three-level directory architecture, and combine the spatio-temporal dimension dynamic expansion mechanism and spatio-temporal composite encoding method to form a unified data logical organization system, providing a directory index basis and refined time-space expression for subsequent data storage and retrieval; Exemplarily, construct a three-level directory architecture of domain-type-source. Create multiple main groups under the root directory to represent the subject area attributes; establish subgroups such as different types of products under each domain group; group by data source within the type subgroup to clarify the data collection device type. As Figure 2 shown in the three-level directory architecture diagram of domain-type-source; Constructing the three-level directory architecture of domain-type-source specifically includes: Domain layer division: Create main groups ` / Meteorology`, ` / Oceanography`, ` / Environmental`, and ` / Geography` under the root directory to represent the attributes of academic disciplines; ` / Meteorology` is the meteorology domain, covering observation, simulation, and forecast data. ` / Oceanography` is the oceanography domain, including measured profiles, numerical simulations, and forecast products. ` / Environmental` is the environmental domain, managing data such as air quality, water quality, and soil. ` / Geography` is the geographic information domain, integrating data such as remote sensing images, GIS analysis, and urban modeling.

[0028] Type layer division: Create subgroups ` / Observation`, ` / Simulation`, ` / Forecast`, ` / Image`, and ` / Video` under each domain group; Observation is observation data (including buoy, shipboard, and satellite observations). Simulation is numerical simulation data (including WRF, ROMS, and WW3). Forecast is forecast products (including numerical models and AI models). Image is image data, and Video is video data.

[0029] Source layer division: Group by sources ` / Satellite`, ` / Buoy`, ` / Radar`, ` / Ship`, ` / AI`, ` / NWP` within the type subgroups to clarify the types of data acquisition devices. Satellite is satellite remote sensing data, Buoy is buoy observation data, Radar is radar observation data, Ship is shipboard data, AI is artificial intelligence forecast data, and NWP is numerical model simulation data.

[0030] Exemplarily, a spatio-temporal dimension dynamic expansion mechanism.

[0031] Create a time series directory within the data source group; for datasets covering the global ocean, divide the space into subgroups according to the longitude and latitude ranges, and each subgroup stores data for the corresponding area; design a spatio-temporal joint coding directory for high-frequency updated data to integrate the spatial location and acquisition time.

[0032] The spatio-temporal dimension dynamic expansion mechanism specifically includes the following: Time axis organization: Create a time series directory nested by / Year / Month / Day within the data source group (four-level expansion of year / month / day / hour). Store the timestamp array (in datetime64 format) in / Index / TimeIndex, and establish the mapping relationship between the time directory and the dataset; Spherical Spacetime Hybrid Index: To address the convergence problem of longitude and latitude grids in high-latitude regions during the partitioning of the Earth's spherical surface, a latitude-dependent resolution function is introduced to dynamically adjust the grid cell size, and a unique grid code is generated through the H3 geospatial grid algorithm. A spatio-temporal R-tree index structure is embedded in the HDF5 file, and the index nodes contain the time range, spatial grid code, and physical address of the data block.

[0033] The latitude-dependent resolution function is: ; In the formula, is the resolution at latitude , is the equatorial resolution, is the polar resolution, is the resolution switching threshold; H3 Hexagonal Grid Encoding: The H3 geospatial grid is adopted to divide the spherical surface into multi-resolution hexagonal cells. The H3 geospatial grid reduces data redundancy in high-latitude regions through the multi-resolution partitioning of hexagonal cells and is deeply adapted to the characteristics of the Earth's spherical surface. The unique code for each grid is generated through the following steps: (1) Convert longitude and latitude to three-dimensional coordinates: ; ; ; In the formula, is the longitude; (2) Determine the resolution level. The resolution level (0 - 15) is determined according to the latitude-dependent resolution function. The higher the level, the finer the grid.

[0034] (3) Calculate the center of the hexagon. Calculate the center position of the hexagon formed by multiple nodes.

[0035] (4) Generate the H3 code. Generate a 64-bit H3 index based on the hexagon center position.

[0036] Hierarchical Spatio-Temporal Encoding: The encoding granularity and dimension are dynamically selected according to the spatio-temporal characteristics of the data. High-frequency data such as satellite observations at the second level adopt the format "Geohash_xxxx_YYYYMMDDHHMMSS", the timestamp is accurate to the second level, and the Geohash accuracy is increased to 6 bits (resolution of about 0.6 km). The timestamp of low-frequency data such as monthly mean simulations is simplified to "YYYYMM", and the Geohash accuracy is reduced to 4 bits (resolution of about 20 km). At the same time, for high-dimensional data such as profile data of the ocean and atmosphere, a vertical dimension identifier is embedded in the encoding to achieve depth association of sea-air profile data.

[0037] S2: Design the storage structure and multi-source adaptation mechanism; Based on the constructed directory and spatio-temporal coding system, design the HDF5 nested grouping storage structure, which supports the topological definition storage of regular and irregular grids, dynamic resolution adaptive storage, and multi-modal data storage of images and videos; Construct a metadata management mechanism to ensure data description consistency and cross-source compatibility; Exemplarily, the grid topology structure definition includes: The regular grid stores the coordinate axes as independent datasets and links them to the main dataset through attributes. The irregular grid stores the three-dimensional coordinate matrix of each node and links it to the main dataset through attributes. The irregular grid creates an `Adjacency_List` dataset and records the adjacent cell indices of each grid cell in a sparse matrix format; Define the grid_type field in the dataset attributes and append the grid type description. Create an independent Group for each time step to store the grid topology and geometric data of that time step. Only store the differences from the previous time step to reduce redundancy.

[0038] Exemplarily, the grid topology structure definition specifically includes: Basic coordinate storage: The regular grid stores the three-dimensional coordinate axes (longitude, latitude, depth / pressure) as independent one-dimensional arrays, namely / Grid / Coordinates / Latitude (latitude array, float32, shape N, where N is the length of the latitude array); / Grid / Coordinates / Longitude (longitude array, float32, shape M, where M is the length of the longitude array); / Grid / Coordinates / Depth (depth array, float32, shape K, where K is the length of the depth array), improving data access efficiency through structured storage and reducing the complexity of multi-dimensional data parsing. The irregular grid stores the three-dimensional coordinates of each node as a matrix (float32, shape Px3, where P is the total number of nodes). The basic coordinates are linked to the main dataset through attributes.

[0039] Adjacency relation description: For the irregular grid, point to the adjacency list dataset through an attribute field in the main dataset. Record the global indices (int32, shape Q, where Q is the total number of all adjacent cell indices in the adjacency list) and the starting positions of the adjacency indices of each cell (int32, shape P+1, where P is the total number of cells) in a sparse matrix format (CSR), reducing redundant space occupancy, quickly locating adjacent cells, and optimizing the topological query performance.

[0040] Metadata mapping: Define the grid_type field (RegularGrid or IrregularGrid) in the dataset attributes and append the grid type description to simplify the data classification and processing logic.

[0041] Dynamic grid deformation: Independent groups are created at each time step to store the grid topology and geometric data for that time step. Only the geometric differences from the previous time step are stored, reducing duplicate storage and lowering the dynamic data storage overhead.

[0042] Exemplarily, dynamic resolution adaptive storage includes storing the base coordinates of grid nodes in the form of a three-dimensional floating-point array, and declaring the coordinate system standard in the attributes. An auxiliary data set is created under the same grouping to store the planar projection coordinates and label the projection method. Regular grids record a fixed resolution, and irregular grids divide the resolution identification by region and record the differences between the equator and the poles in the attributes.

[0043] Specifically, dynamic resolution adaptive storage includes spherical coordinate ontology storage, projection coordinate additional storage, and hierarchical resolution annotation; The spherical coordinate ontology storage includes: 1) Definition of the spherical three-dimensional floating-point array structure. The longitude, latitude, and elevation / pressure of each spherical grid node are stored in sequence as a three-dimensional floating-point array with a shape of (N, 3), where N is the total number of nodes. If the node coordinates are high-precision data (such as satellite positioning data), float64 (double precision) is used; if they are conventional observation data (such as buoys), float32 (single precision) is used, and the precision type is declared through the attribute precision.

[0044] 2) Standard annotation of the spherical coordinate system. The coordinate system standard is declared in the dataset attributes in the format of "standard name (EPSG code)". For example: Attributes: {"crs": "WGS84 (EPSG:4326)"} For special coordinate systems such as CGCS2000, ellipsoid information is extended in the attributes: Attributes: { "crs": "CGCS2000 (EPSG:4490)", "ellipsoid": "GRS80", "semi_major_axis": 6378137.0, "inverse_flattening": 298.257222101 } The projection coordinate additional storage includes: creating an auxiliary data set under the same grouping level as the spherical coordinates, storing planar projection coordinates such as UTM and Web Mercator, annotating the projection method through the attribute "projection_method", and storing the projected planar coordinates (x, y) in one-to-one correspondence with the spherical coordinates.

[0045] The hierarchical resolution annotation includes: recording the fixed resolution in the main data set attributes for regular grids. For irregular grids, the resolution is identified by region division (such as / High_Res_0.1deg, / Low_Res_1.0deg), and the resolution gradient is declared in the attributes (such as 0.1° at the equator and 1.0° at the poles).

[0046] Exemplarily, the multi-source data processing includes: directly storing regular grids as multi-dimensional arrays; linking irregular grids through an independent coordinate data set and attributes; decoding image data into pixel arrays; and treating video data as a multi-dimensional array with a time dimension.

[0047] The specific multi-source data processing includes: Scientific data storage: Regular grids are directly stored as multi-dimensional arrays (float32 or float64), with the data shape being [depth dimension][latitude dimension][longitude dimension], and the coordinate axis data set of the regular grid is pointed to through an attribute field; for irregular grids, the node coordinates are stored as an independent matrix (float32, with a shape of Px3, where P is the total number of nodes), and the node coordinate data set is pointed to through an attribute field.

[0048] Image data storage: Decode the original image into a pixel array and store it in the uint8 format, and declare the image resolution, color type, and compression algorithm in the attributes.

[0049] Video data storage: Store the video data as a multi-dimensional array with a time dimension ([time step][number of channels][height][width]), and declare the frame rate, encoding format, and resolution in the attributes.

[0050] Exemplarily, the metadata management includes: designing the basic metadata of the data set for different types of data.

[0051] Specifically, the metadata management includes: For different types of data, the data set contains the following basic metadata of the data set: scientific data metadata, image data metadata, and video data metadata. See Table 1, Table 2, and Table 3; Table 1 Scientific data metadata

[0052] Table 2 Image data metadata

[0053] Table 3 Video Data Metadata

[0054] An example of the final formed file structure is as follows: # Geospatial Coordinate Ontology Storage / Geospatial / Coordinates (Dataset: float32, shape=1000x3) # Store 3D geospatial coordinates (longitude, latitude, elevation) of 1000 nodes # Attribute declaration of coordinate system standard (WGS84) Attributes: {"crs": "WGS84 (EPSG:4326)"} # Additional storage of planar projected coordinates / Geospatial / Projection (Dataset: float32, shape=1000x2) # Store planar projected coordinates of 1000 nodes # Attribute annotation of projection algorithm Attributes: {"projection_method": "Web Mercator (EPSG:3857)"} # Meteorological observation data (regular / irregular grid) / Meteorology / Observation / Satellite / 2025-05-21 ├── Temperature_Regular (Dataset: float32, shape=10x180x360) │# Store regular grid temperature data (10 depth layers x 180 latitudes x 360 longitudes) │# Metadata declaration of source, unit, grid type, resolution │Attributes: {"source": "GOES-16", "unit": "°C", "grid_type": "regular", "resolution": "0.5°"} │# Coordinate link to an independent axis dataset │Coordinates: {"longitude": / Geospatial / Coordinates_Longitude, "latitude": / Geospatial / Coordinates_Latitude} └── Temperature_Unstructured (Dataset: float32, shape=1000) # Store unstructured grid temperature data (1000 nodes) # Metadata declare source, unit, grid type Attributes: {"source": "Buoy Network", "unit": "°C", "grid_type": "unstructured"} # Coordinates linked to a unified geospatial coordinate dataset Coordinates: {"coordinates": / Geospatial / Coordinates} # Ocean simulation data / Oceanography / Simulation / ROMS / 2025-05-21 ├── Salinity_Regular (Dataset: float32, shape=50x50x10) │# Store regular grid salinity data (50x50 spatial grid x 10 time steps) │# Metadata declare model version, unit, grid type │Attributes: {"model_version": "ROMS v4.0", "unit": "psu", "grid_type": "regular"} └── GridTopology_Unstructured (Dataset: int32, shape=50x50x10) # Store unstructured grid topology (adjacency relationship) # Metadata declare grid type Attributes: {"grid_type": "unstructured"} # Image data / Imaging / Satellite / 2025-05-21 ├── Image_Visible (Dataset: uint8, shape=1000x2000x3) │# Store visible light images (1000x2000 pixels x 3-channel RGB) │# Metadata declares image format, resolution, color space, geospatial coordinate link │Attributes: {"image_format": "PNG", "resolution": "1000x2000", "color_space": "RGB", "geospatial_coords": / Geospatial / Coordinates} └── Image_Infrared (Dataset: uint16, shape=1000x2000) # Store infrared band images (single channel) # Metadata declares image format, bit depth, geospatial coordinate link Attributes: {"image_format": "TIFF", "bit_depth": 16, "geospatial_coords": / Geospatial / Coordinates} # Video data / Video / Drone / 2025-05-21 ├── Video_Drone_Flight (Dataset: uint8, shape=100x1080x1920x3) │# Store aerial video (100 frames x 1080x1920 pixels x 3-channel RGB) │# Metadata declares video format, frame rate, resolution, geospatial coordinate link │Attributes: {"video_format": "H.264", "frame_rate": 30.0, "resolution": "1920x1080", "geospatial_coords": / Geospatial / Coordinates} └── Video_Metadata (Group) # Store video metadata group (optional extension) └── geospatial_coords (SoftLink to / Geospatial / Coordinates) # Index data / Index / TimeIndex (Dataset: datetime64, shape=1000) # Store time series index (e.g., 1000 time points) # For quickly locating data in the time dimension / Index / GeoIndex (Dataset: float32, shape=2x2) # Store geographic extent index (e.g., longitude [-180, 180] x latitude [-90, 90]) # For quickly filtering data in geographic regions S3, build distributed data association and indexing, including performing data standardization preprocessing and cross-source data association.

[0055] Design a standardization preprocessing process to uniformly convert raster data such as formatted satellite images, and resample time series data such as buoys to a unified time base unit system for normalization. Define a standard unit conversion rule library, declared through the `unit` attribute field.

[0056] Exemplarily, the data standardization preprocessing specifically includes: Format unified conversion: Decode raster data (such as satellite images) into a three-dimensional array of `Height×Width×Channels`, and expand multi-band images along the channel dimension; convert the original timestamps of time series data such as buoys (such as local time, GPS time) to UTC time, and interpolate data with non-uniform time steps to generate an equally spaced sequence.

[0057] Unit system normalization: Define a standard unit conversion rule library (such as temperature unified to °C, salinity unified to PSU), declared through the `unit` attribute field.

[0058] Exemplarily, cross-source data association includes performing spatio-temporal alignment on cross-source data and performing spatio-temporal matching of multi-source data. Specifically exemplarily, for spatio-temporal alignment of cross-source data, attach a `time_reference` attribute (UTC timestamp) and a `spatial_extent` attribute (longitude and latitude range) to all datasets, and perform spatio-temporal matching of multi-source data.

[0059] S4, optimize and expand the storage performance, including performing density-sensitive chunking and hierarchical compression.

[0060] Exemplarily, the density-sensitive chunking includes dynamically adjusting the HDF5 chunk size according to the grid node density, and at the same time dynamically adjusting the chunk dimension according to the data access pattern. Time series data is chunked along the time axis. Spatial raster data is chunked by longitude and latitude.

[0061] Exemplarily, the density-sensitive chunking specifically includes: Density-sensitive chunking: Dynamically adjust the HDF5 chunk size according to the grid node density. For high-density areas, use 32×32 chunks, and for sparse areas, use 128×128 chunks. At the same time, dynamically adjust the chunk dimension according to the data access pattern. Time series data is chunked along the time axis (e.g., chunk_shape=(100,1,1)). Spatial raster data is chunked by longitude and latitude (e.g., chunk_shape=(1,512,512)).

[0062] Node density is defined as: ; In the formula, is the total number of nodes in the current area, is the area of the region; Chunk size is defined as: ; Exemplarily, the hierarchical compression includes using a floating-point compression algorithm for floating-point scientific data to maintain numerical accuracy, and using lossless compression for coordinate data; applying JPEG2000 lossy compression to image / video data, and the compression ratio is configurable.

[0063] Another exemplarily, the hierarchical compression specifically includes: using the ZFP floating-point compression algorithm for floating-point scientific data to maintain numerical accuracy; using lossless compression for coordinate data; applying JPEG2000 lossy compression to image / video data, and the compression ratio is configurable; The ZFP algorithm realizes data compression through the following steps: Chunking. Divide the original data into multi-dimensional chunks of a fixed size (4×4×4 for three-dimensional chunks, 4×4 for two-dimensional chunks).

[0064] Quantization. Quantize the data in each chunk, that is, map the data to a smaller range to reduce precision loss.

[0065] Entropy coding. Use run-length coding or Huffman coding to perform lossless compression on the quantized coefficients to generate the final compressed data stream.

[0066] Embodiment 2. The embodiment of the present invention provides a multi-source marine environment data storage system based on HDF5. The system includes: Hierarchical storage framework construction module, used for constructing a three-level directory structure of domain-type-source, implementing a dynamic expansion mechanism for spatio-temporal dimensions, and spatio-temporal composite coding; HDF5 nested grouping storage architecture design module, used for defining grid topology structure, dynamically adaptive storage of resolution, processing multi-source data, and metadata management; Distributed data association and index construction module, used for preprocessing data standardization and cross-source data association; Optimization and expansion module, used for optimizing and expanding storage performance, including density-sensitive block division and hierarchical compression.

[0067] To further illustrate the relevant effects of the embodiments of the present invention, the following experiments are carried out.

[0068] Test example 1: Storage of polar region data. Extract and store the sea surface temperature data of the Arctic region (60°N - 90°N) for one month in the global ocean surface temperature. See Table 4; Table 4 Storage of polar region data

[0069] Aiming at the storage redundancy caused by the convergence of longitude and latitude in high-latitude regions, the grid resolution is dynamically adjusted through hierarchical resolution annotation, significantly reducing the storage size.

[0070] Test example 2: Cross-modal spatio-temporal query efficiency test. Retrieve "Numerical simulation of the equatorial Pacific Ocean in summer 2023 and corresponding satellite remote sensing temperature values", see Table 5; Table 5 Cross-modal spatio-temporal query efficiency test

[0071] The solution of the present invention uses a four-level directory structure to directly connect to data according to the path of / Oceanography / Observation / Satellite, and uses a spatio-temporal R-tree index to filter the target spatio-temporal range at one time, thus significantly reducing the query time and the number of I / O operations.

[0072] The above is only a relatively optimal specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A multi-source marine environment data storage method based on HDF5, characterized in that, The method includes the following steps: S1: Construct a hierarchical logical directory and coding system; taking domain - type - source as the main line, construct a three - level directory structure, and combine the dynamic expansion mechanism of the spatio - temporal dimension and the spatio - temporal composite coding method to form a unified data logical organization system, providing a directory index basis and refined time - space expression for subsequent data storage and retrieval; S2: Design a storage structure and multi - source adaptation mechanism; based on the constructed directory and spatio - temporal coding system, design an HDF5 nested grouping storage structure to support the topological definition storage of regular and irregular grids, dynamic resolution adaptive storage, and multi - modal data storage of images and videos; construct a metadata management mechanism to ensure data description consistency and cross - source compatibility; S3: Construct a data index and association mechanism for multi - source collaboration; aiming at the integration requirements of multi - source heterogeneous data, design a standardized pre - processing process, unify the time base and physical unit system, construct a distributed spatio - temporal index mechanism, and achieve accurate matching and efficient linked query of cross - source data; S4: Carry out performance optimization and storage compression strategies; after completing logical organization, physical storage, and data indexing, introduce density - sensitive block division and multi - level compression strategies, dynamically adjust the block division method according to grid density and access patterns, and adopt compression algorithms of ZFP and JPEG2000 to achieve high - performance and high - compression - ratio optimization of multi - source data storage.

2. The method for storing multi-source marine environmental data based on HDF5 according to claim 1, wherein In step S1, taking domain - type - source as the main line, construct a three - level directory structure, including: create multiple main groups under the root directory to represent the disciplinary domain attributes; establish subgroups of different types of products under each domain group; group by data source within the type subgroup to determine the data acquisition device type.

3. The method for storing multi-source marine environmental data based on HDF5 according to claim 1, characterized in that, In step S1, the dynamic expansion mechanism of the spatio - temporal dimension includes: Time - axis organization, nestedly create time - series directories by / Year / Month / Day within the data source group; store the timestamp array in / Index / TimeIndex, and establish the mapping relationship between the time directory and the data set; Spherical spatio - temporal hybrid index, introduce a latitude - dependent resolution function to dynamically adjust the grid cell size, and generate a unique grid code through the H3 geospatial grid algorithm; embed a spatio - temporal R - tree index structure in the HDF5 file, and the index node contains the time range, spatial grid code, and physical address of the data block; H3 hexagonal grid coding, adopt the H3 geospatial grid to divide the sphere into multi - resolution hexagonal cells; Hierarchical spatio - temporal coding, dynamically select the coding granularity and dimension according to the spatio - temporal characteristics of the data. At the same time, for high - dimensional data such as the profile data of the ocean and atmosphere, embed a vertical dimension identifier in the coding, which is associated with the depth of the ocean - atmosphere profile data.

4. The method for storing multi-source marine environmental data based on HDF5 according to claim 3, wherein The latitude - dependent resolution function is: ; In the formula, is the resolution at latitude , is the equatorial resolution, is the polar resolution, is the resolution switching threshold; Generating a unique grid code through the H3 geospatial grid algorithm includes: (1) Convert longitude and latitude to three - dimensional coordinates; ; ; ; In the formula, is the longitude; (2) Determine the resolution level, and determine the resolution level according to the latitude - dependent resolution function; (3) Calculate the center of the hexagon, calculate the center position of the hexagon formed by multiple nodes; Let the longitude and latitude coordinates of the hexagon vertices be: ; Then the geometric center is ; ; In the formula, and are respectively the longitude and latitude of the th vertex of the hexagon; (4) Generate the H3 code, and generate a 64 - bit H3 index according to the hexagon center position.

5. The multi-source marine environmental data storage method based on HDF5 according to claim 1, characterized in that, In step S2, the topology definition includes: storing the coordinate axes of the regular grid as an independent data set, which is linked to the main data set through attributes; storing the three-dimensional coordinate matrix of each node in the irregular grid, which is linked to the main data set through attributes; creating an `Adjacency_List` data set for the irregular grid, and recording the adjacent cell indices of each grid cell in a sparse matrix format; defining a grid_type field in the data set attributes and attaching a grid type description; creating an independent Group for each time step to store the grid topology and geometric data of that time step.

6. The method for storing multi-source marine environment data based on HDF5 according to claim 1, characterized in that In step S2, the dynamic resolution adaptation storage includes: storing the basic coordinates of the grid nodes in the form of a three-dimensional floating-point array, and declaring the coordinate system standard in the attributes; creating an auxiliary data set under the same group to store the planar projection coordinates and marking the projection method; the regular grid records the fixed resolution, the irregular grid divides the resolution identification by region, and records the differences between the equator and the poles in the attributes; The metadata management includes designing the basic metadata of the data set for different types of data, and the basic metadata of the data set includes scientific data metadata, image data metadata, and video data metadata.

7. The method for storing multi-source marine environmental data based on HDF5 according to claim 1, wherein In step S3, for the integration requirements of multi-source heterogeneous data, a standardized preprocessing process is designed to unify the time reference and physical unit system, including: designing a standardized preprocessing process to uniformly convert the raster data of the formatted satellite images, and resampling the time series data of the buoys to the unified time reference unit system for normalization; defining a standard unit conversion rule library, which is declared through the `unit` field of the attribute; implementing accurate matching and efficient linkage query of cross-source data, including spatio-temporal alignment of cross-source data and spatio-temporal matching of multi-source data.

8. The method for storing multi-source marine environment data based on HDF5 according to claim 1, wherein, In step S4, the density-sensitive block division includes: dynamically adjusting the HDF5 block size according to the grid node density, dynamically adjusting the block dimension according to the data access mode, and dividing the time series data by the time axis; dividing the spatial raster data by longitude and latitude; implementing high-performance and high-compression ratio storage optimization of multi-source data, including using the ZFP floating-point compression algorithm for lossless compression of the coordinate data of the floating-point scientific data; applying JPEG2000 lossy compression to the image / video data.

9. The method for storing multi-source marine environment data based on HDF5 according to claim 8, wherein, Node density is defined as: ; In the formula, is the total number of nodes in the current area, is the area of the area; Chunk size is defined as: 。 10. A multi-source marine environmental data storage system based on HDF5, characterized in that, The system implements the HDF5-based multi-source marine environment data storage method according to any one of claims 1-9, and the system includes: a hierarchical storage framework construction module for constructing a hierarchical logical directory and coding system; taking domain-type-source as the main line, constructing a three-level directory structure, and combining the spatio-temporal dimension dynamic expansion mechanism and the spatio-temporal composite coding method to form a unified data logical organization system, providing a directory index basis and refined time-space expression for subsequent data storage and retrieval; The HDF5 nested grouping storage architecture design module is used to design the storage structure and multi-source adaptation mechanism; based on the constructed directory and spatio-temporal coding system, design the HDF5 nested grouping storage structure, support the topological definition storage of regular and irregular grids, dynamic resolution adaptive storage, and multi-modal data storage of images and videos; construct a metadata management mechanism to ensure data description consistency and cross-source compatibility; The distributed data association and index construction module is used to construct a data index and association mechanism for multi-source collaboration; for the integration requirements of multi-source heterogeneous data, design a standardized preprocessing process, unify the time reference and physical unit system, construct a distributed spatio-temporal index mechanism, and achieve accurate matching and efficient linkage query of cross-source data; The optimization and extension module is used to carry out performance optimization and storage compression strategies; after completing the logical organization, physical storage, and data index, introduce density-sensitive block division and multi-level compression strategies, dynamically adjust the block division method according to grid density and access patterns, and adopt the compression algorithms of ZFP and JPEG2000 to achieve the optimization of multi-source data storage with high performance and high compression ratio.

Citation Information

Patent Citations

  • Efficient distributed organization and management method for mass remote sensing data

    CN101339570A

  • Implement method of lightweight-class global multi-dimensional remote-sensing image network map service

    CN103455624A

  • A three-dimensional data encoding and storing method for massive ocean environment data management

    CN109885572A

  • Sea area port operation condition analysis method and device, and electronic equipment

    CN115563774A

  • Multi-source heterogeneous data management method and system, storage medium and electronic equipment

    CN115935016A

Cited By

  • Multi-modal data enhanced storage system based on Parquet format

    CN121029808A

  • Multi-resolution grid data real-time online visualization method and device

    CN121033329A

  • A method and apparatus for real-time online visualization of multi-resolution grid data

    CN121033329B

  • Block storage method suitable for time domain finite difference method electrically large size grid

    CN121070286A

  • A block storage method suitable for time domain finite difference method electric large size grid

    CN121070286B