A multi-source marine environment data intelligent sharing method

By employing a multi-source intelligent sharing method for marine environmental data, combined with MinIO distributed storage and intelligent subscription and distribution technology, the flexibility and real-time issues of marine data storage and distribution in traditional technologies have been resolved, enabling efficient management and rapid distribution of marine environmental data.

CN119557483BActive Publication Date: 2025-10-24NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411494340.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2025-10-24
Estimated Expiration
2044-10-24

AI Technical Summary

Technical Problem

Traditional marine environmental data sharing technologies are insufficient to meet the needs of efficient management and rapid distribution. Centralized storage systems lack flexibility and real-time performance, and cannot effectively support the diversity and massive volume of marine data.

Method used

A multi-source marine environmental data intelligent sharing method is adopted, including data product adaptation, distributed storage and intelligent subscription distribution. Through the MinIO distributed storage system and intelligent subscription distribution technology, efficient data storage and real-time sharing are achieved.

Benefits of technology

It improves the real-time nature and flexibility of data acquisition, significantly enhances data sharing efficiency and distribution speed, and supports efficient management and personalized subscription of multi-source marine environmental data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557483B_ABST
    Figure CN119557483B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-source marine environment data intelligent sharing method, this method is by obtaining multi-source marine environment data;Multi-source marine environment data is adapted to product, and first multi-source marine environment data is obtained, product adaptation is used to indicate that multi-source marine environment data meets user demand;First multi-source marine environment data is packed, and second multi-source marine environment data is obtained;Second multi-source marine environment data is stored in distribution, and the multi-source marine environment data after storage is obtained;Through intelligent subscription distribution, the data intelligent sharing of the multi-source marine environment data after storage is carried out.The application can improve data sharing efficiency and distribution speed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data intelligent sharing, in particular to a multi-source marine environment data intelligent sharing method. BACKGROUND

[0002] The rapid growth of marine environment data and technological progress have brought an urgent need for efficient data sharing and management. Marine environment data not only includes historical data, but also covers multiple types of data such as marine geology and survey historical data, basic information products, thematic products, and early warning products. These data have important value for marine scientific research, resource development, environmental monitoring, and disaster warning.

[0003] However, the diversity and massiveness of marine data pose challenges to data storage, processing, and sharing. Traditional technologies mainly rely on centralized storage and manual distribution systems. These systems usually use a single data center to store and manage massive data and transmit data through FTP, HTTP, and other protocols. Centralized storage systems usually use large database management systems (such as Oracle, MySQL) to manage data, and file system storage (such as NFS, HDFS) to realize centralized access to data. In addition, data distribution methods usually push data from the center node to the user end or edge node through manual or timed tasks, lacking flexibility and real-time performance. Therefore, traditional data sharing technologies cannot meet the current needs of efficient management and rapid distribution of marine data. SUMMARY

[0004] The present application aims to provide a multi-source marine environment data intelligent sharing method to improve data sharing efficiency and distribution speed.

[0005] In a first aspect, the embodiments of the present application provide a multi-source marine environment data intelligent sharing method, which comprises:

[0006] Obtaining multi-source marine environment data;

[0007] Product adaptation is performed on the multi-source marine environment data to obtain first multi-source marine environment data, wherein the product adaptation is used to indicate multi-source marine environment data that meets user needs;

[0008] The first multi-source marine environment data is packaged to obtain second multi-source marine environment data;

[0009] The second multi-source marine environment data is stored in a distributed manner to obtain stored multi-source marine environment data;

[0010] The stored multi-source marine environment data is intelligently shared through intelligent subscription distribution.

[0011] In some embodiments, the data intelligent sharing of the stored multi-source marine environment data through intelligent subscription distribution comprises:

[0012] obtaining a user request containing a plurality of conditions, and obtaining metadata corresponding to each data record in the stored multi-source marine environment data and the conditions;

[0013] based on the metadata, calculating a matching degree between each data record and each condition in the user request;

[0014] determining a marine environment subscription data set meeting a plurality of conditions in the user request according to the matching degree;

[0015] data intelligent sharing of the marine environment subscription data set through a content distribution strategy.

[0016] In some embodiments, the matching degree between each data record and each condition in the user request is calculated based on the metadata in the following manner:

[0017]

[0018] wherein M j represents the matching degree, Q represents the total number of conditions in the user request, w i represents the importance weight of each condition, q i represents the i-th condition, m ij represents the metadata, and f(·) represents a matching function.

[0019] In some embodiments, the data intelligent sharing of the marine environment subscription data set through the content distribution strategy comprises:

[0020] obtaining a product theme subscribed by a user;

[0021] calculating a correlation between the product theme and the marine environment subscription data set to obtain a correlation calculation result;

[0022] determining a target push data set according to the correlation calculation result;

[0023] data intelligent sharing of the target push data set.

[0024] In some embodiments, the data intelligent sharing of the target push data set comprises:

[0025] data intelligent sharing of the target push data set in a real-time push manner; or, data intelligent sharing of the target push data set in a batch distribution manner.

[0026] In some embodiments, the correlation between the product topic and the marine environment subscription dataset is calculated by:

[0027]

[0028] wherein R new represents the correlation, n represents the total amount of data in the marine environment subscription dataset, a k represents the corresponding weight, s k (·) represents the correlation function, D new represents the marine environment subscription dataset, and S represents the product topic.

[0029] In some embodiments, the distributed storage of the second multi-source marine environment data comprises:

[0030] The second multi-source marine environment data is distributed stored by MinIO.

[0031] In some embodiments, after the data intelligent sharing of the stored multi-source marine environment data by intelligent subscription distribution, the method further comprises:

[0032] If new data is obtained or data is updated, a data update is triggered by an event-driven mechanism;

[0033] The multi-source marine environment data is synchronously updated by incremental synchronization.

[0034] In some embodiments, the multi-source marine environment data is synchronously updated by incremental synchronization by:

[0035] AB = B (t + Δt) - B (t)

[0036]

[0037] wherein AB represents the incremental data, B (t + Δt) represents the data amount at t + Δt, B (t) represents the data amount at the current time t, V sync represents the total amount of synchronous data blocks, N represents the number of data blocks to be synchronized, and AB j represents the jth data block to be synchronized.

[0038] In some embodiments, after the multi-source marine environment data is synchronously updated by incremental synchronization, the method further comprises:

[0039] The data that is synchronously updated is consistency-verified by a hash function.

[0040] Compared with the prior art, the present application has the following beneficial effects:

[0041] The method obtains first multi-source marine environment data by product adaptation on multi-source marine environment data; obtains second multi-source marine environment data by data packaging on the first multi-source marine environment data; obtains stored multi-source marine environment data by distributed storage on the second multi-source marine environment data; and performs data intelligent sharing on the stored multi-source marine environment data by intelligent subscription distribution. In this way, by performing distributed storage on multi-source marine environment data, efficient storage and management of massive marine environment data can be realized, and then by performing data intelligent sharing on multi-source marine environment data by intelligent subscription distribution, the real-time performance and flexibility of data acquisition can be significantly improved. Thus, the data sharing efficiency and distribution speed can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0042] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the description of the embodiments, taken in conjunction with the following drawings in which:

[0043] Figure 1 is a flowchart of an embodiment of the multi-source marine environment data intelligent sharing method provided by the present application;

[0044] Figure 2 is a structural schematic diagram of the best embodiment of the multi-source marine environment data intelligent sharing method provided by the present application;

[0045] Figure 3 is a flowchart of the best embodiment of the multi-source marine environment data intelligent sharing method provided by the present application. DETAILED DESCRIPTION

[0046] The embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar notations represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as a limitation of the present application.

[0047] In the description of the present application, if there is a description of first, second, etc., it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features or the sequence of indicated technical features.

[0048] In the description of the present application, it is to be understood that the orientation description, such as the orientation or position relationship indicated by up, down, etc., is based on the orientation or position relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, and is not to indicate or imply that the indicated device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.

[0049] In the description of the present application, it should be noted that, unless otherwise explicitly defined, the words such as setting, installing, connecting, etc. should be understood broadly, and the person skilled in the art can determine the specific meaning of the above words in the present application in combination with the specific content of the technical solution.

[0050] The rapid growth of marine environmental data and technological progress has brought an urgent need for efficient data sharing and management. Marine environmental data not only includes historical data, but also covers marine geology and mapping historical data, basic information products, thematic products, and early warning products, etc. These data have important value for marine scientific research, resource development, environmental monitoring, and disaster warning, etc.

[0051] However, the diversity and massiveness of marine data pose challenges to data storage, processing, and sharing. Traditional technologies mainly rely on centralized storage and manual distribution systems. These systems usually use a single data center to store and manage massive data, and transmit data through FTP, HTTP, etc. Centralized storage systems usually use large database management systems (such as Oracle, MySQL) to manage data, and use file system storage (such as NFS, HDFS) to realize centralized access of data. In addition, the data distribution method usually pushes data from the center node to the user end or edge node through manual or timing tasks, which lacks flexibility and real-time performance. Therefore, the traditional data sharing technology is difficult to meet the current needs of efficient management and rapid distribution of marine data.

[0052] To solve the problem that the traditional data sharing technology is difficult to meet the current needs of efficient management and rapid distribution of marine data, the present application proposes a multi-source marine environmental data intelligent sharing method.

[0053] Reference Figure 1 The present application provides a multi-source marine environmental data intelligent sharing method, which comprises the following steps:

[0054] Step S100, acquiring multi-source marine environmental data;

[0055] Step S200, product adaptation of multi-source marine environmental data, to obtain first multi-source marine environmental data, product adaptation is used to indicate multi-source marine environmental data meeting user needs;

[0056] Step S300, data packaging of the first multi-source marine environmental data, to obtain second multi-source marine environmental data;

[0057] Step S400, distributed storage of the second multi-source marine environmental data, to obtain stored multi-source marine environmental data;

[0058] In step S500, the stored multi-source marine environment data is intelligently shared through intelligent subscription distribution.

[0059] In this embodiment, the multi-source marine environment data is product-adapted to obtain first multi-source marine environment data; the first multi-source marine environment data is packaged to obtain second multi-source marine environment data; the second multi-source marine environment data is distributed stored to obtain stored multi-source marine environment data; and the stored multi-source marine environment data is intelligently shared through intelligent subscription distribution. In this way, the distributed storage of the multi-source marine environment data can realize efficient storage and management of massive marine environment data, and the intelligent sharing of the multi-source marine environment data through intelligent subscription distribution can significantly improve the real-time performance and flexibility of data acquisition. Thus, the data sharing efficiency and distribution speed can be improved.

[0060] The product adaptation of the multi-source marine environment data can be format adaptation, element adaptation, and regional adaptation, etc.

[0061] The data packaging of the first multi-source marine environment data can be data packaging of the first multi-source marine environment data through compression tools such as ZIP and TAR.GZ, etc., to reduce the data volume and improve the transmission efficiency.

[0062] In some embodiments, the intelligent sharing of the stored multi-source marine environment data through intelligent subscription distribution comprises:

[0063] Obtaining a user request containing multiple conditions, and obtaining metadata corresponding to each data record in the stored multi-source marine environment data and the conditions;

[0064] Based on the metadata, calculating the matching degree between each data record and each condition in the user request;

[0065] According to the matching degree, determining a marine environment subscription data set that meets the multiple conditions in the user request;

[0066] Intelligently sharing the marine environment subscription data set through a content distribution strategy.

[0067] In this embodiment, the matching degree between each data record and each condition in the user request is calculated, then the marine environment subscription data set that meets the multiple conditions in the user request is determined according to the matching degree, and finally the marine environment subscription data set is intelligently shared through a content distribution strategy. In this way, the marine environment subscription data set is determined according to the conditions in the user request, which can better meet the user's needs, and the user can subscribe to data products of specific types or topics.

[0068] In some embodiments, the matching degree between each data record and each condition in the user request is calculated based on the metadata in the following manner:

[0069]

[0070] wherein M j represents the matching degree, Q represents the total number of conditions in the user request, w i represents the importance weight of each condition, q i represents the i-th condition, m ij represents the metadata, and f(·) represents the matching function.

[0071] In some embodiments, the marine environment subscription dataset is intelligently shared by a content distribution strategy, including:

[0072] obtaining a product topic subscribed by a user;

[0073] calculating the correlation between the product topic and the marine environment subscription dataset to obtain a correlation calculation result;

[0074] determining a target push dataset according to the correlation calculation result;

[0075] intelligently sharing the target push dataset.

[0076] In this embodiment, after obtaining the marine environment subscription dataset through the conditions given by the user, the correlation between the product topic and the marine environment subscription dataset is calculated, so that data meeting the user's demand can be further obtained. Then, according to the correlation calculation result, the target push dataset (i.e., the data that needs to be distributed and shared) is determined, so that customized data products can be quickly obtained to meet individual needs.

[0077] In some embodiments, the target push dataset is intelligently shared, including:

[0078] The target push dataset is intelligently shared in a real-time push manner; or, the target push dataset is intelligently shared in a batch distribution manner.

[0079] In this embodiment, the target push dataset is intelligently shared, and the optimal distribution path can be selected according to the geographical location, network condition and data demand of the user. The system automatically manages the push and distribution of data according to the subscription information, supports real-time data push and regular batch distribution, and significantly improves the real-time performance and flexibility of data acquisition.

[0080] In some embodiments, the correlation between the product topic and the marine environment subscription dataset is calculated in the following manner:

[0081]

[0082] wherein R new represents the correlation, n represents the total amount of data in the marine environment subscription dataset, a k represents the corresponding weight, s k (·) represents the correlation function, D new represents the marine environment subscription dataset, and S represents the product theme.

[0083] In some embodiments, the second multi-source marine environment data is distributedly stored, comprising:

[0084] The second multi-source marine environment data is distributedly stored by MinIO.

[0085] In this embodiment, the MinIO distributed storage system is used to realize efficient storage and management of massive marine environment data (i.e., the second multi-source marine environment data). Through the distributed architecture, the system can be horizontally expanded to support parallel storage and access of large-scale data, solving the scalability problem of traditional centralized storage.

[0086] In some embodiments, after the stored multi-source marine environment data is intelligently shared through intelligent subscription distribution, the method further comprises:

[0087] If new data is obtained or data is updated, the data update is triggered through an event-driven mechanism;

[0088] The multi-source marine environment data is synchronously updated through incremental synchronization.

[0089] In this embodiment, it is ensured that when new data is generated or existing data is updated, the synchronization process can be quickly responded and started. Through the event-driven mechanism (such as webhook or message queue event), the data synchronization task can be triggered in real time, ensuring that the data in the system is synchronized with the source data. Incremental synchronization is one of the core technologies of this process, through which the system can identify the newly added or updated data since the last synchronization by comparing the version or timestamp of the data.

[0090] In some embodiments, the multi-source marine environment data is synchronously updated through incremental synchronization in the following manner:

[0091] AB = B (t + At) - B (t)

[0092]

[0093] wherein AB represents the incremental data, B (t + At) represents the data amount at t + At, B (t) represents the data amount at the current time t, V syncN represents the number of data blocks to be synchronized, and ΔB represents the total amount of synchronized data blocks. j The jth data block to be synchronized is represented.

[0094] In some embodiments, after synchronously updating the multi-source marine environmental data through incremental synchronization, the method further comprises:

[0095] Consistency check is performed on the data updated by the hash function.

[0096] In this embodiment, after the data synchronization is completed, consistency check is performed on the synchronized data periodically to ensure that the data between the cloud and the edge node remains consistent.

[0097] For the convenience of those skilled in the art, a set of best embodiments is provided as follows:

[0098] The rapid growth of marine environmental data and technological progress have brought an urgent need for efficient data sharing and management. Marine environmental data not only includes historical data, but also covers marine geology and mapping historical data, basic information products, thematic products, and early warning products. These data have important value for marine scientific research, resource development, environmental monitoring, and disaster warning.

[0099] However, the diversity and massiveness of marine data pose challenges to data storage, processing, and sharing. Traditional data sharing technologies are difficult to meet the current needs of efficient management and rapid distribution of marine data. Traditional technologies mainly rely on centralized storage and manual distribution systems. These systems usually use a single data center to store and manage massive data and transmit data through FTP, HTTP, and other protocols. Centralized storage systems usually use large database management systems (such as Oracle, MySQL) to manage data, and cooperate with file system storage (such as NFS, HDFS) to realize centralized access to data. In addition, the data distribution method is usually to manually or through a timing task to push data from the center node to the user end or edge node, which lacks flexibility and real-time performance.

[0100] Therefore, it is necessary to develop intelligent sharing technology for multi-source massive marine environmental data to achieve efficient integration, secure transmission, and convenient access of data. The core of the technical background is how to build an intelligent sharing platform that can adapt to the characteristics of marine data. This platform needs to be able to process and distribute multi-source data, support secure data exchange, and meet the needs of different users for data access and services. With the development of marine observation technology, the real-time and accuracy requirements of data are becoming higher and higher, which requires the sharing platform not only to be able to quickly respond to user needs, but also to ensure the synchronization and consistency of data. In addition, the marine data sharing platform also needs to have high flexibility and customizability to adapt to the needs of different users and different application scenarios. This includes providing multiple service modes such as data download, online data service, subscription push, customized processing, and emergency service. At the same time, the platform needs to support multiple query methods, including by element type, date, regional range, and combined conditions, to improve the efficiency of user data retrieval. Therefore, the technical background of multi-source massive marine environmental data intelligent sharing technology is multifaceted, involving data storage, processing, sharing, security, and application at multiple levels. These background factors collectively drive the development and innovation of marine data sharing technology to meet the growing demand for data resources in the marine field.

[0101] Current research on multi-source massive marine environmental data sharing technology faces several key problems and challenges, which are particularly prominent in the context of big data and seriously affect the efficiency of data management, sharing, and application. By combining the application of MinIO distributed storage technology, intelligent subscription distribution technology, and data synchronization update technology, the following aspects can be explored to further discuss the problems existing in current research:

[0102] (1) Data sources are diverse and heterogeneous in format.

[0103] Marine environmental data comes from multiple sources, including satellite observations, buoy monitoring, unmanned ship measurements, and weather radars, etc. These data often use different storage formats and standards. For example, satellite data may exist in NetCDF or GeoTIFF format, while buoy data may be stored in time series form, resulting in serious data heterogeneity problems. Current research faces the challenge of how to standardize and manage these multi-source heterogeneous data. Existing solutions are inefficient in handling massive data and lack good support for complex data formats.

[0104] (2) Bottlenecks in massive data storage and management.

[0105] With the development of marine environmental monitoring technology, the amount of data is growing exponentially, and data storage and management have become a core problem. Traditional centralized storage architecture is difficult to meet the current storage needs of marine environmental big data, with poor scalability and often delayed and performance bottlenecks in data access. Especially in the face of large-scale data concurrent reading and writing, the processing capacity of existing systems is limited, which can easily cause system overload or slow response speed, and cannot effectively support real-time data sharing and application.

[0106] (3) Low efficiency of real-time data acquisition and distribution.

[0107] Marine environmental monitoring requires real-time acquisition and analysis of the latest marine state data, such as sea temperature, salinity, wind speed, etc., but existing data distribution systems often use timed updates or batch pushing methods, which cannot efficiently support real-time data subscription and distribution. Especially in the scenario of multiple users requesting data at the same time, the system is prone to congestion, and the delay of data pushing is large, which is difficult to meet the needs of real-time decision support application scenarios.

[0108] (4) Difficulty in maintaining data synchronization and consistency.

[0109] Marine environmental data sharing not only involves a large number of distributed data stored in different places, but also needs to ensure the synchronization and consistency of data between different nodes. Traditional data synchronization methods are inefficient, especially in incremental data updates and data conflict processing, which lack effective mechanisms, making it difficult to guarantee the timeliness and accuracy of data. In the process of multi-source data synchronization, data consistency and integrity are often difficult to maintain effectively, especially in the case of unstable network or node failure, data loss and asynchronization problems are more prominent.

[0110] (5) Technical gap in edge computing and cloud-edge collaboration.

[0111] With the increase in the number of marine data collection points and the wide distribution, massive data needs to be preprocessed and stored locally to reduce the burden on the main data center. However, existing research on the application of edge computing and cloud-edge collaboration is relatively limited, and how to effectively utilize edge nodes for data processing and synchronization remains a problem to be solved. Traditional centralized processing architecture is inefficient in this scenario, and cannot fully utilize the advantages of edge computing to improve the timeliness and sharing efficiency of data processing.

[0112] (6) Lack of effective sharing and collaboration mechanisms.

[0113] In the research of multi-source massive marine environmental data sharing, different users and institutions often have different data needs, which leads to poor flexibility and customizability of data sharing. Existing systems lack intelligent subscription and distribution mechanisms, making it difficult for users to flexibly subscribe to specific data content or topics based on their needs, limiting the efficient use of data. At the same time, cross-institutional collaboration and data sharing face compatibility issues such as data format and interface standards, affecting data flow and cooperation.

[0114] The technical solution of the present embodiment can effectively address the problems existing in current research by combining MinIO distributed storage, intelligent subscription and distribution technology, and data synchronization update technology. The application of these technologies not only solves the heterogeneity and storage bottleneck problems of multi-source data, but also significantly improves the real-time performance and transmission efficiency of data sharing through efficient subscription and distribution and edge computing support. At the same time, the data encryption and access control functions of MinIO enhance the security of the data sharing process, providing comprehensive technical support for the collaborative sharing of multi-source massive marine environmental data. The specific technical solution of the present embodiment includes:

[0115] The multi-source massive marine environmental data intelligent sharing technology (i.e., the multi-source marine environmental data intelligent sharing method) constructs a data sharing exchange directory and a sharing exchange scheduling strategy, and provides sharing services after adapting and processing marine environmental application thematic products. It also supports distributed data synchronization update services. The data sharing exchange directory is used to develop exchange scheduling strategies, ensuring product identification, product configuration processing, and sharing packaging and arrangement. It can realize directory data browsing, data retrieval engine services, demand service processing, subscription distribution management, and content distribution services for various computing and storage environments, supporting marine environmental application thematic product exchange. In addition, it supports distributed data synchronization update services, which can realize data synchronization mechanism construction, synchronization monitoring, and synchronization update functions. The sharing exchange information includes marine environmental historical data, marine geology and surveying historical data, basic information product data, thematic product data, and early warning product data.

[0116] The multi-source massive marine environmental data intelligent sharing technology mainly includes marine environmental directory services, marine environmental data sharing services, and marine environmental data synchronization update services, as shown in Figure 2 .

[0117] The processing flow of the multi-source massive marine environmental data intelligent sharing technology is shown in Figure 3 . The detailed process of the multi-source massive marine environmental data intelligent sharing technology is as follows:

[0118] (1) Multi-source massive marine environmental data (i.e., multi-source marine environmental data) acquisition.

[0119] The sources of multi-source massive marine environmental data and fishery operation data are the marine environment databases in the prior art known to those skilled in the art, for example, the Nation Oceanic and Atmospheric Administration (NOAA) environmental database, the Columbia environmental database, the Copernicus Marine Environment Monitoring Service (CMEMS), and the Western and Central Pacific Fisheries Commission (WCPFC). The types of marine environmental data include chlorophyll concentration, sea surface height, and sea surface temperature, and the fishery operation data is South Pacific longline data, as well as fishery data provided by relevant company groups in provinces and cities. The downloaded data formats are.csv,.xls, and.nc, etc. The data is derived from high-resolution satellite remote sensing, ensuring the accuracy of data acquisition, and providing a prerequisite for the quality of experimental data in the later stage.

[0120] There are various methods and ways to obtain multi-source marine data, involving satellite remote sensing, buoy monitoring, ship observation, unmanned platforms, and marine model simulation. Satellite remote sensing technology is one of the main means to obtain large-scale, long-term continuous marine data. Through the multi-spectral and hyperspectral imaging equipment carried on meteorological and ocean observation satellites, various parameters such as sea surface temperature, sea surface height, sea ice distribution, and marine pigment concentration can be obtained. These satellites include the Fengyun series, marine satellites, and Copernicus satellites, which can provide global marine monitoring data.

[0121] The buoy monitoring network is another important data source. Fixed or drifting buoys deployed in the ocean can collect real-time physical, chemical, and biological data of the ocean surface and deep layers, such as temperature-salinity-depth (CTD), dissolved oxygen, and nutrient salt concentration. Through international buoy networks such as the ARGO program, these buoys can obtain long-term, automated continuous monitoring data of global marine environment. In addition, ship observations in coastal and ocean areas are also an important way to obtain marine data, especially in areas with dense shipping, ships can obtain fine local marine weather, sea waves, and current information through onboard observation equipment.

[0122] The application of unmanned platforms is also increasing, especially the use of unmanned aerial vehicles, unmanned surface vehicles (USVs), and autonomous underwater vehicles (AUVs). These platforms have flexibility and high efficiency, and can penetrate into areas that are difficult for humans to reach or extreme climate conditions to operate, obtaining high-resolution regional marine environmental data. On the other hand, with the advancement of computer technology, marine model simulation has become an important supplement to obtaining marine data.

[0123] By combining meteorological data, historical ocean observation data, and numerical models, future trends in ocean environmental changes such as ocean circulation, tides, and climate change impacts on the ocean can be simulated.

[0124] International ocean data exchange cooperation is also an important way. By participating in global or regional ocean data sharing platforms and plans, such as the Global Ocean Observing System (GOOS) and the Intergovernmental Oceanographic Commission (IOC) of UNESCO's Ocean Data Sharing Plan, rich international ocean monitoring data can be obtained. These data provide basic support for weather forecasting, environmental assessment, and marine resource development.

[0125] Due to the large amount of data and the long period of time involved, manually downloading each piece of data is time-consuming and meaningless. Therefore, a simple data acquisition crawler code is written to automatically download the required data and store it in the local database. The data acquisition method can obtain various environmental data information from the official website, including data acquisition sources, data query formats (grid or table), data provision time periods, and time intervals. The time unit can be in days, weeks, or months.

[0126] (2) Ocean environmental product adaptation.

[0127] According to specific research needs, the obtained multi-source ocean data and products are adapted and structured data is generated and packaged, mainly divided into two stages: product adaptation processing and data packaging arrangement.

[0128] In the product adaptation processing stage, format adaptation, element adaptation, and regional adaptation are required. Since the data obtained comes from different sources such as satellite observations, buoy monitoring, and ship observations, the formats are also different, such as HDF, NetCDF, and GRIB. In order to meet the needs of users, it is necessary to convert the data products into formats specified by users, such as CSV, JSON, or NetCDF. This process is carried out through the use of data conversion tools such as Pandas, NumPy, etc. for format conversion and data preprocessing. According to the specific needs of users, element adaptation is carried out to filter out key marine environmental elements such as temperature, salinity, and flow rate from the data products, and these elements are adapted. Element adaptation not only includes filtering, but also may involve unit conversion (such as conversion between Celsius and Fahrenheit), data aggregation (such as calculation of average values in time or space), and statistical calculations (such as extreme values, coefficient of variation, etc.). Through these operations, it is ensured that users can obtain the key elements they are interested in, and the data meets their use standards. Regional adaptation is also required. Users may only be interested in a specific geographical area, so it is necessary to crop or resample the data to the specified area. This can be done by using geospatial processing tools such as GDAL to crop global or large-scale data in the data product to the area range required by the user. For data involving different spatial resolutions, resampling may also be required to ensure that data from different sources can be effectively integrated.

[0129] After completing the data adaptation processing, the first multi-source marine environmental data is obtained, and then enters the data packaging and arrangement stage. Data packaging is to facilitate the storage and transmission of data, and the data products after data adaptation processing are packaged into a single file or file set to obtain the second multi-source marine environmental data. Specifically, compression tools such as ZIP, TAR.GZ, etc. can be used to reduce data volume and improve transmission efficiency. Especially when transmitting large amounts or large-scale data, compression and packaging can significantly reduce network bandwidth occupancy. In order to ensure good traceability of data during transmission and use, metadata files need to be generated for the packaged data. Metadata records the details of the data source, processing steps, format conversion process, and adaptation operation, ensuring that users can clearly understand the generation and processing process of the data after receiving it, and can correctly use and interpret the data.

[0130] (3) MinIO distributed storage.

[0131] MinIO is a high-performance distributed object storage system compatible with Amazon S3 API, especially suitable for handling massive unstructured data such as marine environmental data. In the context of multi-source massive marine environmental data processing, MinIO provides a reliable, efficient, and flexible distributed storage architecture, data management method, and security guarantee, which can meet the storage and access needs of massive data.

[0132] The distributed storage architecture of MinIO is suitable for the distributed storage and high-concurrency access requirements of marine data. A MinIO cluster is composed of multiple independent storage nodes, each running a MinIO instance, and all nodes are connected through the network to form a complete cluster architecture. Each node can handle read and write operations on data simultaneously, allowing the system to scale horizontally by adding nodes, thus maintaining high availability and performance even as marine environmental data continues to grow. MinIO uses erasure coding technology to shard and redundantly store data across multiple nodes, so that even if some nodes fail, the system can still recover data through the remaining data blocks and redundant blocks to ensure data integrity and availability. For example, a marine observation data object can be divided into multiple data blocks and redundant blocks stored on different nodes, and even if some nodes fail, the data can still be reconstructed through the redundant data on other nodes. This design not only ensures the high availability of data, but also greatly improves the disaster recovery capability and stability of the system.

[0133] In terms of data storage and access, MinIO provides S3-compatible APIs, allowing users to manage and access massive marine environmental data through standard S3 operations. Through S3-compatible PUT, GET, and DELETE operations, users can easily integrate MinIO into existing S3 clients and tools, making it convenient to store and read marine data. MinIO uses an object storage model, which allows each marine data object to be stored as an object. Data objects contain data content and related metadata such as data identifiers, data types, timestamps, and geographic locations. This storage method is particularly suitable for multi-source marine data, as each data object can flexibly store different data information, meeting the management needs of diverse and complex data structures.

[0134] MinIO has excellent data management and expansion capabilities, making it particularly outstanding in handling marine data. Through consistent hashing algorithms, MinIO automatically distributes data evenly across nodes in the cluster, achieving load balancing and avoiding performance bottlenecks on a single node. Suppose a hash function H is used to map nodes and data to a hash space. For a node set N = {n1, n2,..., n k} and a data set D = {d1, d2,..., d m}, the consistent hashing algorithm steps are as follows:

[0135] The first step is to construct the hash ring. The node identifier (such as the node's IP address) is converted into a hash value through a hash function, and the hash value is placed at a certain position in the hash ring. Similarly, the identifier of the data object (such as the data key) is also mapped to the hash ring through the same hash function. The node hash value conversion formula is:

[0136] H(n i )=hash(n i )

[0137] Where H(n i ) is node n i The hash value of .

[0138] The data object hash value conversion formula is:

[0139] H(d j )=hash(d j )

[0140] Where H(d j ) is the data object, d j The hash value of .

[0141] The second step is the data mapping rule. Each data object is assigned to the first node whose hash value is greater than or equal to its hash value by searching clockwise. This means that the data object is stored on the node closest to it.

[0142] The third step is to redistribute data after the node changes. When a new node joins, only the data in front of the new node needs to be redistributed to the node, while other data will not be affected. Similarly, when a node leaves, only the part of the data that the node is responsible for needs to be redistributed to the next node. j , find the first node n in the clockwise direction i , such that:

[0143] H(n i )≥H(d j )

[0144] If no H(d j ) nodes, they are assigned to the first node on the ring. When new nodes join the cluster or old nodes go offline, the system automatically adjusts the data distribution to maintain the overall load balance and performance optimization of the cluster. This automatic load balancing capability is particularly critical for storing data in dynamically changing marine environments. By simply adding nodes, MinIO can easily expand storage capacity and throughput, which is particularly suitable for the deployment needs of marine data centers. MinIO supports deployment across multiple data centers, which provides strong guarantees for geographic redundancy and disaster recovery, ensuring secure storage and timely access to marine data worldwide.

[0145] In terms of data security, MinIO also provides a comprehensive solution to address the sensitivity and high security requirements of ocean data. MinIO supports the TLS / SSL encryption protocol to ensure that data is not tampered with or leaked during transmission. When storing static data, MinIO provides server-side encryption, supporting strong encryption algorithms such as AES-256, and uses external key management services (such as HashiCorp Vault) to manage keys securely. This means that even if the storage device itself is physically damaged or stolen, the data remains encrypted, ensuring data security. In addition, MinIO integrates a role-based access control mechanism (RBAC), allowing users to control access to data through policy files, effectively managing the permissions between users and systems to ensure data compliance and security.

[0146] MinIO has significant advantages in the storage and management of multi-source massive ocean environmental data. Its distributed architecture ensures high availability, scalability, and disaster recovery capabilities, while the S3-compatible API and object storage model provide flexible data management methods. Strong load balancing and data security mechanisms ensure system stability and data security. These features make MinIO an ideal choice for handling and storing massive ocean data, especially in complex and dynamic ocean environmental data application scenarios, providing efficient and reliable support.

[0147] (4) Intelligent subscription distribution technology and data synchronization update.

[0148] After storing multi-source massive ocean data in MinIO, data calling, distribution, and synchronization updates need to be performed according to user instructions and requirements. This embodiment integrates intelligent subscription distribution technology and data synchronization update technology into the entire multi-source massive ocean environmental data intelligent sharing technology system, optimizes user data demand experience, and facilitates the intelligent sharing of massive ocean data.

[0149] In the distributed object storage system MinIO, data products are extracted through Amazon S3-compatible APIs, and the system reads user-requested ocean environmental data objects from different nodes. At the same time, to ensure data integrity and consistency, MinIO combines consistent hashing and erasure coding techniques for data consistency checks. Once data blocks are found to be missing or damaged, the system will automatically start data repair mechanisms to ensure data reliability and integrity. This is crucial for accurate analysis of ocean environmental data, ensuring that scientists and users obtain accurate ocean observation data.

[0150] Assuming that the user request contains Q conditions, each condition can be represented as qi Each data record D j in the database has metadata m ij corresponding to the conditions (such as time range, geographic area, data type, etc.). i The matching function f(q ij ,m j ) is used to determine whether the data record D i satisfies the user condition q j .

[0151] The matching degree M j of the data record D i to the user request can be represented as:

[0152]

[0153] where w j is the importance weight of each condition, usually satisfying The final returned directory C is the collection of data records that satisfy M j >T, where T is a set threshold: C = {D j |M i >T}.

[0154] In data retrieval, the matching function f(q ij ,m j ) is used to determine whether the data record D i satisfies the user condition q ij . The specific matching function can vary based on different data types and retrieval requirements, and the present embodiment does not make specific limitations. Two examples are shown as follows:

[0155] (1) Interval matching function - suitable for matching time ranges, numerical ranges, etc.:

[0156]

[0157] For example, the user queries the time range "2024-01-01 to 2024-01-23", and the data record time is 2024-01-08, then the matching function returns 1, otherwise returns 0. if m i is within the range q ij , then return 1, otherwise return 0. i ij i

[0158] (2) Boolean matching function - suitable for structured data:

[0159] ​​​

[0160] For example, if the user queries the type field of the data record as "temperature", the matching function returns 1, otherwise it returns 0. i is contained in m ij represents the condition q i contains metadata m ij .

[0161] In data distribution management, the subscription mechanism enables users to flexibly subscribe to specific marine data products or topics, such as "real-time ocean temperature" or "ocean salinity changes". By recording user subscription information and setting automatic push time intervals or trigger conditions according to user needs, when new data is generated, the event trigger mechanism automatically pushes the data to the subscriber. In this process, the message queue system (such as Apache Kafka) is used to manage data distribution, ensuring that users can receive the latest data in a timely manner. Data pushing is done through persistent connection technologies such as WebSocket or HTTP / 2, ensuring real-time and high efficiency.

[0162] The content distribution strategy further optimizes the efficiency of data transmission. According to the user's geographic location, network conditions and data needs, the optimal distribution path is selected. For massive marine environmental data, CDN (Content Delivery Network) accelerators can be used to cache data to edge nodes, thereby shortening data transmission distance and improving user access speed. This not only improves response capability, but also provides better services for users who need real-time data updates. For real-time data, real-time pushing can be achieved through persistent connection, while for large-scale historical data, batch data transmission methods are provided, and HTTP download or SFTP is used to achieve more efficient transmission.

[0163] User subscribes to marine environmental data product S. The new data product generated by the data publisher is D new , D new may be a set of marine environmental data obtained by matching the data record D j with the user's request M j , that is, the target marine environmental data set, where D new has a correlation with the user's subscription topic S, which can be represented as P new :

[0164]

[0165] where s k represents the correlation function, and α k represents the corresponding weight. Only when R new > T subscribe (Tsubscribe is the distribution threshold), the data is pushed to the user.

[0166] Correlation coefficient s k is used to calculate the degree of matching between the new data product D new and the user subscription topic S. The specific correlation calculation method can be different based on different data types and retrieval requirements, and the present embodiment does not make specific limitations. For example, the following methods can be used:

[0167] (1) Boolean correlation: used to determine whether the data product meets the subscription topic.

[0168]

[0169] Calculation process: check whether the data product D new meets the conditions of the subscription topic S. For example, if S is "data type = temperature", and the data type of D new is also "temperature", the correlation score is 1, otherwise 0.

[0170] (2) TF-IDF: used to calculate the correlation between the data product and the subscription topic, suitable for text data matching.

[0171] TF: the frequency of the term appearing in the data product D new .

[0172]

[0173] IDF: the inverse document frequency of the term in all data records.

[0174]

[0175] TF-IDF: calculate the correlation of the data product D new to the subscription topic S.

[0176] s k (D new ,S)=TF(q i ,D new )×IDF(q i )

[0177] Calculation process: first text the subscription topic and the data product, then calculate the TF-IDF weight, which represents the importance of the term q i in the data product and the subscription topic. The overall correlation is obtained by weighting the TF-IDF scores of multiple terms.

[0178] (3) Cosine similarity: used for vector space model, the correlation is evaluated by calculating the cosine value of the vector angle between the query condition and the data record.

[0179]

[0180] where: · denotes the dot product of vectors, ‖·‖ denotes the norm (i.e. length) of a vector.

[0181] Calculation process: Convert data product D new and subscription topic S into vector representation, calculate the cosine similarity between them, the value ranges from -1 to 1, the closer the value to 1, the higher the correlation.

[0182] Synthesize correlation score: In subscription management, the correlation score P new is calculated by weighted average.

[0183] The amount of data pushed in real time V realtime should meet the user's network bandwidth B limit:

[0184] V realtime ≤B·t

[0185] where t represents the time interval.

[0186] The amount of data distributed in batches V batch can be divided into m blocks, each block has a size of v i , the total amount is:

[0187]

[0188] Data security is also crucial in marine environmental data management. MinIO ensures data transmission security through TLS / SSL protocol and encrypts static data using symmetric encryption algorithms such as AES-256. The key is securely transmitted through asymmetric encryption algorithms such as RSA, ensuring that data transmission and storage are immune to external attacks. In addition, the present embodiment is based on the Role-Based Access Control (RBAC) model, which restricts user access to data through access policies. Each data request is subject to strict credential verification, ensuring that only authorized users can access specific data.

[0189] Cloud-edge data exchange technology greatly improves the processing efficiency of multi-source marine data. In areas where marine data demand is concentrated, the system using the method of the embodiment (i.e., a multi-source marine environment data intelligent sharing system, which corresponds to a multi-source marine environment data intelligent sharing method setting, and the embodiment does not make specific description) can deploy edge computing nodes to share the load of the cloud main data center. These edge nodes not only can process and store part of the data, but also can provide localized data access services, thereby reducing the pressure on the main data center. The flexibility of the MinIO distributed architecture allows data to be distributed across different edge nodes as needed, ensuring that users can quickly access the required marine environment data. At the same time, the incremental synchronization technology is used to synchronize the changed data of the edge nodes back to the cloud main database, ensuring the consistency of the data between the cloud and the edge nodes. MinIO achieves real-time synchronization of edge nodes and cloud data by automatically handling data conflicts and consistency checks.

[0190] During the data synchronization update process, the system monitors each data source in real time to ensure that when new data is generated or existing data is updated, it can quickly respond and start the synchronization process. Through event-driven mechanisms such as webhooks or message queue events, data synchronization tasks can be triggered in real time to ensure that the data in the system remains synchronized with the source data. Incremental synchronization is one of the core technologies of this process, and by comparing the versions or timestamps of the data, the system can identify the data that has been added or updated since the last synchronization. Using the object versioning feature of MinIO, the system can obtain incremental data and update it to the storage node, while using erasure coding technology to ensure data availability and fault tolerance during the synchronization process.

[0191] Assuming that the data version before time t is B(t), and the data version at time t+Δt is B(t+Δt). The incremental data ΔB is:

[0192] ΔB = B(t+Δt) - B(t)

[0193] The total amount of synchronized data blocks V sync is:

[0194] Where N represents the number of data blocks that need to be synchronized.

[0195] After the data synchronization is completed, the embodiment will also periodically perform consistency checks on the synchronized data to ensure that the data between the cloud and the edge nodes remains consistent. Using the hash function H(x) to check data consistency, assuming that the hash value of data block B j is H(B j ):

[0196] H(B j) = H(source(B j )) = H(destination(B j ))

[0197] where source and destination represent the source node and target node of the data, respectively. The source node refers to the original storage location of the data or the initial source of the data. In data consistency checking, the source node is responsible for generating the initial hash value of the data chunk. The target node refers to the storage target location of the data, which may be the location where the data is synchronized or copied to. In data consistency checking, the target node is responsible for storing the data chunk received from the source node and calculating the hash value of the data chunk to verify consistency.

[0198] If data inconsistency is found, the embodiment will automatically start the data repair process. In addition, the embodiment records detailed logs of each data synchronization, including synchronization time, data source, synchronized data volume, and results, etc. These logs not only can be used for subsequent audit, but also provide strong support for troubleshooting. At the same time, the embodiment sends a synchronization status report to the administrator through email or other message channels, informs the success or failure of synchronization, and attaches detailed log information.

[0199] Subscription distribution technology and data synchronization update play a key role in multi-source massive marine environmental data management. Through data extraction, consistency verification, content distribution, and secure transmission, efficient data acquisition and distribution are ensured. At the same time, the combination of cloud-edge data exchange and real-time synchronization update technology further improves the flexibility and scalability of multi-source marine environmental data intelligent sharing, meeting the needs of massive marine data management and distribution.

[0200] The embodiment realizes efficient sharing and synchronization update of multi-source massive marine environmental data by adopting MinIO distributed storage and subscription distribution technology. Compared with traditional centralized storage systems, the embodiment can better cope with the storage needs of massive data, and at the same time, through distributed architecture, the reliability and speed of data access are improved. Data adaptation and subscription distribution mechanism enable users to quickly obtain customized data products, meeting individual needs. In addition, the combination of edge computing and cloud synchronization technology significantly improves the efficiency and security of data transmission, ensuring the consistency and integrity of data between different nodes, providing a more intelligent and flexible marine data sharing solution. The beneficial effects of the embodiment include:

[0201] (1) MinIO-based distributed storage architecture: This embodiment utilizes the MinIO distributed storage system to achieve efficient storage and management of massive marine environmental data. Through the distributed architecture, the multi-source marine environmental data intelligent sharing method can be horizontally expanded to support parallel storage and access of large-scale data, solving the scalability problem of traditional centralized storage.

[0202] (2) Intelligent data adaptation mechanism: The multi-source marine environmental data intelligent sharing method automatically filters and adapts data according to user needs through intelligent search algorithms, providing customized marine environmental data products. This includes data format conversion, element filtering, and geographic region cropping, allowing users to quickly obtain data that meets specific requirements.

[0203] (3) Data subscription and distribution management: This embodiment introduces a subscription-based distribution mechanism, allowing users to subscribe to specific types or topics of data products. The multi-source marine environmental data intelligent sharing method automatically manages data pushing and distribution based on subscription information, supporting real-time data pushing and periodic batch distribution, significantly improving the real-time and flexibility of data acquisition.

[0204] (4) Automated and intelligent data management: The multi-source marine environmental data intelligent sharing method integrates various intelligent data management functions, such as automated data organization, packaging, and metadata generation, reducing manual intervention and improving data management efficiency and accuracy. At the same time, the multi-source marine environmental data intelligent sharing method supports personalized data recommendations based on user historical behavior, optimizing user experience.

[0205] In summary, compared with the prior art, the technical solution of this embodiment uses MinIO distributed storage and subscription distribution mechanisms to achieve efficient storage, intelligent search, and personalized distribution of massive marine environmental data, significantly improving data access speed and real-time performance. At the same time, through cloud-edge collaboration and secure encrypted transmission, it ensures efficient synchronization and security of data between different nodes, solving the problems of poor scalability, slow data transmission, and lack of intelligent management in traditional centralized storage, and providing a more flexible, reliable, and secure data sharing solution for users.

[0206] The above describes the embodiments of the present application in detail in conjunction with the drawings, but the present application is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the purpose of the present application.

Claims

1. A multi-source marine environment data intelligent sharing method, characterized in that, The method comprises: acquiring multi-source marine environment data; product adaptation is performed on the multi-source marine environment data to obtain first multi-source marine environment data, and the product adaptation is used to indicate multi-source marine environment data meeting user requirements; the first multi-source marine environment data is packaged to obtain second multi-source marine environment data; the second multi-source marine environment data is stored in a distributed manner to obtain stored multi-source marine environment data, comprising, the second multi-source marine environment data is stored in a distributed manner by MinIO; the stored multi-source marine environment data is intelligently shared through intelligent subscription distribution, comprising: acquiring a user request containing multiple conditions, and acquiring metadata corresponding to each data record in the stored multi-source marine environment data and the conditions; based on the metadata, calculating the matching degree between each data record and each condition in the user request; determining a marine environment subscription data set meeting multiple conditions in the user request according to the matching degree; the marine environment subscription data set is intelligently shared through a content distribution strategy, comprising: acquiring a product topic subscribed by a user; calculating the correlation between the product topic and the marine environment subscription data set to obtain a correlation calculation result; determining a target push data set according to the correlation calculation result; the target push data set is intelligently shared. 2.The method of claim 1, wherein, The matching degree between each data record and each condition in the user request is calculated based on the metadata in the following manner: wherein, represents a matching degree, represents a total number of conditions in a user request, represents an importance weight of each condition, represents an condition, represents metadata, represents a matching function. 3.The method of claim 1, wherein, the target push data set is intelligently shared in the following manner: the target push data set is intelligently shared in a real-time push manner; or, the target push data set is intelligently shared in a batch distribution manner. 4.The method of claim 1, wherein, The correlation between the product topic and the marine environment subscription data set is calculated in the following manner: wherein, represents a correlation, represents the total amount of data in the marine environment subscription dataset, represents a corresponding weight, represents a correlation function, represents a marine environment subscription dataset, represents a product topic.

5. The intelligent sharing method of multi-source marine environmental data according to claim 1, characterized in that, after the stored multi-source marine environment data is intelligently shared through intelligent subscription distribution, the method further comprises: if new data is acquired or data is updated, triggering data updating through an event-driven mechanism; the multi-source marine environment data is synchronously updated through incremental synchronization.

6. The intelligent sharing method of multi-source marine environmental data according to claim 5, characterized in that, The multi-source marine environment data is synchronously updated through incremental synchronization in the following manner: wherein, represents incremental data, represents the amount of data at the time instant, represents the amount of data at the current time instant, represents the amount of data at the current time instant, represents the total amount of synchronization data blocks, represents the number of data blocks to be synchronized, represents the first data block to be synchronized.

7. The intelligent sharing method of multi-source marine environmental data according to claim 5, characterized in that, after the multi-source marine environment data is synchronously updated through incremental synchronization, the method further comprises: consistency checking of the data synchronously updated through a hash function.

Citation Information

Patent Citations

  • Processing, storing and sharing method for mass Internet of Things data modeling

    CN113986873A

  • Distributing multi-source push notifications to multiple targets

    US20130067024A1