A profile dataset construction method and system based on multi-source heterogeneous ocean observation

By performing version cleaning, format normalization, and quality control on multi-source heterogeneous ocean observation data, a high-quality ocean profile dataset was constructed, solving the timeliness and consistency problems of multi-source observation data and achieving efficient data processing and storage.

CN120872936BActive Publication Date: 2026-04-28INST OF ATMOSPHERIC PHYSICS CHINESE ACADEMY SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF ATMOSPHERIC PHYSICS CHINESE ACADEMY SCI
Filing Date
2025-06-25
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing ocean observation data suffers from multi-source heterogeneity and redundancy, leading to increased data processing and storage burdens, inconsistent quality, and difficulty in constructing high-quality ocean profile datasets.

Method used

By constructing a method and system for building profile datasets from multi-source heterogeneous ocean observations, including data acquisition, version cleaning, format normalization, quality control, and bias correction, the consistency and accuracy of the data are improved.

Benefits of technology

It enables the rapid construction of high-quality ocean profile datasets, solves the problems of insufficient timeliness, inconsistent formats, difficulty in redundancy control, and inconsistent bias correction of multi-source observation data, and provides reliable data support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872936B_ABST
    Figure CN120872936B_ABST
Patent Text Reader

Abstract

The application provides a profile data set construction method and system based on multi-source heterogeneous ocean observation, wherein the method comprises the following steps: obtaining multi-source heterogeneous original ocean observation profile data and respective description information from a plurality of target ocean data centers / observation agencies in a target time period; performing version cleaning on the original ocean observation profile data according to the unique identifier determined according to the original ocean observation profile data metadata and the description information, and performing high-frequency cleaning on the original ocean observation profile data based on the spatio-temporal joint characteristics to obtain a plurality of target ocean observation profile data; and performing matrix processing on the target ocean observation profile data to obtain corresponding normalized data to construct a profile data set in the target time period. Thus, the multi-source heterogeneous profile data can be sequentially subjected to data cleaning and format normalization, thereby improving the consistency, accuracy and availability of the data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of marine observation and data processing technology, and in particular to a method and system for constructing profile datasets based on multi-source heterogeneous marine observations. Background Technology

[0002] Ocean observation data is a crucial foundation for understanding the internal structure of the ocean and studying its physical properties. Over the past century, more than 18 million ocean observation profiles have been acquired globally. Currently, these profiles are collected, processed, and published by different countries and institutions, resulting in multiple independent datasets of raw ocean observation profiles. Due to differences in processing procedures and coverage areas, each dataset has its own characteristics in terms of accuracy, spatiotemporal distribution, and applicability; there is no single "optimal" dataset.

[0003] Meanwhile, significant differences exist among different institutions and data centers in data storage methods, file structures, and variable naming conventions, severely hindering the unified management and comprehensive utilization of multi-source data. Despite the vast number of observation profiles, the sheer size of the ocean and limited observation capabilities result in significant sparsity and uniqueness in both time and space, giving each profile significant independent value. In practical applications, the same observation equipment may repeatedly collect data within a short period, creating a large amount of redundant observation data. While this redundant information has some reference value in real-time monitoring, it significantly increases the burden of data processing and storage.

[0004] Furthermore, the diverse observation methods and complex types of data errors result in inconsistent data quality. Observational data contains numerous random errors, which need to be identified and eliminated through quality control; it also commonly exhibits systematic biases stemming from instruments or methods, requiring model correction. Therefore, quality control and bias correction are crucial components supporting accurate assessments of marine environmental changes.

[0005] Existing data products and processing workflows largely focus on single data sources, making it difficult to simultaneously adapt to the demands of rapid updates, format fusion, and high-quality management of multi-source observational data. In the absence of a unified processing mechanism, problems such as low data fusion efficiency, difficulties in redundancy control, and inconsistent bias corrections are becoming increasingly prominent, severely hindering the construction and stable production of high-quality ocean profile datasets. There is an urgent need to develop a data processing framework that balances standardization, universality, and automation capabilities. Summary of the Invention

[0006] This application provides a method and system for constructing a profile dataset based on multi-source heterogeneous ocean observations, which can construct a high-quality ocean profile dataset.

[0007] In a first aspect, embodiments of this application provide a method for constructing a profile dataset based on multi-source heterogeneous ocean observations, the method comprising:

[0008] Within a target time period, heterogeneous raw ocean observation profile data, along with their respective descriptive information, are acquired from several target ocean data centers / observation institutions. The raw ocean observation profile data includes metadata and observation data. The descriptive information is used to determine the unique identifier of the raw ocean observation profile data. The metadata is used to describe, locate, or interpret the profile observation process.

[0009] Based on the time information and unique identifiers in the metadata, multiple original ocean observation profile data are cleaned to retain the latest processed version of the original ocean observation profile;

[0010] Based on the time and location information in the metadata, high-frequency cleaning is performed on the data that meet the same spatiotemporal judgment conditions after version cleaning to obtain multiple target ocean observation profile data; high-frequency cleaning includes averaging the observation data that meet the same spatiotemporal judgment conditions;

[0011] The metadata and observation data of the target ocean observation profile data are matrixed to obtain the corresponding normalized data; the normalized data of multiple target ocean observation profile data are used to construct the profile dataset within the target time period.

[0012] Therefore, this application proposes a method for constructing profile datasets based on multi-source heterogeneous ocean observations, which can sequentially perform data cleaning and format normalization on multi-source heterogeneous profile data, thereby improving the consistency, accuracy and usability of the data.

[0013] Secondly, embodiments of this application provide a profile dataset construction system based on multi-source heterogeneous ocean observations, the system comprising:

[0014] The acquisition module is used to acquire multi-source heterogeneous raw ocean observation profile data from several target ocean data centers / observation institutions within a target time period, as well as the descriptive information of each data. The raw ocean observation profile data includes metadata and observation data. The descriptive information is used to determine the unique identifier of the raw ocean observation profile data. The metadata is used to describe, locate, or interpret the profile observation process.

[0015] The processing module is used to perform version cleaning on multiple original ocean observation profiles based on the time information and unique identifiers in the metadata, so as to retain the latest processed version of the original ocean observation profiles.

[0016] The processing module is also used to perform high-frequency cleaning on data that meet the same spatiotemporal judgment conditions in the version-cleaned data based on the time and location information in the metadata, so as to obtain multiple target ocean observation profile data; high-frequency cleaning includes averaging the observation data that meet the same spatiotemporal judgment conditions;

[0017] The processing module is also used to perform matrix processing on the metadata and observation data of the target ocean observation profile data to obtain the corresponding normalized data; the normalized data of multiple target ocean observation profile data are used to construct the profile dataset within the target time period.

[0018] It is understood that the beneficial effects of the second aspect mentioned above can be found in the relevant descriptions of the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This illustration shows a schematic diagram of a profile dataset construction system based on multi-source heterogeneous ocean observations provided in an embodiment of this application;

[0021] Figure 2 This paper illustrates a flowchart of a method for constructing a profile dataset based on multi-source heterogeneous ocean observations, as provided in an embodiment of this application.

[0022] Figure 3 This illustration shows a data cleaning implementation diagram provided in an embodiment of this application;

[0023] Figure 4 This illustration shows a schematic diagram of the high-frequency cleaning implementation provided in an embodiment of this application;

[0024] Figure 5 This application provides an embodiment of the data processing, which includes a comparison diagram of the profiles before and after data processing.

[0025] Figure 6 This paper illustrates a system architecture diagram for constructing profile datasets based on multi-source heterogeneous ocean observations, as provided in an embodiment of this application. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be described below with reference to the accompanying drawings.

[0027] In the description of the embodiments in this application, any embodiment or design that is “exemplary,” “for example,” or “by way of example” should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as “exemplary,” “for example,” or “by way of example” is intended to present the relevant concepts in a concrete manner.

[0028] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized. "Multiple" can refer to one or more, where multiple means two or more.

[0029] To address the problems of inconsistent formats, excessive redundancy, varying quality, and systematic biases in existing multi-source raw ocean observation profile data, this application provides a method and system for constructing a high-quality observation profile dataset based on multi-source ocean observations. This method can sequentially perform data cleaning, format normalization, quality control, and bias correction on multi-source heterogeneous profile data from different data centers and observation institutions, thereby improving data consistency, accuracy, and usability.

[0030] Figure 1 This illustration shows a schematic diagram of a profile dataset construction system based on multi-source heterogeneous ocean observations, provided in an embodiment of this application. For example... Figure 1 As shown, the profile dataset construction system mainly includes the following parts:

[0031] (1) Data acquisition module 110: Acquire multiple raw ocean observation profile data within a target time period from several target ocean data centers / observation institutions.

[0032] (2) Data preprocessing module 120: Generates a profile blacklist and a profile metadata table A. The profile metadata included in table A is auxiliary information used to describe, locate, or interpret the profile observation process and is applied to the entire profile.

[0033] (3) Data cleaning module 130: Cleans multiple original ocean observation profile data using data generated by the data preprocessing module, including: blacklist cleaning: retaining usable original ocean observation profile data; version cleaning: retaining the latest processed version of each profile; high-frequency cleaning: performing high-frequency cleaning on profiles that are sampled in similar time and space and have the same equipment type.

[0034] (4) Data format normalization module 140: count the number of profiles, create and initialize variables, extract and fill metadata, extract and fill observation variables, unify variable naming and data structure, and output data files in the specified standard format.

[0035] (5) Data quality control module 150: scores and marks the quality of standard format marine observation profile data.

[0036] (6) Data bias correction module 160: Cleans up low-quality ocean observation profile data and corrects systematic biases in data caused by designated observation equipment.

[0037] (7) Data output module 170: Outputs the generated high-quality ocean observation profile data.

[0038] Module 110 is used to acquire data, modules 120-160 are used to process data, and module 170 is used to output data.

[0039] Therefore, by constructing a profile dataset construction system including the aforementioned modules 110-170, it is possible to solve problems such as insufficient timeliness caused by differences in the update frequency of multi-source observation data (when high-quality but delayed data has not yet been released, faster-updating data sources can be introduced in advance), missing data from a single data source, inconsistent data formats, difficulties in redundancy control, and inconsistent bias correction. This enables the rapid construction and stable output of high-quality ocean profile datasets, providing reliable data support for marine scientific research and applications.

[0040] Based on the above, in order to make the objectives, content and advantages of the present invention clearer, the following will, in conjunction with embodiments, use WOD (World Ocean Database) data, GTSPP (Global Temperature and Salinity Profile Program) data and CARDC (China Argo Real-time Data Center) data from January 1999 as examples to provide a detailed description of the method for constructing a profile dataset based on multi-source heterogeneous ocean observations proposed in this application.

[0041] Figure 2 A flowchart illustrating a method for constructing a profile dataset based on multi-source heterogeneous ocean observations, as provided in an embodiment of this application, is shown. Figure 2 As shown, the method mainly includes the following execution steps:

[0042] Step S201: Within the target time period, acquire multi-source heterogeneous raw ocean observation profile data and their respective descriptive information from several target ocean data centers / observation institutions; the raw ocean observation profile data includes metadata and observation data; the descriptive information is used to determine the unique identifier of the raw ocean observation profile data; the metadata is used to describe, locate or interpret the profile observation process.

[0043] For example, using Figure 1 The data acquisition module 110 shown acquires raw oceanographic profile data and related descriptive information (such as availability descriptions of the observation profiles, initial quality control markers, etc.) from the data center, observation equipment, and research vessel. "Multi-source" refers to the multiple data sources from which the data is acquired, and "heterogeneous" means that the data structures provided by different data sources may differ.

[0044] This module 110 supports multiple data acquisition methods, including but not limited to: acquiring data after submitting a data request to a designated data center (e.g., acquiring WOD data); directly downloading data via a public network address using remote transmission protocols such as FTP and SSH (e.g., acquiring GTSPP data); acquiring data through manual data copying (e.g., on an oceanographic research vessel); and remotely backing up data through observation equipment (e.g., data from a proprietary oceanographic observation system). In this embodiment, the target time period is one month. It is understood that the target time period could also be one day, one quarter, one year, etc., and is not limited here. After acquiring multiple raw oceanographic observation profile data and their respective descriptive information, the module utilizes... Figure 1 The data preprocessing module 120 shown performs data preprocessing operations.

[0045] First, a blacklist of profiles is generated based on the description information.

[0046] The availability descriptions of the acquired raw oceanographic profile data are integrated and processed to extract the profile identifiers, generating a profile identifier blacklist. Each profile identifier uniquely identifies one observation profile. If the availability description information for an acquired observation profile is empty, the generated profile blacklist will be empty.

[0047] Next, generate the profile metadata table A.

[0048] Information is extracted from the acquired raw observation profile data, and metadata is defined. The metadata of all profiles constitutes the profile metadata list A. Each metadata field corresponds one-to-one with a profile. Each metadata field includes, but is not limited to: observation time, recording time, observation instrument type, profile identifier, latitude and longitude coordinates, and the number of profile samples.

[0049] The raw observation profile data corresponds to at least one observation depth. Each observation depth yields one sampling result; therefore, the number of observation depths / layers equals the number of profile samples. The device can sample at each observation depth in the profile, obtaining one sampling result. A sampling result may include the values ​​of multiple state variables (such as temperature and salinity), which together constitute the observation data at the corresponding observation depth. The observation data from at least one observation depth constitutes the observation data of the raw observation profile data.

[0050] Metadata refers to auxiliary information, in addition to the observation data mentioned above, used to describe, locate, or interpret the profile observation process, and is applied to the entire profile.

[0051] Step S202: Based on the time information and unique identifier in the metadata, perform version cleaning on multiple original ocean observation profile data to retain the latest processed version of the original ocean observation profile.

[0052] For example, using Figure 1 The data cleaning module 130 shown performs version cleaning on the profile metadata table A obtained in step S201 and the multiple original ocean observation profile data.

[0053] Preferably, blacklist cleaning is performed before version cleaning. From multiple original ocean observation profile data sets, those with unique identifiers within the profile blacklist are deleted to retain usable original ocean observation profile data.

[0054] Traverse all metadata in the profile metadata table A. If the identifier of the metadata is in the profile blacklist generated in step S201, delete the corresponding metadata in the profile metadata table A; and delete the ocean observation profile data corresponding to the metadata in the multiple original ocean observation profile data obtained.

[0055] When performing version cleaning on multiple raw ocean observation profile data sets: it is determined whether there is at least one first subset among the multiple raw ocean observation profile data sets; each raw ocean observation profile data set in the first subset has the same unique identifier; if it exists, for each first subset, the first ocean observation profile data of the first subset is determined based on the time information of each raw ocean observation profile data set therein; the first ocean observation profile data is the data with a non-latest record time in the first subset; the first ocean observation profile data of the first subset is deleted from the multiple raw ocean observation profile data sets to perform version cleaning, resulting in version-cleaned raw ocean observation profile data.

[0056] For example, after parsing the blacklist-cleaned profile metadata table A, if there are metadata with the same profile identifier, only the metadata with the latest record time among the same profile identifiers is retained, and the other metadata with earlier record times is deleted. The original ocean observation profile data corresponding to the other metadata with earlier record times is deleted from the available original ocean observation profile data.

[0057] Step S203: Based on the time and location information in the metadata, high-frequency cleaning is performed on the data that meet the same spatiotemporal judgment conditions after version cleaning, resulting in multiple target ocean observation profiles. High-frequency cleaning includes averaging the observation data that meet the same spatiotemporal judgment conditions.

[0058] For example, using Figure 1 The data cleaning module 130 shown performs high-frequency cleaning on data that meet the same spatiotemporal judgment conditions in the version-cleaned data. The purpose of high-frequency cleaning is to effectively reduce redundant observations in the spatiotemporal dimension.

[0059] The high-frequency cleaning process includes: determining whether at least one second subset exists in the cleaned data based on the time and location information in the metadata; ensuring that each original ocean observation profile data in the second subset meets the same spatiotemporal judgment condition; if so, averaging the observation data of each data in the second subset that meets the same spatiotemporal judgment condition to obtain the target observation data corresponding to the second subset; combining the metadata of any original ocean observation profile data with the target observation data to obtain the second ocean observation profile data of the second subset; adding the second ocean observation profile data to the cleaned data and deleting the original ocean observation profile data in the second subset.

[0060] In one implementation, the same spatiotemporal judgment conditions include: each original ocean observation profile data in the second subset corresponds to the same target equipment type; each original ocean observation profile data in the second subset corresponds to the same observation depth; the observation time of each original ocean observation profile data in the second subset is within a preset time resolution range; and the maximum longitude difference and the maximum latitude difference of each original ocean observation profile data in the second subset are respectively less than their respective preset thresholds.

[0061] In another implementation, the spatiotemporal judgment conditions include: each original ocean observation profile data in the second subset corresponds to the same target equipment type; each original ocean observation profile data in the second subset corresponds to the same number of observation depths; the same number is more than one; the observation time of each original ocean observation profile data in the second subset is the same; the longitude and latitude of each original ocean observation profile data in the second subset are the same; and the observation depth corresponding to each original ocean observation profile data in the second subset is the same.

[0062] First, filter the metadata that needs high-frequency cleaning. This includes setting the device types that need high-frequency cleaning, and performing high-frequency cleaning on each device type separately (in this embodiment, high-frequency cleaning is performed on the data of three instrument types: SUR, CTD, and OSD).

[0063] Secondly, the above-mentioned high-frequency cleaning operation is performed on the original oceanographic profile data corresponding to the metadata that requires high-frequency cleaning. For each instrument type that requires high-frequency cleaning, all metadata with a profile sampling count of 1 is extracted from the cleaned profile metadata table A and formed into a single-depth profile metadata table B1, and a single-depth data high-frequency cleaning strategy is performed; the remaining metadata of this instrument type is formed into a multi-depth profile metadata table B2, and a multi-depth high-frequency cleaning strategy is performed.

[0064] This allows for the acquisition of multiple target ocean observation profile data. These data include the original ocean observation profile data that was not subjected to high-frequency cleaning after version cleaning, as well as the ocean observation profile data obtained after high-frequency cleaning.

[0065] Step S204: The metadata and observation data of the target ocean observation profile data are matrixed to obtain the corresponding normalized data; the normalized data of multiple target ocean observation profile data are used to construct the profile dataset within the target time period.

[0066] For example, firstly, using Figure 1 The data format normalization module 140 shown performs matrix processing on the metadata and observation data of the target ocean observation profile data to obtain the corresponding normalized data. The normalized data includes a matrix corresponding to the metadata, a matrix corresponding to the observation data, and a matrix corresponding to the observation depth of the target ocean observation profile data.

[0067] The data format unification module 140 integrates the target ocean observation profile data within a specified time range after high-frequency cleaning in step S203 according to the specified time step (the time step can be, but is not limited to, year, month, day, hour, etc., and is set to month in this embodiment), and converts it into a standardized data file with consistent structure and standardized format, consisting of "one file per time step".

[0068] The specific steps are as follows:

[0069] S11: With a specified time step, traverse the observation data within the specified time range and execute steps S204-2 to S204-7.

[0070] S12: Count the number of profile files included in the current time period, denoted as Tprofile.

[0071] S13: Variable Initialization: Initialize the metadata matrix P to store observation metadata, with the dimension being "metadata dimension × Tprofile". Initialize N+1 observation variable matrices, denoted as M. i (i = 1, ..., N) and D, with dimensions "Tprofile × maximum supported depth layers", where the maximum supported depth layers (i.e., the maximum number of observation depths) are set to 3000. Here, N represents the number of state variables included in the profile (in this embodiment, N is 2, and i ranges from 1 to 2, corresponding to temperature variable M1 and salinity variable M2). All initial matrix values ​​are uniformly filled with missing identifiers, such as "-999".

[0072] S14: Traverse all profile data within the current time period and execute steps S204-5 to S204-7 sequentially.

[0073] S15: Metadata Extraction and Population. Extract the metadata of the current profile and populate it into the metadata matrix P. If the metadata is stored in code form, convert it to standardized text and assign the value; if any metadata item is missing, ignore it.

[0074] S16: State Variable Extraction and Filling. Extract the state variables of the current profile and store them in the corresponding M. i And D. If a state variable is missing, it is ignored. Specifically, if the depth of a state variable does not exceed the maximum supported depth, the extracted state variable is directly filled into the corresponding state variable matrix; if the depth of a state variable exceeds the maximum supported depth, then among the state variable values ​​corresponding to depth layers where the initial quality marker bits are all qualified, the shallowest and deepest observations are retained, and new state variable values ​​are generated by uniformly sampling according to the maximum depth, and filled into the corresponding state variable matrix.

[0075] S17: Standardized Data File Packaging: Package the integrated observation data and metadata matrix P into a standardized data file with a unified structure according to the specified data format (including but not limited to NetCDF, CSV and MAT formats).

[0076] Secondly, utilize Figure 1The data quality control module 150 shown generates corresponding quality control identifiers for each matrix element in the matrix according to a preset quality control scheme. These quality control identifiers include "qualified" and "unqualified".

[0077] The data quality control module 150 iterates through the standardized data obtained in step S204, adopts the CODC-QC quality control scheme, performs quality control one by one, and generates corresponding quality control identifiers.

[0078] Next, using Figure 1 The data deviation correction module 160 shown corrects the deviations of matrix elements that are of substandard quality according to a preset deviation correction scheme.

[0079] The data deviation correction module 160 iterates through the profile data that has undergone quality control processing by module 150, and performs deviation correction processing on the "good data". Specifically, this includes:

[0080] S21: Read the state variable to be corrected, and replace the non-conforming observations in the state variable after quality control processing with the missing identifiers based on the quality control identifiers output by module 150.

[0081] S22: Classify the processed state variables according to the equipment type and process them using the corresponding deviation correction schemes.

[0082] For example: the CH14 scheme is used for XBT data, the GC20 scheme for MBT data, the GC22 scheme for BOT data, and the APB bias correction scheme proposed by the Institute of Atmospheric Physics, Chinese Academy of Sciences, is used for APB data. Corresponding bias correction results are generated for each instrument type.

[0083] Finally, using Figure 1 The data output module 170 shown constructs a profile dataset for the target time period based on the deviation-corrected matrix.

[0084] The data output module 170 outputs high-quality ocean observation profile data according to the specified data format (including but not limited to NetCDF, Mat, CSV, txt, etc.).

[0085] The following will be based on Figure 3 and Figure 4 Taking this example, we will describe in detail the data cleaning process and strategies used in this solution.

[0086] Figure 3 A schematic diagram illustrating a data cleaning implementation provided in an embodiment of this application is shown.

[0087] The data cleaning process employed in this solution mainly includes three steps: blacklist cleaning, version cleaning, and high-frequency cleaning. Each step has specific functions and outputs.

[0088] like Figure 3 As shown, the availability of raw oceanographic profile data first needs to be defined or assessed. Therefore, a profile blacklist needs to be generated in advance based on the data availability definition. This blacklist may contain data that needs to be excluded or specially processed.

[0089] Secondly, metadata table A is generated based on the original ocean observation profile data. This table contains auxiliary information from the original ocean observation profile data used to describe, locate, or interpret the profile observation process.

[0090] Next, the profile blacklist was used as the basis for cleaning, and the metadata table A and the original ocean observation profile data were cleaned to remove or correct the data marked in the blacklist.

[0091] Furthermore, version cleaning can be performed on the data obtained after cleaning the blacklist.

[0092] Furthermore, the data obtained after version cleaning undergoes high-frequency cleaning. This high-frequency cleaning operation includes two parts: single-depth data high-frequency cleaning and multi-depth data high-frequency cleaning.

[0093] Specifically, during high-frequency cleaning of single-depth data, a single-depth profile metadata table B1 is generated from metadata table A. This table contains detailed information about single-depth observations. A high-frequency cleaning strategy for single-depth data is then used to clean the single-depth oceanographic profile data.

[0094] During high-frequency cleaning of multi-depth data, a multi-depth profile metadata table B2 is generated from metadata table A. This table contains ocean observation profile data at multiple depth levels. A high-frequency cleaning strategy for multi-depth data is then used to clean the multi-depth ocean observation profile data.

[0095] Finally, the cleaned target ocean observation profile data is output. This data has undergone multiple cleaning and processing steps and can be used for further analysis and research.

[0096] Figure 4 A schematic diagram illustrating the high-frequency cleaning implementation provided in an embodiment of this application is shown. Figure 4 As shown, after blacklisting and version cleaning of the original ocean observation profile data, a high-frequency cleaning operation is further performed.

[0097] High-frequency cleaning operations consist of two parts: high-frequency cleaning of single-depth data and high-frequency cleaning of multi-depth data.

[0098] The high-frequency cleaning strategy for single-depth data includes the following steps:

[0099] S41: Read the single-depth profile metadata table B1, iterate through all metadata, and extract the single-depth profile observation data specified by the metadata from the original oceanographic observation profile data obtained after version cleaning. Sequentially read the state variables (such as temperature, salinity, dissolved oxygen, etc.) and their initial quality control flags, as well as the depth variables and their initial quality control flags. Expand the corresponding metadata to include observation time, recording time, instrument type, profile identifier, latitude and longitude coordinates, number of profile samples, state variable value, depth variable value, initial quality control flags for the state variables, and initial quality control flags for the depth variables. If the obtained initial quality control flags are empty, the corresponding observation values ​​are considered to be qualified data.

[0100] S42: Group the single-depth profiles. In the single-depth profile metadata table B1, metadata whose initial quality control flags for both the state variable and depth variable are qualified are selected. Grouping is based on two attributes: the depth variable value and a specified time resolution (the time resolution can be, but is not limited to, hourly, daily, weekly, etc.; in this embodiment, the time resolution is set to monthly). Profiles with identical values ​​for both attributes are grouped together. For all groups, steps S43 to S47 are executed to perform single-depth data cleaning.

[0101] S43: Calculate the mean of single-depth data. Calculate the maximum difference in longitude and latitude for all observed profile data within the group. If either is not less than a set threshold (set to 0.01° in this embodiment), the single-depth data cleaning for that group ends directly; otherwise, calculate the mean of all profile state observations within the group and proceed to step S44.

[0102] S44: Clean the single-depth profile data. Read the raw observation profile data corresponding to any metadata in this group, and modify the state variable value to the mean of the state observation values ​​calculated in step S43. Delete the remaining metadata in this group and the raw ocean observation profile data corresponding to the metadata.

[0103] High-frequency cleaning of multi-depth data using a multi-depth data cleaning strategy includes the following steps:

[0104] S51: Group the multi-depth profiles. In the multi-depth profile metadata table B2, use the observation date, latitude and longitude coordinates, and number of profile samples as the basis for grouping the metadata in B2. Groups with the same three attributes are grouped together. For all groups, execute steps S52 to S54 respectively.

[0105] S52: Read multi-depth profile observation information. Read the original oceanographic profile data corresponding to the metadata within the group one by one, and extract the state variables, depth variables, initial quality control flag bits of the state variables, and initial quality control flag bits of the depth variables.

[0106] S53: Calculate the mean of multi-depth data. Compare the depth variables of each observation profile within the group. If the depth variables are completely consistent, select each depth sampling layer based on the initial quality control flag of the state variables, and calculate the mean of the state observation values ​​for each depth sampling layer. Proceed to step S54. If the depth variables are not completely consistent, directly end the multi-depth data cleaning for this group.

[0107] S54: Clean the multi-depth profile data. Read the raw oceanographic profile data corresponding to any metadata in this group, modify the state observation value to the mean of the state observation values ​​calculated in step S53, and modify the initial quality control flag to qualified. Delete the remaining metadata in this group and the raw oceanographic profile data corresponding to the metadata.

[0108] Figure 5 A comparison diagram of the profiles before and after data processing provided in the embodiments of this application is shown. For example... Figure 5 As shown, a comparison of the temperature profiles before and after data processing in January 1999 is presented. The top left image shows the original temperature profiles, displaying multiple original profiles showing temperature variations with depth. It can be seen that some outliers or noise exist in the data, manifesting as some very steep or irregular changes. The top right image shows the high-quality temperature profiles, displaying the processed data. Figure 1 The temperature profile shown is the result of data processing by the system. Compared to the original data, this data appears smoother and more consistent, with outliers and noise removed.

[0109] The bottom left image shows the original salinity profile, displaying multiple raw lines illustrating salinity variations with depth. Outliers or noise can also be observed, manifesting as very steep or irregular changes. The bottom right image shows a high-quality salinity profile, indicating the changes after... Figure 1 The system shown here has undergone data processing to produce salinity profiles. Compared to the original data, these profiles appear smoother and more consistent, with outliers and noise removed.

[0110] It can be seen that, after Figure 1 After data processing, the temperature and salinity profiles of the system shown become smoother and more consistent, and outliers and noise are effectively removed. This demonstrates that the data processing workflow effectively improves the quality and reliability of the data.

[0111] Therefore, this invention discloses a method and system for constructing high-quality ocean profile datasets based on multi-source heterogeneous ocean observations. High-quality ocean observation profile data is generated by performing format normalization, data cleaning, quality control, and bias correction on multi-source heterogeneous ocean profile data from different ocean data centers and observation institutions.

[0112] This invention has the following advantages:

[0113] Based on a cleaning scheme for multi-source heterogeneous data, this invention improves compatibility and automation: It constructs a format normalization scheme to unify variable naming and storage format for observation profile data in various formats such as NetCDF, Mat, CSV, and txt from different marine observation data centers, solving problems such as inconsistent structure and chaotic variable definition between different datasets, significantly reducing the need for manual intervention, and improving the automation level and system compatibility of data processing.

[0114] A multi-version cleaning mechanism based on unique identifiers was constructed. By comparing the record time of the same profile in different versions, the version conflict caused by rolling updates of multi-source data is effectively resolved, ensuring the uniqueness of the profile.

[0115] A high-frequency cleaning mechanism based on the joint characteristics of "time + space" was constructed. By combining key information such as equipment type, observation depth and spatial location, a differentiated averaging strategy was implemented to achieve fine cleaning of high-frequency ocean profile data and effectively reduce redundant observations in the spatiotemporal dimensions.

[0116] It is understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. Furthermore, in some possible implementations, each step in the above embodiments may be selectively executed according to actual circumstances; it may be partially or fully executed, without limitation here. Additionally, all or part of any feature in the above embodiments can be freely and arbitrarily combined without contradiction. The combined technical solutions are also within the scope of this application.

[0117] Figure 6 This paper illustrates a system architecture diagram for constructing profile datasets based on multi-source heterogeneous ocean observations, as provided in an embodiment of this application. Figure 6 As shown, the profile dataset construction system 600 includes:

[0118] The acquisition module 610 is used to acquire multi-source heterogeneous raw ocean observation profile data and their respective descriptive information from several target ocean data centers / observation institutions within a target time period. The raw ocean observation profile data includes metadata and observation data. The descriptive information is used to determine the unique identifier of the raw ocean observation profile data. The metadata is used to describe, locate, or interpret the profile observation process.

[0119] The processing module 620 is used to perform version cleaning on multiple original ocean observation profiles based on the time information and unique identifier in the metadata, so as to retain the latest processed version of the original ocean observation profiles.

[0120] The processing module 620 is also used to perform high-frequency cleaning on data that meet the same spatiotemporal judgment conditions in the version-cleaned data based on the time and location information in the metadata, thereby obtaining multiple target ocean observation profile data. High-frequency cleaning includes averaging the observation data that meet the same spatiotemporal judgment conditions.

[0121] The processing module 620 is also used to perform matrix processing on the metadata and observation data of the target ocean observation profile data to obtain the corresponding normalized data; the normalized data of multiple target ocean observation profile data are used to construct the profile dataset within the target time period.

[0122] The aforementioned system 600 also outputs a profile dataset for the target time period.

[0123] Based on the methods in the above embodiments, this application provides an electronic device. The electronic device may include: at least one memory for storing a program; and at least one processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor executes the methods described in the above embodiments. Exemplarily, the electronic device may be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, server, ultra-mobile personal computer (UMPC), netbook, cellular phone, personal digital assistant (PDA), or artificial intelligence (AI) device. This application does not impose any special limitations on the specific type of the electronic device.

[0124] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of the present invention is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0125] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application. It should be understood that in the embodiments of this application, the order of the process numbers does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0126] The above specific embodiments further illustrate the purpose, technical solution and beneficial effects of this application. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of this application should be included within the scope of protection of this application.

Claims

1. A method for constructing a profile dataset based on multi-source heterogeneous ocean observations, characterized in that, The method includes: Within a target time period, heterogeneous raw ocean observation profile data, along with their respective descriptive information, are acquired from several target ocean data centers / observation institutions. The raw ocean observation profile data includes metadata and observation data. The descriptive information is used to determine the unique identifier of the raw ocean observation profile data. The metadata is used to describe, locate, or interpret the profile observation process. Based on the time information and unique identifier in the metadata, version cleaning is performed on multiple original ocean observation profile data to retain the latest processed version of the original ocean observation profiles. The version cleaning includes: determining whether at least one first subset exists among the multiple original ocean observation profile data; each original ocean observation profile data in the first subset has the same unique identifier; if it exists, for each first subset, based on the time information of each original ocean observation profile data therein, determining that the first subset contains first ocean observation profile data with a non-latest record time; and deleting the first ocean observation profile data of the first subset from the multiple original ocean observation profile data. Based on the time and location information in the metadata, high-frequency cleaning is performed on the data that meets the same spatiotemporal judgment conditions in the cleaned data to obtain multiple target ocean observation profiles; the high-frequency cleaning includes averaging the observation data that meets the same spatiotemporal judgment conditions; wherein, the high-frequency cleaning of the data that meets the same spatiotemporal judgment conditions in the cleaned data includes: Determine whether at least one second subset exists in the cleaned data; the original ocean observation profile data in the second subset meet the same spatiotemporal judgment conditions; the same spatiotemporal judgment conditions include: the original ocean observation profile data in the second subset correspond to the same target equipment type; the original ocean observation profile data in the second subset all correspond to the same observation depth; the observation time of the original ocean observation profile data in the second subset is within a preset time resolution range; and the maximum longitude difference and maximum latitude difference of the original ocean observation profile data in the second subset are respectively less than their respective preset thresholds; if it exists, for each second subset, the observation data of the data that meet the same spatiotemporal judgment conditions are averaged to obtain the target observation data corresponding to the second subset; the metadata of any original ocean observation profile data is combined with the target observation data to obtain the second ocean observation profile data of the second subset; the second ocean observation profile data is added to the cleaned data, and the original ocean observation profile data in the second subset is deleted; The metadata and observation data of the target ocean observation profile data are respectively matrixed to obtain corresponding normalized data; the normalized data includes the matrix corresponding to the metadata, the matrix corresponding to the observation data, and the matrix corresponding to the observation depth of the target ocean observation profile data; the normalized data of multiple target ocean observation profile data are used to construct the profile dataset within the target time period.

2. The method according to claim 1, characterized in that, The method further includes: According to the preset quality control scheme, corresponding quality control identifiers are generated for each matrix element in the matrix; the quality control identifiers include qualified and unqualified.

3. The method according to claim 2, characterized in that, The method further includes: According to the preset deviation revision scheme, the matrix elements in the matrix that are of substandard quality are revised according to the deviation.

4. The method according to claim 1, characterized in that, Before performing version cleaning on multiple raw ocean observation profile data, the method further includes: Generate a blacklist of outlines based on the described information; Among the multiple original ocean observation profile data, the original ocean observation profile data whose unique identifier is within the profile blacklist are deleted in order to retain the usable original ocean observation profile data.

5. A system for constructing profile datasets based on multi-source heterogeneous ocean observations, characterized in that, The system includes: The acquisition module is used to acquire multi-source heterogeneous raw ocean observation profile data and their respective descriptive information from several target ocean data centers / observation institutions within a target time period; the raw ocean observation profile data includes metadata and observation data; the descriptive information is used to determine the unique identifier of the raw ocean observation profile data; the metadata is used to describe, locate, or interpret the profile observation process; A processing module is configured to perform version cleaning on multiple original ocean observation profile data based on the time information and the unique identifier in the metadata, so as to retain the latest processed version of the original ocean observation profile; wherein, the version cleaning includes: determining whether there is at least one first subset among the multiple original ocean observation profile data; each original ocean observation profile data in the first subset has the same unique identifier; if so, for each first subset, based on the time information of each original ocean observation profile data therein, determining that the first subset contains first ocean observation profile data with a non-latest record time; deleting the first ocean observation profile data of the first subset from the multiple original ocean observation profile data; The processing module is further configured to perform high-frequency cleaning on data that meet the same spatiotemporal judgment conditions in the version-cleaned data according to the time and location information in the metadata, to obtain multiple target ocean observation profile data; the high-frequency cleaning includes averaging the observation data that meet the same spatiotemporal judgment conditions; wherein, performing high-frequency cleaning on data that meet the same spatiotemporal judgment conditions in the version-cleaned data includes: Determine whether at least one second subset exists in the cleaned data; the original ocean observation profile data in the second subset meet the same spatiotemporal judgment conditions; the same spatiotemporal judgment conditions include: the original ocean observation profile data in the second subset correspond to the same target equipment type; the original ocean observation profile data in the second subset all correspond to the same observation depth; the observation time of the original ocean observation profile data in the second subset is within a preset time resolution range; and the maximum longitude difference and maximum latitude difference of the original ocean observation profile data in the second subset are respectively less than their respective preset thresholds; if it exists, for each second subset, the observation data of the data that meet the same spatiotemporal judgment conditions are averaged to obtain the target observation data corresponding to the second subset; the metadata of any original ocean observation profile data is combined with the target observation data to obtain the second ocean observation profile data of the second subset; the second ocean observation profile data is added to the cleaned data, and the original ocean observation profile data in the second subset is deleted; The processing module is further configured to perform matrix processing on the metadata and observation data of the target ocean observation profile data to obtain corresponding normalized data; the normalized data includes the matrix corresponding to the metadata, the matrix corresponding to the observation data, and the matrix corresponding to the observation depth of the target ocean observation profile data; the normalized data of multiple target ocean observation profile data are used to construct the profile dataset within the target time period.

Citation Information

Patent Citations

  • Method and system for combining multi-source Rinex (Receiver Independent Exchange Format) observation files

    CN108491454A

  • Quality control method and system based on buoy observation data

    CN119204782A