Global chlorophyll concentration three-dimensional data reconstruction method and system and computer program

Through multi-source data acquisition and stacked integrated learning model, the shortcomings of three-dimensional chlorophyll concentration distribution reconstruction in the existing technology are solved, and high-precision and highly adaptable global chlorophyll concentration three-dimensional data reconstruction is achieved, which is suitable for marine ecological research, fishery resource assessment and global climate change monitoring.

CN120235055AActive Publication Date: 2025-07-01SECOND INST OF OCEANOGRAPHY MNR

Patent Information

Application Number
CN202510703955.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high-precision three-dimensional chlorophyll concentration distribution reconstruction on a global scale, especially in terms of data acquisition range, feature preprocessing, model fusion strategies and independent data set verification, resulting in errors in optically complex areas or extreme water cluster conditions.

Method used

Multi-source data acquisition, stacked integrated learning model and differentiated preprocessing method are used to combine BGC-Argo profile observation, optical factor and mixed layer depth data, and integrated learning model is constructed through random forest, XGBoost, CatBoost, multi-layer perceptron and K nearest neighbor algorithm to carry out three-dimensional spatiotemporal inversion of chlorophyll concentration.

Benefits of technology

It realizes high-precision three-dimensional chlorophyll concentration distribution reconstruction on a global scale, reduces the impact of data loss and observation errors, improves the model's adaptability to complex water environments, and meets the monitoring needs of high spatiotemporal resolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235055A_ABST
    Figure CN120235055A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of marine ecology and water color remote sensing, in particular to a global chlorophyll concentration three-dimensional data reconstruction method and system and a computer program. The method comprises the following steps: acquiring multi-source data; carrying out space-time matching on multi-source and multi-dimensional data; processing and standardizing multi-type data; filtering and preprocessing abnormal data; constructing an inversion model; training a chlorophyll concentration inversion model; checking an independent data set model; making a spatio-temporal data set to be inverted; according to the method, global chlorophyll concentration three-dimensional data reconstruction capacity is remarkably improved, three-dimensional daily data production of global chlorophyll concentration data is supported, and high-reliability data support is provided for marine ecological research and estimation of global three-dimensional chlorophyll concentration data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of marine ecology and ocean color remote sensing, and particularly relates to a method, a system and a computer program for reconstructing three-dimensional data of global chlorophyll concentration. Background Art

[0002] The distribution characteristics of phytoplankton in the ocean are of great significance to marine primary productivity, global carbon cycle, marine fishery resource management, etc. The growth and reproduction degree of phytoplankton is often measured by "chlorophyll concentration". With the continuous development of marine observation means, using satellite remote sensing methods to conduct rapid and wide-coverage monitoring of the global sea area has become the mainstream technical route. However, due to factors such as the observation characteristics of sensors and the limitations of atmospheric correction, satellite remote sensing products can only provide chlorophyll concentration data at the two-dimensional (2D) sea surface level, and it is difficult to directly depict the distribution in the vertical dimension (Depth dimension) of the ocean. In fact, the variation trend of chlorophyll concentration in seawater with depth is jointly affected by multiple factors such as nutrient levels, temperature-salinity structure, mixed layer depth, and optical attenuation characteristics. Therefore, relying solely on two-dimensional chlorophyll concentration data at the sea surface is difficult to accurately reflect the key processes of the marine ecosystem in the vertical direction.

[0003] In order to make up for the deficiency in understanding the three-dimensional chlorophyll concentration distribution in the ocean, the academic and industrial circles have proposed various data interpolation and three-dimensional reconstruction strategies in recent years. Existing research usually adopts the following methods: 1. Single machine learning or deep learning interpolation For large-area missing regions in remotely sensed chlorophyll concentration at the sea surface, some researchers use methods such as neural networks or convolutional autoencoders to perform complementary interpolation on the missing pixels. However, such methods mostly stay at the two-dimensional sea surface level. Although they can improve the integrity of sea surface observation data to a certain extent, they do not incorporate information on vertical depth. For the chlorophyll concentration of deep sea profiles, without measured profile data or other external auxiliary features, it is often impossible to directly complete the filling.

[0004] 2. Numerical model simulation Some research institutions use physical-biological coupled models to simulate the distribution of chlorophyll in three-dimensional space; however, such methods have extremely high requirements for initial conditions, boundary conditions, model parameterization schemes, observation assimilation, etc. They are often time-consuming and expensive and are easily affected by model biases or uncertainties. For the daily monitoring requirements of high resolution and global scale, numerical models alone cannot fully balance computational efficiency and accuracy.

[0005] 3. Preliminary exploration of three-dimensional reconstruction by multi-source fusion In recent years, some emerging studies have begun to consider combining two-dimensional remote sensing products with vertical nutrient characteristics, temperature-salinity profiles, and a small amount of measured three-dimensional chlorophyll concentrations in order to supplement the three-dimensional distribution within a certain sea area and time range. Chinese Patent Application CN115438523A (Publication Date: December 06, 2022) is one such example. This document discloses a method for reconstructing three-dimensional chlorophyll concentration data in the ocean. By obtaining two-dimensional chlorophyll concentration and temperature data at the sea surface, three-dimensional nutrient distribution influencing factors, and some measured profile chlorophyll data, a reconstruction model is constructed to output three-dimensional chlorophyll concentration data. Compared with traditional studies that can only obtain two-dimensional information, the solution of CN115438523A significantly improves the ability to depict chlorophyll concentration in the vertical direction and can better meet the demand for the three-dimensional distribution of chlorophyll.

[0006] However, with the further in-depth study of marine ecology and the need for high-precision, global-scale three-dimensional monitoring, there are still the following deficiencies or limitations: 1. The types of multi-source data are relatively limited CN115438523A has tried to incorporate three-dimensional nutrient factors (such as temperature-salinity data) to enrich vertical information, but important factors such as marine optical characteristics (such as remote sensing reflectance rrs, diffuse attenuation coefficient kd, etc.) and mixed layer depth mlotst have not been deeply utilized. Without these multi-dimensional observations with a relatively high correlation with chlorophyll concentration, the model's depiction of water body optics and marine dynamic environment is still insufficient, and errors are likely to occur in optically complex areas or extreme water mass conditions.

[0007] 2. The strategies for preprocessing and synchronization of multi-source data need to be improved In three-dimensional reconstruction, multi-dimensional data (two-dimensional / three-dimensional, different time periods, different coordinate systems) often need to be strictly matched in time and space, outliers removed, and distribution differentiation processed. Without systematic and differentiated feature preprocessing methods, it is difficult to maximize the complementary information of multi-source data, and even may increase the model deviation due to unbalanced features.

[0008] 3. The model structure needs to be enhanced The neural network or interpolation method used in CN115438523A is a relatively common deep learning idea. However, when facing different sea areas globally, more environmental factors, or higher resolutions, this single model or single structure is prone to overfitting or inaccurate prediction of local complex water bodies. How to use advanced algorithms such as Stacking and multi-base learner fusion to balance the prediction requirements of different waters, different optical conditions, and different depths is still a key technical problem of concern in the industry.

[0009] 4. The verification and generalization evaluation of independent data sets are insufficient Large-scale three-dimensional chlorophyll concentration reconstruction has high requirements for data integrity and model robustness. Without independent datasets (or profile data in more seasons and more sea areas) to evaluate the generalization of the model, it may lead to accurate reconstruction results under specific conditions, but large deviations may occur when environmental changes or geographical regional differences are significant. To be truly applied to long-term monitoring of the global ocean, more rigorous and multi-dimensional tests are needed after training.

[0010] In summary, the existing technologies (including the solution disclosed in CN115438523A) have initially achieved the reconstruction from sea surface two-dimensional data and nutrient factors to three-dimensional chlorophyll concentration distribution, significantly improving the understanding of ocean vertical ecological characteristics. However, in the face of the requirements for higher spatio-temporal resolution, optical characteristics of diverse global sea areas, and multi-source observation fusion, further innovations are still needed in aspects such as data acquisition scope, feature preprocessing, model fusion strategy, and independent dataset testing. Summary of the Invention

[0011] To solve the above technical problems, the objective of the present invention is to provide a method for reconstructing three-dimensional data of global chlorophyll concentration. By incorporating richer multi-source data (such as BGC-Argo profile observations, optical factors rrs and kd, mixed layer depth mlotst, etc.) and using a stacked ensemble learning model, differential preprocessing and unified management of multi-dimensional features are carried out, so as to achieve high-precision reconstruction and dynamic complement of three-dimensional chlorophyll concentration distribution globally, providing key technical support for marine ecological research, fishery resource assessment, and global climate change monitoring.

[0012] To achieve the above objective, the present invention adopts the following technical solutions: A method for reconstructing three-dimensional data of global chlorophyll concentration, the method is based on multi-source and multi-type marine environmental data and uses an ensemble learning model to perform three-dimensional spatio-temporal inversion of chlorophyll concentration, specifically including the following steps: 1) Multi-source data collection: Obtain and integrate temperature, salinity, depth, and chlorophyll concentration data from BGC-Argo profile observations, marine mixed layer depth mlotst data, and remote sensing reflectance rrs and downward irradiance diffuse attenuation coefficient kd data at the remote sensing level; 2) Three-dimensional spatio-temporal matching and fusion: For BGC-Argo three-dimensional profile data and the remaining two-dimensional remote sensing data, perform matching at the same longitude, latitude, and time point based on a daily-scale time window and spatial proximity rules, and fuse BGC-Argo data at different depths at the same point with the corresponding mlotst, rrs, and kd data; 3) Multi-type data processing and standardization: The collected features with different physical meanings and distribution characteristics are processed by cosine transformation, logarithmic transformation and normalization to highlight the differences between extreme high latitudes and low latitudes and stretch the numerical regions of different magnitudes; 4) Abnormal data filtering and preprocessing: Use the quartile method to eliminate outliers and delete invalid filler values ​​in each feature; after completing data cleaning, divide part of the data into independent data sets for model testing, and the rest as training data sets; 5) Inversion model construction: Based on five different principle algorithms, including random forest (RF), XGBoost, CatBoost, multi-layer perceptron (MLP) and K-nearest neighbor (KNN), as base learners, and RF as meta-learner, an integrated learning inversion model is constructed through the stacking strategy; 6) Preparation of spatiotemporal data set to be inverted: The high spatial resolution temperature and salinity external data are fused with the mlotst, rrs, and kd data from the same source as the training data, and the spatiotemporal data set to be inverted is generated using the same feature preprocessing method; 7) Three-dimensional data inversion of chlorophyll concentration: The spatiotemporal data set to be inverted is input into an integrated model that has passed the independent data set verification, the chlorophyll concentration of each depth layer is predicted, and the obtained results are reconstructed into a three-dimensional format with the same dimension as the temperature and salinity data.

[0013] Preferably, in step 1), the multi-source data integration includes BGC-Argo profile observation data, ocean mixed layer depth mlotst data, water body downward irradiance diffuse attenuation coefficient kd_490 at a wavelength of 490 nanometers, remote sensing reflectance rrs_412 and rrs_560 at wavelengths of 412nm and 560nm, and at least one published daily product of sea surface chlorophyll concentration, to improve the adaptability of the model to different regions and seasons.

[0014] Preferably, in step 2), the temporal and spatial matching strategy of the BGC-Argo 3D profile data and the 2D data includes: The time window is 0.5 days before and after, and 1 day in total; In space, the latitude and longitude of the BGC-Argo profile are used as the center to search for the nearest remote sensing grid point for pairing. If the distance exceeds the preset threshold, it is considered invalid data; For each depth in the same profile, the paired surface data of mlotst, rrs, and kd maintain consistent values.

[0015] Preferably, in step 3), the processing methods for different types of features include: The difference between high-latitude and low-latitude regions is enhanced by cosine transformation of latitude; The longitude is normalized to the interval [0,1]; The rrs, kd, and chlorophyll concentration characteristics were first logarithmically transformed and then standardized; Relatively concentrated features such as mlotst, temperature, and salinity are directly standardized.

[0016] Preferably, in step 4), abnormal data filtering and preprocessing includes: Outliers were removed by calculating the interquartile range (IQR) of each eigenvalue; Check the distribution histogram of each feature and remove any specific invalid fill values ​​if found; After preprocessing, at least 5% to 20% of the valid data are randomly selected as independent data sets, and the rest are used as training data sets.

[0017] Preferably, in the step 5), the inversion model is constructed using five algorithms including random forest (RF), XGBoost, CatBoost, MLP, and KNN as base learners to construct a stacked ensemble model, wherein RF is used as a meta-learner; the base learner parameters are set as follows: RF uses a tree depth of 15-20 layers and a Gini coefficient splitting criterion; XGBoost enables the DART mode to prevent overfitting, and the learning rate is dynamically adjusted in the range of 0.05-0.1; CatBoost configures an ordered boosting algorithm to process category features and sets more than 500 iterations; MLP constructs a three-hidden layer structure containing 128, 64, and 32 nodes, and ReLU is used as the activation function; KNN dynamically selects 5-15 neighbors through the elbow rule and adopts a weighted distance voting method; the meta-learner generates a prediction feature matrix through 5-fold cross validation, concatenates the output probability of the base learner with the original feature, and inputs a random forest model containing 200 decision trees, sets the feature screening mode to the square root rule, and enables out-of-bag error evaluation to optimize generalization ability.

[0018] Preferably, the method further comprises the following steps: a) Model training and optimization: Perform multiple cross-validations and parameter searches on each base learner and meta-learner in the integrated model in step 5), and iteratively tune the model based on the comprehensive error index; b) Independent dataset model verification: The retained independent dataset is input into the trained integrated model to evaluate the prediction bias and robustness of the model in different sea areas and depth layers.

[0019] Preferably, in step a), during the training of the base model, stratified sampling is performed based on the optical properties of water bodies, the dataset is divided by 10-fold stratified cross-validation, and hyperparameters are optimized by grid search: for RF and XGBoost, the maximum depth, minimum sample splitting threshold, and subsampling ratio are adjusted; for MLP, the learning rate is optimized in the range of 1e-4 to 1e-2 and the batch size is adjusted; for KNN, the distance weight mode is dynamically adjusted; an early stopping mechanism is set during the training process, and the iteration is terminated when the loss of the validation set does not decrease for 3 consecutive rounds; the meta-model adopts a nested cross-validation framework, the overall performance is evaluated by 5-fold cross-validation in the outer layer, and the parameters of the meta-learner are adjusted by 3-fold cross-validation in the inner layer. At the same time, redundant features with a Gini contribution lower than 1% are removed based on feature importance analysis, and finally, the optimal parameter combination and the preprocessing pipeline including standardization, missing value imputation, and outlier removal threshold are solidified.

[0020] Preferably, in step b), the validation metrics include core accuracy metrics, coefficient of determination R², root mean square error RMSE, mean absolute error MAE, and auxiliary analysis metrics: the standard deviation of the predicted values is calculated by quarter and by sea area to evaluate the spatio-temporal stability, the F1-score is calculated for high-value areas with chlorophyll concentration > 5 mg / m³ to evaluate the extreme value capture ability, at the same time, the distribution consistency between the predicted values and the measured values is analyzed through a scatter density map, a residual spatio-temporal heat map is drawn to locate systematic deviation areas, and a Taylor diagram is generated to comprehensively display the collaborative performance of multiple models.

[0021] Preferably, in step 6), when making the spatio-temporal dataset to be inverted, temperature and salinity can use external high-spatial-resolution datasets, and the remaining mlotst, rrs, and kd data have the same source as the training dataset, and prediction is performed using the same processing method and standardization parameters as in the training stage.

[0022] Preferably, in step 7), the preprocessed dataset to be inverted is input into the validation model, and batch calculation is implemented using the Dask library for memory optimization; the output results are filtered by physical thresholds, and the chlorophyll concentration is limited to the range of 0-100 mg / m³, and out-of-limit values are assigned invalid values; the inversion results are reconstructed into a four-dimensional array, time × latitude × longitude × depth, and its vertical dimension is strictly aligned with the GLORYS temperature and salinity data, and the vertical profile data is saved in the NetCDF4 format, supporting multi-dimensional visualization analysis using tools such as Panoply and Matplotlib.

[0023] Furthermore, the present invention also provides a system for reconstructing three-dimensional data of global chlorophyll concentration, which implements the described method, including: A data acquisition module for collecting and integrating multi-source ocean environment data such as BGC-Argo profile observation data, ocean mixed layer depth mlotst, remote sensing reflectance rrs, and diffuse attenuation coefficient kd; A spatio-temporal matching module for matching BGC-Argo three-dimensional profile data with two-dimensional remote sensing data in a daily-scale time window and on spatially adjacent grids; A feature processing module for performing cosine transform, logarithmic transformation, and normalization on multi-type data respectively to highlight the differences between high and low latitudes and stretch the numerical range; An anomaly filtering and preprocessing module for removing anomalies and outliers through quartile method and invalid value detection means, and dividing the training set and the independent data set; A model construction and training module for constructing a stacked ensemble model based on random forest, XGBoost, CatBoost, multi-layer perceptron, and K-nearest neighbor algorithms, and performing multiple cross-validations and parameter optimizations; A model testing module for inputting the independent data set into the trained ensemble model to evaluate the prediction performance of the model at different regions and depths; A data set preparation module for generating a spatio-temporal data set to be inverted in the same standardization and preprocessing manner as in the training stage; A three-dimensional inversion and reconstruction module for applying the qualified ensemble model to the data set to be inverted to obtain the three-dimensional distribution result of chlorophyll concentration, and outputting a result file with the same spatial dimension format as the temperature and salinity data.

[0024] Preferably, the data acquisition module is further connected to external high-spatial-resolution temperature and salinity data sources for uniformly introducing temperature and salinity information into the training set and the data set to be inverted, so as to improve the spatio-temporal accuracy and coverage of the three-dimensional reconstruction of chlorophyll concentration.

[0025] Preferably, the feature processing module includes: A latitude and longitude processing unit for performing cosine transform on latitude and normalizing longitude in the range of [0,1]; An optical and biological feature processing unit for performing logarithmic transformation on rrs, kd, and chlorophyll concentration features and then normalizing them; A mixed layer depth processing unit for directly normalizing mlotst.

[0026] Preferably, the model construction and training module includes: Base learner units for running random forest (RF), XGBoost, CatBoost, multi-layer perceptron (MLP), and KNN models respectively, and outputting prediction results; A meta-learner unit for receiving the prediction results of each base learner, stacking them into new input features for training; A cross-validation and parameter search unit is used to perform grid search or Bayesian optimization under K-fold cross-validation and comprehensively select the optimal parameter combination based on indicators such as MSE, MAE, and R².

[0027] Preferably, the three-dimensional chlorophyll concentration results output by the three-dimensional inversion and reconstruction module are consistent with the temperature and salinity data in terms of spatial resolution and vertical profile, and can be directly used in application fields such as marine ecological environment assessment, global climate change monitoring, and aquaculture fishery resource management.

[0028] Furthermore, the present invention also provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the method is implemented.

[0029] Furthermore, the present invention also provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the method is implemented.

[0030] Due to the adoption of the above technical solution, the present invention incorporates richer multi-source data (such as BGC-Argo profile observations, optical factors rrs and kd, mixed layer depth mlotst, etc.) and uses a stacked integrated learning model to perform differential preprocessing and unified management of multi-dimensional features, thereby achieving high-precision reconstruction and dynamic complementation of the three-dimensional chlorophyll concentration distribution globally, providing key technical support for marine ecological research, fishery resource assessment, and global climate change monitoring. It has the following technical effects: 1. Significantly improve the reconstruction accuracy of three-dimensional chlorophyll data: The present invention introduces multi-source observation information at the data level, including two-dimensional sea surface chlorophyll and temperature, BGC-Argo profile observations, mixed layer depth mlotst, and optical factors such as remote sensing reflectance rrs and diffuse attenuation coefficient kd. Compared with traditional schemes that only rely on sea surface or a small amount of nutrient profile data, it can more accurately simulate and restore the distribution law of chlorophyll concentration in the vertical depth and multi-sea area environment.

[0031] 2. Effectively reduce the impact of data missing and observation errors on the reconstruction results: By adopting differential data cleaning, preprocessing, and standardization methods for different features (temperature and salinity, optical data, mixed layer depth, etc.), the present invention can reduce the interference of noise points and outliers on model training; at the same time, using daily-scale spatio-temporal matching and interpolation methods can significantly alleviate the problem of satellite observation data missing caused by cloud cover, extreme weather conditions, etc., strengthening the spatio-temporal coverage and completeness of the reconstructed three-dimensional chlorophyll concentration.

[0032] 3. The stacked integrated learning model improves the adaptability to complex water environments: The present invention adopts a multi-algorithm fusion at the model level (such as random forest, XGBoost, CatBoost, MLP, KNN) and a stacking integration strategy with RF as the meta-learner. Compared with a single neural network or a single interpolation model, it can more fully explore the non-linear associations between multi-source and multi-dimensional features, reduce the dependence on the limitations of a single algorithm, and help maintain a more stable prediction effect in different sea areas, different seasons or different depth scenarios.

[0033] 4. Achieving high spatio-temporal resolution chlorophyll inversion in the global or large-scale sea areas: By applying the trained integrated learning model to high-resolution temperature, salinity data and other data sources consistent with the training, the present invention can generate a three-dimensional chlorophyll concentration field at the daily scale and even higher spatio-temporal frequencies. Compared with the method relying on numerical model assimilation, the present invention achieves a better balance between computational efficiency and prediction accuracy, and can meet the needs of multi-faceted aspects such as marine ecology, fishery resource assessment and global climate research for high-resolution three-dimensional chlorophyll data.

[0034] 5. Providing more complete inputs for subsequent data visualization and ecological analysis: The three-dimensional chlorophyll concentration data output after reconstruction maintains the same spatial resolution and time series as the observed data such as temperature, salinity and depth, and mixed layer depth, which is convenient for visualization and further analysis. For example, in three-dimensional profiles, time series graphs or marine ecological process simulations, the dynamic changes of chlorophyll in different sea areas and depth layers can be presented more comprehensively and intuitively, providing key technical support for marine environmental monitoring and ecological assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a flowchart of the method for reconstructing three-dimensional global chlorophyll concentration data.

[0036] Figure 2 is a schematic diagram of the model verification results.

[0037] Figure 3 is a schematic diagram of the model inversion result Figure 1 .

[0038] Figure 4 is a schematic diagram of the model inversion result Figure 2 .

[0039] Figure 5 is a comparison schematic diagram of the model inversion result and the reference data. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0041] like Figure 1 As shown, the global chlorophyll concentration three-dimensional data reconstruction method proposed in the present invention mainly includes the following nine steps: 1. Multi-source data acquisition; 2. Multi-source and multi-dimensional data spatiotemporal matching; 3. Multi-type data processing and standardization; 4. Abnormal data filtering and preprocessing; 5. Inversion model construction; 6. Model training and optimization; 7. Independent data set verification; 8. Preparation of data sets to be inverted; 9. Spatiotemporal data inversion and post-processing. Each step will be described in detail below.

[0042] Step 1: Multi-source data collection; In the process of marine environmental research and data assimilation, accurate acquisition of chlorophyll concentration often requires consideration of the coupling relationship between multiple physical, chemical and optical characteristics. The present invention uses the following data sources to achieve effective reconstruction of the three-dimensional chlorophyll concentration of the ocean: 1. BGC-Argo profile data Temperature, salinity, depth and chlorophyll fluorescence: BGC-Argo floats collect underwater profile information in an unevenly distributed manner in the global ocean. For each BGC-Argo profile, data on several depth layers at the corresponding longitude and latitude and time point can be obtained, including temperature, salinity, and chlorophyll concentration measured by fluorescence sensors.

[0043] Depth information: The depth sampling of BGC-Argo profiles can usually extend to 1000 meters or even 2000 meters, depending on the float hardware configuration and the intended observation target. In the present invention, a depth range that can reflect the main characteristics of vertical chlorophyll distribution is selected, such as the interval of 0 to 1000 meters.

[0044] 2. Ocean mixed layer depth (mlotst) This example obtains ocean mixed layer depth data from the CMEMS (Copernicus Marine Environment Monitoring Service) platform. The mixed layer depth has an important impact on the transport of nutrients in the surface and subsurface layers, light conditions, and phytoplankton growth, and is indispensable in the reconstruction of three-dimensional chlorophyll distribution.

[0045] 3. Water optical data (kd, rrs) kd_490: Diffuse attenuation coefficient of downwelling irradiance at 490 nm wavelength, which can quantitatively describe the optical transparency of water body and the light attenuation rate; rrs_412 / rrs_560: Remote sensing reflectance at 412 nm and 560 nm wavelengths respectively, which can reflect the scattering and absorption characteristics of seawater in the blue and green spectral bands.

[0046] In this embodiment, the OC-CCI (Ocean Colour Climate Change Initiative) multi-sensor fusion product is preferably used to improve the coverage and consistency of data in the global ocean area.

[0047] 4. Sea surface chlorophyll concentration benchmark (Modis data) Although BGC-Argo provides in-situ chlorophyll profile data, there are often richer satellite remote sensing data available as a control or reference in the sea surface layer. The present invention uses the sea surface chlorophyll concentration product of the Modis (Aqua / Terra) satellite as a benchmark to test or assist in interpolating missing data, so as to effectively utilize a large amount of sea surface information in the training stage.

[0048] Through the acquisition and integration of the above multi-source data, the present invention obtains relatively complete ocean environmental elements, laying a foundation for the subsequent three-dimensional chlorophyll inversion. It should be noted that these data have different spatio-temporal resolutions and data formats, and need to be aligned, interpolated and standardized in the subsequent steps.

[0049] Step 2: Spatio-temporal matching of multi-source and multi-dimensional data Since the BGC-Argo data is in the form of a three-dimensional profile (longitude, latitude, depth), while other parameters (such as mlotst, rrs, kd, etc.) are mostly two-dimensional (longitude, latitude) or sea surface layer data, reasonable spatio-temporal matching must be carried out. The following method is adopted in this embodiment: 1. Setting of time window Select 1 day as the time matching window. If the observation time of the BGC-Argo profile is 12:00 on a certain day, the remote sensing or mixed layer depth data within that day (0:00 - 24:00) is regarded as the same day's observation. If the cross-day difference does not exceed 12 hours, the influence is generally considered negligible.

[0050] 2. Spatial proximity criterion The BGC-Argo profiles have specific longitude and latitude (loni, lati) (\text{lon}_i, \text{lat}_i) (loni, lati). For two-dimensional grids (such as mlotst, kd, etc.) on the same day, the nearest neighbor or bilinear interpolation is used to find the match with the profile coordinate points. If the distance exceeds 1 / 4° or 4 km (depending on the resolution), there is no available supporting data for this profile point on that day.

[0051] 3. Depth homogenization Each depth layer of the three-dimensional profile of BGC-Argo corresponds to the actually measured temperature, salinity, and chlorophyll values; for the same profile, parameters such as mlotst, rrs, kd, etc. at different depths have the same values (equivalent to assuming that there are no significant differences vertically, only mapping the surface data). This allows each depth layer to have the same external environmental characteristic information during subsequent modeling, but the chlorophyll concentration may vary with depth.

[0052] Through the above steps of time matching, spatial proximity, and depth mapping, a unified multi-dimensional sample record can be formed in the dataset, that is: , where each dimension element is aligned on the same record, facilitating subsequent feature processing and training.

[0053] Step 3: Processing and standardization of multi-type data Different elements vary greatly in numerical range, physical meaning, and distribution characteristics. To enable the machine learning model to better capture potential patterns, the present invention performs differential processing and standardization for each feature: 1. Latitude cosine transformation In oceanographic research, there are significant differences in climate, ocean dynamics, and optical conditions between low latitudes (tropics) and high latitudes (polar regions). To highlight the clustering characteristics between the polar and tropical regions and also considering the transition zone in the mid-latitude region, the present invention converts the latitude Lat to cos(Lat×Π / 180) to make the difference between high and low latitudes more obvious.

[0054] 2. Longitude normalization The longitude Lon is normalized to the interval [0,1]: , This avoids discontinuity when crossing the 180° prime meridian and is also conducive to subsequent model reading.

[0055] 3. Logarithmic transformation and standardization For features such as rrs, kd, and chlorophyll concentration that often exhibit power-law or exponential distributions, a logarithmic stretch is first performed: X′ = log(X + ∈), where ∈ is a very small positive number (to prevent taking the logarithm when X = 0). Then, perform Z - score standardization on X: , where μ X′ is the sample mean and σ X′ is the sample standard deviation.

[0056] 4. Standardization of the mixed layer depth (mlotst) Since most of the values of mlotst are concentrated between dozens of meters and hundreds of meters, its distribution does not have a large - scale exponential change like rrs and kd. It can be directly normalized using the Z - score or min - max method.

[0057] If mlotst max < 2000, the following can be used , or a similar linear stretching method, which can be determined according to the data distribution.

[0058] Through the above steps, multi - source and multi - type features can be mapped to a relatively "balanced" numerical range, reducing the impact of different dimensions and distribution forms on the machine learning model, thereby improving the modeling effect.

[0059] Step 4: Abnormal data filtering and pre - processing In the scenario of multi - source data fusion, the original data usually has different degrees of missing values, noise values, and even outliers. If not processed, the abnormal values may introduce large biases in the subsequent training and inversion stages, resulting in a decline in the generalization ability of the model. Therefore, systematic abnormal filtering and pre - processing of the data are required. The specific contents are as follows: 1. Removing outliers by the interquartile method Principle: The interquartile method (Interquartile Range, IQR) mainly calculates the first quartile Q1 and the third quartile Q3, and the interquartile range IQR = Q3 - Q1, so as to set a threshold range for distinguishing normal data from abnormal values. The common practice is that data outside [Q1 - 1.5×IQR, Q3 + 1.5×IQR] is regarded as an outlier.

[0060] Implementation details: 1) For each key feature (such as sea - surface chlorophyll concentration, temperature, salinity, mixed layer depth mlotst, optical properties rrs and kd, etc.), calculate its quartiles respectively, and judge whether it is an outlier according to the IQR principle; 2) For extreme outliers, such as certain deep chlorophyll concentrations soaring to hundreds of mg / m³ or temperatures falling below physically unreasonable negative ranges, additional review can be conducted in combination with oceanographic prior knowledge or historical distribution ranges; 3) If determined to be an outlier, it is removed from the dataset to avoid bringing such extreme anomalies into model training.

[0061] 2. Checking for invalid values in the data distribution histogram Purpose: Some data providers use placeholders (such as -999, 9999, NaN) to fill in data when observation fails or algorithm anomalies occur. If such invalid values are not screened out, it will seriously affect the training process.

[0062] Steps: 1) Draw a distribution histogram or box plot for each feature, focusing on whether there are abnormal peaks or sharp breaks in the distribution; 2) If there is a non-physical dense accumulation in certain numerical ranges, such as a large number of -999 or 9999, it is regarded as an invalid value; 3) Delete or mark and remove the records of invalid values to ensure that the remaining samples are all data points with reasonable physical meanings.

[0063] 3. Dataset division After filtering out abnormal data, the cleaned dataset needs to be divided into a training set and an independent dataset: 1) Training set: Used for model training and cross-validation. It occupies most of the total data volume and is the main source for the model to learn the distribution laws of the ocean environment and chlorophyll concentration; 2) Independent dataset: Used for final performance verification, does not participate in any training and parameter tuning of the model, and ensures an objective and fair evaluation of the model.

[0064] Division strategy: In this method, generally, a part (recommended not less than 20% - 30%) is randomly selected from the overall data as the independent dataset, or strictly segmented according to spatio-temporal characteristics (such as different years, different geographical regions) to ensure that there is no overlap in time and space between training and verification. This can truly reflect the generalization performance of the model in unknown regions.

[0065] Through the above preprocessing, most outliers and invalid observation points can be removed, obtaining higher-quality and more consistent samples, laying a solid foundation for model construction and training in the subsequent steps.

[0066] Step Five: Inversion model construction To fully explore the potential relationship between multi-source ocean data and chlorophyll concentration and improve the prediction accuracy in different water masses and spatio-temporal environments, the present invention adopts the Stacking Ensemble Learning method to converge the prediction capabilities of multiple algorithm models. Specifically, five types of algorithms, namely Random Forest (RF), XGBoost, CatBoost, Multi-Layer Perceptron (MLP), and KNN, are selected as the Base Learners, and Random Forest (RF) is used as the MetaLearner. The key algorithms and parameter configurations are described as follows: 1. Random Forest (RF) Base Learner Tree depth: 15 - 20 layers; Splitting criterion: Gini coefficient; Explanation: Deep trees help capture complex non-linear features but may cause overfitting; therefore, subsequent cross-validation and parameter search are needed to balance the model depth and generalization ability.

[0067] 2. XGBoost DART mode: Enable the Dropouts meet Multiple Additive Regression Trees algorithm with a dropout mechanism to help suppress overfitting; Learning rate: Dynamically adjusted between 0.05 and 0.1 for fine-tuning in different training stages; Explanation: XGBoost has advantages in dealing with sparse data and feature importance screening. The DART can further avoid overfitting to small-scale data or local features.

[0068] 3. CatBoost Core features: Ordered Boosting and Gradient Bias minimization; Number of iterations: More than 500 times; Explanation: CatBoost has good adaptability to categorical features, missing values, and imbalanced data. It can also have a certain degree of fault tolerance and robustness for the hierarchical features or optical category conversions that may occur in ocean observations.

[0069] 4. Multi-Layer Perceptron (MLP) Network structure: Consists of 3 hidden layers with 128, 64, and 32 nodes respectively; Activation function: ReLU; Description: MLP is suitable for mining the non-linear relationships of high-dimensional continuous features; in the present invention, temperature-salinity, optics, geographical coordinates, etc. can all be used as input features. With multiple hidden layers, the network can abstract potential laws layer by layer.

[0070] 5. KNN Number of neighbors: Dynamically selected between 5 and 15 through the "elbow method"; Voting method: Weighted distance method (the closer the distance, the higher the weight); Description: KNN has an intuitive advantage in small sample or local distribution determination and is suitable for making up for local details that may be ignored by other complex models. It can effectively supplement some heterogeneous regions in ocean data (such as the boundary between the inshore and the open ocean).

[0071] After building the five base learners, they are combined to form the first-layer model; immediately afterwards, the predicted outputs of each base learner (usually the predicted values or probability distributions in the regression scenario) are generated through 5-fold cross-validation. Then these prediction results are concatenated with the original features to form a new feature matrix, which is then input into the random forest (RF) serving as the meta-learner.

[0072] Meta-learner (RF): 1) It contains 200 decision trees; 2) Feature screening mode max_features='sqrt', that is, only randomly select within the range of the square root of the number of features during each split; 3) Enable out-of-bag error evaluation (OOB), which can dynamically evaluate the generalization error during the training process and assist in tuning.

[0073] This stacked structure can better take into account the advantages of multiple algorithms compared to a single model: for example, XGBoost and CatBoost have strong performance in the direction of gradient boosting, RF is good at feature randomization and integration, while MLP and KNN can capture complex non-linear and local neighbor relationships. By the meta-learner re-synthesizing the outputs of the first layer, the adaptability of the model to the large-scale and multi-factor ocean environment can be effectively improved.

[0074] Step Six: Model Training and Optimization After the model architecture is determined, each base learner (RF, XGBoost, CatBoost, MLP, KNN) and the meta-learner (RF) need to be trained and parameter-tuned. To ensure that the model has sufficient generalization ability in terms of time, region, etc., the present invention emphasizes the cross-validation and parameter search strategies, and designs measures such as stratified sampling and early stopping mechanism according to the characteristics of ocean data: 1. 10-fold stratified cross-validation for the base learners Stratification criteria: Stratify by combining water optical characteristics (such as kd, rrs) and chlorophyll concentration segments (such as 0 - 1 mg / m³, 1 - 5 mg / m³, 5 - 10 mg / m³, >10 mg / m³). This can avoid uneven distribution of high or low concentration areas after splitting the training samples; Training - validation division: Divide the training set into 10 equal parts. Select 1 part as the validation set in each round, and the remaining 9 parts as the training set; After performing this 10 times in turn, comprehensively measure the average performance; Grid search: 1) For Random Forest (RF) and XGBoost: Adjust the maximum depth (such as 6, 8, 10, 12, 15), minimum sample split threshold, and subsampling ratio (0.6 - 1.0); 2) For MLP: Search for the learning rate between 1e - 4 and 1e - 2, and try batch sizes (batch_size) of 32, 64, and 128; 3) For KNN: Set two modes of equal weight or inverse distance weighting for comparison, and find the optimal number among 5 - 15 neighbors; Early stopping mechanism: When the validation set loss or error does not significantly decrease for 3 consecutive rounds, stop training to prevent overfitting in mini - batch updates.

[0075] 2. Nested cross - validation of the meta - learner Outer - layer 5 - fold evaluation: Based on the first - layer 10 - fold validation, perform another round of 5 - fold cross - validation for the meta - learner to ensure balanced evaluation of the final performance when combining predictions of different base learners; Inner - layer 3 - fold adjustment of meta - learner parameters: Include the number of decision trees, maximum depth, etc.; Feature importance screening: Based on the trained Random Forest meta - learner, statistically calculate the Gini contribution of each feature, and eliminate redundant features with a contribution less than 1% to further simplify the model and reduce the risk of overfitting; Solidify the optimal parameter combination: Based on the validation results of different folds, obtain the optimal hyperparameters (such as tree depth, learning rate, number of iterations, number of neighbors) of each base learner and the meta - learner, and retain the configuration of the corresponding data pre - processing pipeline (standardization, missing value imputation, outlier removal threshold, etc.) to form the final deployable model.

[0076] Through the above training and optimization process, the model can form a more robust understanding of chlorophyll concentration changes in multiple dimensions (optical properties, depth, space - time), and can avoid overfitting as much as possible with limited training samples, achieving accurate prediction of the three - dimensional distribution of chlorophyll in a large range and even the global sea area.

[0077] Step Seven: Independent dataset validation After the model obtains the optimal hyperparameters, it is necessary to conduct a final performance evaluation on the previously excluded independent dataset. This independent dataset is strictly separated from the training set in terms of time (2020 - 2022) and space, and does not contain any training samples to ensure the objectivity of the evaluation. The specific evaluation methods are as follows: 1. Core accuracy metrics Coefficient of determination R 2 : Measures the overall goodness of fit between the predicted values and the measured values of the model. The closer it is to 1, the more accurate it indicates; Root Mean Square Error (RMSE): Can better reflect the sensitivity of extreme value errors and is suitable for evaluating the prediction of high - concentration chlorophyll regions; Mean Absolute Error (MAE): More intuitively measures the average deviation. The lower the value, the smaller the overall error.

[0078] 2. Spatiotemporal stability analysis Calculate the standard deviation of the predicted values by quarter (spring, summer, autumn, winter) or by sea area (coastal, open ocean, low - latitude, high - latitude, etc.) to evaluate the sensitivity of the model to seasonal and regional changes. If the standard deviation is too large, it indicates that the prediction is unstable in some spatiotemporal intervals, and the reasons need to be further analyzed.

[0079] 3. Accuracy of high - value regions (F1 - score) Conduct discriminant statistics on high - value regions where the chlorophyll concentration > 5 mg / m³, and calculate the F1 - score as a comprehensive index of precision and recall to test the model's recognition ability in eutrophic or algal bloom areas.

[0080] 4. Consistency of the distribution of predicted values and measured values Check whether the predicted values - measured values are distributed near the diagonal through a scatter density plot; if there is an obvious deviation or large - area dispersion, it indicates that the model has systematic errors or poor local fitting.

[0081] 5. Residual spatiotemporal heat map Draw a heat map of the residuals obtained by subtracting the measured values from the predicted values in the time - space dimension, which can quickly locate whether there are continuous deviations in certain months or sea areas, helping for subsequent targeted improvement.

[0082] 6. Taylor Diagram Comprehensively compare the standard deviation, correlation coefficient, and mean square error of multiple models or the same model under different stratification settings, visually display the performance of the model in terms of spatiotemporal distribution consistency, and also help analyze the collaborative performance between base learners and meta - learners.

[0083] If in this independent dataset, the overall R of the model 2 is at a relatively high level (e.g., above 0.8), both RMSE and MAE are small, and the F1-score in the high-value area also has good performance, then it can be determined that the model has feasible generalization ability in real unknown scenarios and can enter the next stage of ocean data inversion and large-scale application.

[0084] Step Eight: Preparation of the dataset to be inverted After completing model training and independent dataset verification, predictions can be made on larger-scale or completely new sea area spatio-temporal data. At this time, it is necessary to prepare the dataset to be inverted and input it into the verified model. To avoid a decline in prediction performance caused by inconsistent data distributions, the present invention strictly maintains consistency with the training stage in terms of data sources and preprocessing processes, including the following details: 1. Data sources and spatio-temporal resolutions Temperature and salinity data: Use the GLORYS12V1 product of the CMEMS platform, with a horizontal resolution of 1 / 12° (about 8 km) and 50 layers in the vertical direction; in this embodiment, interpolation or scaling is used to meet the subsequent 4 km resolution requirements.

[0085] Optical parameters (a, kd): From the MODIS Aqua L3 satellite product, with a general time resolution of daily scale; to be consistent with the model input features, interpolation is also required to a 4 km spatial resolution and a 1-day time resolution.

[0086] Mixed layer depth mlotst: The mixed layer depth field corresponding to the time and space resolutions can be directly obtained from the CMEMS database. If the resolution differences are large, unification is also required.

[0087] 2. Data resampling and matching Spatial resampling: Use bilinear interpolation to unify data with a resolution of 1 / 12° or coarser to 4 km. If the data source resolution is already less than or equal to 4 km, downsampling or nearest neighbor interpolation can be used; Time alignment: Set the time window to 1 day. Whether it is optical satellite data or temperature and salinity reanalysis data, it needs to be mapped to the same day. If there is a slight time difference (such as night or afternoon orbits), it can be regarded as the same observation period within 24 hours to reduce data gaps.

[0088] 3. Consistent preprocessing Including coordinate scaling (normalizing longitude and latitude to [0, 1] or [-1, 1]), logarithmic transformation (if log(chl-a), log(SST), etc. have been used), outlier rejection threshold (consistent with training), and feature standardization (such as mean-variance normalization), etc. At the same time, retain the missing value imputation mode and feature selection scheme determined during training to ensure that the data distribution is as consistent as possible with the training data when input into the model.

[0089] 4. Spatial matching window To ensure the coordinate correspondence between ocean profile data and surface / satellite data, the present invention usually sets the spatial window to 4 km; that is, if there are no valid grid points within 4 km, it is regarded as invalid data and the inversion of this point is skipped. After completing this step, an input feature set highly matching the training process can be obtained, covering the temperature, salinity data, optical parameters, longitude, latitude, time information, etc. of each depth layer, providing a basis for subsequent spatio-temporal data inversion.

[0090] Step Nine: Spatio-temporal data inversion and post-processing When the prepared dataset to be inverted is standardized and matched through the above steps, it can be input into the verified stacked integrated model for three-dimensional prediction of chlorophyll concentration. To ensure the efficiency and reasonable memory usage in large-scale data calculations, the present invention recommends using Dask or other distributed computing frameworks for batch prediction. The main process is as follows: 1. Batch block prediction Split according to the spatial or temporal dimension of the dataset, such as in units of days or longitude-latitude grids, and cut the data into several chunks, and call the trained model for prediction chunk by chunk; this can make full use of parallel computing and avoid memory bottlenecks caused by loading too much data at once.

[0091] 2. Physical threshold filtering Set a reasonable bio-optical range for the output chlorophyll concentration, such as 0 - 100 mg / m³; if it exceeds this range, assign an invalid value or fallback to the interpolation threshold; this threshold can be based on oceanographic literature or historical observation statistics to ensure physical compliance and eliminate possible numerical explosion phenomena.

[0092] 3. Data reconstruction and storage Store the prediction results in the form of a four-dimensional array:

[0093] where t is time, lat and lon are latitude and longitude respectively, and depth is the vertical profile layer; Align strictly with GLORYS thermohaline data in the vertical direction to ensure that subsequent analyses can conveniently compare chlorophyll with thermohaline characteristics at the same layer; it is recommended that the output file adopt the NetCDF4 format, which can not only maintain the multi-dimensional array structure but also attach variable names, attribute descriptions, coordinate information, etc.

[0094] 4. Visualization and Further Analysis Use general tools (such as Python plotting libraries like Panoply or Matplotlib) to visualize the results through 3D or 2D slicing; users can draw sea surface chlorophyll distribution maps, depth profile cross-section diagrams, and time series animations according to their needs to visually evaluate the dynamic changes of the reconstructed chlorophyll concentration in different regions and at different times.

[0095] The finally output 3D chlorophyll concentration data can be used in multiple fields such as marine ecological research, fishery resource assessment, algal bloom early warning, and global climate change monitoring, providing key ecological indicators with high precision and high spatio-temporal resolution for scientists and decision-making departments. In the example of this article, Figure 2 the verification results of the model on an independent dataset can be shown, Figure 3 , Figure 4 while presenting an example of the inversion results of the model at a typical moment or region, and Figure 5 comparing the inversion results with reference data (such as measured profiles) to help readers more intuitively understand the distribution of 3D chlorophyll and the model accuracy.

[0096] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0097] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1means for the functions specified in one or more blocks.

[0098] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one Figure 1 one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks.

[0099] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one Figure 1 one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks.

[0100] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0101] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0102] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0103] The foregoing is a description of embodiments of the present invention. Through the above description of the disclosed embodiments, those skilled in the art can implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for reconstructing three-dimensional data of global chlorophyll concentration, characterized in that The method is based on multi-source and multi-type marine environmental data and uses an integrated learning model to perform three-dimensional spatiotemporal inversion of chlorophyll concentration, and specifically includes the following steps: 1) Multi-source data acquisition: Acquire and integrate temperature, salinity, depth and chlorophyll concentration data from BGC-Argo profile observations, ocean mixed layer depth mlotst data, and remote sensing reflectance rrs and downward irradiance diffuse attenuation coefficient kd data at the remote sensing level; 2) Three-dimensional spatiotemporal matching and fusion: For BGC-Argo three-dimensional profile data and other two-dimensional remote sensing data, matching is performed at the same longitude and latitude and time point based on the daily time window and spatial neighbor rule, and the BGC-Argo data at different depths at the same point are fused with the corresponding mlotst, rrs, and kd data; 3) Multi-type data processing and standardization: The collected features with different physical meanings and distribution characteristics are processed by cosine transformation, logarithmic transformation and normalization to highlight the differences between extreme high latitudes and low latitudes and stretch the numerical regions of different magnitudes; 4) Abnormal data filtering and preprocessing: Use the quartile method to eliminate outliers and delete invalid filler values ​​in each feature; after completing data cleaning, divide part of the data into independent data sets for model testing, and the rest as training data sets; 5) Inversion model construction: Based on five different principle algorithms, including random forest (RF), XGBoost, CatBoost, multi-layer perceptron (MLP) and K-nearest neighbor (KNN), as base learners, and RF as meta-learner, an integrated learning inversion model is constructed through the stacking strategy; 6) Preparation of spatiotemporal data set to be inverted: The high spatial resolution temperature and salinity external data are fused with the mlotst, rrs, and kd data from the same source as the training data, and the spatiotemporal data set to be inverted is generated using the same feature preprocessing method; 7) Three-dimensional data inversion of chlorophyll concentration: The spatiotemporal data set to be inverted is input into an integrated model that has passed the independent data set verification, the chlorophyll concentration of each depth layer is predicted, and the obtained results are reconstructed into a three-dimensional format with the same dimension as the temperature and salinity data.

2. The three-dimensional data reconstruction method of global chlorophyll concentration according to claim 1, wherein: In the step 1), the multi-source data integration includes BGC-Argo profile observation data, ocean mixed layer depth mlotst data, water body downward irradiance diffuse attenuation coefficient kd_490 at a wavelength of 490 nanometers, remote sensing reflectance rrs_412 and rrs_560 at wavelengths of 412nm and 560nm, and at least one published daily product of sea surface chlorophyll concentration, to improve the adaptability of the model to different regions and seasons.

3. The three-dimensional data reconstruction method of global chlorophyll concentration according to claim 1, wherein: In step 2), the temporal and spatial matching strategies of the BGC-Argo 3D profile data and the 2D data include: The time window is 0.5 days before and after, and 1 day in total; In space, the latitude and longitude of the BGC-Argo profile are used as the center to search for the nearest remote sensing grid point for pairing. If the distance exceeds the preset threshold, it is considered invalid data; For each depth under the same profile, the paired mlotst, rrs, and kd surface layer data maintain consistent values; And / or, in step 3), the processing methods for different types of features include: Enhancing the difference between high-latitude and low-latitude regions for latitude through cosine transformation; Normalizing longitude to the interval [0,1]; First performing logarithmic transformation on the rrs, kd, and chlorophyll concentration features, and then standardizing them; Directly performing standardization processing on features such as mlotst, temperature, and salinity with relatively concentrated distributions; And / or, in step 4), the abnormal data filtering and preprocessing include: Removing outliers by calculating the interquartile range (IQR) of each feature value; Checking the distribution histograms of each feature, and removing them if specific invalid filling values are found; After preprocessing, at least 5% - 20% of the valid data is randomly selected as the independent dataset, and the rest is used as the training dataset; And / or, in step 5), the inversion model is constructed using five algorithms, namely Random Forest (RF), XGBoost, CatBoost, MLP, and KNN, as the base learners to construct a stacked ensemble model, where RF is used as the meta-learner; The parameters of the base learners are set as follows: RF uses a tree depth of 15 - 20 layers and the Gini coefficient splitting criterion; XGBoost enables the DART mode to prevent overfitting, and the learning rate is dynamically adjusted in the interval of 0.05 - 0.1; CatBoost configures the ordered boosting algorithm to process categorical features and sets more than 500 iterations; MLP constructs a three-hidden layer structure containing 128, 64, and 32 nodes, and the ReLU activation function is selected; KNN dynamically selects 5 - 15 nearest neighbors through the elbow method and uses the weighted distance voting method; the meta-learner generates a prediction feature matrix through 5-fold cross-validation. After splicing the output probabilities of the base learners with the original features, they are input into a random forest model containing 200 decision trees. The feature screening mode is set to the square root rule, and at the same time, the out-of-bag error evaluation is enabled to optimize the generalization ability.

4. The three-dimensional data reconstruction method of global chlorophyll concentration according to claim 1, characterized in that: The method further includes the following steps: a) Model training and optimization: Perform multiple cross-validations and parameter searches on each base learner and meta-learner in the ensemble model in step 5), and iteratively optimize the model based on the comprehensive error index; b) Independent dataset model testing: Input the reserved independent dataset into the trained ensemble model to evaluate the prediction deviation and robustness of the model in different sea areas and depth layers.

5. The three-dimensional data reconstruction method of global chlorophyll concentration according to claim 4, characterized in that: In step a), during the training of the base model, stratified sampling is performed based on the optical properties of water bodies. The dataset is divided through 10-fold stratified cross-validation, and hyperparameters are optimized by combining grid search: for RF and XGBoost, the maximum depth, minimum sample split threshold, and subsampling ratio are adjusted; for MLP, the learning rate is optimized within the range of 1e-4 to 1e-2 and the batch size; for KNN, the distance weight mode is dynamically adjusted; an early stopping mechanism is set during the training process, and the iteration is terminated when the loss of the validation set does not decrease for 3 consecutive rounds; the meta-model adopts a nested cross-validation framework, with 5-fold evaluation of the overall performance in the outer layer and 3-fold adjustment of the meta-learner parameters in the inner layer. At the same time, redundant features with a Gini contribution of less than 1% are removed based on feature importance analysis, and finally, the optimal parameter combination and the preprocessing pipeline including standardization, missing value imputation, and outlier removal threshold are solidified. And / or, in step b), the validation metrics include core accuracy metrics, the coefficient of determination R², root mean square error RMSE, mean absolute error MAE, and auxiliary analysis metrics: the standard deviation of the predicted values is calculated by quarter and by sea area to evaluate the spatio-temporal stability, the F1-score is calculated for high-value regions with chlorophyll concentration > 5 mg / m³ to evaluate the extreme value capture ability, at the same time, the distribution consistency between the predicted values and the measured values is analyzed through a scatter density map, a residual spatio-temporal heat map is drawn to locate systematic deviation regions, and a Taylor diagram is generated to comprehensively display the collaborative performance of multiple models.

6. The three-dimensional data reconstruction method of global chlorophyll concentration according to claim 1, characterized in that: In step 6), when making the spatio-temporal dataset to be inverted, temperature and salinity can use external high-spatial-resolution datasets, and the remaining mlotst, rrs, and kd data have the same source as the training dataset, and prediction is performed using the same processing method and standardization parameters as in the training stage; And / or, in step 7), the preprocessed dataset to be inverted is input into the validation model, and the Dask library is used to achieve batch calculation for memory optimization; The output results are filtered by physical thresholds, and the chlorophyll concentration is limited to the range of 0 - 100 mg / m³. Values exceeding the limit are assigned invalid values; the inversion results are reconstructed into a four-dimensional array, time × latitude × longitude × depth, and its vertical dimension is strictly aligned with the GLORYS temperature and salinity data. The vertical profile data is saved in the NetCDF4 format, supporting multi-dimensional visualization analysis using tools such as Panoply and Matplotlib.

7. A system for reconstructing three-dimensional data of global chlorophyll concentration, characterized in that, This system implements the method described in any one of claims 1 - 6, including: A data acquisition module for collecting and integrating multi-source ocean environment data such as BGC-Argo profile observation data, ocean mixed layer depth mlotst, remote sensing reflectance rrs, and diffuse attenuation coefficient kd; A spatio-temporal matching module for matching BGC-Argo three-dimensional profile data and two-dimensional remote sensing data on a daily-scale time window and a spatially adjacent grid; A feature processing module for performing cosine transformation, logarithmic transformation, and normalization on different types of data respectively to highlight the differences between high and low latitudes and stretch the numerical span; An anomaly filtering and preprocessing module for removing anomalies and outliers through quartile method and invalid value detection means, and dividing the training set and the independent dataset; A model construction and training module, which is used to construct a stacked ensemble model based on random forest, XGBoost, CatBoost, multi-layer perceptron and K-nearest neighbor algorithm, and perform multiple cross-validations and parameter optimizations; A model verification module, which is used to input an independent data set into the trained ensemble model to evaluate the prediction performance of the model at different regions and depths; A data set preparation module, which is used to generate a spatio-temporal data set to be inverted in the same standardization and preprocessing manner as in the training stage; A three-dimensional inversion and reconstruction module, which is used to apply the verified ensemble model to the data set to be inverted to obtain the three-dimensional distribution result of chlorophyll concentration, and output a result file with the same spatial dimension format as the temperature and salinity data; 8. The system according to claim 7, wherein: The data acquisition module is further connected to an external temperature and salinity data source with high spatial resolution, and is used to uniformly introduce temperature and salinity information into the training set and the data set to be inverted, so as to improve the spatio-temporal accuracy and coverage of the three-dimensional reconstruction of chlorophyll concentration; And / or, the feature processing module includes: A latitude and longitude processing unit, which is used to perform cosine transformation on latitude and normalize longitude in the interval [0,1]; An optical and biological feature processing unit, which is used to perform logarithmic transformation on rrs, kd, and chlorophyll concentration features and then normalize them; A mixed layer depth processing unit, which is used to directly standardize mlotst; And / or, the model construction and training module includes: A base learner unit, which is respectively used to run random forest (RF), XGBoost, CatBoost, multi-layer perceptron (MLP) and KNN models, and output prediction results; A meta-learner unit, which is used to receive the prediction results of each base learner, stack them into new input features for training; A cross-validation and parameter search unit, which is used to perform grid search or Bayesian optimization under K-fold cross-validation, and comprehensively select the optimal parameter combination based on indicators such as MSE, MAE, and R².

9. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instruction is executed by a processor, it implements the method described in any one of claims 1-6.

10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by a processor, it implements the method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Ocean three-dimensional chlorophyll concentration data reconstruction method, device, equipment and medium

    CN115438523A

  • Ocean chlorophyll concentration three-dimensional distribution inversion method, terminal and medium

    CN116008267A

  • Lake water chlorophyll concentration inversion method based on Sentinel-2 image

    CN118070237A

  • Inversion method of subsurface chlorophyll concentration based on 1D-DSCAM

    CN119229299A

  • Global chlorophyll remote sensing inversion optimization method and device covering high latitude area

    CN119249901A

Cited By

  • Chlorophyll concentration inversion method and system based on hierarchical feature embedding

    CN120721687A