Method and system for reconstructing sea surface pH remote sensing migration constrained by in-situ pCO2
By constructing pH labels through on-site pCO2 observations and introducing reliability weights, combined with the source domain basic model and target domain residual correction, the problem of insufficient reliability in sea surface pH remote sensing reconstruction was solved, and high-precision and traceable regional sea surface pH prediction was achieved.
Patent Information
- Application Number
- CN202611114464.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-27
- Publication Date
- 2026-08-25
AI Technical Summary
Existing technologies lack a complete closed loop in remote sensing reconstruction of sea surface pH, including label construction, uncertainty quantification, source domain pre-training, target domain adaptation, and result output. This results in insufficient reliability of sea surface pH reconstruction in marginal sea areas, mixed label sources, and unrefined uncertainty processing. Furthermore, there is a lack of effective bias correction when transferring global models to regional areas.
pH tags are constructed through on-site pCO2 observations. A tag reliability weighting system and target domain residual correction are introduced to establish a four-level tag source reliability level. A two-layer structure of source domain basic model and target domain residual correction model is adopted to ensure that tag generation and running variables are isolated, and the credibility information of prediction results is output synchronously.
It achieves more accurate, reliable and traceable sea surface pH reconstruction, improves the prediction accuracy of the target domain, and reduces RMSE by about 23.9%. It is suitable for high-confidence pH reconstruction in marginal seas, shelf seas and open sea areas.
Smart Images

Figure CN122631855A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of marine environmental monitoring and remote sensing data processing technology, and in particular relates to a method and system for remote sensing migration reconstruction of sea surface pH under on-site pCO2 constraints. Background Technology
[0002] The marine carbonate system is a key factor influencing the global carbon cycle and ocean acidification, and sea surface pH is an important indicator of the system's state. Traditional sea surface pH observations rely on on-site shipborne measurements or buoy observations, but due to limitations in observation costs and coverage, global ocean pH observation data exhibits a severe imbalance in its spatiotemporal distribution, particularly in marginal seas and shelf areas.
[0003] To fill observational gaps, existing research has attempted to utilize satellite remote sensing data and reanalysis data to establish global or regional inversion models of sea surface pH using machine learning methods. Jiang et al. (2022, Remote Sensing of Global SeaSurface pH Based on Massive Underway Data and Machine Learning) disclosed a technical solution that uses large-scale underway pCO2 data and total alkalinity (TA) estimated from parameters such as SST, SSS, and location to construct near-field pH labels through carbonate system calculations, and then trains a global sea surface pH model using remote sensing and reanalysis variables. This solution initially achieves the basic route of "pCO2 + TA estimation to calculate pH labels + machine learning inversion," but it still has the following technical shortcomings: First, the sources of labels are mixed and there is a lack of management on the reliability levels of the labels. Current technology mixes direct field pH measurements, derived pH values from pCO2, pH values from gridded products, and model-predicted soft labels for model training and validation, without distinguishing the reliability of labels from different sources. Direct field pH measurements are limited by instrument calibration and field conditions; pCO2-derived pH values are affected by TA estimation errors and the selection of carbonate calculation constants; and gridded products are themselves model reconstruction results rather than independent true values. The reliability of labels from different sources varies fundamentally. Using them with equal weight leads to low-reliability labels contaminating the training effect of high-reliability labels, or high-reliability labels being diluted by low-reliability labels.
[0004] Secondly, there is a lack of uncertainty quantification and weighting mechanisms for each label. While existing technologies have conducted sensitivity analyses on SST and Chl-a input errors, they do not synthesize pCO2 observation errors, TA estimation errors, and SST and SSS matching errors into personalized uncertainties for each label. When the pCO2 observation error corresponding to a label is large or the environmental matching distance is far, the reliability of that label should decrease accordingly. However, existing technologies use a uniform processing method, which fails to achieve refined control over label quality.
[0005] Third, there is a lack of effective bias correction mechanisms when migrating global models to regional target domains. Global source domain data and marginal sea target domain data exhibit distributional differences in carbonate system characteristics, leading to systematic biases when directly extrapolating global models to the target domain. Existing technologies do not establish residual fine-tuning mechanisms for the target domain, making it impossible to regionalize global models using limited in-situ pCO2 data from the target domain.
[0006] Fourth, the lack of effective isolation between label generation variables and runtime variables poses a risk of inflated accuracy. Some studies have used variables such as pCO2, TA, and DIC as input features during model training, resulting in models with seemingly high accuracy. However, in actual product deployment, these variables cannot be continuously acquired through remote sensing, leading to a severe disconnect between training and runtime accuracy. Current technologies do not explicitly isolate these two types of variables, nor do they record the actual usable input patterns for model operation in the output.
[0007] Fifth, the pH reconstruction process lacks a closed-loop, traceable output system. Existing processes typically output only a single predicted value, without simultaneously outputting label uncertainty, sample weights, target domain corrections, and operating mode identifiers. This makes it impossible for users to determine the reliability and applicability of the prediction results.
[0008] In summary, existing technologies for sea surface pH remote sensing reconstruction lack a complete closed loop from label construction, uncertainty quantification, source domain pre-training, target domain adaptation to result output, making it difficult to meet the application requirements for high-reliability pH reconstruction in marginal sea areas. Summary of the Invention
[0009] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for remote sensing migration reconstruction of sea surface pH under on-site pCO2 constraints. It aims to achieve more accurate, reliable and traceable regional sea surface pH reconstruction by constructing highly reliable pCO2-derived pH tags, introducing a tag reliability weighting system and target domain residual correction.
[0010] To achieve the above-mentioned objectives, the first objective of this invention is to provide a method for remote sensing migration reconstruction of sea surface pH under in-situ pCO2 constraint, comprising: Step 1: Obtain on-site pCO2 observation records for the target sea area, match sea surface temperature (SST) and sea surface salinity (SSS) based on the spatiotemporal information of each observation record, and record the environmental matching distance. Step 2: Train the regional total alkalinity (TA) estimation model using regional carbonate observation data isolated from the target testing period, estimate TA using field pCO2 observation records, and obtain the TA estimation error, which serves as the TA input uncertainty in S3. ; Step 3: Input the field pCO2, estimated TA, SST and SSS into the carbonate system to solve for the field pCO2 derived pH label, and determine the sensitivity of pH to field pCO2, estimated TA, SST and SSS. Propagate the TA estimation error to the label uncertainty of the pCO2 derived pH label. Step 4: Classify the available pH labels according to their source reliability level and label them as pH label sources; and determine the sample purpose and sample weight based on the label source reliability level, label uncertainty and environmental matching distance. Step 5: Establish a whitelist of product operation characteristics. Based on the continuous availability of each input variable in different operation modes, exclude the pCO2 and TA variables used for tag generation from operation modes without corresponding continuous input guarantees. Train the source domain basic model using global source domain samples and the whitelisted environmental variables that have been filtered and retained. Determine the operation mode based on the input variables actually used in the inference stage and output the corresponding operation mode identifier. Step 6: Based on the target domain label constrained by sample weights, calculate the residual between the target domain label and the predicted value of the source domain basic model, and train the target domain residual correction model. Combine the predicted value of the source domain basic model with the correction amount output by the target domain residual correction model to obtain the predicted value of sea surface pH of the target domain, and generate an adaptation source associated with the predicted value of sea surface pH of the target domain. The adaptation source is used to characterize that the predicted value of sea surface pH of the target domain is obtained by the target domain residual correction method. Step 7: Associate the label uncertainty, pH label source, sample weight, adaptation source, operating mode identifier, and target domain sea surface pH prediction value with the same sample or grid index and output them.
[0011] Preferably, S1 includes surface screening and deduplication at the same location: records with a depth greater than a preset threshold are removed, and when multiple records exist at the same time and location, the shallowest layer record is retained, wherein the preset threshold is 5m.
[0012] Preferably, the inputs to the TA estimation model include: year center value, month sine and cosine, latitude and longitude, SST, SSS, SST squared, SSS squared, SST-SSS interaction term, and latitude-longitude and salinity interaction term; the TA estimation model is trained and validated by time, flight or spatial partition, and the data of the target test period are not involved in the model fitting.
[0013] Preferably, the formula for calculating the label uncertainty is:
[0014] in, The sensitivity of pH to pCO2, The sensitivity of pH to TA, The sensitivity of pH to SST, The sensitivity of pH to SSS. Due to the uncertainty of pCO2 at the site, Uncertainty estimated for on-site TA Due to the uncertainty of on-site SST, Due to the uncertainty of SSS on site, This is to address the uncertainty of derived pH labeling.
[0015] Preferably, the formula for calculating the sample weight is:
[0016] in, For sample weights, To retain the lowest weight, For the reliability level coefficient of the label source, To address the uncertainty of derived pH labels, The parameter for the scale of label uncertainty decay. To match the distance to the environment, To match the distance attenuation scale parameters to the environment.
[0017] Preferably, the usage constraints of the reliability level of the label source are as follows: direct field pH is used for the highest level training or validation, field pCO2 derived pH is used for target domain training and time hold evaluation, weak product labels are only used for background constraints or pre-training, and soft model labels are only used for distillation or consistency constraints.
[0018] Preferably, the whitelist of product operation characteristics includes at least one or more of the following variables: latitude and longitude, month, SST, SSS, chlorophyll a, wind field components, sea surface height, mixing layer depth and their derived variables; the operation mode includes at least the operation remote sensing mode, pCO2 constraint mode and carbonate consistency mode, and each mode records the available inputs, applicable scope and evaluation indicators respectively.
[0019] Preferably, the fields output by S7 include one or more of the following: model version, target domain correction amount, environment matching distance, input data source, spatial and temporal resolution, and missing measurement processing identifier.
[0020] A second objective of this invention is to provide a remote sensing migration reconstruction system for sea surface pH constrained by in-situ pCO2, comprising: The observation and labeling layer is used for quality screening, environmental matching, and carbonate system calculation of multi-source observations, and converts the on-site pCO2 observations of the target sea area into pCO2-derived pH labels and their uncertainties. The label source reliability classification and weighting layer is used to receive the pCO2-derived pH label and its uncertainty and environmental matching distance, classify the label source reliability, and determine the sample weight. The migration reconstruction and output layer is used to build a source domain basic model using global source domain data, and to complete the regional adaptation of the source domain basic model by target domain labels that have undergone label source reliability classification and sample weight constraints, and finally outputs the sea surface pH prediction results.
[0021] A third objective of this invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method for remotely sensing and reconstructing sea surface pH under field pCO2 constraints.
[0022] A fourth objective of this invention is to provide a computer program product comprising a computer program that, when executed by a processor, implements the aforementioned method for remotely sensing and reconstructing sea surface pH under field pCO2 constraints.
[0023] The advantages and positive effects of this application are: This invention uses in-situ pCO2 observations as a constraint and employs a five-stage closed-loop process—label construction, label source reliability grading, source domain pre-training, target domain adaptation, and isolated output—to achieve high-reliability remote sensing reconstruction of sea surface pH in marginal sea areas. Specifically: This invention transforms in-situ pCO2 observations into pH labels through environmental matching and regional TA estimation, avoiding complete reliance on gridded products as the master truth value. This gives the target domain labels in-situ observation constraints, allowing them to more realistically reflect the characteristics of the regional carbonate system.
[0024] This invention establishes a four-level label source reliability rating system, which distinguishes between direct pH, pCO2-derived pH, weak product labels, and soft model labels. Through label-by-label uncertainty propagation and weighting mechanisms, high-confidence labels play a greater role in training, while the weight of low-confidence labels is reduced, thus avoiding model bias caused by mixing different label source reliability levels.
[0025] This invention employs a two-layer structure of "source domain pre-training + target domain residual correction." The source domain basic model learns globally applicable carbonate environmental relationships, while the target domain residual model utilizes regional field constraints to correct for distribution differences. The synergy between the two significantly improves prediction accuracy. Examples show that, using the technical solution of this invention, the RMSE of target domain pH prediction is reduced by approximately 23.9%.
[0026] This invention generates variables (pCO2, TA, DIC, ...) by running a feature whitelist. The variables are strictly isolated from the product's operating variables to ensure consistency of input between the model training phase and the product's operating phase, thus avoiding a severe decrease in operating accuracy due to the use of variables that cannot be continuously obtained during training.
[0027] This invention simultaneously outputs predicted values, label sources, label uncertainty, sample weights, adaptation sources, and operating mode identifiers. Users can trace the basis and credibility of each prediction result, facilitating product quality control and subsequent analysis.
[0028] This invention can be adapted to other marginal seas, shelf seas, or open seas where direct pH observations are insufficient but pCO2 observations are abundant, simply by replacing the regional pCO2 observations, environmental field, and regional TA empirical model. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1 A flowchart of a preferred embodiment of the present invention is shown; Figure 2 A system block diagram of a preferred embodiment of the present invention is shown. Detailed Implementation
[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] Explanation of the name: TA: Total Alkalinity; SST: Sea Surface temperature; SSS: Sea Surface Salinity; DIC: Dissolved Inorganic Carbon; Aragonite Saturation: The saturation level of aragonite.
[0033] This invention can be applied to marginal seas, shelf seas, or open seas where direct pH observations are insufficient but pCO2 observations are abundant. The basic idea is to avoid directly using gridded pH products as the sole true value. Instead, it prioritizes utilizing regional in-situ pCO2 observations, constructing target domain pH labels with source, error, and weights through environmental variable matching and carbonate system calculations. Then, it trains a basic model using global source domain data, performing residual correction or weighted joint adaptation using target domain samples and a weakly labeled background field. The system operation model uses only long-term, continuously available remote sensing, temperature-salinity, dynamic, and spatiotemporal variables. Label generation variables such as pCO2, TA, DIC, and saturation are only used for label construction, carbonate consistency checks, or constrained modes, and must not be mixed into operation modes without corresponding input guarantees.
[0034] To better understand the technical solution of the present invention, the following is in conjunction with... Figure 1 The method of the present invention will be explained in a non-limiting manner.
[0035] Please see Figure 1 A method for remote sensing migration reconstruction of sea surface pH constrained by in-situ pCO2, mainly including: S1: Read global source domain data and target domain observation data. Acquire global source domain samples, in-situ pCO2 observations of the target sea area, the target domain environmental field, and optional weak labels for gridded products. Global source domain samples are used to train the source domain base model; in-situ pCO2 observations of the target sea area are used to construct target domain pCO2-derived pH labels; the target domain environmental field includes at least sea surface temperature (SST), sea surface salinity (SSS), and location and time information; gridded products are only used as weak labels or background constraints and are not used as independent in-situ validation data. When reading data, record the time, longitude, latitude, sampling depth, parameter type, parameter unit, data source, and quality label for each observation, and distinguish between direct in-situ pH, in-situ pCO2, calculated derived values, product data, and model output.
[0036] S2, perform surface screening and deduplication at the same site. The validity of in-situ pCO2 observations in the target area is checked. Valid observations refer to records with valid sampling time, latitude and longitude, sampling depth, and corresponding carbonate parameters, and that pass the data source quality markup, missing value, and physical extent checks. Sea surface observations are screened according to a preset surface depth threshold, with records having a sampling depth of no more than 5 m being preferred. When multiple records at different depths exist at the same time and location, the shallowest record is retained; when duplicate records exist at the same time, location, and depth, deduplication is performed based on the data source identifier and observation record identifier. Records exceeding the surface depth range, lacking necessary fields, or failing the quality check are not included in the subsequent label construction process. This step outputs the screened and deduplicated surface pCO2 observation records for the target area.
[0037] S3 matches environmental variables corresponding to on-site pCO2 observations. Based on the time and spatial location of each surface pCO2 observation, SST, SSS, and location-time information are matched within preset time and spatial windows. For each match, the time difference, spatial distance, and overall environmental matching distance between the observation and the environmental field are recorded. When the matching result exceeds the preset time or spatial range, the record is marked as a failed match or a low-quality match sample to avoid forcing environmental variables that are too far away to be assigned to the field observation. This step outputs the SST, SSS, location and time information, environmental matching distance, and matching quality identifier corresponding to each field pCO2 observation.
[0038] S4, Estimate the Total Alkalinity (TA) and its error in the target region. Train the regional TA estimation model using regional carbonate observation data isolated from the target test period. Model inputs may include SST, SSS, latitude and longitude, month, monthly cycle term, and temperature-salinity and spatial interaction terms. The regional TA estimation model divides training and validation data according to time, cruise, or spatial partitions; data from the target test period are not included in model fitting. Input the environmental variables output from S3 into the regional TA estimation model to obtain the TA estimate for each field pCO2 observation. Determine the TA estimation error based on time-leaving validation, residual statistics, or prediction intervals. This step outputs the TA estimate, TA estimation error, TA model version, and data isolation identifier.
[0039] S5, calculate the pCO2-derived pH label and propagate input errors. Input the field pCO2, estimated TA, SST, and SSS into the carbonate system solution process, and obtain the pCO2-derived pH label according to a unified pH scale, equilibrium constant, and calculation parameter combination. This step can also simultaneously calculate DIC and... This is used to check the internal consistency of the carbonate system, but the variables and their low-error results are not a substitute for the accuracy of long-term operating remote sensing models. The local sensitivity of pH to pCO2, TA, SST, and SSS is determined separately, and the four types of input errors are propagated as uncertainties for each derived pH label:
[0040] in, , , , The sensitivity of pH to the corresponding input. , , , For uncertainties in on-site pCO2, TA estimation, SST, and SSS, The uncertainty of the i-th pH label.
[0041] Local sensitivity can be determined through finite difference input perturbation or by obtaining the label uncertainty distribution through Monte Carlo perturbation. This step outputs the pCO2-derived pH label, label uncertainty, source of label generation variables, carbonate calculation parameters, and calculation status identifier. The label should be marked as "field-constrained pCO2-derived pH," and should not be marked as direct field pH.
[0042] S6, Determine the label source reliability level. Based on the label source and independence, classify available pH labels into different label source reliability levels and determine the permissible uses for each type of label. Direct field pH labels are used for training or independent validation at the highest label source reliability level; field pCO2-derived pH labels are used for target domain training, adaptation, and time hold-out evaluation, but not for independent field pH validation; weak product labels are used for background constraints, source domain pre-training, or product controls, and are not included in the independent field validation set; soft model labels are used for knowledge distillation, coverage supplementation, or consistency constraints, and do not replace direct observations or derived hard labels. This step outputs the label source, label source reliability level, and permissible uses for each label.
[0043] S7, determine sample weights and sample selection status. Based on the label source reliability level determined in S6, the label uncertainty obtained in S5, and the environment matching distance obtained in S3, determine the training weight of each sample. The sample weights can be determined according to the following formula:
[0044] in, Let be the weight of the i-th sample. For the reliability level coefficient of the label source, To address the uncertainty of derived pH labels, For the tag uncertainty attenuation scale parameter, To match the distance to the environment, Environmental matching distance attenuation scale parameter, The minimum retention weight.
[0045] Samples with high label reliability, low label uncertainty, and close environmental matching distance are given higher weights. Samples below the preset weight threshold or that do not meet the requirements for their evidentiary use are marked as not entering the corresponding training or validation set. This step outputs the sample weights and sample selection status.
[0046] S8, Train the source domain basic model and determine the operating mode. Establish a whitelist of product operating characteristics. This whitelist records the name, data source, temporal resolution, spatial resolution, continuous availability, and missing test handling method for each input variable. Train the source domain basic model using global source domain samples and the filtered and retained whitelisted environmental variables. pCO2, TA, DIC, and [other parameters] are used for label generation or carbonate consistency checks. Variables must be specified, and the system must not enter a long-term operating mode without corresponding continuous input guarantees. The operating mode is determined based on the input variables actually used during the inference phase. Operating modes may include remote sensing mode, pCO2-constrained mode, and carbonate consistency mode. Different operating modes record their input variables, applicable scope, and evaluation indicators. This step outputs the source domain basic model, the predicted values of the source domain basic model, and the operating mode identifier.
[0047] S9 performs target domain adaptation and generates adaptation sources. The weighted target domain labels output from S7 are input into the target domain adaptation process. For the i-th target domain training sample, the residual between its target domain label and the source domain base model prediction value is calculated:
[0048] in, For the target domain residual; The target domain is labeled with pH. This represents the predicted value from the source domain basic model.
[0049] Using product operational characteristics as input and target domain residuals as training targets, a target domain residual correction model is trained using sample weights determined by S7. The predicted values from the source domain basic model are combined with the target domain residual correction values to obtain the predicted sea surface pH value for the target domain.
[0050] in, This represents the predicted sea surface pH value for the target area. This represents the correction amount output by the target domain residual correction model.
[0051] Samples from the target testing period must not be used to fit the source domain base model, regional TA estimation model, or target domain residual correction model. This step generates the adaptation source associated with the predicted value. In the above implementation, the adaptation source is labeled "target domain residual correction". As an alternative implementation, source domain samples, high-level labels of the target domain, and low-weight weak labels can be jointly adapted according to label reliability weights; in this case, the adaptation source is labeled "evidence-weighted joint adaptation".
[0052] S10, synchronously output prediction results and evidence metadata. Following the same sample or grid index, associate the predicted sea surface pH value of the target domain with the label source, label uncertainty, sample weight, adaptation source, and operating mode identifier. The synchronous output may also include the source domain base model prediction value, target domain residual correction, environmental matching distance, model version, input data source, spatial and temporal resolution, and missing data handling identifier. When the operating mode is pCO2-constrained, the data source and applicable scope of pCO2 are output synchronously. For long-term operating grids without corresponding target domain labels, the label source, label uncertainty, and sample weights are recorded as "no corresponding label" or "not applicable" to avoid mistakenly writing model prediction results as direct field observations or pCO2-derived labels.
[0053] In this embodiment, global source domain data, target sea area field pCO2 data, and gridded carbonate background products are used as the basis. The regional TA empirical model is separated into training and testing periods by time, with a TA time-delay test RMSE of approximately 5.08 μmol / kg. The system generates 2290 pCO2-constrained pH tags from field pCO2 data over 34 observation months, and divides them by time into 1822 target domain training samples and 468 target domain test samples.
[0054] When using only long-term available remote sensing and environmental variables, the pHRMSE of the global source domain model directly extrapolated to the aforementioned target test set is 0.03932. After adding target domain constraints for adaptation, the RMSE of the fully covered remote sensing candidate model is 0.02992, a relative reduction of approximately 23.9%. The above evaluation target is the derived pH from in-situ pCO2 constraints and does not belong to independent in-situ pH measurement verification.
[0055] Includes on-site pCO2, TA, DIC or The low-error model is only used for label construction or carbonate consistency description and is not used as the accuracy of long-term running remote sensing models.
[0056] Please see Figure 2 A remote sensing migration reconstruction system for sea surface pH constrained by in-situ pCO2, comprising: The observation and labeling layer is used for quality screening, environmental matching, and carbonate system calculation of multi-source observations, converting in-situ pCO2 observations of the target sea area into derived pH labels with clear sources, calculation conditions, and uncertainties. This layer outputs surface observation records, environmental matching variables, environmental matching distances, TA estimates and their errors, and pCO2-derived pH labels and their uncertainties, providing input for the second layer's label source reliability grading and sample weight determination. The workflow of the observation and labeling layer includes: Surface screening and deduplication: This module is used to generate valid observation samples from raw observation data that meet the requirements for sea surface pH reconstruction. Valid observations refer to observation records that possess valid sampling time, latitude and longitude, sampling depth, and corresponding carbonate parameters, and pass data source quality marking, missing value checks, and physical extent checks. This module filters surface records according to a preset depth threshold; when multiple observation records at different depths exist at the same time and location, the shallowest record is retained or merged according to preset rules; records that do not meet surface conditions or quality control requirements are discarded, and the module outputs surface observation records that have undergone quality screening and deduplication.
[0057] Environmental Variable Matching: This module is used to match the environmental variables required for label calculation and subsequent modeling of screened field pCO2 observations. Based on the time and spatial location of field pCO2 observations, this module matches sea surface temperature (SST), sea surface salinity (SSS), and location-time information within preset time and spatial windows, and records the matching time difference, spatial distance, or comprehensive environmental matching distance. Records exceeding the preset matching range can be discarded or marked as low-match quality samples. The module outputs the matched SST, SSS, location-time information, and environmental matching distance for use in determining sample weights.
[0058] Regional TA Estimation: This module estimates regional TA for pCO2 records lacking contemporaneous in-situ total alkalinity (TA) observations. It builds a TA estimation model using regional carbonate observation data isolated from the target testing period, estimating the TA at the location of in-situ pCO2 observations based on SST, SSS, latitude and longitude, month, and corresponding derived variables. Model training and validation implement data isolation according to time, cruise, or spatial partitions to avoid data from the target testing period participating in model fitting. It outputs the TA estimate and TA estimation error for each record. The TA estimation error is further used as input for pH label uncertainty propagation.
[0059] Tag Calculation and Error Propagation: This module calculates pCO2-derived pH tags based on field pCO2, estimated TA, SST, and SSS, and quantifies the tag uncertainty. It solves for sea surface pH using a unified carbonate system parameter setting, pH scale, and calculation constant combination, while retaining the input data source, calculation parameters, and calculation status identifier. The module determines the sensitivity of pH to pCO2, TA, SST, and SSS, and uses error propagation, input perturbation, or Monte Carlo calculations to synthesize the four types of input errors into a tag-by-tag uncertainty. The module outputs the pCO2-derived pH tag, tag uncertainty, tag generation variable source, and carbonate calculation label.
[0060] The label source reliability grading and weighting layer is used to evaluate the evidence strength, applicability, and training contribution of pH labels from different sources. This layer does not recalculate pH labels; instead, it receives the label source, label uncertainty, and environmental matching distance from the first layer's output to form the label source reliability level, sample purpose, sample weight, and sample selection criteria. The workflow of the label source reliability grading and weighting layer includes: Tag Source Reliability Grading and Usage Determination: This module is used to classify tag source reliability levels based on tag source and independence, and to limit the usage scope of different tag source reliability levels. This module must distinguish at least the following four tag categories, outputting the tag source reliability level, tag source, and permitted usage for each tag, providing a tag source reliability level coefficient for the sample weighting module. The four tag categories are: Direct field pH tags are used for training or independent validation at the highest tag source reliability level. Field pCO2-derived pH labels are used for target domain training, adaptation, and time-leave evaluation, but are not presented as independent field pH validation. Weak labels for products are used for background constraints, source domain pre-training, or trend comparison, but are not used as independent field validation labels. Model soft labels are used for knowledge distillation, coverage supplementation, or consistency constraints, but do not replace hard labels and independent verification.
[0061] Sample weight determination: This module determines the training contribution of samples based on the strength of label evidence and sample reliability. It comprehensively considers the reliability level of the label source, label uncertainty, environment matching distance, and the degree of observational support in the target domain to determine sample weights. Samples with higher label source reliability, lower label uncertainty, and closer environment matching distance receive higher weights; samples with weak product labels, soft model labels, high uncertainty labels, or poor matching quality receive lower weights. For samples below a preset weight threshold or that do not meet the usage conditions, this module can generate an exclusion flag to prevent them from entering the corresponding training or validation set. This module outputs sample weights and sample selection flags, primarily used to constrain target domain adaptation.
[0062] The transfer reconstruction and output layer is used to build a basic model using global source domain data, and to perform regional adaptation by using target domain labels that have undergone source reliability grading and sample weight constraints. The final output is a sea surface pH prediction result with source and operational status identifiers. This layer has two types of input paths: a whitelist of global source domain samples and operational characteristics for training the source domain basic model; and the target domain labels, sample weights, and selection identifiers output by the second layer for target domain adaptation. The workflow of the transfer reconstruction and output layer includes: Source Domain Pre-training and Runtime Isolation: This module trains the source domain base model using global source domain samples and ensures that the model input matches the actual product's operating conditions. It establishes a whitelist of product operating features, determining available features based on the continuous availability of different input variables during the inference phase. Variables such as pCO2 and TA used for label generation must not enter operating modes without corresponding continuous input guarantees. Includes pCO2, TA, DIC, or... The model is used only as a carbonate constraint or consistency model and is saved and evaluated separately from the long-term operating remote sensing model. The output includes the source domain basic model, source domain predicted values and operating mode identifier.
[0063] Target Domain Adaptation: This module is used to correct the systematic bias of the source domain base model in the target area using weighted labels of the target sea area. In the residual correction implementation, this module calculates the residual between the target domain label and the predicted value of the source domain base model. It then establishes a residual correction model using target domain training samples constrained by sample weights. Finally, it combines the source domain predicted value with the target domain correction to obtain the predicted sea surface pH value for the target area. The output includes the predicted sea surface pH value, the target domain correction, and the adaptation source. The adaptation source indicates that the predicted value was obtained using a residual correction method, an evidence-weighted joint adaptation method, or other defined adaptation methods.
[0064] Synchronous Output: This module is used to associate, save, and output prediction results with corresponding evidence, adaptation information, and operational status. Based on the same sample or grid index, it associates and outputs sea surface pH predictions, label sources, label uncertainties, sample weights, adaptation sources, and operational mode identifiers. It can further output model version, environmental matching distance, target domain correction amount, input data source, and missing data handling identifiers.
[0065] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for remotely sensing and reconstructing sea surface pH under field pCO2 constraints.
[0066] A computer program product includes a computer program that, when executed by a processor, implements the above-described method for remotely sensing and reconstructing sea surface pH under field pCO2 constraints.
[0067] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented, in whole or in part, as a computer program product, the computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line, or wireless (e.g., infrared, wireless, microwave, etc.) means). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0068] The above description is only a preferred embodiment of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for remote sensing migration reconstruction of sea surface pH under in-situ pCO2 constraint, characterized in that, include: Step 1: Obtain on-site pCO2 observation records for the target sea area, match sea surface temperature (SST) and sea surface salinity (SSS) based on the spatiotemporal information of each observation record, and record the environmental matching distance. Step 2: Train the regional total alkalinity (TA) estimation model using regional carbonate observation data isolated from the target testing period, estimate TA using field pCO2 observation records, and obtain the TA estimation error, which serves as the TA input uncertainty in S3. ; Step 3: Input the field pCO2, estimated TA, SST and SSS into the carbonate system to solve for the field pCO2 derived pH label, and determine the sensitivity of pH to field pCO2, estimated TA, SST and SSS. Propagate the TA estimation error to the label uncertainty of the pCO2 derived pH label. Step 4: Classify the available pH labels according to their source reliability level and label them as pH label sources; and determine the sample purpose and sample weight based on the label source reliability level, label uncertainty and environmental matching distance. Step 5: Establish a whitelist of product operation characteristics. Based on the continuous availability of each input variable in different operation modes, exclude the pCO2 and TA variables used for tag generation from operation modes without corresponding continuous input guarantees. The source domain basic model is trained using global source domain samples and a whitelist of filtered and retained environment variables. The operating mode is determined based on the input variables actually used during the inference phase, and the corresponding operating mode identifier is output. Step 6: Based on the target domain label constrained by sample weights, calculate the residual between the target domain label and the predicted value of the source domain basic model, train the target domain residual correction model, combine the predicted value of the source domain basic model with the correction amount output by the target domain residual correction model to obtain the predicted value of sea surface pH of the target domain, and generate an adaptation source associated with the predicted value of sea surface pH of the target domain. The adaptation source is used to characterize that the predicted value of sea surface pH of the target domain is obtained by the target domain residual correction method. Step 7: Associate the label uncertainty, pH label source, sample weight, adaptation source, operating mode identifier, and target domain sea surface pH prediction value with the same sample or grid index and output them.
2. The method for remote sensing migration reconstruction of sea surface pH under pCO2 constraint according to claim 1, characterized in that, S1 includes surface screening and deduplication at the same location: records with a depth greater than a preset threshold are removed, and when multiple records exist at the same time and location, the shallowest layer record is retained. The preset threshold is 5m.
3. The method for remote sensing migration reconstruction of sea surface pH under pCO2 constraint according to claim 1, characterized in that, The inputs to the TA estimation model include: year center value, month sine and cosine, latitude and longitude, SST, SSS, SST squared, SSS squared, SST-SSS interaction term, and latitude-longitude and salinity interaction term; the TA estimation model is trained and validated by time, flight or spatial partition, and the data of the target test period are not involved in the model fitting.
4. The method for remote sensing migration reconstruction of sea surface pH under pCO2 constraint according to claim 1, characterized in that, The formula for calculating the label uncertainty is: in, The sensitivity of pH to pCO2, The sensitivity of pH to TA, The sensitivity of pH to SST, The sensitivity of pH to SSS. Due to the uncertainty of pCO2 at the site, Uncertainty estimated for on-site TA Due to the uncertainty of on-site SST, Due to the uncertainty of SSS on site, This is to address the uncertainty of derived pH labeling.
5. The method for remote sensing migration reconstruction of sea surface pH under pCO2 constraint according to claim 1, characterized in that, The formula for calculating the sample weight is: in, For sample weights, To retain the lowest weight, For the reliability level coefficient of the label source, To address the uncertainty of derived pH labels, For the tag uncertainty attenuation scale parameter, To match the distance to the environment, To match the distance attenuation scale parameters to the environment.
6. The method for remote sensing migration reconstruction of sea surface pH under pCO2 constraint according to claim 1, characterized in that, The usage constraints of the reliability level of the label source are as follows: direct field pH is used for the highest level training or validation, field pCO2 derived pH is used for target domain training and time hold-out evaluation, weak product labels are only used for background constraints or pre-training, and soft model labels are only used for distillation or consistency constraints.
7. The method for remote sensing migration reconstruction of sea surface pH under in-situ pCO2 constraint according to claim 1, characterized in that, The whitelist of product operation characteristics includes at least one or more of the following variables: latitude and longitude, month, SST, SSS, chlorophyll a, wind field components, sea surface height, mixing layer depth and their derived variables; the operation mode includes at least the operation remote sensing mode, pCO2 constraint mode and carbonate consistency mode, and each mode records the available inputs, applicable scope and evaluation indicators respectively.
8. The method for remote sensing migration reconstruction of sea surface pH under in-situ pCO2 constraint according to claim 1, characterized in that, The fields output by S7 include one or more of the following: model version, target domain correction, environment matching distance, input data source, spatial and temporal resolution, and missing data handling identifier.
9. A remote sensing migration reconstruction system for sea surface pH constrained by in-situ pCO2, characterized in that, The system for implementing the sea surface pH remote sensing migration reconstruction method with in-situ pCO2 constraint as described in any one of claims 1-8 comprises: The observation and labeling layer is used for quality screening, environmental matching, and carbonate system calculation of multi-source observations, and converts the on-site pCO2 observations of the target sea area into pCO2-derived pH labels and their uncertainties. The label source reliability classification and weighting layer is used to receive the pCO2-derived pH label and its uncertainty and environmental matching distance, classify the label source reliability and determine the sample weight; The migration reconstruction and output layer is used to build a source domain basic model using global source domain data, and to complete the regional adaptation of the source domain basic model by target domain labels that have undergone label source reliability classification and sample weight constraints, and finally outputs the sea surface pH prediction results.
10. A computer-readable storage medium storing a computer program, characterized in that, When executed by the processor, the program implements the remote sensing migration reconstruction method for sea surface pH under field pCO2 constraints as described in any one of claims 1-8.