Method and system for ocean acidification remote sensing reconstruction mechanism attribution and confidence ranking

CN122654873APending Publication Date: 2026-08-28STATE OCEAN TECH CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611114462.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-27
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0008]本发明提供一种海洋酸化遥感重建机制归因与置信度分级方法及系统,以解决现有技术中置信度来源不清、机制解释混淆、发布规则粗放及观测规划依据不足的问题

Benefits of technology

[0026]本申请具有的优点和积极效果是:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122654873A_ABST
    Figure CN122654873A_ABST
Patent Text Reader

Abstract

The application discloses a marine acidification remote sensing reconstruction mechanism attribution and confidence grading method and system, and belongs to the technical field of marine environment monitoring. The method comprises the following steps: obtaining marine acidification variable reconstruction prediction, target domain label, label source and label error, product operation characteristics, and space-time information; calculating divergence risk, feature distance risk, space-time observation deficiency risk and label risk; normalizing the divergence risk, feature distance risk, space-time observation deficiency risk and label risk; establishing a carbonate system checking channel and a remote sensing environment attribution channel; mapping the target domain residual error, traceable risk vector, sample density and mechanism group contribution to the same space grid or process partition; determining the product release level; and outputting the product release level, risk component, mechanism contribution and observation priority area. The application can solve the problems of unclear confidence source, confused mechanism explanation, rough release rules and insufficient observation planning basis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of marine environmental monitoring technology, and in particular relates to a method and system for attribution and confidence grading of marine acidification remote sensing reconstruction mechanism. Background Technology

[0002] Ocean acidification is a key indicator of global climate change, and reconstructing sea surface pH, pCO2, and other acidification variables using remote sensing data has become a primary means of obtaining large-scale, long-term time series data. However, the reliability assessment and mechanistic interpretation of the reconstruction results are crucial for both scientific research and operational applications.

[0003] In existing technologies, such as those described by Chau et al. (2022), sea surface pCO2 is reconstructed using ensemble neural networks, and the ensemble standard deviation is used as the uncertainty to identify priority observation areas. However, existing technologies have the following significant drawbacks: The confidence assessment dimension is too singular: Existing technologies mostly use a single comprehensive score or set uncertainty to characterize product accuracy, making it difficult to trace the specific source of low confidence results (such as model divergence, feature extrapolation, observation sparsity or label error itself), which makes it impossible for users to judge the cause of risk.

[0004] Confusion exists in mechanism explanation: When performing mechanism attribution, existing methods often mix "carbonate variables involved in pH tag calculation" (such as pCO2, TA) with "remote sensing variables available during product operation" (such as SST, Chl-a), which can easily lead to misjudging the consistency of tag calculation as a mechanism contribution from remote sensing observation data, resulting in distorted mechanism explanation.

[0005] The product release rules are too rudimentary: existing geobiochemical reanalysis model products lack refined release level rules and have not established a triple constraint mechanism of "overall confidence level, single risk upper limit, and minimum spatiotemporal observation support", which may lead to high-risk areas of local errors being incorrectly classified into high-confidence products.

[0006] The observation planning lacks a basis: existing methods fail to effectively combine error hotspots, observation gaps and mechanism hypotheses, making it difficult to generate observation priority areas with clear scientific objectives and failing to effectively guide the subsequent deployment of ocean acidification monitoring.

[0007] Therefore, there is an urgent need for a method and system for attributing and classifying the confidence level of ocean acidification remote sensing reconstruction mechanisms, which can assess confidence level in multiple dimensions, separate mechanism channels, release information in a refined manner and guide observation planning. Summary of the Invention

[0008] This invention provides a method and system for attributing and classifying the mechanism of ocean acidification remote sensing reconstruction, in order to solve the problems of unclear confidence sources, confusing mechanism interpretations, coarse release rules, and insufficient basis for observation planning in the prior art.

[0009] To achieve the above-mentioned objectives, the first objective of this invention is to provide a method for attribution and confidence grading of ocean acidification remote sensing reconstruction mechanisms, comprising: S1. Obtain the ocean acidification variable reconstruction prediction of the target sea area, the target domain label and its label source and label error, product operation characteristics, spatiotemporal information of training samples and spatiotemporal information of the grid to be evaluated; S2. For each sample or grid to be evaluated, calculate the divergence risk of prediction results from multiple teacher models, the feature distance risk of the sample to be evaluated relative to the training sample in the standardized product operation feature space, the risk of insufficient spatiotemporal observation based on neighboring observation distance and local sample density, and the label risk based on label source and label error. S3. Normalize and retain the divergence risk, feature distance risk, spatiotemporal observation support risk and label risk as traceable risk vectors, and form a comprehensive confidence level based on the risk vectors; S4. Establish separate verification channels for the carbonate system and remote sensing environmental attribution channels, and save the evaluation results of the two channels separately. S5. Map the target domain residual, the traceable risk vector, the sample density, and the mechanism group contribution of the remote sensing environment attribution channel to the same spatial grid or process partition to form a spatial joint evaluation result. S6. The product release level is determined by three conditions: overall confidence level, individual risks, and minimum spatiotemporal observation support. When any individual risk exceeds the corresponding upper limit or the spatiotemporal observation support is lower than the minimum requirement, the highest release level is prohibited. The reconstruction error and coverage of the high-confidence retained subset are output synchronously for the results screened by confidence level. S7. Based on low confidence, existing sample reconstruction error and spatial observation gaps, determine the priority of supplementary observations, and output the product release level, risk component, mechanism contribution and observation priority area.

[0010] The low confidence level refers to the overall confidence level being lower than a preset low confidence threshold.

[0011]

[0012] in, To assess the overall confidence level, This is the minimum preset low confidence threshold. The target domain validation set is used to determine the threshold by comparing the error and coverage of the retained subset under different candidate thresholds, and selecting the threshold that simultaneously satisfies the condition that the error is not higher than a preset upper limit and the coverage is not lower than a preset lower limit. .

[0013] Preferably, the teacher model includes at least two regression models with different structures, and the risk of divergence is determined by the standard deviation, range, or quantile interval of each teacher's predicted value.

[0014] Preferably, the running feature distance risk is determined by multiple nearest neighbor distances from the sample to be evaluated to the training sample, and the pCO2, TA, DIC, and [other parameters] used for label generation are [also considered]. The variable is not included in the calculation of the running feature distance.

[0015] Preferably, the risk of insufficient spatiotemporal observation is determined by at least two of the following: spatial distance to the target domain, temporal distance, local sample density, and regional coverage.

[0016] Preferably, the label risk is determined based on the evidence level and corresponding uncertainty of direct field pH, field pCO2-derived pH, weak product labeling, and model soft labeling.

[0017] Preferably, the carbonate system verification channel uses the variables involved in solving the carbonate system to evaluate the consistency of the calculation, while the remote sensing environment attribution channel only uses the spatiotemporal, temperature-salinity, dynamic, and bio-optical variables available during the product operation phase to evaluate the contribution of the mechanism group.

[0018] Preferably, the overall confidence level is obtained by a weighted sum of four types of non-negative risks through a monotonically decreasing transformation, while each risk component is retained in the output record instead of being replaced by the overall confidence level.

[0019] Preferably, the mechanism group of the remote sensing environment attribution channel includes at least a spatial-temporal group, a thermo-salinity group, a dynamic mixing group, and a bio-optical group, and the carbonate system verification channel separately includes pCO2, TA, DIC, and Group.

[0020] Preferably, the product release levels include at least high confidence, conditional confidence, low confidence and missing test levels, and the highest release level simultaneously meets the conditions of comprehensive confidence, individual risk and spatiotemporal observation support.

[0021] Preferably, the coverage rate is the ratio of the number of samples or grids that enter the specified release level or high confidence subset to the total number of evaluable samples or grids, and is output simultaneously with the corresponding root mean square error, mean absolute error, or bias.

[0022] Preferably, the priority of supplementary observations increases with increasing low confidence, normalized existing sample reconstruction error, and observation gap, and is associated with outputting suggested supplementary carbonate, nutrient, turbidity, water color, or dynamic observation variables.

[0023] A second objective of this invention is to provide an attribution and confidence grading system for remote sensing reconstruction mechanisms of ocean acidification, comprising: The multi-source data access module is used to acquire reconstruction prediction, target domain labels and label errors, product operation characteristics, training samples and the spatial and temporal locations of the grid to be evaluated; The multi-source risk calculation module is used to calculate the teacher model divergence risk, running feature distance risk, insufficient spatiotemporal observation risk, and label risk respectively, and retains the four risk components. The overall confidence module is used to generate an overall confidence score based on the four risk components. The dual-channel mechanism attribution module is used to separate the carbonate system verification channel and the remote sensing environment attribution channel and output the mechanism group contributions separately. The spatial joint evaluation module is used to jointly preserve the mechanism group contributions of residuals, risk components, sample density, and remote sensing environmental attribution channels within the same spatial grid or process partition. The product release module is used to generate release levels based on comprehensive confidence, single risk upper limit, and minimum spatiotemporal observation support, and simultaneously outputs coverage when outputting high-confidence subsets; The observation priority module is used to generate observation priority areas based on low confidence, existing sample reconstruction errors, and observation gaps.

[0024] A third objective of this invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method for attribution and confidence grading of ocean acidification remote sensing reconstruction mechanisms.

[0025] A fourth objective of this invention is to provide a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method for attribution and confidence grading of ocean acidification remote sensing reconstruction mechanisms.

[0026] The advantages and positive effects of this application are: This invention simultaneously records teacher model divergence, extrapolation of operational features, spatiotemporal observation support, and label error, enabling the tracing of the specific source of low-confidence results.

[0027] This invention separates the verification interpretation of the carbonate system from the attribution interpretation of the remote sensing environment, thus avoiding the misinterpretation of the high contribution of the label-generating variable as mechanistic evidence that can be obtained through long-term remote sensing.

[0028] This invention combines confidence screening with spatial error attribution, and can output a high-confidence reconstruction subset, a global product release level, and a low-confidence risk area.

[0029] This invention transforms error hotspots, observation gaps, and mechanism hypotheses into supplementary observation priorities, providing a basis for subsequent deployment of ocean acidification monitoring.

[0030] This invention requires that the screening results be reported synchronously to ensure coverage, thus avoiding the substitution of a small number of highly reliable subsets for overall performance.

[0031] This invention is applicable to sea surface pH, pCO2, Reconstruction quality control of various ocean acidification variables. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 A flowchart of the first embodiment of the present invention is shown; Figure 2 A system block diagram of the present invention is shown; Figure 3 A flowchart of the second embodiment of the present invention is shown. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] Explanation of the name: SST: Sea Surface temperature; SSS: Sea Surface Salinity; DIC: Dissolved Inorganic Carbon; Aragonite Saturation: The saturation level of aragonite.

[0036] TA: Total Alkalinity; MLD: Mixed Layer Depth; SSH: Sea Surface Height; Chl-a: Chlorophyll-a (chlorophyll a); Kd490: Diffuse attenuation coefficient at a wavelength of 490 nm.

[0037] The term "teacher model" as used in this specification refers to one or more trained reconstruction models used to generate prediction results, model soft labels, or model divergence information; "student model" refers to a model trained using reference labels and the output of the teacher model to generate continuous reconstruction results. The training method of passing prediction information from the teacher model to the student model is an implementation method of knowledge distillation.

[0038] This invention is applicable to sea surface pH, pCO2, Quality control of grid reconstruction results for other ocean acidification variables. The core idea is to evaluate the model not solely using a single accuracy index, but by separately calculating model divergence, feature extrapolation, spatiotemporal observation support, and label error, forming a traceable confidence vector. Simultaneously, the mechanism analysis is divided into a carbonate system verification channel and a remote sensing environment attribution channel to avoid interpretative confusion caused by label-generated variables. Furthermore, confidence, residuals, mechanism evidence, and sample density are converted into product release levels, low-confidence-risk zones, and supplementary observation priority zones.

[0039] Please see Figure 1 The first embodiment, a method for attribution and confidence grading of ocean acidification remote sensing reconstruction mechanisms, mainly includes: S1. Obtain the ocean acidification variable reconstruction prediction of the target sea area, the target domain label and its label source and label error, product operation characteristics, spatiotemporal information of training samples and spatiotemporal information of the grid to be evaluated; S2. For each sample or grid to be evaluated, calculate the divergence risk of prediction results from multiple teacher models, the feature distance risk of the sample to be evaluated relative to the training sample in the standardized product operation feature space, the risk of insufficient spatiotemporal observation based on neighboring observation distance and local sample density, and the label risk based on label source and label error. S3. Normalize and retain the divergence risk, feature distance risk, insufficient spatiotemporal observation risk and label risk as traceable risk vectors, and form a comprehensive confidence level based on the risk vectors; S4. Establish separate verification channels for the carbonate system and remote sensing environmental attribution channels, and save the evaluation results of the two channels separately. S5. Map the target domain residual, the traceable risk vector, the sample density, and the mechanism group contribution of the remote sensing environment attribution channel to the same spatial grid or process partition to form a spatial joint evaluation result. S6. The product release level is determined by three conditions: overall confidence level, individual risks, and minimum spatiotemporal observation support. When any individual risk exceeds the corresponding upper limit or the spatiotemporal observation support is lower than the minimum requirement, the highest release level is prohibited. The reconstruction error and coverage of the high-confidence retained subset are output synchronously for the results screened by confidence level. S7. Based on low confidence, existing sample reconstruction error and spatial observation gaps, determine the priority of supplementary observations, and output the product release level, risk component, mechanism contribution and observation priority area.

[0040] The low confidence level refers to the overall confidence level being lower than a preset low confidence threshold.

[0041]

[0042] in, To assess the overall confidence level, This is the minimum preset low confidence threshold. The target domain validation set is used to determine the threshold by comparing the error and coverage of the retained subset under different candidate thresholds, and selecting the threshold that simultaneously satisfies the condition that the error is not higher than a preset upper limit and the coverage is not lower than a preset lower limit. The low confidence level refers to the overall confidence level. Below the preset low confidence threshold .

[0043] To better understand the method of the present invention, the following non-limiting explanation is provided: The teacher model comprises at least two regression models with different structures. In one implementation, the teacher model ensemble includes a histogram gradient boosting regression model, a random forest regression model, and an extreme random tree regression model. Each teacher model is trained using the same product performance characteristics and the same data partitioning, and outputs predicted values ​​for the same sample to be evaluated.

[0044] The risk of discrepancy is determined by the standard deviation, range, or quantile interval of the predicted values ​​from each teacher model. Specifically, for example, when n teacher models output predicted values ​​for the same sample... , … At that time, the average predicted value of the teacher model set for the sample to be evaluated for:

[0045] This is the teacher model's ID, with a value of [value]. ; Indicates the first The predicted values ​​of the teacher models for the same sample to be evaluated; This represents the total number of teacher models.

[0046] when At that time, the predicted standard deviation of the teacher model set for the samples to be evaluated was used. This indicates discrepancies in predictions among different teacher models:

[0047] Establish a cumulative distribution function based on the empirical distribution of teacher standard deviations in training data or independent validation data. Risk of divergence in normalized teacher models for:

[0048] in, To prefit and fix the empirical cumulative distribution function on the training data or independent validation data, The value range of is [0,1]. The more inconsistent the teachers' predictions, the... The larger.

[0049] The operational feature distance risk is determined by multiple nearest neighbor distances from the sample to be evaluated to the training samples, and is used for label generation by pCO2, TA, DIC, and Variables are not included in the calculation of the distance to the operational features. For example: first, the product operational features are standardized using the mean and standard deviation fitted from the training samples, and then the distance to the samples to be evaluated is calculated. With training samples The k nearest training samples are selected based on the Euclidean distance between them. , … The average distance is taken as the feature distance:

[0050] in, The distance to the running features of the sample or grid to be evaluated; This is the ID of the nearest neighbor training sample, with a value of [value]. ; Represents the distance to the sample to be evaluated in the running feature space. No. The feature vector of the nearest training sample; This is the preset number of nearest neighbor samples.

[0051] Establish a cumulative distribution function based on the empirical distribution of feature distances in the training data or independent validation data. Normalized operating characteristics distance risk for:

[0052] in, To prefit and fix the empirical cumulative distribution function on the training data or independent validation data, The value range of is [0,1]; the larger the feature distance, the better. The larger the value, the better. The reference distribution and its version are saved with the model and are not redefined using test data or the batch to be evaluated.

[0053] The risk of insufficient spatiotemporal observations is determined by at least two of the following: spatial distance to the target domain observation, temporal distance, local sample density, and regional coverage. For example, the nearest spatial distance, nearest temporal distance, local sample density, and regional coverage from the grid to be evaluated to the target domain observation are calculated separately, and each component is normalized to [0,1]. Since larger spatial and temporal distances indicate lower spatiotemporal observation support, while larger sample density and regional coverage indicate higher spatiotemporal observation support, the direction of each component should be unified first, and then the spatiotemporal observation support should be formed according to the following formula. :

[0054] in, To normalize the nearest spatial distance, To normalize the nearest time distance, To normalize the local sample density, To normalize the regional coverage; Spatial distance weights As a time distance weight, For local sample density weights, This represents the weighting for regional coverage. Unused indicators have a weight of 0, and at least two indicators must be used, with the sum of their weights greater than 0. This addresses the risk of insufficient normalized spatiotemporal observations. for:

[0055] When the grid to be evaluated is far from the target domain observation, the time interval is long, or the local sample density is low, the spatiotemporal observation support decreases. Increase. The above relationship is based on the premise that all components are normalized to [0,1] and have the same direction; the normalization parameters and weights should be calibrated by training data or validation data, or determined according to preset rules, and saved with the model version, without using test data to determine them.

[0056] The label risk is determined based on the evidence levels and corresponding uncertainties of direct field pH, field pCO2-derived pH, weak product labeling, and model soft labeling. For example, evidence level risk coefficients are set for direct field pH, field pCO2-derived pH, weak product labeling, and model soft labeling, and normalized label risk is formed by combining the corresponding label uncertainties. :

[0057] in, The level of evidence risk corresponding to the source of the label. To normalize label uncertainty, Weights for label uncertainty, used to adjust Contribution to label risk. For the boundary truncation operator, Limited to [0,1]. Direct on-site pH has the lowest risk of evidence level, while model-based soft labels have a higher risk of evidence level.

[0058] The carbonate system verification channel uses the variables involved in solving the carbonate system to evaluate the consistency of calculations, while the remote sensing environment attribution channel only uses the spatiotemporal, temperature-salinity, dynamic, and bio-optical variables available during the product operation phase to evaluate the contributions of the mechanism group.

[0059] The overall confidence level is obtained by weighting four types of non-negative risks and performing a monotonically decreasing transformation. At the same time, each risk component is retained in the output record instead of being replaced by the overall confidence level.

[0060]

[0061] in, The overall confidence level ranges from [0,1], with a higher value indicating a lower overall risk. For the teacher model disagreement risk weights, To run the feature distance risk weight, Risk weighting for insufficient spatiotemporal observation. The risk weights for the labels are a, b, c, and d, all of which are non-negative and control the degree of influence of the four types of risks on the overall confidence level, respectively. This represents the lower confidence level. All four risk categories must be saved simultaneously during output; not just one. .

[0062] The mechanism group of the remote sensing environment attribution channel includes at least the space-time group, the thermo-salinity group, the dynamic mixing group, and the bio-optical group. The carbonate system verification channel separately includes pCO2, TA, DIC, and Group.

[0063] The product release levels include at least high confidence, conditional confidence, low confidence, and missing test levels, with the highest release level simultaneously satisfying the conditions of comprehensive confidence, individual risk, and spatiotemporal observation support. For example: High trust level: Each individual risk does not exceed its corresponding upper limit, and the spatiotemporal observation support meets the minimum requirements.

[0064] Conditional confidence level: The prediction result and necessary inputs are complete, but a single condition does not meet the high confidence requirement, and .

[0065] Low confidence level: Prediction results can still be generated, but It also outputs the main sources of risk in a timely manner.

[0066] Missing test level: Required input is missing, the model cannot form an effective prediction, or the risk component cannot be calculated.

[0067] in To set the maximum preset low confidence threshold, To minimize the preset low confidence threshold, satisfy The individual risks mentioned include , , and The corresponding upper limits are determined based on training data, independent validation data, or preset business rules, respectively; the minimum requirement for spatiotemporal observation support is... ,in This is the preset minimum spatiotemporal observation support threshold.

[0068] For any selected evaluation results, the coverage rate should be output simultaneously:

[0069] in, The screening coverage rate is within the evaluable sample or spatial grid range, with a value range of [0,1], and is multiplied by 100% when reported as a percentage; The number of evaluable samples or grids to retain after applying specified confidence thresholds, single-item risk caps, minimum spatiotemporal observation support, or product release level rules; The total number of evaluable samples or grids before screening is defined as follows: evaluable means having complete necessary inputs, capable of forming effective predictions, and having calculable required risk components. The numerator and denominator use the same time range, spatial range, and counting unit. .

[0070] The priority of supplementary observations increases with low confidence, normalized existing sample reconstruction error, and the degree of observation gaps, and is associated with the output of suggested supplementary observation variables such as carbonate, nutrients, turbidity, water color, or dynamics.

[0071] Please see Figure 3 The second embodiment, a method for attribution and confidence grading of ocean acidification remote sensing reconstruction mechanisms, mainly includes: Step 1, Read Data: Read the reconstruction prediction, target domain label, label source, label error, product operation characteristics, training sample location, and grid location to be evaluated, and check the consistency of fields, units, and indexes.

[0072] Step 2, Calculate teacher dissent risk: Generate predictions from multiple real teacher models, and calculate teacher mean, standard deviation, and normalized dissent risk.

[0073] Step 3, calculate the running feature distance risk: standardize using only the features available during the running phase, calculate and normalize the nearest neighbor distance from the sample to be evaluated to the training domain.

[0074] Step 4, calculate the risk of insufficient spatiotemporal observation and the risk of labeling: form two risk components based on spatial and temporal distance, local density, regional coverage, label source and labeling error respectively.

[0075] Step 5, retain the four types of risk vectors and form a comprehensive confidence score: retain the four types of risks, use a monotonically decreasing transformation to form C, and record the weights and normalization parameters.

[0076] Step 6: Separate the carbonate system verification channel and the remote sensing environment attribution channel: The carbonate system verification channel uses labels to generate relevant variables; the remote sensing environment attribution channel only uses product runtime variables, and the results of the two channels are saved separately.

[0077] Step 7, perform spatial joint evaluation: map the mechanism group contributions of residuals, risk components, sample density, and remote sensing environmental attribution channels to the same grid or process partition.

[0078] Step 8: Apply three types of release conditions: simultaneously check the overall confidence level, the upper limit of individual risk, and the minimum spatiotemporal observation support. If any one of them is not met, the highest level cannot be entered.

[0079] Step 9, Generate Release Levels and Low Trust Regions: Generate high trust, conditional trust, low trust, and missing test levels according to the release rules, and merge adjacent low trust grids.

[0080] Step 10, synchronously report the reconstruction error and coverage of the high-confidence retained subset: calculate RMSE, MAE or bias for the high-confidence retained subset with target labels, and synchronously report the coverage; no reconstruction error is reported for unlabeled product grids.

[0081] Step 11, Generate observation priority areas: According to The algorithm normalizes existing sample reconstruction errors and the degree of observation gaps to calculate the priority of supplementary observations, and outputs the priority region and suggested supplementary variables.

[0082] Calculation of confidence level for multi-source evidence: For each sample, the system obtains multiple teacher model predictions. The average predicted value of the teacher model set for the samples to be evaluated. for:

[0083] The predictive standard deviation of the teacher model ensemble for the samples to be evaluated for:

[0084] Based on the empirical distribution of teachers' predicted standard deviations in training data or independent validation data, an empirical cumulative distribution function is pre-established and fixed. The normalized teacher model disagreement risk was calculated. Simultaneously, the average distance from a sample to its nearest neighbor in the training set is calculated in the standardized running feature space and normalized to... ; Calculate based on the observation distance and density of the target domain ;Calculated based on label source and error The final confidence level is obtained from the four risk components:

[0085] When a certain risk component exceeds the upper limit of a single item, even if the overall confidence level is high, it shall not be allowed to enter the highest release level.

[0086] Progressive ablation and mechanism contribution Construct two sets of ablation levels that cannot be mixed and interpreted. The carbonate system verification pathway may include: A0: Spatial and temporal characteristics; A1: Add SST and SSS; A2: Increase pCO2 and TA; A3: Increase DIC and ; A4: Increase the carbonate ratio and sensitivity derived variables.

[0087] Remote sensing environmental attribution pathways may include: R0: Spatial-temporal characteristics; R1: Add SST and SSS; R2: Adds MLD, SSH, and wind farm; R3: Increase Chl-a, Kd490 and remote sensing reflectance; R4: Added temperature-salt interaction, upstream proxy, and regional partitioning; R5: Added target domain adaptation; R6: Filter high-confidence subsets by confidence level.

[0088] The error variation between adjacent levels can be expressed as:

[0089] in, This represents the root mean square error of the reconstructed model on the evaluation sample set after adding the k-th level mechanism variable; This represents the root mean square error of the reconstructed model on the evaluation sample set after adding the (k-1)th level mechanism variable; This represents the error variation between adjacent levels. A value greater than 0 indicates that the newly added mechanism group reduces the error; a value less than or equal to 0 indicates that the mechanism group does not provide effective gain under the current label and sample conditions. The low error of the carbonate system verification channel cannot replace the product accuracy of the remote sensing environment attribution channel.

[0090] Spatial error attribution and observation recommendations The system statistically analyzes residuals, RMSE, bias, four risk components, sample density, and mechanism group contributions according to preset sea areas or spatial grids. If a region simultaneously exhibits high residuals, high model divergence, high feature extrapolation risk, or low spatiotemporal observational support, it is marked as a low-confidence region. Supplementary observation priority is also implemented. It can be represented as:

[0091] in, To normalize the reconstruction error of existing samples, To normalize the degree of spatial observation gaps, The existing sample reconstruction error weights are used to control the normalization of the existing sample reconstruction error. Impact on supplementary observation priority The observation gap weight is used to control the normalized observation gap degree. Impact on the priority of supplementary observations. If the region exhibits low salinity, high Chl-a, anomalous MLD / SSH, or other process characteristics, the system provides the mechanism hypotheses to be verified and the corresponding observational variables. For example, in nearshore estuaries, supplementary observations of pH, TA, DIC, nutrients, and turbidity may be suggested; in upwelling regions, supplementary observations of profile DIC, wind field, and MLD may be suggested.

[0092] Key technical parameters 1. Number of teacher models: no less than 2, preferably 3 or more.

[0093] 2. Confidence level range: can be limited to 0.05-1.00, and simultaneously save... , , and .

[0094] 3. Feature distance nearest neighbor number: preferably 3-10 nearest neighbors.

[0095] 4. Confidence weight , , , All values ​​are non-negative and can be calibrated by the validation set or preset according to the level of evidence.

[0096] 5. High-confidence screening method: Fixed threshold, quantile or multi-condition release rules can be used; any screening result must be reported with coverage.

[0097] 6. Mechanism Group: The remote sensing environmental attribution pathway should include at least the spatial-temporal, thermohaline, dynamic mixing, and bio-optical groups; the carbonate system verification pathway should separately include pCO2, TA, DIC, and... Group.

[0098] 7. Release Level: Can be set to four levels: High Trust, Conditional Trust, Low Trust, and Missing Tests, or can be adjusted according to business needs.

[0099] 8. Spatial attribution indices: RMSE, MAE, Bias, residual standard deviation, sample density, model divergence, feature distance, label error, spatiotemporal observation support, and mechanism group contribution.

[0100] In one embodiment, the system trains a teacher ensemble consisting of different structural regression models, using derived pH labels constrained by pCO2 in the South China Sea and multi-source remote sensing environmental variables as the objects. Both the teacher and students use models without pCO2, TA, DIC, and... The operational remote sensing characteristics. The pH RMSE of the fully covered hard-labeled student model on 468 time-leaved target test samples was 0.03113.

[0101] The system retains 234 test samples based on the median teacher confidence level, representing a 50% coverage rate. The RMSE of the high-confidence subset is 0.02027. This result must be expressed as "subset precision at 50% coverage" and cannot replace full-coverage precision.

[0102] The combined diagnostic risk index showed only a weak positive correlation with the absolute error, with a correlation coefficient of approximately 0.184. This indicates that a single comprehensive score is insufficient to explain all the errors, and components such as model divergence, feature distance, spatiotemporal observation support, and label error need to be retained. In the remote sensing environmental attribution channel, the RMSE increments of the grouping permutations for the spatial-temporal, thermohaline, bio-optical, and dynamic groups were approximately 0.00655, 0.00247, 0.00090, and 0.00072, respectively. The target labels mentioned above are all derived pH values ​​constrained by in-situ pCO2 and do not belong to independent in-situ pH validation.

[0103] A third embodiment provides an attribution and confidence grading system for remote sensing reconstruction mechanisms of ocean acidification, comprising: The multi-source data access module is used to acquire reconstruction prediction, target domain labels and label errors, product operation characteristics, training samples and the spatial and temporal locations of the grid to be evaluated; The multi-source risk calculation module is used to calculate the teacher model divergence risk, running feature distance risk, insufficient spatiotemporal observation risk, and label risk respectively, and retains the four risk components. The overall confidence module is used to generate an overall confidence score based on the four risk components. The dual-channel mechanism attribution module is used to separate the carbonate system verification channel and the remote sensing environment attribution channel and output the mechanism group contributions separately. The spatial joint evaluation module is used to jointly preserve the mechanism group contributions of residuals, risk components, sample density, and remote sensing environmental attribution channels within the same spatial grid or process partition. The product release module is used to generate release levels based on comprehensive confidence, single risk upper limit, and minimum spatiotemporal observation support, and simultaneously outputs coverage when outputting high-confidence subsets; The observation priority module is used to generate priority observation areas based on low confidence, existing sample reconstruction errors, and observation gaps.

[0104] To better understand the system of the present invention, the following non-limiting explanations are provided: Please see Figure 2 The system in this invention mainly includes: an input and multi-source risk layer, a mechanism and spatial joint layer, and a release and observation layer; wherein: The input and multi-source risk layer includes a multi-source data access module, a multi-source risk calculation module, and a comprehensive confidence module; the mechanism and spatial joint layer includes a dual-channel mechanism attribution module and a spatial joint evaluation module; and the release and observation layer includes a product release module and an observation priority module.

[0105] The multi-source data access module provides the multi-source risk calculation module with reconstruction predictions, label evidence, operational characteristics, and spatial-temporal information. The multi-source risk calculation module generates teacher disagreement risk, operational characteristic distance risk, insufficient spatiotemporal observation risk, and label risk. The comprehensive confidence module retains the four risk vectors and generates a comprehensive confidence score. The dual-channel mechanism attribution module generates evaluation results for the carbonate system verification channel and the remote sensing environment attribution channel, respectively. The spatial joint evaluation module maps residuals, risk components, sample density, and the mechanism group contribution of the remote sensing environment attribution channel to a unified spatial unit. The product release module generates release levels based on comprehensive confidence, single-item risk upper limits, and minimum spatiotemporal observation support. The observation priority module generates observation priority areas based on low confidence, existing sample reconstruction errors, and observation gaps.

[0106] The multi-source data access module acquires remote sensing reconstruction predictions of ocean acidification, target domain labels and their sources and errors, product operational characteristics, training sample information, and spatial and temporal information of samples or grids to be evaluated. Target domain labels include one or more of the following: direct in-situ pH labels, in-situ pCO2-constrained derived pH labels, weak product labels, and soft model labels. Product operational characteristics are remote sensing, reanalysis, or auxiliary environmental variables that can be continuously obtained during long-term product operation.

[0107] The multi-source risk calculation module includes a teacher model divergence risk calculation unit, a runtime feature distance risk calculation unit, a spatiotemporal insufficient observation risk calculation unit, and a label risk calculation unit.

[0108] The teacher model divergence risk calculation unit calculates the teacher model divergence degree based on the prediction differences of multiple teacher models for the same sample or grid to be evaluated, and normalizes the teacher model divergence degree to obtain the teacher model divergence risk. The more dispersed the teacher model's predictions, the higher the risk of divergence in the teacher model.

[0109] The runtime feature distance risk calculation unit compares the distance between the runtime feature vector of the sample or grid to be evaluated and the runtime feature distribution of the training samples, and normalizes the distance to obtain the runtime feature distance risk. The greater the deviation of the sample to be evaluated or the grid from the distribution of the running features of the training samples, the higher the risk of the running features.

[0110] The spatiotemporal observation insufficiency risk calculation unit is used to calculate the spatiotemporal observation support based on one or more of the following: the number of valid observations around the sample to be evaluated or the grid, spatial distance, and temporal distance. And calculate the risk of insufficient spatiotemporal observations according to the following formula:

[0111] in, The normalized spatiotemporal observation support is defined with a value range of [0,1]. This poses a risk of insufficient spatiotemporal observations. The denser the effective observations and the closer they are to the sample or grid being evaluated in terms of spatiotemporal distance, the greater the risk. The larger, The smaller.

[0112] The label risk calculation unit calculates label risk based on the label evidence level and label uncertainty. The direct on-site pH label, the on-site pCO2-constrained derived pH label, the product weak label, and the model soft label each correspond to a preset evidence risk; the greater the label uncertainty or the lower the evidence level, the higher the label risk. The multi-source risk calculation module stores four types of risk components respectively, and does not use a single risk component to replace the comprehensive evaluation result.

[0113] The comprehensive confidence module is used to fuse evidence regarding teacher model divergence risk, runtime feature distance risk, insufficient spatiotemporal observation risk, and label risk to form a comprehensive confidence score. In one implementation, the overall confidence level is calculated according to the following formula:

[0114] in, The higher the overall confidence level, the lower the overall risk of the sample or grid being evaluated. Risk of divergence in normalized teacher models; To normalize the distance risk of operational characteristics; Risks associated with insufficient normalized spatiotemporal observations; Risks associated with normalized labeling; For the teacher model disagreement risk weights, To run the feature distance risk weight, Risk weighting for insufficient spatiotemporal observation. The risk weights for the labels are a, b, c, and d, all of which are non-negative and control the degree of influence of the four types of risks on the overall confidence level, respectively. This is the preset lower limit of the overall confidence level; This indicates that the calculation result will be restricted to [ Within the range of [1]. While outputting the overall confidence level, the system retains and outputs four risk components to identify the specific sources of low confidence.

[0115] The dual-channel mechanism attribution module separates the carbonate system verification channel from the remote sensing environment attribution channel, and outputs the mechanism group contributions separately for each. The carbonate system verification channel utilizes pCO2, TA, DIC, and... One or more of the following are used to verify the pH reconstruction results using the carbonate system; the remote sensing environmental attribution channel only utilizes operational characteristics that can be continuously obtained during the long-term product operation phase to evaluate the grouped perturbations, grouped ablation, or contributions of feature groups such as spatial-temporal, thermohaline, bio-optical, and ocean dynamics. The label-generated variables in the carbonate system verification channel are not considered as long-term product operational characteristics without continuous input guarantees, and the mechanism group contributions are used to explain the model response relationship, not as separate proof of causal mechanisms.

[0116] The spatial joint evaluation module maps reconstruction residuals, teacher model divergence risk, operational feature distance risk, insufficient spatiotemporal observation risk, label risk, sample density, and mechanism group contributions from the remote sensing environment attribution channel to the same spatial grid or process partition, forming a spatial joint evaluation result. By comparing residuals, risk sources, spatiotemporal observation support, and mechanism group contributions in different regions, it identifies regions with high errors, high extrapolation risk, low spatiotemporal observation support, or anomalous mechanism responses.

[0117] The product release module generates a product release level based on the overall confidence level, the upper limit of each individual risk, and the minimum spatiotemporal observation support. Only when the overall confidence level reaches a preset threshold, each individual risk does not exceed its corresponding upper limit, and the spatiotemporal observation support meets the minimum requirement, will the sample or grid to be evaluated be included in the corresponding high-level release range. When outputting the high-confidence screening results, the system simultaneously outputs the coverage rate of the retained samples or grids relative to all valid samples or grids, avoiding substituting the accuracy of a low-coverage subset for the accuracy of a fully covered product.

[0118] The observation priority module generates observation priority areas by integrating low-confidence regions, existing sample reconstruction errors, and observation gaps. For regions with low overall confidence, large existing sample reconstruction errors, and low spatiotemporal observation support, the priority of supplementary observations is increased, and the module outputs the priority observation areas, main sources of risk, and the spatial range of recommended supplementary observations, providing a basis for subsequent field observation deployment and product updates.

[0119] Fourth embodiment: A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for attribution and confidence grading of ocean acidification remote sensing reconstruction mechanisms.

[0120] Fifth embodiment, a computer program product, including a computer program that, when executed by a processor, implements the above-described method for attribution and confidence grading of ocean acidification remote sensing reconstruction mechanisms.

[0121] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented, in whole or in part, as a computer program product, the computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line, or wireless (e.g., infrared, wireless, microwave, etc.) means). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0122] The above description is only a preferred embodiment of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for attribution and confidence grading of ocean acidification remote sensing reconstruction mechanisms, characterized in that, include: S1. Obtain the ocean acidification variable reconstruction prediction of the target sea area, the target domain label and its label source and label error, product operation characteristics, spatiotemporal information of training samples and spatiotemporal information of the grid to be evaluated; S2. For each sample or grid to be evaluated, calculate the divergence risk of prediction results from multiple teacher models, the feature distance risk of the sample to be evaluated relative to the training sample in the standardized product operation feature space, the risk of insufficient spatiotemporal observation based on neighboring observation distance and local sample density, and the label risk based on label source and label error. S3. Normalize and retain the divergence risk, feature distance risk, insufficient spatiotemporal observation risk and label risk as traceable risk vectors, and form a comprehensive confidence level based on the risk vectors; S4. Establish separate verification channels for the carbonate system and remote sensing environmental attribution channels, and save the evaluation results of the two channels separately. S5. Map the target domain residual, the traceable risk vector, the sample density, and the mechanism group contribution of the remote sensing environment attribution channel to the same spatial grid or process partition to form a spatial joint evaluation result. S6. The product release level is determined by three conditions: overall confidence level, individual risks, and minimum spatiotemporal observation support. When any individual risk exceeds the corresponding upper limit or the spatiotemporal observation support is lower than the minimum requirement, the highest release level is prohibited. The reconstruction error and coverage of the high-confidence retained subset are output synchronously for the results screened by confidence level. S7. Based on low confidence, existing sample reconstruction error and spatial observation gaps, determine the priority of supplementary observations, and output the product release level, risk component, mechanism contribution and observation priority area.

2. The method for attribution and confidence grading of ocean acidification remote sensing reconstruction mechanisms according to claim 1, characterized in that, The teacher model includes at least two regression models with different structures, and the risk of divergence is determined by the standard deviation, range, or quantile interval of each teacher's predicted value.

3. The method for attribution and confidence grading of ocean acidification remote sensing reconstruction mechanisms according to claim 1, characterized in that, The feature distance risk is determined by multiple nearest neighbor distances from the sample to be evaluated to the training samples, and is used for label generation by pCO2, TA, DIC, and The variables are not included in the calculation of the running feature distance; the risk of insufficient spatiotemporal observation is determined by at least two of the following: spatial distance to the target domain observation, temporal distance, local sample density, and regional coverage; the label risk is determined based on the evidence level and corresponding uncertainty of direct field pH, field pCO2 derived pH, weak product label, and soft model label.

4. The method for attribution and confidence grading of ocean acidification remote sensing reconstruction mechanisms according to claim 1, characterized in that, The carbonate system verification channel uses the variables involved in solving the carbonate system to evaluate the consistency of calculations, while the remote sensing environment attribution channel only uses the spatiotemporal, temperature-salinity, dynamic, and bio-optical variables available during the product operation phase to evaluate the contributions of the mechanism group.

5. The method for attribution and confidence grading of ocean acidification remote sensing reconstruction mechanisms according to claim 1, characterized in that, The mechanism group of the remote sensing environment attribution channel includes at least the space-time group, the thermo-salinity group, the dynamic mixing group, and the bio-optical group. The carbonate system verification channel separately includes pCO2, TA, DIC, and Group.

6. The method for attribution and confidence grading of ocean acidification remote sensing reconstruction mechanisms according to claim 1, characterized in that, The product release levels include at least high confidence, conditional confidence, low confidence, and missing data levels, with the highest release level simultaneously satisfying the conditions of comprehensive confidence, individual risk, and spatiotemporal observation support. The coverage rate is the ratio of the number of samples or grids entering the specified release level or high confidence subset to the total number of evaluable samples or grids, and is output simultaneously with the corresponding root mean square error, mean absolute error, or bias. The priority of supplementary observations increases with the increase of low confidence, normalized existing sample reconstruction error, and observation gap, and is associated with the output of suggested supplementary carbonate, nutrient, turbidity, water color, or dynamic observation variables.

7. An attribution and confidence grading system for remote sensing reconstruction of ocean acidification mechanisms, characterized in that, include: The multi-source data access module is used to acquire reconstruction prediction, target domain labels and label errors, product operation characteristics, training samples and the spatial and temporal locations of the grid to be evaluated; The multi-source risk calculation module is used to calculate the teacher model divergence risk, running feature distance risk, insufficient spatiotemporal observation risk, and label risk respectively, and retains the four risk components. The overall confidence module is used to generate an overall confidence score based on the four risk components. The dual-channel mechanism attribution module is used to separate the carbonate system verification channel and the remote sensing environment attribution channel and output the mechanism group contributions separately. The spatial joint evaluation module is used to jointly preserve the mechanism group contributions of residuals, risk components, sample density, and remote sensing environmental attribution channels within the same spatial grid or process partition. The product release module is used to generate release levels based on comprehensive confidence, single risk upper limit, and minimum spatiotemporal observation support, and simultaneously outputs coverage when outputting high-confidence subsets; The observation priority module is used to generate observation priority areas based on low confidence, existing sample reconstruction errors, and observation gaps.

8. The attribution and confidence grading system for remote sensing reconstruction mechanism of ocean acidification according to claim 7, characterized in that, The spatial joint evaluation module summarizes risk, residual, sample density, and mechanism contribution at two scales: preset process partitioning and fixed spatial grid.

9. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the attribution and confidence grading method for ocean acidification remote sensing reconstruction mechanism as described in any one of claims 1-6.

10. A computer-readable storage medium storing a computer program, characterized in that, When executed by the processor, the program implements the attribution and confidence grading method for ocean acidification remote sensing reconstruction mechanism as described in any one of claims 1-6.