A glacier mass balance prediction method and system based on uncertainty perception
Patent Information
- Application Number
- CN202611264027.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-20
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]针对现有技术的以上缺陷或改进需求,本发明提供了一种基于不确定度感知的冰川物质平衡预测方法及系统,由此解决现有冰川物质平衡预测技术存在难以深度挖掘不确定度多维结构和样本间关系、不具备物理先验引导的模型架构、在小样本微调中易引发灾难性遗忘、预测输出形式单一、缺乏可靠性评估的技术问题
(1)本发明基于分项不确定度为每个冰川样本生成用于表征置信度的权重,训练时使用权重矩阵对损失函数进行加权,深度挖掘不确定度多维结构和样本间关系,并引入物理先验引导的注意力掩码构建模型架构。引入层自适应低秩适配参数对基础模型进行参数微调,避免引发灾难性遗忘。现有模型的预测输出为单一点值,缺乏可靠性评估。本发明采用多随机种子集成推断策略,通过不同随机种子训练得到多个模型,对同一输入获得多个预测值并计算其均值和标准差,生成附带置信区间的冰川物质平衡预测结果,可以评估预测可靠性,为水资源管理决策提供了可量化的风险边界。当 K小于5时,样本量过小,计算出的标准差极易受到单次预测中偶然出现的极端值的影响,导致不确定性被严重高估或低估。K≥5保证了研究者在有限的计算资源下,能够以最低限度的统计可靠性,客观地呈现模型预测的置信区间,从而科学地评估冰川物质平衡产品的可信度。
Smart Images

Figure CN122817883A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of glaciology and artificial intelligence, and more specifically, relates to a method and system for predicting glacial mass balance based on uncertainty perception. Background Technology
[0002] The high-altitude regions of Asia possess the world's largest glacial reserves outside the polar regions, and changes in their mass balance are crucial to regional water security and sea-level rise prediction. Due to the harsh natural environment and complex topography of this region, traditional field observation methods can only cover a very small number of glaciers, resulting in a lack of data. Current technologies for estimating large-scale glacial mass balance largely rely on physical model simulations or satellite remote sensing inversion. Physical models are limited by uncertainties in their parameterization schemes and are prone to systematic biases in areas lacking field calibration; while remote sensing inversion products, although covering a wide area, introduce complex uncertainties through downscaling and error propagation processes.
[0003] Currently, the core challenge in this field is how to effectively utilize widely covered but uncertain inversion data, combined with sparse but high-precision measured data, to construct a model that combines generalization ability and local accuracy. While existing technologies attempt to introduce deep learning models such as Transformer and cross-attention mechanisms for glacier mass balance prediction, they suffer from the following shortcomings: Training sample quality is treated equally, making the model highly sensitive to noise. Remote sensing inversion of mass balance data covers a wide range but varies in quality. Inversion results from different glaciers, years, and regions may exhibit different uncertainties. Existing methods assign equal weights to all training samples, leading to the gradient contribution of high-noise samples being treated as equivalent to that of high-quality samples. This results in noise interfering with the direction of model parameter updates, leading to insufficient robustness. The model architecture lacks physical prior guidance. Existing methods employ the standard Transformer architecture, whose multi-head self-attention mechanism enables undifferentiated global interactions among all features. However, in glacier mass balance modeling, not all feature combinations have physical meaning. For example, there may be a known physical relationship between a certain topographic feature and a certain climate feature, while other feature combinations have no physical association. The standard Transformer cannot distinguish between effective and ineffective interactions and is prone to learning spurious correlations when data is sparse. Fine-tuning with small sample measured data can easily lead to catastrophic forgetting. When only a very limited amount of field observation data is available, traditional full-parameter fine-tuning strategies involve changing all model parameters. While adapting the model to local regional characteristics, this irreversibly overwrites the generalization rules learned from massive amounts of data during the pre-training phase. Existing models provide single-point predictions, lacking reliability assessment. Current methods output only a single, definitive value for any given glacier and year input, making it impossible for users to determine the reliability of the prediction.
[0004] It is evident that existing glacier mass balance prediction technologies suffer from several technical problems: difficulty in deeply exploring the multidimensional structure of uncertainty and the relationships between samples; lack of a model architecture guided by physical priors; susceptibility to catastrophic amnesia during small-sample fine-tuning; limited predictive output formats; and a lack of reliability assessment. Summary of the Invention
[0005] In response to the above-mentioned deficiencies or improvement needs of existing technologies, this invention provides a method and system for predicting glacier mass balance based on uncertainty perception. This solves the technical problems of existing glacier mass balance prediction technologies, such as difficulty in deeply mining the multidimensional structure of uncertainty and the relationship between samples, lack of a model architecture guided by physical priors, susceptibility to catastrophic forgetting in small sample fine-tuning, single prediction output form, and lack of reliability assessment.
[0006] To achieve the above objectives, according to a first aspect of the present invention, a method for predicting glacier mass balance based on uncertainty perception is provided, comprising: a training phase and a prediction phase; During the training phase, K different random seeds are set, and the training process is executed independently for each random seed to obtain K target region adaptation models, where K ≥ 5. The training process executed independently for each random seed includes the following steps: (1) Extract the feature vector of each glacier sample from the multi-source dataset of glaciers, including: static topographic features, dynamic meteorological features, true value of glacier mass balance and partial uncertainty. Combine the feature vectors of multiple glacier samples into a sample feature matrix. Divide the training set composed of multiple sample feature matrices into a pre-training set and a fine-tuning set. Generate weights for each glacier sample to characterize confidence based on partial uncertainty. The weights of all glacier samples in the sample feature matrix form a weight matrix. (2) The deep learning model is trained using a pre-training set. The initial parameters of the deep learning model are controlled by random seeds. The deep learning model performs nonlinear regression mapping on the features after the fusion of static terrain features and dynamic meteorological features to predict the simulation results of glacier mass balance. The error between the simulation results of glacier mass balance and the true value of glacier mass balance in the sample feature matrix is used as the loss function. The loss function is weighted using a weight matrix to obtain the total loss. The parameters are updated by backpropagation and trained until convergence to obtain the basic model. (3) When training the base model using the fine-tuning set, introduce the layer adaptive low-rank adaptation parameter to fine-tune the parameters of the base model and obtain the target region adaptation model; In the prediction phase, the static topographic features and dynamic meteorological features of the target glacier are input into the K target area adaptation models to obtain K predicted values. The mean of the K predicted values is used as the predicted value of glacier mass balance, and the standard deviation of the K predicted values is used to construct the confidence interval of the predicted value of glacier mass balance.
[0007] Furthermore, the weights in step (1) are generated as follows: The partial uncertainties of each glacier sample are extracted, including outlier interpolation uncertainty, elevation variation uncertainty, and density transformation uncertainty. The square root of the sum of the squares of the outlier interpolation uncertainty, elevation variation uncertainty, and density transformation uncertainty is taken as the total uncertainty. An exponential decay function is used to map the total uncertainty to the weight of each glacier sample.
[0008] Furthermore, the weights in step (1) are generated as follows: Each glacier sample is treated as a node, and the feature similarity between glacier samples is used as the edge weight to construct an uncertainty propagation graph; Extract the partial uncertainties for each glacier sample, including outlier interpolation uncertainty, elevation variation uncertainty, and density transformation uncertainty. Take the square root of the sum of the squares of outlier interpolation uncertainty, elevation variation uncertainty, and density transformation uncertainty as the total uncertainty. Use the total uncertainty and partial uncertainties of each glacier sample as the initial state of the node, and calculate the initial confidence score of the node. Multiple rounds of iterative information transmission are performed on the uncertainty propagation graph. In each round of iterative information transmission, each node aggregates the confidence scores of its neighboring nodes to obtain the confidence scores of each node in each round of iterative information transmission. The information entropy of each node in each round of iterative information transmission is calculated. Based on the information entropy, the confidence scores of each node in each round of iterative information transmission are weighted and fused to obtain the weight of each glacier sample.
[0009] Furthermore, the deep learning model includes: a feature grouping embedding module, a multi-head self-attention module, a physical information-guided attention mask module, a hierarchical feature fusion module, and a regression simulation head; The feature grouping embedding module groups the input data according to static terrain features and dynamic meteorological features, and then embeds and encodes them through independent linear mapping layers to obtain grouped feature representations. The physical information-guided attention masking module generates an attention mask matrix based on the known physical relationships between features; The multi-head self-attention module includes L stacked multi-head self-attention layers, where L≥2. Each layer of multi-head self-attention multiplies the grouped feature representation with the query projection matrix, key projection matrix, and value projection matrix respectively to obtain the query matrix, key matrix, and value matrix. The query matrix and key matrix are first subjected to dot product matching and then scaled for stabilization. Then, they are added to the attention mask matrix to obtain the intermediate matrix. The intermediate matrix is processed by the Softmax function and then multiplied with the value matrix to obtain the feature representation output by each layer of multi-head self-attention. The hierarchical feature fusion module fuses the feature representations output by the multi-head self-attention output of L stacked layers one by one to obtain fused features; The regression simulation head uses a multilayer perceptron to perform nonlinear regression mapping on the fused features to predict the simulation results of glacier mass balance.
[0010] Furthermore, before training the base model with the fine-tuning set, a low-rank decomposition matrix is added to the weight matrix of the base model, specifically as follows: Add a low-rank decomposition matrix to the weight matrix corresponding to the linear mapping layer of the feature grouping embedding module; Add low-rank decomposition matrices to the query projection matrix, key projection matrix, and value projection matrix of the multi-head self-attention module, respectively; Add a low-rank decomposition matrix to the weight matrix used in the layer-by-layer fusion process of the hierarchical feature fusion module; Add a low-rank decomposition matrix to the weight matrix corresponding to the multilayer perceptron of the regression simulation head.
[0011] Furthermore, the specific method for fine-tuning the parameters is as follows: Set the biases of the low-rank decomposition matrix and the regression simulation head as fine-tuning parameters, freeze all other parameters of the base model, train the base model using the fine-tuning set, and update only the fine-tuning parameters during training.
[0012] Furthermore, the rank of the low-rank decomposition matrix added in the feature grouping embedding module is less than the rank of the low-rank decomposition matrix added in the multi-head self-attention module, which is less than the rank of the low-rank decomposition matrix added in the hierarchical feature fusion module and the regression simulation head.
[0013] To achieve the above objectives, according to a second aspect of the present invention, a glacier mass balance prediction system based on uncertainty perception is provided, comprising: a training module and a prediction module; The training module is used to set K different random seeds, independently execute the training process for each random seed, and obtain K target region adaptation models, where K ≥ 5. The following training is performed independently for each random seed: Feature vectors for each glacier sample are extracted from the multi-source dataset of glaciers, including: static topographic features, dynamic meteorological features, true values of glacier mass balance and partial uncertainties. Feature vectors of multiple glacier samples are combined to form a sample feature matrix. The training set composed of multiple sample feature matrices is divided into a pre-training set and a fine-tuning set. Weights for each glacier sample are generated based on partial uncertainties to characterize confidence. The weights of all glacier samples in the sample feature matrix are combined to form a weight matrix. The deep learning model is trained using a pre-training set. The initial parameters of the deep learning model are controlled by random seeds. The deep learning model performs nonlinear regression mapping on the features after fusing static terrain features and dynamic meteorological features to predict the simulation results of glacier mass balance. The error between the simulation results of glacier mass balance and the true values of glacier mass balance in the sample feature matrix is used as the loss function. The loss function is weighted using a weight matrix to obtain the total loss. The parameters are updated by backpropagation and trained until convergence to obtain the basic model. When training the base model using fine-tuning sets, layer adaptive low-rank adaptation parameters are introduced to fine-tune the parameters of the base model, resulting in a target region adaptation model. The prediction module is used to input the static topographic features and dynamic meteorological features of the target glacier into K target area adaptation models to obtain K predicted values. The mean of the K predicted values is used as the predicted value of glacier mass balance, and the standard deviation of the K predicted values is used to construct the confidence interval of the predicted value of glacier mass balance.
[0014] To achieve the above objectives, according to a third aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of a method for predicting glacier mass balance based on uncertainty awareness.
[0015] To achieve the above objectives, according to a fourth aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of a method for predicting glacier mass balance based on uncertainty perception.
[0016] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: (1) This invention generates weights to characterize the confidence level for each glacier sample based on the partial uncertainty. During training, the weight matrix is used to weight the loss function, deeply explore the multidimensional structure of uncertainty and the relationship between samples, and introduce a physical prior-guided attention mask to construct the model architecture. Layer adaptive low-rank fitting parameters are introduced to fine-tune the parameters of the basic model to avoid catastrophic forgetting. The prediction output of the existing model is a single point value, which lacks reliability assessment. This invention adopts a multi-random seed ensemble inference strategy, which obtains multiple models through training with different random seeds, obtains multiple prediction values for the same input, and calculates their mean and standard deviation, generating glacier mass balance prediction results with confidence intervals. This can assess the reliability of the prediction and provide a quantifiable risk boundary for water resource management decisions. When K is less than 5, the sample size is too small, and the calculated standard deviation is easily affected by the extreme values that occur accidentally in a single prediction, resulting in a serious overestimation or underestimation of uncertainty. K≥5 ensures that researchers can objectively present the confidence interval of the model prediction with the minimum statistical reliability under limited computing resources, thereby scientifically assessing the credibility of glacier mass balance products.
[0017] (2) This invention provides two schemes for generating weights to characterize confidence. The basic weight scheme extracts the total uncertainty of each sample and maps it to a sample weight value using an exponential decay function. This scheme is simple to calculate and easy to implement, and is suitable for scenarios with high computational efficiency requirements. The composite weight scheme constructs a multi-source uncertainty propagation graph to realize the mutual propagation of confidence between samples. It jointly models the total uncertainty, the uncertainty of each component, and the similarity between samples, and considers both the level of total uncertainty and the structure of the uncertainty of each component. Compared with the single total uncertainty scheme, it is more comprehensive. It generates weights through iterative information propagation and entropy weighting. This iterative process aggregates the confidence scores of neighboring nodes, enabling the confidence of high-confidence samples to propagate to neighboring low-confidence samples in the feature space, and the confidence of low-confidence samples is corrected and improved. The composite weight scheme has higher accuracy and is suitable for scenarios with high computational accuracy requirements.
[0018] (3) Existing methods are standard Transformer or cross-attention mechanisms; this invention introduces feature grouping embedding and physically-guided attention masks into the standard Transformer framework. The purpose of the grouping embedding strategy is to ensure that features with different physical meanings are differentiated before being fed into the attention mechanism. The physically-guided attention mask module generates an attention mask matrix based on the known physical relationships between features, and selectively constrains the attention weights of the multi-head self-attention module. This design enables the model to interact only between physically meaningful features, reduces spurious correlation learning between unrelated feature pairs, and improves the model's generalization ability when data is sparse.
[0019] (4) Existing technologies, such as full parameter fine-tuning or three-stage progressive unfreezing, cannot avoid catastrophic forgetting when there is very little measured data. This invention introduces layer adaptive low-rank fitting parameters to fine-tune the parameters of the basic model to achieve effective fitting with <1% of trainable parameters. The large-scale glacier change patterns learned from massive data during the pre-training stage are completely preserved. At the same time, high-precision fitting capability to the measured area is obtained through high-level larger-rank fitting, fundamentally solving the problem of catastrophic forgetting. The bias of the low-rank decomposition matrix and the regression simulation head are set as fine-tuning parameters, and all other parameters of the basic model are frozen to ensure that the large-scale generalization patterns learned during the pre-training stage are completely preserved. Differentiated low-rank fitting dimensions are set for different functional layers of the model: smaller ranks are used in the bottom feature embedding layer because these layers encode general physical laws and do not require large adjustments; medium ranks are used in the multi-head self-attention layer; and larger ranks are used in the layered feature fusion module and the regression simulation head. Because these layers are directly oriented to the prediction output, more degrees of freedom are needed to fit local features. Attached Figure Description
[0020] Figure 1 This is a flowchart of the method provided in an embodiment of the present invention.
[0021] Figure 2 This is a simulation result of the mass balance of the RGI60-13.08624 glacier provided in an embodiment of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0023] like Figure 1 As shown, a method for predicting glacier mass balance based on uncertainty perception includes a training phase and a prediction phase. During the training phase, K different random seeds are set, and the training process is executed independently for each random seed to obtain K target region adaptation models, where K ≥ 5. The training process executed independently for each random seed includes the following steps: (1) Extract the feature vector of each glacier sample from the multi-source dataset of glaciers, including: static topographic features, dynamic meteorological features, true value of glacier mass balance and partial uncertainty. Combine the feature vectors of multiple glacier samples into a sample feature matrix. Divide the training set composed of multiple sample feature matrices into a pre-training set and a fine-tuning set. Generate weights for each glacier sample to characterize confidence based on partial uncertainty. The weights of all glacier samples in the sample feature matrix form a weight matrix. (2) The deep learning model is trained using a pre-training set. The initial parameters of the deep learning model are controlled by random seeds. The deep learning model performs nonlinear regression mapping after fusing the static terrain features and dynamic meteorological features of each glacier sample in the sample feature matrix. The predicted value of glacier mass balance of each glacier sample in the sample feature matrix is output. The error between the predicted value of glacier mass balance of each glacier sample in the sample feature matrix and its true value of glacier mass balance is used as the loss function of each glacier sample in the sample feature matrix. The loss function of each glacier sample in the sample feature matrix is weighted using a weight matrix to obtain the total loss. The parameters are updated by backpropagation and trained until convergence to obtain the basic model. (3) When training the base model using the fine-tuning set, introduce the layer adaptive low-rank adaptation parameter to fine-tune the parameters of the base model and obtain the target region adaptation model; In the prediction phase, the static topographic features and dynamic meteorological features of the target glacier are input into the K target area adaptation models to obtain K predicted values. The mean of the K predicted values is used as the predicted value of glacier mass balance, and the standard deviation of the K predicted values is used to construct the confidence interval of the predicted value of glacier mass balance.
[0024] Step (1) includes: The geometric boundaries of the target glacier were extracted based on global glacier inventory data (RGI 6.0). Static topographic features and dynamic meteorological features of each glacier were extracted using meteorological reanalysis data (ERA5-Land) and digital elevation models. Z-score normalization was used to normalize the continuous features and construct a basic feature vector. The static topographic features include glacier area, median elevation, slope and aspect, and the dynamic meteorological features include monthly mean temperature, solid precipitation and positive accumulated temperature. Obtain the annual material balance values and total uncertainty ( ), outlier interpolation uncertainty components ( ), Elevation variation uncertainty components ( and density transformation uncertainty components ( The remote sensing inversion dataset is obtained and connected with the basic feature vector to generate a standardized sample feature matrix.
[0025] Preferably, the generation of the weight matrix for characterizing sample confidence is performed using either of the following two schemes: Option 1: Basic Weighting Scheme Extract the total uncertainty for each sample The weights are mapped using an exponential decay function:
[0026] in The hyperparameter used to control the decay rate typically has the following range: The weights of all samples are normalized.
[0027] This solution is simple to calculate and easy to implement, making it suitable for scenarios with high computational efficiency requirements.
[0028] Option 2: Composite Weighting Scheme This scheme achieves the mutual propagation of confidence levels between samples by constructing a multi-source uncertainty propagation graph, which specifically includes the following sub-steps: (a) Constructing an uncertainty propagation diagram: Using each sample as a node ( A graph structure is constructed using the feature similarity between samples as edge weights. The edge weight is calculated as follows:
[0029] in , For the sample and samples eigenvectors, This is the bandwidth hyperparameter. The graph structure establishes connections between similar samples in the feature space, facilitating subsequent confidence propagation.
[0030] (b) Calculation of initial confidence level of nodes: The total uncertainty and individual uncertainties of each sample are combined as the initial state of the node, and the initial confidence score is calculated:
[0031] in Let be the coefficient of variation of the three individual uncertainties. , This is a hyperparameter. The initial score considers both the total uncertainty level and the structure of individual uncertainties, making it more comprehensive than a single total uncertainty scheme. The standard deviation of the partial uncertainty is... This represents the mean of the individual uncertainties.
[0032] (c) Information propagation based on graph structure: Perform on the uncertainty propagation graph Round-by-round information transmission ( In each round, each node aggregates the confidence information of its neighboring nodes:
[0033] in For the propagation coefficient, For iteration rounds, For nodes The set of neighboring nodes, This is the normalized edge weight. This iterative process allows the confidence of high-confidence samples to propagate to neighboring low-confidence samples in the feature space, thereby correcting and improving the confidence of low-confidence samples.
[0034] (d) Entropy-weighted fusion: Calculate the information entropy of the confidence estimate for each node in each round:
[0035] in . For the traversal index in the summation process, an adaptive weighted fusion is performed on the confidence estimates of each round based on the entropy value, with rounds having higher weights for lower entropy values:
[0036] in This is the temperature hyperparameter. Finally, the composite weights are normalized to obtain... .
[0037] Step (2) includes: The deep learning model is a hierarchical feature interaction deep learning model based on the Transformer architecture, including: a feature grouping and embedding module, a multi-head self-attention module, a physical information-guided attention mask module, a hierarchical feature fusion module, and a regression simulation head; The feature grouping embedding module groups the input data according to their physical meaning. Static terrain feature groups include glacier area, median elevation, slope, and aspect, while dynamic meteorological feature groups include monthly average temperature, solid precipitation, and positive accumulated temperature. These are embedded and encoded using independent linear mapping layers.
[0038]
[0039] Obtain grouped feature representation and Furthermore, learnable positional encodings are superimposed. The purpose of this grouping embedding strategy is to ensure that features with different physical meanings are differentiated before being fed into the attention mechanism.
[0040] The multi-head self-attention module performs global feature interaction modeling on the grouped feature representation. For the input sequence , respectively with the query projection matrix Key projection matrix Sum projection matrix Multiplication to calculate the query, key, and value matrix:
[0041]
[0042]
[0043] Attention output is:
[0044] in Attention mask matrix guided by physical information represents the feature dimension of the key matrix. The multi-head self-attention module consists of L stacked multi-head self-attention modules and a feedforward network.
[0045] The physical information-guided attention masking module generates an attention mask matrix based on the known physical relationships between features, selectively constraining the attention weights of the multi-head self-attention module. Specifically, a physical correlation matrix between features is defined based on knowledge from the field of glaciology. ,in Representation of features and characteristics There are known physical correlations between them, where the characteristics refer to glacier area, median elevation, slope, aspect, monthly average temperature, solid precipitation, or positive accumulated temperature. This indicates no direct physical connection or no explicit physical meaning. The attention mask matrix is as follows:
[0046] For feature pairs with no physical association, their attention weights are masked (set to zero). (After softmax, the value approaches 0), and the model no longer allocates attention between these feature pairs. This design allows the model to interact only between physically meaningful features, reducing spurious correlation learning between physically unrelated feature pairs and improving the model's generalization ability when data is sparse.
[0047] Compared with existing technical solutions that use cross-attention mechanisms to achieve bidirectional interaction between meteorology and terrain, this invention introduces the following design into the standard Transformer architecture: (1) grouping features according to their physical meaning before sending them into the attention mechanism; (2) selectively shielding features with no physical association based on domain knowledge. These designs add physical prior constraints and are fundamentally different.
[0048] The hierarchical feature fusion module fuses feature representations at different levels of abstraction layer by layer:
[0049] in For the first The output representation of the layer, This is for the splicing operation. This layered fusion strategy preserves information at different levels of abstraction.
[0050] The regression simulation head uses a multilayer perceptron to perform nonlinear regression mapping on the final fused features to predict the simulation results of glacier mass balance:
[0051] The uncertainty-weighted mean square error is used as the loss function during the pre-training phase.
[0052] in, and These are the actual value and predicted value of glacier mass balance for the i-th glacier sample (node) in the pre-training set, respectively.
[0053] Step (3) includes: Freeze all the original parameter matrices of the base model to ensure that the large-scale generalization rules learned in the pre-training phase are fully preserved. Differentiated low-rank adaptation dimensions are set for different functional layers of the model: lower rank is used in the bottom feature embedding layers because these layers encode general physical laws and do not require significant adjustments; medium rank is used in the multi-head self-attention layers; and higher rank is used in the hierarchical feature fusion module and regression simulation head because these layers directly face the prediction output and require more degrees of freedom to adapt to local features. Add low-rank decomposition matrices to the weight matrices of each functional layer. and The original weights are updated to (in The dimension is , The dimension is , For the aforementioned differentiated rank), only the layers , Set the matrix and output layer parameters to a trainable state; Based on sparse field observation data, a mean squared error loss function with weighted decay is adopted:
[0054] Perform backpropagation, updating only the low-rank adaptation matrix and output layer parameters. Let M be the sample summary of the fine-tuning set. This is the weight decay rate. and These are the actual and predicted values of glacier mass balance for the j-th glacier sample (node) in the fine-tuning set, respectively.
[0055] The advantage of this module is that the large-scale glacier change patterns learned from massive remote sensing inversion data during the pre-training stage are fully preserved. At the same time, it achieves high-precision fitting of the measured area through high-level large-rank adaptation, fundamentally solving the problem of catastrophic forgetting.
[0056] This invention systematically solves the problems of existing technologies being sensitive to training sample noise and prone to catastrophic forgetting during small sample fine-tuning through the synergistic effect of the following technical means: the basic weight scheme or composite weight scheme suppresses the gradient of low-quality samples from the source, and the contribution of high-noise samples to the loss function is greatly compressed, so that they do not dominate the direction of parameter update; the attention mask module guided by physical information shields the attention calculation between feature pairs that are not physically related, and prevents noisy features from interfering with the data representation inside the model through attention weights; the weighted loss function enables the gradient signal of high-quality samples to dominate during backpropagation.
[0057] The multi-random seed deep ensemble inference strategy used in this invention specifically includes: (a) Multi-random seed ensemble during the training phase: set up Each seed has a different random seed (e.g., 2021, 2022, ..., 2027). A complete pre-training and fine-tuning process is performed independently for each seed.
[0058] In deep learning model training, the initial parameters are controlled by random seeds. Different random seeds cause the model parameters to be at different initial positions in the loss function at the start of training. As training progresses, these parameters will converge to different local optima of the loss function. Specifically, in the glacier mass balance modeling scenario of this invention: The model trained on seed A may focus more on the physical path of "temperature to ablation" because its initialization makes the weight of monthly average temperature greater than other features. The model trained on seed B may focus more on the physical path of "precipitation to accumulation" because its initialization makes solid precipitation have a greater weight than other features. The model trained with seed C may focus more on the modulating effect of terrain on climate because its initialization makes the weight of static terrain features greater than that of dynamic meteorological features.
[0059] Thus obtained Adaptation model for each target region These models differ in both parameter space and functional preferences, collectively constituting an effective sampling of model initialization uncertainty.
[0060] (b) Integrated inference in the prediction phase: Input the target glacier and the year sample simultaneously. Each model undergoes one forward propagation to obtain... Predicted material balance values .calculate: mean , as the final predicted value of material balance; Standard deviation As a measure of uncertainty, it quantifies the cognitive uncertainty of the model caused by different random seeds.
[0061] (c) Confidence interval generation: Based on the mean and standard deviation, a confidence interval for the predicted material balance values is constructed. This corresponds to a 95% confidence level, meaning that, assuming the prediction error follows a normal distribution, this interval has a 95% probability of containing the true mass balance value. Verification shows that the distribution of the predicted values from the K models basically satisfies the normality assumption.
[0062] Existing technologies output only a single, specific point value for any given glacier-year input. Users cannot judge the reliability of the prediction results. This invention trains multiple models by setting multiple random seeds, obtains a set of predicted values, and calculates their mean and standard deviation, thereby achieving more accurate results.
[0063] Specifically, in a water resource management scenario, if a model predicts an annual mass balance of -0.52 mwe for a glacier, the manager cannot determine the reliability of this value. However, this invention provides a predicted value of -0.52 mwe with a 95% confidence interval of [-0.70, -0.34] mwe. A narrower interval indicates... The models' predictions for the glacier are highly consistent, indicating high confidence and can be directly used for decision-making. However, a wider prediction range suggests significant discrepancies between models trained with different random seeds, indicating high uncertainty. In such cases, managers should consider other information for a comprehensive assessment or prioritize on-site monitoring. This quantifiable reliability feedback is a key capability for big data-driven models to achieve operational applications.
[0064] The glacier mass balance prediction method based on uncertainty perception provided by this invention can be widely applied in fields such as artificial rain enhancement, snow enhancement effect evaluation, glacier protection engineering scheme formulation, and long-term water resource evolution prediction. In terms of weather modification, this invention's method uses changes in meteorological elements under different operational timing, areas, and dosages as input to simulate the degree of mass balance restoration in glacier accumulation zones. It then quantifies the reliability of the restoration effect based on prediction confidence intervals, providing a basis for quantitative optimization of operational schemes. In terms of water resource prediction, this invention's method can utilize future scenario data output from global climate models to drive models, simulating the evolution of upstream glacier mass balance. The output results can be further used for meltwater runoff trend analysis, providing scientific support for hydropower project site selection and installed capacity design.
[0065] Example 1 The implementation of the technical solution of this invention will be clearly and completely described below using the specific example of glacier RGI60-13.08624. This glacier is located in the RGI 13 region of the Asian high-altitude area, with complete remote sensing inversion records and sparse field observation data, which can fully demonstrate the technical effectiveness of this invention in addressing noise sensitivity and catastrophic forgetting issues. Finally, the effectiveness of the method is verified by comparing the predicted and measured data of the target glacier from 2004 to 2023.
[0066] Step 1: Obtain multi-source datasets of glaciers and construct a standardized sample feature matrix. First, based on the RGI 6.0 global glacier inventory data, the geometric boundaries of all glaciers corresponding to RGI zones 13, 14, and 15 were extracted. Static topographic features were then statistically analyzed using a digital elevation model: minimum elevation, maximum elevation, median elevation, average slope, average aspect, and glacier area. These features constitute a static topographic feature vector. .
[0067] Simultaneously, monthly temperature and precipitation corresponding to glaciers with remote sensing inversion mass balance data were extracted using ERA5-Land meteorological reanalysis data. After elevation reduction correction and spatial weighting, these were aggregated into annual-scale dynamic climate variables, and monthly mean temperature standard deviation and positive accumulated temperature were supplemented to form a dynamic meteorological feature vector. .
[0068] Regarding mass balance labeling, a global remote sensing inversion dataset integrating geodetic and glaciological observations was used to obtain glacier-by-glacier, year-by-year mass balance values and partial uncertainties. The partial uncertainties are: outlier interpolation uncertainty (…). ), Elevation variation uncertainty ( ), density transformation uncertainty ( The total uncertainty is synthesized from independent components:
[0069] The above operations were performed on glaciers in RGI 13–15 zones of the Asian highlands, aligning them by RGI ID and year, and retaining all years that simultaneously possess complete climate characteristics and effective mass balance labels. The final result is a standardized sample feature matrix, with each sample's fields being: [RGI_ID, Year, , , , , , The dataset was divided as follows: within the RGI 13–15 partition, with the centroid of glacier RGI60-13.08624 as the origin, glaciers beyond 50 km were included in the pre-training set, and neighboring glaciers within 50 km formed the fine-tuning set; the observation data of glacier RGI60-13.08624 from 2004 to 2023 were used as an independent test set.
[0070] Step 2: Analyze the individual uncertainty components to generate a weight matrix representing the sample confidence level. This embodiment employs a more accurate composite weighting scheme, using the 2022 sample with RGI60-15.03586 as the node. Detailed explanation. The selected hyperparameters are: , ,bandwidth Propagation coefficient Number of iterations Temperature parameters .
[0071] (a) Constructing an uncertainty propagation diagram In the standardized feature space, select and Euclidean distance of the 4 nearest neighbors Calculate edge weights:
[0072] Similarly, we can obtain And so on, eventually forming a graph structure. .
[0073] (b) Calculation of initial confidence level of nodes The three component uncertainties [0.14, 0.11, 0.08] have a mean of Standard deviation Coefficient of variation:
[0074] The initial confidence score is calculated using the following formula:
[0075] Substitute parameters for calculation:
[0076] (c) Information propagation based on graph structure After normalizing the edge weights, the first round of propagation iterations is calculated as follows:
[0077] Substitute the values:
[0078] The confidence sequence for each round is obtained by iterating for 5 consecutive rounds.
[0079] (d) Entropy-weighted fusion Calculate the proportional probability and information entropy for each round. ,by The composite weight of the sample is obtained by weighting the scores of each round with respect to the weights. After min-max normalization of all samples, the final training weights are... Due to significant variations in the uncertainty of the three terms in the 2022 sample of this glacier, the weights were reduced from an equal weight of 1.0 to 0.84, achieving a refined weight reduction.
[0080] Step 3: Construct a hierarchical feature interaction deep learning model based on the Transformer architecture and perform weighted pre-training. The constructed model includes a feature grouping embedding module, a multi-head self-attention module, a physical information-guided attention masking module, a hierarchical feature fusion module, and a regression simulation head.
[0081] Feature grouping embedding module The input features are divided into static terrain groups according to their physical meaning. ) and dynamic meteorological group ( ), respectively through independent linear mapping layers and Projecting onto the 256-dimensional latent space, we obtain and And superimposed with learnable positional codes to form the input sequence. .
[0082] Multi-head self-attention module The model includes Multi-head self-attention with stacked layers (head=8, For each layer, calculate:
[0083]
[0084]
[0085] Attention output formula:
[0086] Physical information-guided attention masking module Physical correlation matrix defined based on knowledge in the field of glaciology For example, there is a direct physical correlation between summer average temperature and positive accumulated temperature. Average slope aspect has no direct physical correlation with total annual precipitation. Generate a mask based on this. ,when The corresponding position is set to This forces the attention weights to be 0 after softmax. This design ensures that the model allocates attention only between feature pairs that have physical meaning, avoiding spurious correlations.
[0087] Layered feature fusion module The hidden states of the 4-layer output (i.e., the outputs of the 4-layer multi-head self-attention system) splicing along the passage, via Compressed to 512 dimensions .
[0088] Regression simulation head It outputs a single-substance balance prediction value.
[0089] During pre-training, all pre-training samples are loaded, and the uncertainty-weighted mean square error loss function is used:
[0090] in The composite weight scheme from step two is pre-calculated and fixed. The AdamW optimizer is used with an initial learning rate of 1e. -4 Batch size 2048, early stop strategy monitoring and validation set loss.
[0091] Step 4: Perform layer-adaptive low-rank fine-tuning using sparse field observation data The fine-tuning data consisted of neighboring glaciers within 50 km of the centroid of the RGI60-13.08624 glacier. First, all original parameters of the pre-trained model from step three were completely frozen. Then, low-rank decomposition matrices with differentiated ranks were added to different functional layers. Specific settings: Low-level feature embedding layer ( (weight matrix): rank ,Right now , ; Projection matrix of the middle multi-head self-attention layer :rank ; Weight matrix of high-level feature fusion module and regression output layer: rank .
[0092] Original weight update formula:
[0093] Only The output layer bias was set to trainable, and all other parameters were frozen. Fine-tuning was performed using the SGD optimizer (learning rate 5e). -3 (Momentum 0.9), loss function with weight decay:
[0094] Step 5: Generate a prediction product with confidence intervals using a multi-random seed ensemble inference strategy. (a) Multi-random seed ensemble during the training phase set up 10 random seeds (2021, 2022, …, 2027). For each seed, starting from random initialization, a complete pre-training and fine-tuning process is independently executed to obtain 7 functional target region adaptation models. Different seeds cause the model to converge to different local optima in the parameter space, effectively sampling initialization uncertainty.
[0095] (b) Integrated inference in the prediction phase Taking any target year of the RGI60-13.08624 glacier as an example, its climate and topographic features were input into seven models respectively, resulting in seven mass balance predictions. The mean was calculated. As the final predicted value for material balance, the standard deviation As a cognitive uncertainty of the model.
[0096] (c) Confidence interval generation Construct a 95% confidence interval:
[0097] The width of this interval directly reflects the reliability of the prediction. The narrower the interval, the higher the consensus among multiple models and the stronger the confidence in decision-making.
[0098] like Figure 2 As shown, to verify the practical performance of the method of this invention, the multi-random seed ensemble model, after training and fine-tuning as described above, was used to retrospectively predict 20 years of the RGI60-13.08624 glacier from 2004 to 2023, and the predictions were compared with the measured mass balance values. The predicted values and confidence intervals are the results of the ensemble of 7 models. All values are in millimeters of water equivalent (mmwe) to maintain consistency with the measured values. The confidence intervals cover the measured values for most years, with almost all measured values falling within the 95% confidence interval, showing a high degree of consistency with the theoretical values, thus verifying the reliability of the uncertainty output by the multi-random seed ensemble.
[0099] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for predicting glacier mass balance based on uncertainty perception, characterized in that, include: Training phase and prediction phase; During the training phase, K different random seeds are set, and the training process is executed independently for each random seed to obtain K target region adaptation models, where K ≥ 5. The training process executed independently for each random seed includes the following steps: (1) Extract the feature vector of each glacier sample from the multi-source dataset of glaciers, including: static topographic features, dynamic meteorological features, true value of glacier mass balance and partial uncertainty. Combine the feature vectors of multiple glacier samples into a sample feature matrix. Divide the training set composed of multiple sample feature matrices into a pre-training set and a fine-tuning set. Generate weights for each glacier sample to characterize confidence based on partial uncertainty. The weights of all glacier samples in the sample feature matrix form a weight matrix. (2) The deep learning model is trained using a pre-training set. The initial parameters of the deep learning model are controlled by random seeds. The deep learning model performs nonlinear regression mapping on the features after the fusion of static terrain features and dynamic meteorological features to predict the simulation results of glacier mass balance. The error between the simulation results of glacier mass balance and the true value of glacier mass balance in the sample feature matrix is used as the loss function. The loss function is weighted using a weight matrix to obtain the total loss. The parameters are updated by backpropagation and trained until convergence to obtain the basic model. (3) When training the base model using the fine-tuning set, introduce the layer adaptive low-rank adaptation parameter to fine-tune the parameters of the base model and obtain the target region adaptation model; In the prediction phase, the static topographic features and dynamic meteorological features of the target glacier are input into the K target area adaptation models to obtain K predicted values. The mean of the K predicted values is used as the predicted value of glacier mass balance, and the standard deviation of the K predicted values is used to construct the confidence interval of the predicted value of glacier mass balance.
2. The glacier mass balance prediction method based on uncertainty perception as described in claim 1, characterized in that, The weights in step (1) are generated as follows: The partial uncertainties of each glacier sample are extracted, including outlier interpolation uncertainty, elevation variation uncertainty, and density transformation uncertainty. The square root of the sum of the squares of the outlier interpolation uncertainty, elevation variation uncertainty, and density transformation uncertainty is taken as the total uncertainty. An exponential decay function is used to map the total uncertainty to the weight of each glacier sample.
3. The glacier mass balance prediction method based on uncertainty perception as described in claim 1, characterized in that, The weights in step (1) are generated as follows: Each glacier sample is treated as a node, and the feature similarity between glacier samples is used as the edge weight to construct an uncertainty propagation graph; Extract the partial uncertainties for each glacier sample, including outlier interpolation uncertainty, elevation variation uncertainty, and density transformation uncertainty. Take the square root of the sum of the squares of outlier interpolation uncertainty, elevation variation uncertainty, and density transformation uncertainty as the total uncertainty. Use the total uncertainty and partial uncertainties of each glacier sample as the initial state of the node, and calculate the initial confidence score of the node. Multiple rounds of iterative information transmission are performed on the uncertainty propagation graph. In each round of iterative information transmission, each node aggregates the confidence scores of its neighboring nodes to obtain the confidence scores of each node in each round of iterative information transmission. The information entropy of each node in each round of iterative information transmission is calculated. Based on the information entropy, the confidence scores of each node in each round of iterative information transmission are weighted and fused to obtain the weight of each glacier sample.
4. A method for predicting glacier mass balance based on uncertainty perception as described in any one of claims 1-3, characterized in that, The deep learning model includes: a feature grouping embedding module, a multi-head self-attention module, a physical information-guided attention mask module, a hierarchical feature fusion module, and a regression simulation head; The feature grouping embedding module groups the input data according to static terrain features and dynamic meteorological features, and then embeds and encodes them through independent linear mapping layers to obtain grouped feature representations. The physical information-guided attention masking module generates an attention mask matrix based on the known physical relationships between features; The multi-head self-attention module includes L stacked multi-head self-attention layers, where L≥2. Each layer of multi-head self-attention multiplies the grouped feature representation with the query projection matrix, key projection matrix, and value projection matrix respectively to obtain the query matrix, key matrix, and value matrix. The query matrix and key matrix are first subjected to dot product matching and then scaled for stabilization. Then, they are added to the attention mask matrix to obtain the intermediate matrix. The intermediate matrix is processed by the Softmax function and then multiplied with the value matrix to obtain the feature representation output by each layer of multi-head self-attention. The hierarchical feature fusion module fuses the feature representations output by the multi-head self-attention output of L stacked layers one by one to obtain fused features; The regression simulation head uses a multilayer perceptron to perform nonlinear regression mapping on the fused features to predict the simulation results of glacier mass balance.
5. The glacier mass balance prediction method based on uncertainty perception as described in claim 4, characterized in that, Before training the base model with the fine-tuned set, a low-rank decomposition matrix is added to the weight matrix of the base model, specifically as follows: Add a low-rank decomposition matrix to the weight matrix corresponding to the linear mapping layer of the feature grouping embedding module; Add low-rank decomposition matrices to the query projection matrix, key projection matrix, and value projection matrix of the multi-head self-attention module, respectively; Add a low-rank decomposition matrix to the weight matrix used in the layer-by-layer fusion process of the hierarchical feature fusion module; Add a low-rank decomposition matrix to the weight matrix corresponding to the multilayer perceptron of the regression simulation head.
6. The glacier mass balance prediction method based on uncertainty perception as described in claim 5, characterized in that, The specific method for fine-tuning the parameters is as follows: Set the biases of the low-rank decomposition matrix and the regression simulation head as fine-tuning parameters, freeze all other parameters of the base model, train the base model using the fine-tuning set, and update only the fine-tuning parameters during training.
7. A method for predicting glacier mass balance based on uncertainty perception as described in claim 5 or 6, characterized in that, The rank of the low-rank decomposition matrix added in the feature grouping embedding module is less than the rank of the low-rank decomposition matrix added in the multi-head self-attention module, which is less than the rank of the low-rank decomposition matrix added in the hierarchical feature fusion module and the regression simulation head.
8. A glacier mass balance prediction system based on uncertainty perception, characterized in that, include: Training module and prediction module; The training module is used to set K different random seeds, independently execute the training process for each random seed, and obtain K target region adaptation models, where K ≥ 5. The following training is performed independently for each random seed: Feature vectors for each glacier sample are extracted from the multi-source dataset of glaciers, including: static topographic features, dynamic meteorological features, true values of glacier mass balance and partial uncertainties. Feature vectors of multiple glacier samples are combined to form a sample feature matrix. The training set composed of multiple sample feature matrices is divided into a pre-training set and a fine-tuning set. Weights for each glacier sample are generated based on partial uncertainties to characterize confidence. The weights of all glacier samples in the sample feature matrix are combined to form a weight matrix. The deep learning model is trained using a pre-training set. The initial parameters of the deep learning model are controlled by random seeds. The deep learning model performs nonlinear regression mapping on the features after fusing static terrain features and dynamic meteorological features to predict the simulation results of glacier mass balance. The error between the simulation results of glacier mass balance and the true values of glacier mass balance in the sample feature matrix is used as the loss function. The loss function is weighted using a weight matrix to obtain the total loss. The parameters are updated by backpropagation and trained until convergence to obtain the basic model. When training the base model using fine-tuning sets, layer adaptive low-rank adaptation parameters are introduced to fine-tune the parameters of the base model, resulting in a target region adaptation model. The prediction module is used to input the static topographic features and dynamic meteorological features of the target glacier into K target area adaptation models to obtain K predicted values. The mean of the K predicted values is used as the predicted value of glacier mass balance, and the standard deviation of the K predicted values is used to construct the confidence interval of the predicted value of glacier mass balance.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of a glacier mass balance prediction method based on uncertainty awareness as described in any one of claims 1 to 7.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of a glacier mass balance prediction method based on uncertainty awareness as described in any one of claims 1 to 7.