A minute-level mountainous terrain near-rainfall prediction method and system based on improved Vision Transformer terrain physical constraints
By improving the Vision Transformer network and combining it with the physical constraints of terrain-forced precipitation, the problem of inaccurate heavy rainfall location forecast in complex mountainous environments is solved, and efficient and accurate minute-level nowcasting precipitation forecast is achieved.
Patent Information
- Application Number
- CN202511029929.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-25
AI Technical Summary
In complex mountainous environments, existing nowadays precipitation forecasting methods cannot fully consider the effects of terrain forcing, resulting in inaccurate forecasts of heavy precipitation locations. Traditional methods are computationally expensive and unstable, and deep learning methods lack the constraints of terrain physical mechanisms, resulting in low forecast accuracy.
A terrain physical constraint method based on the improved Vision Transformer is constructed. Through a position-adaptive dynamic convolutional network and a rain-free mask feature extraction strategy, combined with a combined loss function of terrain-forced precipitation physical constraints, model training is optimized to achieve efficient extraction of multi-source meteorological data and terrain features and capture of spatiotemporal dependencies, and output minute-level precipitation forecasts.
It improves the accuracy of heavy rainfall forecasts in complex mountainous environments, reduces random and meaningless rainfall forecasts, and improves the model's forecasting skills in complex mountainous environments.
Smart Images

Figure CN120541622B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of meteorological nowcasting precipitation, and in particular to a minute-level mountain nowcasting precipitation forecasting method and system based on improved VisionTransformer terrain physical constraints. Background Art
[0002] The disaster-causing rate of heavy rainfall within the near-term (0-2h) is very high. In mountainous areas with complex terrain, it often causes urban and rural waterlogging, mountain torrents, mudslides and other disasters, seriously threatening industrial and agricultural production and people's lives and property safety in economically fragile areas. Therefore, the study of high-quality near-term precipitation forecasts in complex mountainous environments is of great importance and significance for reducing casualties and property losses, improving disaster prevention and mitigation capabilities, and promoting social and economic development.
[0003] Nowcasting methods primarily include numerical models and extrapolation. Numerical models are computationally expensive and require long optimization cycles, resulting in instability in the 0-2 hour timeframe and limited application value. Traditional extrapolation methods, such as single-unit centroid tracking, cross-correlation, and optical flow, fail to fully exploit the characteristic information of the massive amount of high-dimensional meteorological data available. Their accuracy for complex, nonlinear precipitation processes is low, and decreases rapidly with increasing forecast time. Extrapolation methods, such as deep learning, have become a key technology for nowcasting in recent years due to their robust nonlinear computational capabilities. However, in complex mountainous environments, nowcasting is significantly influenced by topography. Current research has not fully considered the impact of orographically forced precipitation and lacks the constraints of the physical mechanisms of orographic precipitation in complex mountainous environments, resulting in limited modeling quality and heavy precipitation forecast accuracy. Summary of the Invention
[0004] Purpose of the invention: The purpose of the present invention is to provide a minute-level mountain near-fall precipitation forecasting method and system based on improved Vision Transformer terrain physical constraints, so as to solve the problems of inaccurate forecast of the near-fall location and low heavy precipitation forecasting skills caused by the influence of terrain forcing in complex mountainous environments.
[0005] Technical solution: The method of the present invention comprises the following steps:
[0006] Preprocess and normalize multi-source meteorological observation station data and terrain feature data in complex mountainous areas to obtain multi-source meteorological observation grid data and terrain feature grid data, and construct a model training dataset;
[0007] An improved Vision Transformer network based on a position-adaptive dynamic convolutional network and a rain-free mask feature extraction strategy is constructed. The network structure includes a feature extraction layer, a spatiotemporal feature fusion layer, and a mapping output layer. The feature extraction layer uses a position-adaptive dynamic convolutional network and a rain-free feature extraction strategy to achieve efficient dynamic independent extraction of multi-source meteorological observation grid data and terrain feature grid data, obtaining a classification feature vector. The classification feature vector is input into the spatiotemporal feature fusion layer, where the global and local spatiotemporal dependencies of the data features are captured through the Vision Transformer encoder, and the classification feature vectors are concatenated and fused. Finally, the mapping output layer maps the fused feature vector to the minute-level precipitation field for the next two hours, outputting a precipitation forecast.
[0008] Construct a combined loss function that can provide physical constraints on orographic forced precipitation, and constrain the physical consistency of model forecasts of orographic forced precipitation in complex mountainous environments;
[0009] The improved Vision Transformer network is trained based on the training dataset, and the relevant parameters are adjusted to obtain the optimal nowcasting precipitation forecast model of the improved Vision Transformer network.
[0010] Real-time multi-source meteorological observation grid data and terrain feature grid data are input into the optimal nowcasting model of the improved VisionTransformer network to obtain minute-level nowcasting forecasts.
[0011] Furthermore, the multi-source meteorological observation station data in complex mountainous areas include: radar combined reflectivity factor and ground meteorological automatic station minute-level precipitation observation station data, and the terrain feature data include: terrain slope, terrain angle and terrain gradient.
[0012] Furthermore, the multi-source meteorological observation station data and terrain feature data in complex mountainous areas are preprocessed and normalized to obtain multi-source meteorological observation grid data and terrain feature grid data, and the model training dataset is constructed; specifically:
[0013] Missing values are supplemented and singular values are replaced based on the climatological average value; then the inverse distance weighted method is used to unify the horizontal spatial resolution to the set grid field, and the temporal resolution is synchronized to the set interval. Data normalization is then performed to form multi-source meteorological observation grid data and terrain feature grid data. Finally, the model training dataset is constructed using the sample paradigm of "multi-source meteorological observation grid data and terrain feature grid data for the first 2 hours as input → precipitation forecast for the next 2 hours as output".
[0014] Furthermore, the position-adaptive dynamic convolutional network includes matrix diversification, position relationship learning and combined convolution kernel function. Matrix diversification is a matrix set , matrix set Included A weight matrix, that is , ;
[0015] The calculation formula for position relationship learning is:
[0016] ,
[0017] in, is a nonlinear function, represent Normalization function, is the position relationship vector of the input point, is the position adaptive weight coefficient matrix;
[0018] The combined convolution kernel function is obtained by combining the weight matrix in matrix diversification and the position adaptive weight coefficient of position relationship learning to obtain a dynamic convolution kernel of spatial information. :
[0019] ,
[0020] in, For the Position adaptive weight coefficient.
[0021] Furthermore, the spatiotemporal feature fusion layer is based on the Vision Transformer encoder. The number of heads is set according to the data type to capture the global and local spatiotemporal dependencies of data features, achieve high-quality spatiotemporal feature extraction, and enhance feature expression capabilities through nonlinear transformation of the feedforward neural network. The calculation formula is as follows:
[0022] First, perform spatiotemporal feature projection:
[0023] , , ,
[0024] in, is the query matrix, is the bond matrix, is the value matrix, is the number of long positions, is the input feature, is the time step, is the spatial position, is the feature dimension; 、 、 They are 、 、 The learnable projection matrix of
[0025] Then perform multi-head attention mechanism calculation:
[0026] ,
[0027] in, is the multi-head attention classification feature vector, Its function is to convert any real number into a probability distribution. is the bond matrix The transposed matrix of The projected dimensions for the query and key in each header;
[0028] After outputting the multi-head attention classification feature vector, perform splicing and fusion:
[0029] ,
[0030] in, is the feature matrix after splicing and fusion, For splicing operation, is the multi-head attention weight coefficient, and finally the feedforward neural network is used for nonlinear transformation:
[0031] ,
[0032] ,
[0033] ,
[0034] in, is the feature vector output by the regularization layer, Output feature vector for feedforward neural network, is the eigenvector after nonlinear transformation, is a nonlinear activation function, 、 is the weight coefficient of the feedforward neural network, 、 is the bias term of the feedforward neural network, is the regularization layer, It is a feed-forward neural network.
[0035] Furthermore, the mapping output layer uses a fully connected layer to flatten the feature vector and map it to the minute-level precipitation for the next two hours, outputting precipitation forecast data. The specific calculation formula is as follows:
[0036] Compress the spatiotemporal features and flatten them into vectors:
[0037] ,
[0038] in, is the flattened eigenvector, For the flattening operation, is the eigenvector after nonlinear change;
[0039] Finally, the map outputs the precipitation forecast:
[0040] ,
[0041] in, 、 , is the output dimension, is the original dimension of the input feature vector, is the dimension of the middle hidden layer, is the output normalized precipitation value, 、 is the bias term of the fully connected layer, is a nonlinear activation function; then, through anti-normalization restoration, the output is the precipitation forecast :
[0042] ,
[0043] in, 、 Represent the maximum and minimum values in the input data samples respectively.
[0044] Furthermore, the constructed combined loss function The expression is:
[0045] ,
[0046] in, is the ETS precipitation score loss item, is the orographic forced precipitation loss term, 、 are the weight coefficients of the ETS precipitation score loss term and the terrain forced precipitation loss term, respectively;
[0047] The calculation formula is:
[0048] ,
[0049] ,
[0050] in, is the time number of forecast accuracy, is the number of empty report events, is the number of missed events, is the number of correct negations, is the expected number of hits of the random forecast;
[0051] The loss function for terrain features is as follows:
[0052] ,
[0053] in, Indicates the grid points, is the total number of grid points, 、 represent the weight coefficients of precipitation gradient and terrain gradient respectively, 、 represent the precipitation on the windward and leeward slopes, respectively. 、 Represent the terrain gradient and precipitation gradient grid data respectively.
[0054] The system of the present invention comprises:
[0055] The dataset construction unit is used to preprocess and normalize the multi-source meteorological observation station data and terrain feature data in complex mountainous areas to obtain multi-source meteorological observation grid data and terrain feature grid data and construct a model training dataset;
[0056] The network construction unit is used to construct an improved Vision Transformer network based on a position-adaptive dynamic convolutional network and a rain-free mask feature extraction strategy. The network structure includes a feature extraction layer, a spatiotemporal feature fusion layer, and a mapping output layer. The feature extraction layer uses a position-adaptive dynamic convolutional network and a rain-free feature extraction strategy to achieve efficient dynamic independent extraction of multi-source meteorological observation grid data and terrain feature grid data to obtain a classification feature vector. The classification feature vector is input into the spatiotemporal feature fusion layer, and the global and local spatiotemporal dependencies of the data features are captured through the Vision Transformer encoder. The classification feature vectors are then concatenated and fused. Finally, the fused feature vector is mapped to the minute-level precipitation field for the next two hours through the mapping output layer to output a precipitation forecast.
[0057] The combined loss function construction unit is used to construct a combined loss function that can provide physical constraints on terrain-forced precipitation, constraining the physical consistency of the model forecast of terrain-forced precipitation in complex mountainous environments;
[0058] A training unit, used to train the improved Vision Transformer network based on the training data set, adjust relevant parameters, and obtain an optimal nowcasting precipitation forecast model of the improved Vision Transformer network;
[0059] The prediction unit is used to input real-time multi-source meteorological observation grid data and terrain feature grid data into the optimal nowcasting model of the improved Vision Transformer network to obtain minute-level nowcasting forecasts.
[0060] The electronic device of the present invention includes a memory, a processor, and a computer program / instruction stored in the memory and executable on the processor. When the computer program / instruction is executed by the processor, the steps of the method for minute-level mountainous area nowcasting based on improved Vision Transformer terrain physical constraints are implemented.
[0061] The computer-readable storage medium of the present invention stores computer instructions, which, when called, are used to execute the steps of the method for minute-level mountain nowcasting forecasting based on improved Vision Transformer terrain physical constraints.
[0062] Beneficial effects: Compared with the existing technology, the significant technical effects of the present invention are: (1) the position-adaptive dynamic convolutional network is introduced to improve the Vision Transformer network, each data feature is extracted independently, and a rain-free mask feature extraction strategy is designed, which improves the data feature extraction quality and computational efficiency; (2) through the physical mechanism of terrain-forced precipitation and the ETS precipitation score, a customized combined loss function is designed to impose consistency constraints on the physical law of terrain-forced precipitation in precipitation forecast, reduce random and meaningless precipitation forecasts, make the heavy precipitation forecast location more accurate, and improve the model's heavy precipitation forecasting skills in complex mountainous environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 Flow chart of the method of the present invention;
[0064] Figure 2 This is an example diagram of the two-dimensional grid data distribution of radar combination reflectivity factors;
[0065] Figure 3 To improve the Vision Transformer network structure diagram;
[0066] Figure 4 Schematic diagram of the position dynamic convolutional network structure. DETAILED DESCRIPTION
[0067] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0068] like Figure 1 As shown, the present invention provides a minute-level mountain nowcasting method based on improved Vision Transformer terrain physical constraints, comprising the following steps:
[0069] Step 1: Preprocess and normalize the multi-source meteorological observation station data and terrain feature data in complex mountainous areas to obtain multi-source meteorological observation grid data and terrain feature grid data, and construct a model training dataset;
[0070] Multi-source meteorological observation site data include radar combined reflectivity factors and ground meteorological automatic station minute-level precipitation observation site data. Terrain feature data include: terrain slope, terrain angle and terrain gradient. Multi-source meteorological observation grid data include: radar combined reflectivity factors and ground meteorological automatic station minute-level precipitation observation grid data. Terrain feature grid data include grid data of terrain slope, terrain angle and terrain gradient. Multi-source meteorological observation grid data and terrain feature grid data constitute the model training data set.
[0071] Collect radar combined reflectivity factor, ground meteorological automatic station minute-level precipitation observation site data, terrain slope, terrain angle and terrain gradient data in complex mountainous areas. Radar combined reflectivity factor is the maximum reflectivity factor of each vertical height layer in meteorological radar observation projected onto the coordinate grid point to form two-dimensional distribution data, which can comprehensively reflect the precipitation particle density, cloud structure and storm intensity characteristics, and is widely used in severe convective weather monitoring and other scenarios. Figure 2 As shown in the figure, different colors represent the intensity of the radar reflectivity factor, measured in decibels (dBz). A higher intensity indicates a greater probability of heavy precipitation and severe convective weather during the actual weather process. The data is corrected and supplemented using the climatological mean (generally the average over the past three decades) to improve data quality. Specifically, the climatological mean is used to fill in missing values and replace outliers. Subsequently, the inverse distance weighted interpolation method is used to unify the horizontal spatial resolution to a set grid field (a 0.125° × 0.125° grid field in this example), and the temporal resolution is synchronized to a set interval (a 10-minute interval in this example). Data normalization is then performed to form a dataset containing multi-source meteorological observation grid data (radar reflectivity factor and minute-level precipitation observation grid data from ground-based automatic meteorological stations) and terrain feature grid data (grid data on terrain slope, terrain angle, and terrain gradient). The model training dataset is then constructed using the sample paradigm of "multi-source meteorological observation grid data and terrain feature grid data for the previous 2 hours as input, and precipitation for the next 2 hours as output."
[0072] Terrain slope, terrain angle, and terrain gradient are calculated as follows:
[0073] (1),
[0074] (2),
[0075] (3),
[0076] in, is the terrain slope, is the terrain angle, is the terrain gradient, and They are 、 The elevation change rate in the direction. The terrain angle result ranges from 0° to 360°, which means the angle rotated clockwise from the north direction. and When it is 0, it is a flat area and can be marked as an invalid value.
[0077] The inverse distance weighted interpolation method is used to preprocess the observation data of the station to form a 10-minute grid field with a horizontal spatial resolution of 0.125°×0.125°. The calculation formula is as follows:
[0078] (4),
[0079] (5),
[0080] in, is the inverse distance weight coefficient of the discrete site, For the The Euclidean distance from the discrete station to the grid point to be interpolated, For the The Euclidean distance from the discrete station to the grid point to be interpolated, is the number of nearest points around the required interpolation point, is the meteorological element observation value at discrete stations, For the surrounding discrete sites, is the grid point value that needs to be interpolated in the end.
[0081] The data normalization calculation formula is as follows:
[0082] (6),
[0083] in, 、 Represent the maximum and minimum values of the entire sample of input data, respectively. is the dataset sample after normalization, is the grid value.
[0084] Step 2: Construct an improved VisionTransformer network with a position-adaptive dynamic convolutional network and a rain-free mask feature extraction strategy, and write a loss function that can provide physical constraints for terrain-forced precipitation. The improved VisionTransformer network is divided into a feature extraction layer, a spatiotemporal feature fusion layer, and a mapping output layer, as shown in the following example: Figure 3 As shown. Among them:
[0085] Feature extraction layer: A position-adaptive dynamic convolutional network is introduced to realize feature extraction of multi-source grid data (i.e., the unified grid field in the training data set in step 1: ground meteorological automatic station minute-level precipitation observation grid data, radar combined reflectivity factor grid data, and terrain slope, terrain angle, and terrain gradient grid data) to obtain classification feature vectors. At the same time, a rain-free feature extraction strategy is designed. For moments when no precipitation occurs, special masking is performed on the ground meteorological automatic station minute-level precipitation observation grid data and radar combined reflectivity factor grid data. The specific approach is: when the ground meteorological automatic station minute-level precipitation observation grid data are all 0, the grid values of the ground meteorological automatic station minute-level precipitation observation grid data and radar combined reflectivity factor grid data are marked as the normalized special value -1. No feature extraction is performed on such data in the model to avoid invalid feature extraction and improve computational efficiency and feature extraction quality. The processing method is as follows:
[0086] when season: ;
[0087] in, is the minute-level precipitation observation grid data of the ground meteorological automatic station at a certain moment, The radar combination reflectivity factor grid data at a certain moment.
[0088] The position-adaptive dynamic convolutional network is obtained by matrix diversification, position relationship learning and combined convolution kernel function calculation.
[0089] Matrix diversification is a matrix set , matrix set Each of represents a weight matrix, , is the number of weight matrices, set to 8 or 16.
[0090] (7),
[0091] The weight matrix of each neighborhood point obtained ,Right now The weighted sum of weight matrices, (such as Figure 4 (a)).
[0092] The calculation formula for position relationship learning is:
[0093] (8),
[0094] in, is a nonlinear function (such as logarithmic function, cosine function, etc.), represent Normalization function, is the position relationship vector of the input point, is the position adaptive weight coefficient (such as Figure 4 (b)).
[0095] Combined convolution kernel function: A dynamic convolution kernel with spatial information is obtained by combining the weight matrix in matrix diversification with the weight coefficient learned from positional relationship. (like Figure 4 In (c)) is:
[0096] (9),
[0097] in, For the Position adaptive weight coefficient.
[0098] Spatiotemporal feature fusion layer: Based on the Vision Transformer encoder, the classification feature vector is positionally encoded and the number of multiple heads is set according to the data type to capture the global and local spatiotemporal dependencies of the classification feature vector, thus achieving high-quality spatiotemporal feature extraction (such as Figure 2 ), and enhance the feature expression ability through nonlinear transformation of feedforward neural network. The calculation formula is as follows:
[0099] To prevent the computational complexity from exploding, we first perform spatiotemporal feature projection:
[0100] , , (10),
[0101] in, is the query matrix, is the bond matrix, is the value matrix, is the number of long positions, is the input feature, is the time step, is the spatial position, is the feature dimension; 、 、 They are 、 、 The learnable projection matrix.
[0102] Then perform multi-head attention mechanism calculation:
[0103] (11),
[0104] in, is the multi-head attention classification feature vector, Its function is to convert any real number into a probability distribution, highlighting important information and suppressing minor information through exponential operations. is the query matrix, is the bond matrix The transposed matrix of Projected dimensions for the query and key in each header.
[0105] After outputting the multi-head attention classification feature vector, perform splicing and fusion:
[0106] (12),
[0107] in, is the feature matrix after splicing and fusion, For splicing operation, is the multi-head attention weight coefficient matrix, and finally the feedforward neural network is used for nonlinear transformation:
[0108] (13),
[0109] (14),
[0110] (15),
[0111] in, is the feature vector output by the regularization layer, Output feature vector for feedforward neural network, is the characteristic item after nonlinear change, is a nonlinear activation function, 、 is the weight coefficient of the feedforward neural network, 、 is the bias term of the feedforward neural network, is the regularization layer, It is a feed-forward neural network.
[0112] Mapping output layer: The mapping output layer uses a fully connected layer to flatten the feature vector and map the feature vector to the next 2 hours of minute-level precipitation, outputting the precipitation forecast. The specific calculation formula is as follows:
[0113] Flattened to a vector:
[0114] (16),
[0115] in, is the flattened eigenvector, Flattening operation.
[0116] Finally, the map outputs the precipitation forecast:
[0117] (17),
[0118] in, 、 , is the output dimension, is the original dimension of the input feature vector, is the dimension of the middle hidden layer, is the output normalized precipitation value, 、 is the bias term of the fully connected layer, is a nonlinear activation function; then, through anti-normalization restoration, the output is the precipitation forecast .
[0119] (18),
[0120] in, 、 They represent the maximum and minimum values in the entire sample of input data respectively.
[0121] Step 3: Construct a combined loss function that can provide physical constraints on terrain-forced precipitation, and impose physical consistency constraints on terrain-forced precipitation in complex mountainous environments on model precipitation forecasts.
[0122] In order to maintain the physical consistency of terrain-forced precipitation in nowcasting under complex mountainous environments, reduce random and meaningless precipitation forecasts, make the heavy precipitation forecast location more accurate, and improve the model's heavy precipitation forecasting skills under complex mountainous environments, a combined loss function is constructed by combining the physical mechanism of terrain-forced precipitation and the ETS precipitation score. , the calculation formula is as follows:
[0123] (19),
[0124] in, is the ETS precipitation score loss item, is the orographic forced precipitation loss term, 、 are the weight coefficients of the ETS precipitation score loss term and the terrain forced precipitation loss term, respectively.
[0125] The calculation formula is:
[0126] (20),
[0127] (twenty one),
[0128] in, is the time number of forecast accuracy, is the number of empty report events, is the number of missed events, is the number of correct denials (no precipitation is forecasted, no precipitation is observed), is the expected number of hits of the random forecast.
[0129] is the terrain forced loss function, and the formula is as follows:
[0130] (twenty two),
[0131] in, Indicates the grid points, is the total number of grid points, 、 represent the weight coefficients of precipitation gradient and terrain gradient respectively, 、 represent the precipitation on the windward and leeward slopes, respectively. 、 The first term on the right side of the equation penalizes situations where the predicted precipitation on the leeward side exceeds that on the windward side; the second term forces the terrain gradient to align with the precipitation gradient, thereby achieving a physical consistency constraint on terrain-forced precipitation.
[0132] Step 4: Based on the training data set and the improved Vision Transformer network, continuously adjust the weight coefficient of the combined loss function 、 , optimizer learning rate, batch size, training rounds and other parameters, and trained the optimal precipitation forecast model of the improved Vision Transformer network. The details are as follows:
[0133] Weight coefficient of combined loss term: as initially set 、 , which is dynamically adjusted according to the importance of the terrain constraints.
[0134] Optimizer: Use Adam optimizer with initial learning rate lr=1e-4.
[0135] Batch size: Set according to GPU memory, such as batch_size=32.
[0136] Training rounds: You can set the maximum number of rounds to max_epochs=20, and dynamically terminate training using the early stopping method.
[0137] Dynamic parameter adjustment: If 5 consecutive rounds of validation sets If it does not decrease, increase the terrain constraint weight (such as += 0.1, upper limit 0.8).
[0138] Early stopping mechanism: If the validation loss does not decrease for 20 consecutive rounds, the training is terminated and the weight coefficient of the loss term is further adjusted to start training again to avoid overfitting.
[0139] Step 5: Input the real-time multi-source meteorological observation grid data and terrain feature grid data into the optimal nowcasting model of the improved VisionTransformer network to obtain the nowcasting data;
[0140] The radar combined reflectivity factor of the mountainous area, minute-level observation site data of the ground automatic meteorological station, as well as terrain slope, angle and gradient data are collected in real time. After preprocessing, the unified grid data is formed and input into the optimal nowcasting forecast model of the improved Vision Transformer network. The output is a nowcasting forecast with terrain physical constraints.
[0141] The embodiment of the present invention further provides a minute-level mountain nowcasting precipitation forecasting system based on improved Vision Transformer terrain physical constraints, including:
[0142] The dataset construction unit is used to preprocess and normalize the multi-source meteorological observation station data and terrain feature data in complex mountainous areas to obtain multi-source meteorological observation grid data and terrain feature grid data and construct the model training dataset;
[0143] The network construction unit is used to construct an improved Vision Transformer network based on a position-adaptive dynamic convolutional network and a rain-free mask feature extraction strategy. The network structure includes a feature extraction layer, a spatiotemporal feature fusion layer, and a mapping output layer. The feature extraction layer uses a position-adaptive dynamic convolutional network and a rain-free feature extraction strategy to achieve efficient dynamic independent extraction of multi-source meteorological observation grid data and terrain feature grid data to obtain a classification feature vector. The classification feature vector is input into the spatiotemporal feature fusion layer, and the global and local spatiotemporal dependencies of the data features are captured through the Vision Transformer encoder. The classification feature vectors are then concatenated and fused. Finally, the fused feature vector is mapped to the minute-level precipitation field for the next two hours through the mapping output layer to output a precipitation forecast.
[0144] The combined loss function construction unit is used to construct a combined loss function that can provide physical constraints on terrain-forced precipitation, constraining the physical consistency of the model forecast of terrain-forced precipitation in complex mountainous environments;
[0145] A training unit, used to train the improved Vision Transformer network based on the training data set, adjust relevant parameters, and obtain an optimal nowcasting precipitation forecast model of the improved Vision Transformer network;
[0146] The prediction unit is used to input real-time multi-source meteorological observation grid data and terrain feature grid data into the optimal nowcasting model of the improved Vision Transformer network to obtain minute-level nowcasting forecasts.
[0147] An embodiment of the present invention also provides an electronic device, including a memory, a processor, and a computer program / instruction stored in the memory and executable on the processor. When the computer program / instruction is executed by the processor, the steps of the method for minute-level mountainous nowcasting prediction based on improved Vision Transformer terrain physical constraints are implemented.
[0148] An embodiment of the present invention also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute the steps of the method for minute-level mountain precipitation forecasting based on improved Vision Transformer terrain physical constraints.
Claims
1. A method for minute-level mountain nowcasting based on terrain physical constraints of an improved Vision Transformer, characterized by: The following steps are involved: Preprocess and normalize multi-source meteorological observation station data and terrain feature data in complex mountainous areas to obtain multi-source meteorological observation grid data and terrain feature grid data, and construct a model training dataset; An improved Vision Transformer network based on a position-adaptive dynamic convolutional network and a rain-free mask feature extraction strategy is constructed. The network structure includes a feature extraction layer, a spatiotemporal feature fusion layer, and a mapping output layer. The feature extraction layer uses a position-adaptive dynamic convolutional network and a rain-free feature extraction strategy to achieve efficient dynamic independent extraction of multi-source meteorological observation grid data and terrain feature grid data, obtaining a classification feature vector. The classification feature vector is input into the spatiotemporal feature fusion layer, where the global and local spatiotemporal dependencies of the data features are captured through the Vision Transformer encoder, and the classification feature vectors are concatenated and fused. Finally, the mapping output layer maps the fused feature vector to the minute-level precipitation field for the next two hours, outputting a precipitation forecast. Construct a combined loss function that can provide physical constraints on terrain forced precipitation, and constrain the physical consistency of model forecasts of terrain forced precipitation in complex mountainous environments; construct a combined loss function The expression is: , in, is the ETS precipitation score loss item, is the orographic forced precipitation loss term, 、 are the weight coefficients of the ETS precipitation score loss term and the terrain forced precipitation loss term, respectively; The calculation formula is: , in, Indicates the grid points, is the total number of grid points, 、 represent the weight coefficients of precipitation gradient and terrain gradient respectively, 、 represent the precipitation on the windward and leeward slopes, respectively. 、 Represent the terrain gradient and precipitation gradient grid data respectively; The improved Vision Transformer network is trained based on the training dataset, and the relevant parameters are adjusted to obtain the optimal nowcasting precipitation forecast model of the improved Vision Transformer network. Real-time multi-source meteorological observation grid data and terrain feature grid data are input into the optimal nowcasting model of the improved Vision Transformer network to obtain minute-level nowcasting forecasts.
2. The method for minute-level mountain nowcasting based on improved Vision Transformer terrain physical constraints according to claim 1 is characterized in that: The multi-source meteorological observation station data in complex mountainous areas include: radar combined reflectivity factor and minute-level precipitation observation station data of ground meteorological automatic stations; the terrain feature data include: terrain slope, terrain angle and terrain gradient.
3. The method for minute-level mountain nowcasting based on improved Vision Transformer terrain physical constraints according to claim 1 is characterized in that: The multi-source meteorological observation station data and terrain feature data in complex mountainous areas are preprocessed and normalized to obtain multi-source meteorological observation grid data and terrain feature grid data, and the model training dataset is constructed. Specifically: Missing values are supplemented and singular values are replaced based on the climatological average value. The inverse distance weighted method is then used to unify the horizontal spatial resolution to the set grid field, and the temporal resolution is synchronized to the set interval. Data normalization is then performed to form multi-source meteorological observation grid data and terrain feature grid data. Finally, a model training dataset is constructed using the sample paradigm of "multi-source meteorological observation grid data and terrain feature grid data for the first two hours as input → precipitation forecast for the next two hours as output." 4. The method for minute-level mountain nowcasting based on improved Vision Transformer terrain physical constraints according to claim 1 is characterized in that: Position-adaptive dynamic convolutional network includes matrix diversification, position relationship learning and combined convolution kernel function. Matrix diversification is a matrix set , matrix set Included A weight matrix, that is , ; The calculation formula for position relationship learning is: , in, is a nonlinear function, represent Normalization function, is the position relationship vector of the input point, is the position adaptive weight coefficient matrix; The combined convolution kernel function is obtained by combining the weight matrix in matrix diversification and the position adaptive weight coefficient of position relationship learning to obtain a dynamic convolution kernel of spatial information. : , in, For the Position adaptive weight coefficient.
5. The method for minute-level mountain nowcasting based on improved Vision Transformer terrain physical constraints according to claim 1 is characterized in that: The spatiotemporal feature fusion layer is based on the Vision Transformer encoder. The number of heads is set according to the data type to capture the global and local spatiotemporal dependencies of data features, achieve high-quality spatiotemporal feature extraction, and enhance feature expression capabilities through nonlinear transformations of the feedforward neural network. The calculation formula is as follows: First, perform spatiotemporal feature projection: , , , in, is the query matrix, is the bond matrix, is the value matrix, is the number of long positions, is the input feature, is the time step, is the spatial position, is the feature dimension; 、 、 They are 、 、 The learnable projection matrix of Then perform multi-head attention mechanism calculation: , in, is the multi-head attention classification feature vector, Its function is to convert any real number into a probability distribution. is the bond matrix The transposed matrix of The projected dimensions for the query and key in each header; After outputting the multi-head attention classification feature vector, perform splicing and fusion: , in, is the feature matrix after splicing and fusion, For splicing operations, is the multi-head attention weight coefficient, and finally the feedforward neural network is used for nonlinear transformation: , , , in, is the feature vector output by the regularization layer, Output feature vector for feedforward neural network, is the eigenvector after nonlinear transformation, is a nonlinear activation function, 、 is the weight coefficient of the feedforward neural network, 、 is the bias term of the feedforward neural network, is the regularization layer, It is a feed-forward neural network.
6. The method for minute-level mountain nowcasting based on improved Vision Transformer terrain physical constraints according to claim 1 is characterized in that: The mapping output layer uses a fully connected layer to flatten the feature vector and map it to the minute-level precipitation for the next two hours, outputting precipitation forecast data. The specific calculation formula is as follows: Compress the spatiotemporal features and flatten them into vectors: , in, is the flattened eigenvector, For the flattening operation, is the eigenvector after nonlinear change; Finally, the map outputs the precipitation forecast: , in, 、 , is the output dimension, is the original dimension of the input feature vector, is the dimension of the middle hidden layer, is the output normalized precipitation value, 、 is the bias term of the fully connected layer, is a nonlinear activation function; then, through anti-normalization restoration, the output is the precipitation forecast : , in, 、 Represent the maximum and minimum values in the input data samples respectively.
7. The method for minute-level mountain nowcasting based on improved Vision Transformer terrain physical constraints according to claim 1 is characterized in that: The calculation formula is: , , in, is the time number of forecast accuracy, is the number of empty report events, is the number of missed events, is the number of correct negations, is the expected number of hits of the random forecast.
8. A minute-level mountain nowcasting system based on improved Vision Transformer terrain physical constraints, characterized by: include: The dataset construction unit is used to preprocess and normalize the multi-source meteorological observation station data and terrain feature data in complex mountainous areas to obtain multi-source meteorological observation grid data and terrain feature grid data and construct the model training dataset; The network construction unit is used to construct an improved Vision Transformer network based on a position-adaptive dynamic convolutional network and a rain-free mask feature extraction strategy. The network structure includes a feature extraction layer, a spatiotemporal feature fusion layer, and a mapping output layer. The feature extraction layer uses a position-adaptive dynamic convolutional network and a rain-free feature extraction strategy to achieve efficient dynamic independent extraction of multi-source meteorological observation grid data and terrain feature grid data to obtain a classification feature vector. The classification feature vector is input into the spatiotemporal feature fusion layer, and the global and local spatiotemporal dependencies of the data features are captured through the Vision Transformer encoder. The classification feature vectors are then concatenated and fused. Finally, the fused feature vector is mapped to the minute-level precipitation field for the next two hours through the mapping output layer to output a precipitation forecast. The combined loss function construction unit is used to construct a combined loss function that can provide physical constraints on terrain forced precipitation, and to constrain the physical consistency of the model forecast of terrain forced precipitation in complex mountainous environments; the constructed combined loss function The expression is: , in, is the ETS precipitation score loss item, is the orographic forced precipitation loss term, 、 are the weight coefficients of the ETS precipitation score loss term and the terrain forced precipitation loss term, respectively; The calculation formula is: , in, Indicates the grid points, is the total number of grid points, 、 represent the weight coefficients of precipitation gradient and terrain gradient respectively, 、 represent the precipitation on the windward and leeward slopes, respectively. 、 Represent the terrain gradient and precipitation gradient grid data respectively; A training unit, used to train the improved Vision Transformer network based on the training data set, adjust relevant parameters, and obtain an optimal nowcasting precipitation forecast model of the improved Vision Transformer network; The prediction unit is used to input real-time multi-source meteorological observation grid data and terrain feature grid data into the optimal nowcasting model of the improved VisionTransformer network to obtain minute-level nowcasting forecasts.
9. An electronic device, characterized in that: The invention comprises a memory, a processor, and a computer program / instruction stored in the memory and executable on the processor, wherein when the computer program / instruction is executed by the processor, the steps of the method for minute-level mountain nowcasting prediction based on improved Vision Transformer terrain physical constraints are implemented according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which, when called, are used to execute the steps of the method for minute-level mountain nowcasting prediction based on improved VisionTransformer terrain physical constraints as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Lightweight rapid forecasting system and method for typhoon and storm surge
CN119805625A
Near-surface meteorological field downscaling method based on topographic constraint Transform model
CN120045923A