A self-supervised neural network architecture search method for post-processing of rainfall forecasts
Through the self-supervised neural network architecture search method, a highly accurate neural network model is automatically designed, which solves the problems of small rainfall areas, high similarity and diversified geographical environment in the existing rainfall forecasting technology, and improves the accuracy of rainfall forecasting.
Patent Information
- Application Number
- CN202210029296.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-12
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-01-12
AI Technical Summary
The existing rainfall forecasting technology faces problems such as small rainfall areas, highly similar rainfall to backgrounds, and diversified geographical environments, making it time-consuming and difficult to effectively solve these challenges.
The self-supervised neural network architecture search method is adopted, and the high-accuracy neural network model is automatically designed to improve the accuracy of rainfall forecasts by introducing self-supervised search strategies and continuous perceived regularization functions (HSS).
Automatically design of high-accuracy neural network models is achieved, which improves the accuracy of rainfall forecasts, especially in terms of rainfall rating classification.
Smart Images

Figure CN114492960B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer prediction models, and particularly relates to a self-supervised neural network architecture search method for post-processing of rainfall forecasts. Technical Background
[0002] Limited by complex environmental observations and high requirements for expert domain knowledge, rainfall forecasting is one of the most difficult tasks in weather forecasting. Previous rainfall forecasting techniques mainly relied on numerical statistics-based models, such as numerical weather prediction (NWP) model systems, objective interpretation applications, and downscaling grid post-processing techniques. The development of NWP models has laid the foundation for rainfall forecasting. However, cloud dynamics and microphysical processes during rainfall occurrence cannot be effectively represented by a single model (such as the NWP model). Inevitably, biases will occur based on single-model predictions. To avoid such biases, ensemble forecasting methods combine multiple outputs of model variables. Before the booming development of deep learning, extracting useful rainfall forecasting information was a commonly used method.
[0003] Existing learning-based network design methods are limited by high requirements for expert domain knowledge, and rainfall forecasting faces challenges such as small rain areas, high similarity between rainfall and background, and diverse geographical environments. Manually designing networks to address these main challenges is very time-consuming.
[0004] The goal of neural architecture search (NAS) is to search for a robust and well-performing neural architecture by automatically selecting and combining different basic operations from a predefined search space to maximize the accuracy of a given task. These NAS-based methods can be divided into three aspects according to architecture optimization methods: 1. Evolutionary algorithms; 2. Reinforcement learning methods; 3. Gradient descent methods. They all have their own advantages, among which the differentiable NAS algorithm optimized by the gradient descent method has a fast search speed. Summary of the Invention
[0005] In view of the defects of the prior art, the present invention provides a self-supervised neural network architecture search method for post-processing of rainfall forecasts to reduce the manual design cost in rainfall prediction work; the present invention introduces a rainfall-aware search space with a self-supervised search strategy and a new continuous perception regularization function (HSS) to calculate the loss during training, so as to achieve the purpose of improving the performance of the final rainfall prediction task.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A self-supervised neural network architecture search method for post-processing rainfall forecasts. The method includes the NWP model processing the original meteorological data to obtain the input data graph for post-processing rainfall forecasts, inputting the input data graph for post-processing rainfall forecasts into the self-supervised neural network architecture search model. The search model uses the target network and the online network for contrastive learning, calculates the output difference between the target network and the online network and calculates the MSE loss, updates the target network and the online network after gradient, and obtains the online model after the search is completed. Then use the training data set to continue training the model, calculate the loss using the HSS loss function and update the gradient, and finally obtain the fully trained prediction model.
[0008] It should be noted that the initial parameters of the target network are the same as those of the online network. During the search process, the target network will be periodically updated by exponential moving average (EMA) according to the parameters of the online network; each time input data is input, the input data will be randomly cropped into x 1 , x 2 , x' 1 and x' 2 four seeds, and each network inputs two seeds; the backbone of the entire architecture contains multiple blocks, x 1 and x 2 will choose different paths when passing through the backbone to search for the optimal architecture; the online network performs gradient update according to the contrast loss between the output data and the output of the target network.
[0009] It should be noted that first, two online networks and target networks with the same initial parameters will be initialized. By iteratively searching the data set, a piece of data x will be randomly split into four seeds and divided into two groups (x 1 , x 2 ), and each network inputs two groups of seeds; after obtaining the network output, calculate the mean square error (MSE) of the two groups of inputs, perform loss calculation and backpropagation, and finally update the online network parameters, and use the exponential moving update (EMA) strategy to periodically update the target network parameters until the data set iteration is completed.
[0010] It should be noted that the search space adopted during the search process is a custom block-based operation block suitable for post-processing rainfall forecasts; the search space includes appropriate CNN and transformer operations, including residual blocks (RB), spatial awareness blocks (SAB), and channel awareness blocks (CAB); the channel awareness module is represented as follows:
[0011]
[0012] where z is the input data; z’ is the expectation of z; Sum(·) represents the internal summation; n = w × h - 1, and w and h are the width and height of the input data; Lamda = 10 -4 is a hyperparameter.
[0013] It should be noted that after the search is completed, the searched online model is obtained, and the search model is used to enter the training stage; in the training stage, the rainfall prediction post-processing data is still used, and the HSS loss function is used for loss calculation; the HSS loss function is expressed as follows:
[0014]
[0015] where loss MSE and loss HSS are the precipitation loss and intensity classification respectively; c H is a coefficient; ε is a constant, and its value can be 10 -10 .
[0016] The beneficial effects of the present invention are as follows: The self-supervised neural network architecture search model automatically designs a neural network model with high accuracy, and designs a regularization function by introducing the intensity classification index HSS to improve the accuracy of rainfall level classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Schematic diagrams for describing the technical background of the rainfall prediction method; among them, (a) data characteristics of rainfall forecasts. (b) Normal RGB images in object detection. (c) Comparison of previous methods and AdaNAS in post-processing of ensemble rainfall forecasts;
[0018] Figure 2 Schematic diagram of the overall process for describing the AdaNAS method to process rainfall forecast input data;
[0019] Figure 3 Schematic diagrams for describing the details of the channel-aware block (a) and the spatial-aware block (b);
[0020] Figure 4 Schematic diagram of the detailed architecture of a certain search of AdaNAS;
[0021] Figure 5 Schematic diagram of the implementation of the channel-aware block and the spatial-aware block in the AdaNAS search space;
[0022] Figure 6 Schematic diagram of the detailed architecture of a certain search of AdaNAS;
[0023] Figure 7 Multilevel association table for describing rainfall classification;
[0024] Figure 8 Implementation expression of the loss function for describing HSS. SPECIFIC EMBODIMENTS
[0025] The present invention will be further described below in conjunction with the accompanying drawings. It should be noted that this embodiment is based on the present technical solution, and detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to this embodiment.
[0026] The present invention is a self-supervised neural network architecture search method for post-processing rainfall forecasts. The method includes the NWP model processing the original meteorological data to obtain the input data graph for post-processing rainfall forecasts, inputting the input data graph for post-processing rainfall forecasts into the self-supervised neural network architecture search model. The search model uses the target network and the online network for contrastive learning, calculates the output difference between the target network and the online network and performs MSE loss calculation. After gradient updating the target network and the online network, the searched online model is obtained, and the training data set is used to continue training the model. After calculating the loss using the HSS loss function and gradient updating, the finally trained prediction model is obtained.
[0027] It should be noted that the initial parameters of the target network are the same as those of the online network. During the search process, the target network will be periodically updated by exponential moving average (EMA) according to the parameters of the online network; each time the input data is input, the input data will be randomly cropped into x 1 , x 2 , x' 1 and x' 2 four seeds, and two seeds are input into each network; the backbone of the entire architecture contains multiple blocks, and x 1 and x 2 will select different paths when passing through the backbone to search for the optimal architecture; the online network performs gradient update according to the contrast loss between the output data and the output of the target network.
[0028] It should be noted that first, two online networks and target networks with the same initial parameters will be initialized. By iteratively searching the data set, a piece of data x is randomly split into four seeds and divided into two groups (x 1 , x 2 ), and two groups of seeds are input into each network; after obtaining the network output, the mean square error (MSE) of the two groups of inputs is calculated, loss calculation and backpropagation are performed, and finally the online network parameters are updated, and the target network parameters are periodically updated using the exponential moving update (EMA) strategy until the data set iteration is completed.
[0029] It should be noted that the search space adopted during the search process is a custom block-based operation block suitable for post-processing rainfall forecasts; the search space includes appropriate CNN and transformer operations, including residual blocks (RB), spatial awareness blocks (SAB), and channel awareness blocks (CAB); the channel awareness module is represented as follows:
[0030]
[0031] Where z is the input data; z' is the expected value of z; Sum(·) represents the internal summation; n = w × h-1, w and h are the width and height of the input data; Lamda = 10 -4 is a hyperparameter.
[0032] It should be noted that after the search is completed, the searched online model is obtained, and the search model is used to enter the training phase; the training phase still uses the rainfall prediction post-processing data, and the HSS loss function is used for loss calculation; the HSS loss function is expressed as follows:
[0033]
[0034] Among them, loss MSE and loss HSS are the precipitation loss and intensity classification respectively; c H is a coefficient; ε is a constant, which can be 10 -10 .
[0035] Example
[0036] In order to fully understand the technical solution of the present invention, it is necessary to briefly explain the abbreviations and key terms used in the following embodiments.
[0037] NAS (neural architecture search): A neural network architecture search method that defines the search space, search strategy, and evaluation indicators to achieve the purpose of automatically searching for the optimal neural network architecture, reducing the cost of manual design, and achieving better performance.
[0038] NWP (numerical weather prediction) is a method of predicting the state of atmospheric movement and weather phenomena in a certain period of time in the future by solving the set of fluid mechanics and thermodynamics equations describing the weather evolution process based on the actual atmospheric conditions and under certain initial and boundary conditions through numerical calculations on large computers.
[0039] EMA (exponential moving average) can be used to estimate the local mean of a variable, so that the update of the variable is related to the historical value over a period of time. The moving average can be regarded as the average of the values of the variable over a period of time. Compared with the direct assignment of the variable, the value obtained by the moving average is smoother and less jittery on the graph, and the sliding average will not fluctuate greatly due to an abnormal value at a certain time.
[0040] The HSS score, one of the quantitative test methods for prediction accuracy, measures the prediction accuracy by excluding the cases where random predictions are correct. The value range of HSS is from negative infinity to 1, with the highest being 1.
[0041] Figure 1 Describes the differences between the input image data for rainfall prediction and the general RGB image input data, as well as the common techniques of current rainfall prediction methods. It includes (a) the input diagram of data features for rainfall prediction, (b) the ordinary RGB image for object detection, and (c) the comparison between previous methods and AdaNAS in the post - processing of ensemble rainfall prediction.
[0042] Figure 1 The rainfall prediction image shown in (a) is completely different from the ordinary RGB image containing clear and obvious objects in terms of image classification and object detection ( Figure 1 (b)). Because the images for rainfall prediction usually contain small rain areas and different geographical locations. Therefore, it is unreasonable to apply the existing custom networks based on ordinary RGB images to the rainfall prediction task of the present invention.
[0043] There are two main aspects in the post - processing process of rainfall prediction, as Figure 1 (c) shows, including traditional statistic - based methods and the latest learning - based methods. The traditional post - processing methods are statistic - based and use simple operations such as numerical averaging and sorting. The learning - based methods use multi - layer neural networks to post - process the ensemble rainfall prediction, and the learning - based methods perform better than the statistic - based methods. However, due to the high requirements of network design for expert domain knowledge. As known to the present invention, rainfall prediction faces challenges such as small rainfall areas, high similarity between rainfall / background, and diverse geographical environments. Manually designing a network to solve these main challenges is very time - consuming.
[0044] Based on the above problems, the present invention designs a neural network architecture search method based on self - supervised contrastive learning, called AdaNAS. Figure 2 It is the process of a search model and a training model for AdaNAS.
[0045] First, the original meteorological data needs to be processed by the NWP model to obtain the post - processing dataset for rainfall prediction. Then it enters the search stage, and the input pictures for post - processing of rainfall prediction are used as the input of the search stage.
[0046] Figure 3It is a schematic diagram of the architecture of the AdaNAS contrast learning process. Self-supervised contrast learning uses auxiliary tasks to mine information from unsupervised data to improve the quality of learned representations. In this way, supervised information is constructed to search for networks in order to learn representations valuable for downstream tasks, such as rainfall forecasting. Specifically, before the search process, a target network with the same structure as the online network is copied. As Figure 3 shown, the initial parameters of the target network are the same as those of the online network. During the search process, the target network is periodically updated using exponential moving average (EMA) according to the parameters of the online network. Each time input data is provided, the input data is randomly cropped into x 1 , x 2 , x' 1 and x' 2 four seeds, and two seeds are input into each network. The backbone of the entire architecture contains multiple blocks, and x 1 and x 2 will choose different paths when passing through the backbone to search for the optimal architecture. The online network performs gradient updates according to the contrast loss between the output data and the output of the target network. The pseudocode of the search process is as Figure 4 shown. That is, first, an online network and a target network with the same initial parameters are initialized. By iteratively searching the dataset, a piece of data x is randomly split into four seeds and divided into two groups (x 1 , x 2 ), and two groups of seeds are input into each network. After obtaining the network output, the mean squared error (MSE) of the two groups of inputs is calculated, loss calculation and backpropagation are performed, and finally the parameters of the online network are updated, and the parameters of the target network are periodically updated using the exponential moving update (EMA) strategy until the dataset iteration is completed.
[0047] The search space adopted during the search process is a custom block-based operation block suitable for post-processing work of rainfall prediction. It includes a search space with appropriate CNN and transformer operations, including residual blocks (RB), spatial awareness blocks (SAB), and channel awareness blocks (CAB). The channel awareness module is a more robust transformer-based module that determines the importance of neurons through the suppression effect on the surrounding space. The expression is as follows:
[0048]
[0049] where z is the input data; z' is the expectation of z; Sum(·) represents internal summation; n = w × h - 1, where w and h are the width and height of the input data; Lamda = 10 -4 is a hyperparameter. The composition details of the perception blocks of SAB and CAB of the present invention are as Figure 5 shown.
[0050] The search process uses a block-based self-supervised contrastive learning method to search for network architectures, rather than using shared weights as in one-shot NAS. To reduce the computational cost, most NAS methods adopt the weight-sharing rating scheme in one-shot NAS methods. The evaluation method with shared weights has low accuracy, and reducing shared weights can effectively improve the accuracy of evaluating the architecture ranking. The block method solves the prediction error of shared weights by dividing the network depth and retaining the original search space. Each block of the supernetwork is trained separately before being connected to the overall search. Therefore, each block of the present invention has a separate structure instead of being stacked on the same structure, which makes the network structure of the present invention more extensible. After a complete search phase, the possible search result network models are as Figure 6 shown. That is, the previous feature extractor and classifier are fixed. In the middle search backbone, it is divided into several blocks, and each block contains several operations. Each operation corresponds to selecting one from the custom search space perception blocks (RB, SAB, CAB). The operation with the largest priority weight is obtained through the gradient descent method, and finally the structure of the entire block is determined.
[0051] As Figure 2 shown, after the search is completed, a searched online model is obtained, and the search model is used to enter the training phase. In the training phase, the post-processed data of rainfall prediction is still used, and the HSS loss function is used for loss calculation. HSS is a classification criterion that measures the accuracy of prediction by excluding the cases where random predictions are correct. The loss function formula is as follows:
[0052]
[0053] where loss MSE and loss HSS are the precipitation loss and intensity classification respectively; c H is a coefficient; ε is a very small constant, which is taken as 10 -10 .
[0054] The reason why the present invention introduces HSS instead of Accuracy (ACC) is that when the meteorological center reports the forecast results to the public, it does not need to report the specific rainfall amount, but the level of rainfall amount. To verify the practical applicability of the post-processing method of the present invention, the rainfall prediction accuracy is introduced to prove whether it can correctly predict the rainfall level. The present invention divides rain into five categories according to the size of the rain: no [0.0, 0.1) mm / day, light rain [0.1, 10.1) mm / day, moderate rain [10.1, 25.1) mm / day, heavy rain [25.1, 50.1) mm / day, and heavy rainfall [50.1, ∞) mm / day. The present invention creates a multi-category association table according to the levels of all prediction maps, as Figure 7As shown. L = 5 is the number of rainfall categories; n i,j represents the number of cases where the actual observation is at level i but the prediction is at level j; N i ’ represents the total count of predicting rainfall at level i; N j represents the total count of observing at level j, N T represents the total prediction count. The general ACC score cannot truly reflect the prediction ability of the model for a dataset with a severely uneven class distribution. For example, the proportion of light precipitation events reaches 76.4%. The model only needs to predict any input light to ensure an ACC of 76.4%. However, such a model does not have sufficient prediction ability. HSS excludes this possibility, so the present invention introduces HSS into the regularization function. The calculation formulas of ACC and HSS are shown in reference Figure 7 after, and are implemented as Figure 8 shown. After iteratively inputting all the data, a trained prediction model is finally obtained, which can be used for rainfall forecasting.
[0055] For those skilled in the art, various corresponding changes can be made according to the above technical solutions and concepts, and all such changes should be included within the protection scope of the claims of the present invention.
Claims
1. A self-supervised neural network architecture search method for post-processing rainfall forecasts, characterized in that, the method includes the NWP model processing the original meteorological data to obtain an input data map for post-processing rainfall predictions, inputting the input data map for post-processing rainfall predictions into a self-supervised neural network architecture search model. The search model uses a target network and an online network for contrastive learning, calculates the output difference between the target network and the online network and calculates the MSE loss, updates the gradients of the target network and the online network to obtain an online model after the search is completed, continues to train the model using the training data set, calculates the loss using the HSS loss function and then updates the gradients, and finally obtains a fully trained prediction model; The initial parameters of the target network are the same as those of the online network. During the search process, the target network will be periodically updated by exponential moving average (EMA) according to the parameters of the online network; Each time data is input, the input data will be randomly cropped into x 1 , x 2 , x' 1 and x' 2 Four seeds, with two seeds input to each network; The backbone of the entire architecture contains multiple blocks, x 1 and x 2 will choose different paths when passing through the backbone to search for the optimal architecture; The online network performs gradient updates according to the comparison loss between the output data and the output of the target network; First, two online networks and a target network with the same initial parameters are initialized. By iteratively searching the dataset, a piece of data x is randomly split into four seeds and divided into two groups (x 1 , x 2 ). Each network is input with two groups of seeds. After obtaining the network outputs, the mean squared error (MSE) of the two groups of inputs is calculated, loss calculation and backpropagation are performed, and finally the online network parameters are updated, and the target network parameters are updated regularly using the exponential moving average (EMA) strategy until the dataset iteration is completed; The search space adopted during the search process is a custom block-based operation block suitable for post-processing rainfall prediction work; the search space includes appropriate CNN and transformer operations, including residual blocks (RB), spatial awareness blocks (SAB), and channel awareness blocks (CAB); the channel awareness block is represented as follows: Among them, z is the input data; z' is the expectation of z; Sum(·) represents the internal summation; n = w × h - 1, where w and h are the width and height of the input data; Lamda = 10 -4 is a hyperparameter; After the search is completed, a searched online model is obtained, and the search model is used to enter the training stage; the training stage still uses the post-processing data of rainfall predictions and calculates the loss using the HSS loss function. The HSS loss function is represented as follows: where loss MSE and loss HSS are precipitation loss and intensity classification respectively; c H is a coefficient; ε is a constant.
Citation Information
Patent Citations
Driver longitudinal car-following behavior model construction method based on deep reinforcement learning
CN112201069A
Neural network acceleration and embedding compression systems and methods with activation sparsification
IN202127001230A