Intelligent estimation method for rainfall of local area
By integrating historical rainfall data with atmospheric circulation factors, the initial population range of the optimization algorithm is dynamically determined, and the hyperparameters of the neural network model are optimized. This solves the problem of local optimum solutions in rainfall prediction in climate transition zones, achieves accurate estimation of rainfall in local areas, and improves the reliability and adaptability of prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-28
- Publication Date
- 2026-04-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing optimization algorithms are prone to getting stuck in local optima during rainfall prediction, leading to inaccurate estimates of rainfall during the flood season, especially in seasonal and stochastic rainfall data in climate transition zones.
By integrating historical rainfall data from multiple monitoring stations in the target area with atmospheric circulation factors, the fluctuation and periodic characteristics of rainfall data are accurately analyzed. The initial population range of the optimization algorithm is dynamically determined, and the hyperparameters of the neural network model are optimized by combining the atmospheric circulation factors of the period to be predicted, so as to achieve accurate estimation of rainfall in local areas.
It effectively adapts to the complex variation characteristics of rainfall in climate transition zones, improves the reliability and pertinence of rainfall estimation, and achieves accurate estimation of rainfall for the forecast period.
Smart Images

Figure CN121786513A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and specifically to a method for intelligent estimation of rainfall in a local area. Background Technology
[0002] In the transitional zone between warm temperate and subtropical climates, the region is controlled by the westerly atmospheric circulation, resulting in distinct climatic characteristics and large fluctuations in rainfall. It exhibits both obvious seasonal features and random fluctuations, leading to the coexistence of flood and drought risks. This not only affects daily life and agricultural production but also seriously threatens the lives and property of local residents. Therefore, intelligent estimation and early forecasting of rainfall in this region has become a key requirement for reducing related economic losses.
[0003] In these transitional climate zones, a flood season lasting several months occurs each year, during which rainfall increases dramatically, often reaching the highest annual rainfall. Meanwhile, rainfall data exhibits repetitive, ambiguous, and highly random characteristics over time, posing a significant challenge to traditional forecasting models.
[0004] Existing technologies have attempted to integrate long short-term memory networks (LSTM) with sparrow optimization algorithms or back propagation (BP) neural networks with particle swarm optimization algorithms for rainfall prediction. However, these optimization algorithms require uniform distribution during population initialization. When faced with rainfall data that exhibits significant seasonality and randomness, their fixed initialization patterns can lead to the algorithms getting trapped in local optima during the optimization process, especially when dealing with the transition from non-flood season to flood season or vice versa. This significantly reduces the accuracy and reliability of key rainfall estimation during the flood season. Summary of the Invention
[0005] To address the technical problem in existing technologies where the initialization method of optimization algorithms easily gets trapped in local optima during rainfall prediction, leading to inaccurate rainfall estimates during the flood season, the present invention aims to provide an intelligent rainfall estimation method for local areas. The specific technical solution adopted is as follows: Firstly, a method for intelligently estimating rainfall in a local area is provided, comprising: acquiring atmospheric circulation factors and historical rainfall data from multiple monitoring stations in the target area; analyzing the fluctuation and periodic characteristics of the rainfall data based on the historical rainfall data and atmospheric circulation factors to obtain the complexity during the flood season and the complexity during the non-flood season; determining the initial population range of the optimization algorithm based on the complexity during the flood season, the complexity during the non-flood season, and the atmospheric circulation factors; optimizing the hyperparameters of the neural network model using the optimization algorithm based on the atmospheric circulation factors and the initial population range, and training the neural network model based on the hyperparameters obtained from the optimization; and inputting the atmospheric circulation factors of the period to be predicted into the trained neural network model to obtain the predicted rainfall value of the target area during the period to be predicted.
[0006] Based on the above technical solution, the intelligent estimation method for local rainfall provided by this invention integrates historical rainfall data from multiple monitoring stations in the target area with atmospheric circulation factors. It accurately analyzes the fluctuation and periodic characteristics of rainfall data to obtain the complexity during the flood season and the complexity during the non-flood season. Based on this, it dynamically determines the initial population range of the optimization algorithm, avoiding the limitations of a fixed initialization mode. Then, combined with the atmospheric circulation factors of the period to be predicted, the optimal hyperparameters are obtained through optimization algorithm and the neural network model is trained. This enables accurate estimation of rainfall in the local area during the period to be predicted, effectively adapting to the complex variation characteristics of rainfall in the climate transition zone and improving the reliability and pertinence of rainfall estimation.
[0007] In conjunction with the first aspect mentioned above, in one possible implementation, the method for analyzing the fluctuation and periodic characteristics of rainfall data based on historical rainfall data and atmospheric circulation factors to obtain the complexity of the flood season and the complexity of the non-flood season specifically includes: analyzing the degree of correlation between historical rainfall data and atmospheric circulation factors at monitoring stations within the target area, and classifying multiple monitoring stations into different levels; clustering the historical rainfall data of each monitoring station to obtain rainfall data during the flood season and rainfall data during the non-flood season; determining the complexity value of each monitoring station based on its historical rainfall data and atmospheric circulation factors; and determining the complexity of the flood season and the complexity of the non-flood season in the target area based on the complexity value of each monitoring station, the weight of its level, and the clustering results of the rainfall data during the flood season and the rainfall data during the non-flood season.
[0008] In conjunction with the first aspect above, in one possible implementation, the method for determining the complexity value of each monitoring station based on historical rainfall data and atmospheric circulation factors specifically includes: obtaining a first parameter based on the rank correlation between historical rainfall data and atmospheric circulation factors at each monitoring station; obtaining a second parameter based on the degree of random fluctuation in historical rainfall data at each monitoring station; obtaining a third parameter based on the sequence similarity of historical rainfall data at each monitoring station in adjacent periods; and determining the complexity value of each monitoring station based on the first, second, and third parameters.
[0009] In conjunction with the first aspect mentioned above, in one possible implementation, the method for determining the flood season complexity and non-flood season complexity of a target area based on the complexity value and weight of each monitoring station's level, combined with the clustering results of flood season and non-flood season rainfall data, specifically includes: analyzing the impact of each rainfall data point on the complexity value of its respective monitoring station based on historical rainfall data for each monitoring station, and obtaining a first flood season complexity component and a first non-flood season complexity component based on the clustering results; analyzing the impact of each rainfall data point on the global information entropy based on historical rainfall data and weight of multiple monitoring stations' levels, and obtaining a second flood season complexity component and a second non-flood season complexity component based on the clustering results; obtaining the flood season complexity based on the first flood season complexity component and the second flood season complexity component; and obtaining the non-flood season complexity based on the first non-flood season complexity component and the second non-flood season complexity component.
[0010] In conjunction with the first aspect mentioned above, in one possible implementation, the method of analyzing the impact of historical rainfall data on the complexity value of each monitoring station, and obtaining the first flood season complexity component and the first non-flood season complexity component based on clustering results, specifically includes: for each monitoring station, comparing the complexity value of the monitoring station with the masked complexity value of each rainfall data to determine the first degree of impact of each rainfall data on the complexity value; the masked complexity value is the complexity value re-determined after masking one rainfall data; statistically analyzing the average first degree of impact of flood season rainfall data and the average first degree of impact of non-flood season rainfall data for each monitoring station, and normalizing the average first degree of impact to obtain the first flood season complexity component and the first non-flood season complexity component.
[0011] In conjunction with the first aspect mentioned above, in one possible implementation, the method described above, which analyzes the impact of historical rainfall data from multiple monitoring stations and the weights of their respective levels on the global information entropy, and obtains the second flood season complexity component and the second non-flood season complexity component based on clustering results, specifically includes: performing multi-scale decomposition processing on the historical rainfall data of each monitoring station to obtain approximate and detailed components of each rainfall data, and constructing an enhanced feature vector based on atmospheric circulation factors; determining the global information entropy of the historical rainfall data based on the enhanced feature vectors corresponding to the historical rainfall data from multiple monitoring stations, and analyzing the second impact degree of each rainfall data on the global information entropy based on the weights of the monitoring station's level; statistically analyzing the mean of the second impact degree of the flood season rainfall data and the mean of the second impact degree of the non-flood season rainfall data for each monitoring station, and normalizing the mean of the second impact degree to obtain the second flood season complexity component and the second non-flood season complexity component.
[0012] In conjunction with the first aspect mentioned above, in one possible implementation, the method for determining the initial population range of the optimization algorithm based on the complexity of the flood season, the complexity of the non-flood season, and atmospheric circulation factors specifically includes: statistically analyzing the global statistical characteristics of historical rainfall data and calculating the mean value of atmospheric circulation factors; determining dynamic adjustment coefficients based on the atmospheric circulation factors and the mean value of atmospheric circulation factors for the period to be predicted; and mapping the initial population range of the hyperparameter space to generate based on the dynamic adjustment coefficients and rainfall statistical characteristics.
[0013] In conjunction with the first aspect mentioned above, in one possible implementation, the method for classifying multiple monitoring stations by analyzing the correlation between historical rainfall data and atmospheric circulation factors at the target area specifically includes: constructing a spatiotemporal correlation matrix between historical rainfall data and atmospheric circulation factors for each monitoring station; determining the likelihood function for each monitoring station based on its historical rainfall data and atmospheric circulation factors, and constructing a likelihood function matrix; and classifying each monitoring station by level using a dynamic Bayesian algorithm based on the spatiotemporal correlation matrix and the likelihood function matrix.
[0014] In conjunction with the first aspect mentioned above, in one possible implementation, the method of clustering historical rainfall data of each monitoring station to obtain flood season rainfall data and non-flood season rainfall data specifically includes: constructing a data matrix of historical rainfall data of each monitoring station according to time series; dividing the historical rainfall data in the data matrix into two clusters using a preset clustering algorithm; and determining whether the rainfall data at each time point is flood season rainfall data or non-flood season rainfall data based on the difference in the average rainfall of the two clusters.
[0015] In conjunction with the first aspect mentioned above, in one possible implementation, after obtaining historical rainfall data and atmospheric circulation factors, the method further includes: performing stationarity tests and denoising on the historical rainfall data; and normalizing the atmospheric circulation factors.
[0016] Secondly, a localized rainfall estimation device is provided, comprising: a processor and a storage medium; the storage medium includes instructions, and the processor is configured to execute the instructions to perform the actions described in the first aspect and any possible implementation thereof. This localized rainfall estimation device may be an electronic device or a chip within an electronic device.
[0017] Thirdly, a computer-readable storage medium is provided, in which instructions are stored, which, when executed on a local area intelligent rainfall estimation device, cause the local area intelligent rainfall estimation device to perform the actions described in the first aspect and any possible implementation thereof.
[0018] Fourthly, a computer program product containing instructions is provided, which, when run on a local area intelligent rainfall estimation device, causes the local area intelligent rainfall estimation device to perform the actions described in the first aspect and any possible implementation thereof.
[0019] The present invention has the following beneficial effects: By integrating historical rainfall data and atmospheric circulation factors from multiple monitoring stations in the target area, the fluctuation and periodic characteristics of rainfall data are accurately analyzed to obtain the complexity during the flood season and the complexity during the non-flood season. Based on this, the initial population range of the optimization algorithm is dynamically determined, avoiding the limitations of a fixed initialization mode. Then, combined with the atmospheric circulation factors of the period to be predicted, the optimal hyperparameters are obtained through optimization algorithm and the neural network model is trained, thereby achieving accurate estimation of rainfall in the local area during the period to be predicted. This effectively adapts to the complex variation characteristics of rainfall in the climate transition zone and improves the reliability and pertinence of rainfall estimation. Attached Figure Description
[0020] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A flowchart illustrating a method for intelligent estimation of rainfall in a local area, as provided in one embodiment of the present invention; Figure 2 A flowchart illustrating another method for intelligent estimation of rainfall in a local area, provided in one embodiment of the present invention; Figure 3 A flowchart illustrating another method for intelligent estimation of rainfall in a local area, provided in one embodiment of the present invention; Figure 4 This is a schematic diagram of the hardware structure of an intelligent rainfall estimation device for a local area, provided as an embodiment of the present invention. Detailed Implementation
[0022] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a localized intelligent rainfall estimation method proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0024] The following description, in conjunction with the accompanying drawings, details the specific scheme of the intelligent estimation method for local rainfall provided by the present invention.
[0025] Please see Figure 1 The diagram illustrates a flowchart of a method for intelligently estimating rainfall in a local area according to an embodiment of the present invention. This method includes: S1. Obtain atmospheric circulation factors and historical rainfall data from multiple monitoring stations in the target area.
[0026] In some implementation methods, monitoring stations covering different terrains and regions of the target area are selected to ensure that the data can comprehensively reflect the rainfall distribution characteristics of the target area. At the same time, atmospheric circulation factors corresponding to the time dimension of historical rainfall data are obtained from authoritative meteorological data release platforms (such as the official website of the National Atmospheric Administration Bureau) to provide comprehensive and matching basic data for subsequent rainfall characteristic analysis and complexity calculation.
[0027] In the transitional zone between warm temperate and subtropical climates, atmospheric circulation factors can include any one of the following parameters: westerly wind index, latitude and western extension point of the subtropical high-pressure ridge, polar vortex area index, southwest airflow intensity index, and sea surface temperature anomaly in the central and eastern equatorial Pacific. These atmospheric circulation factors can reflect the circulation conditions that influence rainfall formation in the target area from different dimensions. The westerly index characterizes the strength of atmospheric circulation in the westerly belt. Its numerical changes directly affect the frequency and intensity of cold air activity in the target area, thereby altering the spatiotemporal distribution characteristics of rainfall and providing key circulation background parameters for the analysis of rainfall fluctuation characteristics. The latitude of the subtropical high-pressure ridge is represented by specific latitude values, and the western extension point of the ridge is represented by specific longitude values. Both are normalized and weighted to form a comprehensive factor, with the weights determined based on the dominant factors influencing rainfall in the target region. Changes in the position and extent of the subtropical high-pressure system regulate the moisture transport path and convergence conditions in the target region. When the ridge latitude is further north and the western extension point is further west, it is easier to guide warm, moist air currents northward to converge with cold air, increasing the probability of rainfall. Therefore, the latitude and western extension point data of the subtropical high-pressure ridge can provide core evidence for the analysis of rainfall cycle characteristics (such as the start and end times of the flood season). As an important source of cold air in the Northern Hemisphere, the size of the polar vortex reflects the activity level of cold air. When the polar vortex area index is large, the frequency of southward influence of cold air increases, which will change the intensity and duration of rainfall in the target area and provide circulation driving parameters for the analysis of the random fluctuation of rainfall in complexity calculation. The southwest airflow is an important water vapor transport carrier in the target area. The southwest airflow intensity index directly determines the abundance of water vapor supply and is significantly positively correlated with rainfall. This factor provides a key quantitative indicator for the correlation analysis between atmospheric circulation factors and rainfall. Sea surface temperature anomalies affect global atmospheric circulation through air-sea interaction. El Niño (positive sea surface temperature anomaly) or La Niña (negative sea surface temperature anomaly) events can indirectly change the atmospheric circulation pattern of the target region, thereby affecting the total amount of rainfall during the flood season and non-flood season. Therefore, sea surface temperature anomalies in the central and eastern equatorial Pacific can provide a macroscopic circulation reference for complexity calculation and determination of the initialization range of optimization algorithms.
[0028] Furthermore, to ensure the accuracy and effectiveness of subsequent analyses, after obtaining historical rainfall data and atmospheric circulation factors, targeted preprocessing of both types of data is required: The stationarity of historical rainfall data is tested using methods such as time series plot analysis, correlation plot analysis, or d-order differencing to determine whether there are trend or periodic fluctuations in the data. Data that does not meet the stationarity requirements are adjusted to ensure that the historical rainfall data meets the core premise of time series analysis and avoids deviations in subsequent feature analysis and complexity calculations caused by non-stationary data.
[0029] To address the noise generated during the acquisition of historical rainfall data due to equipment interference such as sensor vibration, algorithms such as Gaussian filtering or bilateral filtering are used for noise reduction to filter out abnormal interference signals in the data and improve the purity and reliability of historical rainfall data.
[0030] The atmospheric circulation factors are processed using a linear minimum-maximum (min-max) normalization algorithm, which maps atmospheric circulation factors of different magnitudes and units to the interval [0, 1], eliminating the influence of differences in data magnitude and enabling atmospheric circulation factors and historical rainfall data to have a basis for collaborative analysis.
[0031] S2. Based on historical rainfall data and atmospheric circulation factors, analyze the fluctuation and periodic characteristics of rainfall data to obtain the complexity during the flood season and the complexity during the non-flood season.
[0032] In some implementations, a spatiotemporal correlation matrix and a likelihood function matrix are first constructed using preprocessed historical rainfall data and normalized atmospheric circulation factors. A dynamic Bayesian algorithm is then used to classify monitoring stations into different levels and assign corresponding weights. Simultaneously, a pre-defined clustering algorithm distinguishes between flood season and non-flood season based on the mean rainfall. Next, the complexity value of each individual station is calculated using rank correlation, random fluctuation, and temporal similarity. Then, by masking data, the first level of influence in the local dimension is obtained; and by decomposing and updating data components and combining them with station weights, the second level of influence in the global dimension is obtained. Subsequently, after normalizing the two levels of influence, the flood season complexity and non-flood season complexity of the target area are synthesized using pre-defined weights.
[0033] S3. Determine the initial population range of the optimization algorithm based on the complexity during the flood season, the complexity during the non-flood season, and the atmospheric circulation factor.
[0034] In some implementations, the first extreme value range during the flood season and the second extreme value range outside the flood season are statistically analyzed from historical rainfall data. The mean of all atmospheric circulation factors is calculated, and then the atmospheric circulation factors for the period to be predicted are compared with this mean. If the atmospheric circulation factors for the period to be predicted are greater than or equal to the mean, the first extreme value range is dynamically adjusted in conjunction with the complexity of the flood season to determine the initial population range of the optimization algorithm; if they are less than the mean, the second extreme value range is dynamically adjusted in conjunction with the complexity of the non-flood season to determine the initial population range.
[0035] S4. Based on the atmospheric circulation factor and the initial population range, use the optimization algorithm to optimize the hyperparameters of the neural network model, and train the neural network model according to the hyperparameter configuration obtained by optimization.
[0036] In some implementations, the predetermined initial population range is first used as one of the search space boundaries of the optimization algorithm (such as particle swarm optimization or genetic algorithm). This range dynamically adapts to the circulation characteristics (flood season or non-flood season) and the complexity of rainfall data during the period to be predicted, ensuring the rationality and diversity of the starting point for hyperparameter optimization, and effectively avoiding the problem of traditional fixed initialization mode easily getting trapped in local optima.
[0037] Next, the optimization objective of the optimization algorithm is set as minimizing the prediction error of the neural network model on the training set (for example, using mean squared error or mean absolute error as the fitness function). Preprocessed historical rainfall data is used as training labels, and the normalized atmospheric circulation factor of the corresponding time series is used as training features, together forming the training sample set of the neural network.
[0038] Subsequently, an optimization algorithm is initiated for iterative optimization. In each iteration, the algorithm generates a set of candidate solutions for neural network hyperparameters (including but not limited to: learning rate, number of hidden layer nodes, regularization coefficient, batch size, etc.) within the initialized population range and other preset spaces. This set of hyperparameters is used to temporarily configure the neural network model (e.g., BP neural network, LSTM network), and a rapid training and validation round is performed using the entire training sample set. The error between the model's predicted values and the actual historical rainfall is calculated and used as the fitness value of this set of candidate hyperparameter solutions.
[0039] The optimization algorithm, guided by minimizing prediction error, continuously updates the hyperparameter population through multiple iterations (e.g., 50-100 rounds), ultimately selecting a set of hyperparameter combinations that optimize the model's prediction performance. This process utilizes a dynamically adapted initial population range, allowing hyperparameter optimization to focus more on configuration ranges that fit the current circulation background and rainfall complexity characteristics, thereby improving the efficiency of model training and final performance.
[0040] Finally, the optimal hyperparameter combination obtained through optimization is used to formally configure the neural network model. The complete training sample set is input into the model for thorough training (e.g., 100-500 iterations) until the model loss converges to a preset accuracy threshold. This yields a neural network model with optimized hyperparameters that has fully learned the complex nonlinear mapping relationship between historical rainfall and atmospheric circulation factors.
[0041] S5. Input the atmospheric circulation factor of the period to be predicted into the trained neural network model to obtain the predicted rainfall value of the target area during the period to be predicted.
[0042] In some implementations, the normalized atmospheric circulation factor for the period to be predicted is input into a trained neural network model. Based on the learned correlation patterns, the model performs forward propagation calculations on the input circulation features and finally outputs the predicted rainfall value for the target area during the period to be predicted. This enables intelligent and accurate estimation of rainfall in local areas, providing reliable forecast support for responding to disasters such as floods and droughts.
[0043] Atmospheric circulation factors (such as the position of the subtropical high and the westerly index) are parameters of large-scale weather systems with more stable variation patterns. Existing numerical weather prediction models (such as the Global Forecast System (GFS)) can accurately predict them 7-10 days in advance. Therefore, the atmospheric circulation factors for the forecast period are not measured values, but rather predicted in advance by mature meteorological models. In contrast, rainfall is greatly affected by local small-scale factors (such as valley winds and sudden convection), exhibiting strong randomness and cannot be directly obtained through simple prediction.
[0044] Based on the above technical solution, by integrating historical rainfall data and atmospheric circulation factors from multiple monitoring stations in the target area, the fluctuation and periodic characteristics of rainfall data are accurately analyzed to obtain the complexity during the flood season and the complexity during the non-flood season. Based on this, the initial population range of the optimization algorithm is dynamically determined, avoiding the limitations of a fixed initialization mode. Then, combined with the atmospheric circulation factors of the period to be predicted, the optimal hyperparameters are obtained through optimization algorithm and the neural network model is trained, thereby achieving accurate estimation of rainfall in the local area during the period to be predicted. This effectively adapts to the complex variation characteristics of rainfall in the climate transition zone and improves the reliability and pertinence of rainfall estimation.
[0045] In one possible implementation, combining Figure 1 ,like Figure 2 As shown, the method in S2 described above can be specifically implemented through the following steps S21 to S24, which are explained in detail below: S21. Analyze the correlation between historical rainfall data and atmospheric circulation factors at monitoring stations in the target area, and classify multiple monitoring stations into different levels.
[0046] In some implementations, a spatiotemporal correlation matrix is constructed for historical rainfall data and atmospheric circulation factors at each monitoring station. This matrix uses the monitoring station as the core dimension and integrates correlation data from the time dimension (layered by year and month). Each element in the matrix corresponds to the combination information of historical rainfall data and corresponding atmospheric circulation factors at a specific time point for a single monitoring station. This achieves a structured presentation of the spatiotemporal matching relationship between station data and circulation factors, providing a clear data carrier for subsequent quantitative analysis of correlation and ensuring that the analysis process is not divorced from the spatiotemporal context.
[0047] For example, taking 5 monitoring stations as an example, the rainfall data R (i.e., the average rainfall in the kth month) of each monitoring station is combined with the atmospheric circulation factor C (i.e., the average atmospheric circulation factor in the kth month) into a tuple (R, C). Then, taking each monitoring station as the channel, the monthly rainfall data of each monitoring station in each year is used as the rows of the matrix, and the rainfall data of the same month of each monitoring station in the 5 years is used as the columns, to construct a 5×5×12 spatiotemporal correlation matrix. Each element stores the tuple (R, C) of the corresponding monitoring station, corresponding year, and corresponding month.
[0048] Secondly, the likelihood function for each monitoring station is determined based on historical rainfall data and atmospheric circulation factors. Referring to the example above, the likelihood function for the i-th monitoring station in year t is... It can be represented as: In the formula, and Let represent the rainfall data and atmospheric circulation factor of the i-th monitoring station in the k-th month of year t, respectively. and Let represent the rainfall data and atmospheric circulation factor of the j-th monitoring station in the k-th month of the t-th year, respectively.
[0049] The molecular weighting of the synergistic contribution of rainfall and atmospheric circulation factors within the 12 months of year t at the i-th monitoring station reflects the correlation strength between the monitoring station's own data and circulation factors.
[0050] The denominator quantifies the total collaborative contribution of all five stations over the 12 months of year t, serving as a global benchmark for correlation strength. In real-world scenarios, there is no rainfall data for each month within five years, nor is there atmospheric circulation factor that is zero (extreme cases of five years without rainfall are not applicable for rainfall prediction), therefore the above formula does not need to consider the case where the denominator is zero.
[0051] Dividing the two values yields the likelihood function for the i-th monitoring station in year t, which reflects the proportion of the correlation between its data and atmospheric circulation factors relative to the global correlation. The larger the value, the higher the correlation.
[0052] Subsequently, the likelihood function values of all monitoring stations in each year were arranged according to the dimension of monitoring station and time to construct a likelihood function matrix. This matrix transforms the degree of correlation between the stations into a directly comparable quantitative indicator, providing accurate numerical support for the classification and avoiding bias caused by subjective judgment.
[0053] Based on the above example, we construct a 5×5 likelihood function matrix by using the likelihood function of each monitoring station over 5 years as rows and the likelihood functions of the 5 monitoring stations in the same year as columns.
[0054] Finally, based on the spatiotemporal correlation matrix and the likelihood function matrix, a dynamic Bayesian algorithm is used to classify each monitoring station into different levels. The dynamic Bayesian algorithm integrates spatiotemporal matching information from the spatiotemporal correlation matrix and quantitative correlation indicators from the likelihood function matrix, while considering both data correlation over time and differences between stations, to evaluate and rank the comprehensive reference value of each monitoring station. The algorithm outputs the level result for each monitoring station (e.g., level 1-5, where higher levels indicate a stronger correlation between the station's data and atmospheric circulation factors, and higher reference value), thus achieving differentiated distinction between different monitoring stations.
[0055] S22. Cluster the historical rainfall data of each monitoring station to obtain rainfall data during the flood season and rainfall data during the non-flood season.
[0056] In some implementations, the historical rainfall data of each monitoring station is first used to construct a data matrix based on time series. This matrix uses time points as the row dimension and rainfall values as the column dimension. Each element precisely corresponds to the historical rainfall data of a single monitoring station at a specific time point, achieving structured integration of scattered rainfall data. This preserves the continuity of the time dimension and makes the data more regular, providing a well-organized and computable data foundation for subsequent cluster analysis and avoiding clustering bias caused by disordered data.
[0057] For example, the monthly rainfall data collected by each monitoring station over 5 years can be used as the rows of a matrix, and the rainfall data of the same month from the 5 monitoring stations over 5 years can be used as the columns of the matrix, and then the data can be combined into a 5×60 two-dimensional matrix.
[0058] Secondly, a pre-defined clustering algorithm divides the historical rainfall data in the data matrix into two clusters. The algorithm calculates the similarity (e.g., Euclidean distance) between rainfall data at different time points in the data matrix, aggregating data with similar rainfall values and convergent temporal distribution characteristics into the same cluster, thus achieving a preliminary classification of the historical rainfall data. The clustering process is based on the inherent distribution patterns of the data, ensuring that the classification results better reflect the actual rainfall characteristics of the target region, thereby improving the objectivity and accuracy of the classification.
[0059] Finally, based on the difference in the mean rainfall between the two clusters, the median of the two means was used as a classification threshold. All rainfall data in clusters with rainfall exceeding this median were labeled as flood season rainfall data, corresponding to periods of high and concentrated rainfall. Conversely, all rainfall data in clusters with rainfall less than or equal to this median were labeled as non-flood season rainfall data, corresponding to periods of low and dispersed rainfall. This method of using the mean difference accurately defines the two types of data and clearly delineates the cyclical characteristics of rainfall.
[0060] S23. Determine the complexity value of each monitoring station based on historical rainfall data and atmospheric circulation factors.
[0061] In some implementations, the first step is to determine the weight of each monitoring station based on its classification. A higher station classification indicates a stronger correlation between its historical rainfall data and atmospheric circulation factors, resulting in higher reference value and a larger corresponding weight. This differentiated weight allocation allows data from high-reliability monitoring stations to contribute more significantly to complexity calculations, ensuring that the complexity value accurately reflects the rainfall data quality characteristics of the target station.
[0062] Secondly, the first parameter is obtained based on the rank correlation between historical rainfall data and atmospheric circulation factors at each monitoring station. Specifically, using preprocessed historical rainfall data and normalized atmospheric circulation factors as input, a preset rank correlation analysis algorithm (such as Spearman's rank correlation algorithm or Kendall's rank correlation algorithm) is used to calculate the correlation between the rainfall data sequence and the corresponding atmospheric circulation factor sequence at each monitoring station, thus obtaining the first parameter. The first parameter quantifies the degree of monotonic correlation between the two types of sequences; the larger the absolute value of the first parameter, the stronger the correlation between the rainfall data and atmospheric circulation factors at the monitoring station.
[0063] Next, based on the degree of random fluctuation in historical rainfall data for each monitoring station, the second parameter is obtained. Specifically, for the historical rainfall data sequence of each monitoring station, a preset permutation entropy algorithm is used to analyze the degree of random fluctuation in the sequence: setting the embedding dimension m (with values [3, 7]) and the time delay parameter τ (empirical value of 1), the one-dimensional time series data is transformed into multiple high-dimensional phase space vectors, each vector consisting of m data points selected every τ points in the original sequence. Then, each phase space vector is sorted by element size to generate a unique permutation index sequence (i.e., permutation pattern). Subsequently, the occurrence frequency of each permutation pattern in all vectors is counted, and its probability of occurrence is calculated. Then, according to the information entropy formula, the probability of each pattern is summed with the negative value of the corresponding logarithmic product with base 2 to calculate the original permutation entropy value, which is then divided by the theoretical maximum value. After normalization, the second parameter is obtained. The second parameter intuitively represents the random fluctuation characteristics of the rainfall data. The larger the value of the second parameter, the more drastic the fluctuation and the higher the randomness of the rainfall data at the monitoring station in the time dimension.
[0064] Then, the third parameter is obtained based on the sequence similarity of historical rainfall data from each monitoring station across adjacent periods. Specifically, the historical rainfall data from each monitoring station is divided into multiple independent time-series subsequences according to the principles of temporal continuity, consistent length, and period alignment, ensuring that the similarity of adjacent subsequences effectively reflects the differences in rainfall trend changes. Annual periods are preferred (as they best reflect the interannual fluctuation characteristics of rainfall). If the data volume is insufficient (e.g., less than 3 years), seasonal periods can be selected. A preset sequence similarity analysis algorithm (such as cosine similarity) is used to calculate the similarity between adjacent time-series subsequences one by one, and the average of all adjacent sequence similarities is taken as the third parameter. The third parameter reflects the consistency of the periodic trend of rainfall data from the monitoring stations; the smaller the value of the third parameter, the greater the difference in rainfall trend between adjacent periods.
[0065] Finally, weight coefficients are set for the three parameters (the sum of the three coefficients is 1), and a weighted fusion calculation is performed on the first, second, and third parameters. The complexity value of the i-th monitoring station is then calculated. It can be represented as: In the formula, The first parameter represents the rank correlation between the rainfall data sequence and the atmospheric circulation factor sequence at the i-th monitoring station. The larger the absolute value, the stronger the correlation (whether positive or negative) between the rainfall data and the atmospheric circulation factor sequence, and the greater the influence of the atmospheric circulation factor on the rainfall data of the i-th monitoring station, and the lower the complexity. This reflects the strength of the correlation between rainfall data and atmospheric circulation factors. The weaker the correlation, the larger the value of this item and the higher the complexity.
[0066] This represents the normalized permutation entropy of the rainfall data sequence at the i-th monitoring station. The larger the value, the greater the randomness and complexity of the rainfall data at the i-th monitoring station.
[0067] This represents the mean similarity between rainfall data sequences of all adjacent periods within the i-th monitoring station, with a value range of [0, 1] (if bounded indices such as cosine similarity are used, they need to be normalized to the [0, 1] interval through linear transformation). The larger the value, the more similar the changing trends of rainfall data in adjacent periods at the i-th monitoring station, the smaller the difference in randomness, and the lower the complexity. It reflects the difference in the randomness of rainfall between adjacent cycles; the greater the difference, the higher the complexity.
[0068] , , This represents the weighting coefficient, used to adjust the contribution of each parameter, and the sum of these coefficients is 1.
[0069] By combining the complex quantitative results from three dimensions, the comprehensive and complex characteristics of rainfall data from a single monitoring station are fully characterized. The larger the value, the more significant the randomness and overall complexity of the rainfall data at the i-th monitoring station.
[0070] Specifically, the weighting coefficients can be determined through empirical setting, historical data backtesting, and regional scenario adaptation: First, initial basic weights are set based on the importance of three dimensions: domain experience, factor correlation, random fluctuation, and annual differences; then, historical data is selected, and backtesting is performed to optimize the weighting from the perspectives of the correlation between complexity and extreme rainfall, station level consistency, and cross-time stability; finally, scenario adaptation is performed based on the terrain of the target area (such as complex or flat terrain) and the dominant atmospheric circulation type (such as large-scale circulation or local micro-circulation), ultimately determining a weighting combination that both conforms to the theoretical importance of dimensions and adapts to the actual data characteristics and regional scenario. For example, taking a plain area in the transition zone between the warm temperate and subtropical zones as an example, the initial weights... =0.3、 =0.4、 =0.3. Historical data backtesting revealed that rainfall during the flood season in this region is greatly influenced by the position of the subtropical high-pressure ridge (atmospheric circulation factor), and extreme rainfall events are related to... The correlation in =0.35、 =0.35、 The value reaches its maximum at 0.3, and the final weight of this region is determined as follows: =0.35、 =0.35、 =0.3.
[0071] S24. Based on the complexity value and weight of each monitoring station's level, and combined with the clustering results of flood season rainfall data and non-flood season rainfall data, determine the flood season complexity and non-flood season complexity of the target area.
[0072] In some implementations, the method described in S24 above can be specifically implemented using the following S241 to S243, which are explained in detail below: S241. For the historical rainfall data of each monitoring station, analyze the degree of influence of each rainfall data on the complexity value of the monitoring station. Combine the clustering results to obtain the complexity component of the first flood season and the complexity component of the first non-flood season.
[0073] First, for each monitoring station, based on its corresponding complexity value, individual data points in the historical rainfall data of that station are individually masked. For each masking operation, the station's complexity value is recalculated based on the remaining historical rainfall data, the corresponding atmospheric circulation factor, and the station's weight; this value is defined as the post-masking complexity value. By comparing the monitoring station's complexity value with the post-masking complexity value, the degree of influence of each rainfall data point on the complexity value is quantified, i.e., the first degree of influence. The first degree of influence of the c-th rainfall data point within the i-th monitoring station is denoted as: ,in, This represents the complexity value of the i-th monitoring station. This represents the complexity value after masking at the i-th monitoring station, excluding the c-th rainfall data point, and recalculating the masked complexity.
[0074] Subsequently, all primary impact levels corresponding to the flood season rainfall data of each monitoring station were selected, and their average values were calculated to obtain the mean primary impact level during the flood season. Similarly, all primary impact levels corresponding to the non-flood season rainfall data were selected, and their average values were calculated to obtain the mean primary impact level during the non-flood season.
[0075] Finally, the min-max normalization algorithm was used to normalize the mean of the first impact degree during the flood season and the mean of the first impact degree outside the flood season for all monitoring stations, mapping the values uniformly to the interval [0, 1]. This operation eliminated the differences in data magnitude between different monitoring stations, making the mean impact degree of each station at different time periods horizontally comparable, and obtaining the complexity component of the first flood season. and the complexity component of the first non-flood season .
[0076] S242. Based on the historical rainfall data and the weights of the levels of multiple monitoring stations, analyze the impact of each rainfall data point on the global information entropy. Combine the clustering results to obtain the second flood season complexity component and the second non-flood season complexity component.
[0077] First, for the historical rainfall data of each monitoring station, a preset multi-scale decomposition algorithm is used for hierarchical processing. The decomposition basis function is set to four-level wavelet decomposition (daubechies 4, db4), and the number of decomposition levels is set (e.g., 3). Each rainfall data point is decomposed into one approximate component and multiple detail components. The approximate component is used to characterize the overall trend of rainfall data, while the detail components are used to characterize the local fluctuations in rainfall data. This decomposition enables the accurate extraction of features at different scales from the rainfall data.
[0078] Next, to comprehensively analyze the information structure of precipitation data in the global context, an enhanced global data matrix integrating multi-scale precipitation characteristics and circulation driving factors is constructed. Specifically, for each precipitation data point at each monitoring station, all components obtained from its multi-scale decomposition (one approximate component and multiple detail components) are concatenated with the corresponding normalized atmospheric circulation factor to form an enhanced feature vector. The enhanced feature vectors of all monitoring stations and all precipitation data points are then arranged row-wise to construct the global data matrix. Each row of this matrix represents the joint characteristics of a data point in both the multi-scale space and the circulation factor space.
[0079] Subsequently, a pre-defined dimensionality reduction algorithm (such as principal component analysis (PCA)) is used to process the enhanced global data matrix, extract the core feature values of the matrix, and obtain the global information entropy by calculating the sum of the Shannon entropy of the probability distribution of the principal component scores in each dimension, which is used to quantify the overall complexity and uncertainty of the fused feature space.
[0080] Based on the global information entropy and the weight of each monitoring station, the contribution of each rainfall data point to the feature fusion of the global information structure is calculated, i.e., the second degree of influence. One feasible calculation method is to assess the change in global information entropy after removing the data point: In the formula, This represents the contribution of the c-th rainfall data point to the feature fusion of the global information structure. Represents global information entropy. This represents the global information entropy recalculated after removing the row corresponding to the c-th rainfall data point from the enhanced global data matrix.
[0081] The larger the value, the more significant the impact of the data point on the complexity of global feature fusion.
[0082] Finally, the mean of the second degree of influence of rainfall data during the flood season and the mean of the second degree of influence of rainfall data outside the flood season are calculated for each monitoring station. Then, the min-max normalization algorithm is used to normalize the two types of means for all stations to eliminate differences between stations and obtain the second flood season complexity component. Second non-flood season complexity component .
[0083] S243. Obtain the flood season complexity based on the first flood season complexity component and the second flood season complexity component; obtain the non-flood season complexity based on the first non-flood season complexity component and the second non-flood season complexity component.
[0084] First, weight coefficients are set for the first and second components (the sum of the two coefficients is 1). These weight coefficients are determined based on the rainfall characteristics and data reliability requirements of the target area, and are used to balance the contribution of local station influences to the overall data correlation. For example, for scenarios with uniform spatial distribution of rainfall and strong representativeness of monitoring station data, such as plains, the weight of the first component can be set to 0.6 and the weight of the second component to 0.4; for scenarios with large spatial differences in rainfall and strong global correlation, such as hilly areas, the weight of the first component can be set to 0.4 and the weight of the second component to 0.6. Further optimization can be achieved through backtesting of historical data (such as comparing the complexity under different weights with the correlation of actual extreme rainfall events) to select the weight combination most suitable for the target area.
[0085] For flood season scenarios, the complexity components of the first flood season are... Complexity component of the second flood season The flood season complexity of the target region is obtained by performing a weighted fusion calculation based on preset weights (i.e., weighted multiplication followed by summation). Similarly, for non-flood season scenarios, the complexity components of the first non-flood season are... Complexity component of the second non-flood season The non-flood season complexity of the target region is obtained by performing weighted fusion calculations with equal weights. .
[0086] By weighted fusion of two-dimensional components, the final complexity value fully reflects the impact of data within a single site on local complexity, while also taking into account the role of global data correlation features on overall complexity. This comprehensively and accurately characterizes the combined complexity of rainfall data during the flood season and non-flood season in the target area, providing a scientific and reliable quantitative basis for the dynamic determination of the population range for subsequent optimization algorithms.
[0087] In one possible implementation, combining Figure 2 ,like Figure 3 As shown, the method in S3 above can be specifically implemented through the following steps S31 to S33, which are explained in detail below: S31. Calculate the global statistical characteristics of historical rainfall data and calculate the average value of atmospheric circulation factors.
[0088] In some implementations, the first step is to extract global statistical features that reflect the overall distribution and fluctuation characteristics of the preprocessed historical rainfall data. These features include at least the global maximum, global minimum, global mean, and global standard deviation of the historical rainfall data.
[0089] Secondly, the mean of atmospheric circulation factors corresponding to all historical rainfall data was calculated. As a reference benchmark for judging the tendency of circulation conditions during the period to be predicted, it is compared with the atmospheric circulation factors during the period to be predicted. By comparing the data, we can initially identify the type of rainfall period corresponding to the period to be predicted (whether it is more likely to be the flood season or the non-flood season), providing a directional basis for the selection of the target area.
[0090] S32. Determine the dynamic adjustment coefficient based on the atmospheric circulation factor and the mean value of the atmospheric circulation factor for the period to be predicted.
[0091] In some implementations, if the atmospheric circulation factor for the period to be predicted is greater than or equal to the mean atmospheric circulation factor, it indicates that the circulation conditions at this time are closer to the circulation characteristics of historical flood seasons; if the atmospheric circulation factor for the period to be predicted is less than the mean atmospheric circulation factor, it indicates that the circulation conditions are closer to the circulation characteristics of historical non-flood seasons. Specifically, this means: Regarding the westerly index, the mid-latitude westerlies during the flood season are typically of moderate intensity and fluctuate actively (index values tend to be high), which facilitates the southward movement of cold air and their convergence with warm, moist air currents, leading to rainfall. If the westerly index for the forecast period is greater than or equal to the historical average, it indicates that the intensity or fluctuation characteristics of the westerlies are more consistent with the typical value range of historical flood seasons, and the circulation conditions are conducive to the formation and maintenance of rainfall systems, more closely resembling the circulation characteristics of the flood season. Conversely, if the westerly index for the forecast period is less than the historical average, it is closer to the circulation characteristics of the non-flood season.
[0092] Regarding the latitude and westward extension of the subtropical high-pressure ridge, during the flood season, the ridge shifts northward (to a higher latitude) and the westward extension extends westward (to a greater longitude), guiding warm and humid air currents northward, which converge with cold air to form a large-scale rain belt. If the latitude and westward extension of the subtropical high-pressure ridge are both greater than or equal to the historical averages during the forecast period, it indicates that the north-south position and westward extension of the subtropical high are more consistent with the typical distribution during historical flood seasons, and the circulation conditions are conducive to the formation of the rain belt during the flood season, more closely resembling the circulation characteristics of the flood season; conversely, if the latitude and westward extension of the subtropical high-pressure ridge are less than the historical averages during the forecast period, it is closer to the circulation characteristics of the non-flood season.
[0093] Regarding the polar vortex area index, during the flood season, polar vortices typically split and affect the mid-latitudes (resulting in a higher area index), which facilitates the southward movement of cold air and interaction with warm, moist air currents, triggering rainfall. If the polar vortex area index for the forecast period is greater than or equal to the historical average, it indicates that the range of influence of the polar vortex is more consistent with the typical value range of historical flood seasons, and the intensity and extent of cold air activity are conducive to the development of rainfall systems, more closely resembling the circulation characteristics of the flood season. Conversely, if the polar vortex area index for the forecast period is less than the historical average, it is closer to the circulation characteristics of the non-flood season.
[0094] Regarding the southwest airflow intensity index, the southwest airflow intensity is relatively high during the flood season, serving as a key channel for the transport of moisture from the South China Sea and the Indian Ocean to the land, providing ample moisture for rainfall. If the southwest airflow intensity index for the forecast period is greater than or equal to the historical average, it indicates that the moisture transport intensity is more consistent with the typical range of historical flood season values, and the moisture conditions are conducive to the formation of heavy rainfall, more closely resembling the circulation characteristics of the flood season. Conversely, if the southwest airflow intensity index for the forecast period is less than the historical average, it is closer to the circulation characteristics of the non-flood season.
[0095] Regarding sea surface temperature anomalies in the central and eastern equatorial Pacific, during El Niño years (a global climate phenomenon caused by abnormally high sea surface temperatures in the central and eastern equatorial Pacific), the western Pacific subtropical high is often stronger and abnormally positioned. Rainfall during the flood season tends to exhibit extreme distributions, with more precipitation in the south and less in the north, or vice versa, closely correlated with typical flood season circulation responses. If the sea surface temperature anomaly in the central and eastern equatorial Pacific during the forecast period is greater than or equal to the historical average (i.e., leaning towards El Niño-type sea surface temperature characteristics), it indicates that the type of sea surface temperature anomaly is more consistent with the typical circulation response of historical flood seasons, indirectly reflecting that atmospheric circulation conditions are closer to flood season characteristics. Conversely, if the sea surface temperature anomaly in the central and eastern equatorial Pacific during the forecast period is less than the historical average, it is closer to non-flood season circulation characteristics.
[0096] Based on this, the atmospheric circulation factors for the period to be predicted can be compared with historical averages to determine which precipitation pattern the current forecasting task favors, and the corresponding complexity value can be selected to calculate the dynamic adjustment coefficient accordingly. : In the formula, and These represent the complexity during the flood season and the non-flood season, respectively. This represents the atmospheric circulation factors for the period to be predicted. This represents the mean of the atmospheric circulation factor corresponding to all rainfall data.
[0097] S33. Based on the dynamic adjustment coefficient and rainfall statistical characteristics, map the initial population range of the hyperparameter space.
[0098] In some implementations, the goal of the optimization algorithm (such as particle swarm optimization) is to find the optimal combination of neural network hyperparameters (including learning rate, number of hidden layer neurons, etc.). Each hyperparameter has its theoretical or empirical feasible region; for example, the feasible region of the learning rate can be
[10] . -5 10 -1 The feasible region for the number of neurons can be [1, 100]. However, this method does not uniformly apply a fixed feasible region across all hyperparameters. Instead, it dynamically determines the initial search range for key hyperparameters that significantly impact model performance based on rainfall characteristics and complexity. The specific method is as follows: First, select one or more continuous hyperparameters that are sensitive to model capacity and fitting ability as the objects of dynamic adjustment, and set a relatively wide empirical baseline range for each hyperparameter. (i.e., feasible region). Then, use the dynamic adjustment coefficient. and the coefficient of variation of rainfall The rainfall uncertainty reflected by the ratio of the global standard deviation to the global mean (representing the relative fluctuation of rainfall) narrows the baseline range of this hyperparameter to a more focused initial population range. , represented as: In the formula, and These represent the empirical lower and upper bounds of the hyperparameters, respectively.
[0099] Calculate half the width of the baseline range, which represents the maximum possible shrinkage.
[0100] This represents the amount of contraction, specifically the distance that the price can be simultaneously reduced from both ends of the reference range inwards. The maximum possible contraction is used as the upper limit for adjustment, utilizing a dynamic adjustment coefficient. (Characterizing the stochastic volatility and forecasting difficulty of rainfall patterns) and the coefficient of variation of rainfall. (Characterizing the relative fluctuation of historical rainfall data) is used for comprehensive adjustment.
[0101] Through the above methods, the initial population range for the key hyperparameters of the optimization algorithm is finally obtained. This range matches the dimensions of the hyperparameters themselves and logically inherits and utilizes the spatiotemporal complexity characteristics of rainfall data, realizing an intelligent mapping from data features to model optimization configuration.
[0102] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0103] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0104] In this embodiment of the invention, the intelligent rainfall estimation device for a local area can be divided into functional units according to the above method example. For example, each function can be divided into its own functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0105] This invention also provides a schematic diagram of the hardware structure of a local rainfall intelligent estimation device, see [link / reference]. Figure 4 The local rainfall intelligent estimation device 400 includes a processor 401, and optionally, a memory 402 connected to the processor 401.
[0106] In the first possible implementation, see Figure 4 The intelligent rainfall estimation device 400 for localized areas also includes a transceiver 403. The processor 401, memory 402, and transceiver 403 are connected via a bus. The transceiver 403 is used to communicate with other devices or communication networks. Optionally, the transceiver 403 may include a transmitter and a receiver. The device in the transceiver 403 that implements the receiving function can be considered as a receiver, which is used to perform the receiving steps in the embodiments of the present invention. The device in the transceiver 403 that implements the transmitting function can be considered as a transmitter, which is used to perform the transmitting steps in the embodiments of the present invention.
[0107] Based on the first possible implementation method Figure 4 The schematic diagram shown can be used to illustrate the structure of the intelligent rainfall estimation device for local areas involved in the above embodiments.
[0108] in, Figure 4 This can also be illustrated by the system chip in the intelligent rainfall estimation device for a localized area. In this case, the actions performed by the aforementioned intelligent rainfall estimation device for a localized area can be implemented by this system chip. The specific actions performed can be found above and will not be repeated here.
[0109] In implementation, each step of the method provided in this embodiment can be completed by integrated logic circuits in the processor or by instructions in software form. The steps of the method disclosed in this embodiment can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor.
[0110] The processor in this invention may include, but is not limited to, at least one of the following: a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a microcontroller unit (MCU), or an artificial intelligence processor, etc., which are various computing devices that run software. Each computing device may include one or more cores for executing software instructions to perform calculations or processing. The processor may be a standalone semiconductor chip or integrated with other circuits into a single semiconductor chip. For example, it may be integrated with other circuits (such as encoding / decoding circuits, hardware acceleration circuits, or various bus and interface circuits) to form a System-on-a-Chip (SoC), or it may be integrated as a built-in processor within an ASIC. The ASIC with the integrated processor may be packaged separately or together with other circuits. In addition to the cores for executing software instructions to perform calculations or processing, the processor may further include necessary hardware accelerators, such as field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), or logic circuits that implement dedicated logic operations.
[0111] The memory in the embodiments of the present invention may include at least one of the following types: read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions; random access memory (RAM) or other types of dynamic storage devices capable of storing information and instructions; or electrically erasable programmable read-only memory (EEPROM). In some scenarios, the memory may also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.
[0112] This invention also provides a computer-readable storage medium including instructions that, when run on a computer, cause the computer to perform any of the methods described above.
[0113] This invention also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform any of the methods described above.
[0114] This invention also provides a chip, which includes a processor and an interface circuit. The interface circuit is coupled to the processor. The processor is used to run computer programs or instructions to implement the above-described method. The interface circuit is used to communicate with other modules outside the chip.
[0115] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).
[0116] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings and the disclosure, will understand and implement other variations of the disclosed embodiments in carrying out the claimed invention. In this invention, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several of the functions listed in this invention.
[0117] Although the invention has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made therein without departing from the spirit and scope of the invention. Accordingly, this specification and drawings are merely illustrative of the invention and are to be considered as covering any and all modifications, variations, combinations, or equivalents within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if such modifications and modifications of the invention fall within the scope of the invention and its equivalents, the invention is also intended to include such modifications and modifications.
Claims
1. A method for intelligently estimating rainfall in a local area, characterized in that, include: Acquire atmospheric circulation factors and historical rainfall data from multiple monitoring stations in the target area; Based on the analysis of historical rainfall data and atmospheric circulation factors, the fluctuation and periodic characteristics of rainfall data are obtained, and the complexity during the flood season and the complexity during the non-flood season are obtained. The initial population range of the optimization algorithm is determined based on the flood season complexity, non-flood season complexity, and atmospheric circulation factor. Based on the atmospheric circulation factor and the initial population range, the hyperparameters of the neural network model are optimized using an optimization algorithm, and the neural network model is trained based on the hyperparameters obtained from the optimization. By inputting the atmospheric circulation factor for the period to be predicted into the trained neural network model, the predicted rainfall value for the target area during the period to be predicted is obtained.
2. The intelligent rainfall estimation method according to claim 1, characterized in that, Based on the analysis of historical rainfall data and atmospheric circulation factors, the fluctuation and periodic characteristics of rainfall data are determined, resulting in the complexity during the flood season and the complexity outside the flood season, including: The correlation between historical rainfall data and atmospheric circulation factors at monitoring stations in the target area was analyzed, and multiple monitoring stations were classified into different levels. Clustering the historical rainfall data of each monitoring station yields rainfall data during the flood season and rainfall data during the non-flood season. The complexity value for each monitoring station is determined based on historical rainfall data and atmospheric circulation factors. Based on the complexity value and weight of each monitoring station's level, and combined with the clustering results of flood season rainfall data and non-flood season rainfall data, the flood season complexity and non-flood season complexity of the target area are determined.
3. The intelligent rainfall estimation method according to claim 2, characterized in that, Based on historical rainfall data and atmospheric circulation factors for each monitoring station, the complexity value for each monitoring station is determined, including: The first parameter is obtained based on the rank correlation between historical rainfall data and atmospheric circulation factors at each monitoring station; The second parameter is obtained based on the degree of random fluctuation in historical rainfall data for each monitoring station; The third parameter is obtained based on the sequence similarity of historical rainfall data from each monitoring station in adjacent periods; Based on the first parameter, the second parameter, and the third parameter, the complexity value of each monitoring station is determined.
4. The intelligent rainfall estimation method according to claim 3, characterized in that, Based on the complexity value and weight of each monitoring station's level, and combined with the clustering results of flood season and non-flood season rainfall data, the flood season complexity and non-flood season complexity of the target area are determined, including: For each monitoring station's historical rainfall data, the impact of each rainfall data point on the complexity value of its respective monitoring station is analyzed. Combined with the clustering results, the complexity components for the first flood season and the first non-flood season are obtained. Based on the historical rainfall data and the weights of the different levels at multiple monitoring stations, the influence of each rainfall data point on the global information entropy is analyzed. Combined with the clustering results, the complexity components of the second flood season and the second non-flood season are obtained. The flood season complexity is obtained based on the first flood season complexity component and the second flood season complexity component; The non-flood season complexity is obtained based on the first non-flood season complexity component and the second non-flood season complexity component.
5. The intelligent rainfall estimation method according to claim 4, characterized in that, For each monitoring station's historical rainfall data, the impact of each rainfall data point on the complexity value of its respective monitoring station is analyzed. Combined with clustering results, the complexity components for the first flood season and the first non-flood season are obtained, including: For each monitoring station, the complexity value of the monitoring station is compared with the complexity value of each rainfall data after masking to determine the first degree of influence of each rainfall data on the complexity value; the complexity value after masking is the complexity value re-determined after masking a rainfall data. The average first degree of influence of rainfall data during the flood season and the average first degree of influence of rainfall data during the non-flood season are statistically analyzed for each monitoring station. The average first degree of influence is then normalized to obtain the first flood season complexity component and the first non-flood season complexity component.
6. The intelligent rainfall estimation method according to claim 4, characterized in that, Based on historical rainfall data from multiple monitoring stations and their respective weights, the impact of each rainfall data point on the global information entropy is analyzed. Combined with clustering results, the complexity components for the second flood season and the second non-flood season are obtained, including: The historical rainfall data of each monitoring station is decomposed into approximate and detailed components to obtain the approximate components and detailed components of each rainfall data, and then combined with atmospheric circulation factors to construct an enhanced feature vector. Based on the enhanced feature vectors corresponding to historical rainfall data from multiple monitoring stations, the global information entropy of the historical rainfall data is determined, and the degree of secondary influence of each rainfall data on the global information entropy is analyzed in conjunction with the weight of the monitoring station's level. The mean of the second degree of influence of rainfall data during the flood season and the mean of the second degree of influence of rainfall data during the non-flood season are statistically analyzed for each monitoring station. The mean of the second degree of influence is then normalized to obtain the second flood season complexity component and the second non-flood season complexity component.
7. The intelligent rainfall estimation method according to claim 2, characterized in that, The initial population range for the optimization algorithm is determined based on the aforementioned flood season complexity, non-flood season complexity, and atmospheric circulation factors, including: The global statistical characteristics of historical rainfall data were analyzed, and the mean values of atmospheric circulation factors were calculated. The dynamic adjustment coefficient is determined based on the atmospheric circulation factor and the mean atmospheric circulation factor for the period to be predicted. Based on the dynamic adjustment coefficient and rainfall statistics, the initial population range of the hyperparameter space is generated by mapping.
8. The intelligent rainfall estimation method according to claim 2, characterized in that, The correlation between historical rainfall data and atmospheric circulation factors at monitoring stations within the target area was analyzed, and multiple monitoring stations were classified into different levels, including: Construct a spatiotemporal correlation matrix between historical rainfall data and atmospheric circulation factors for each monitoring station; The likelihood function for each monitoring station is determined based on historical rainfall data and atmospheric circulation factors, and a likelihood function matrix is constructed. Based on the spatiotemporal correlation matrix and likelihood function matrix, each monitoring station is classified into different levels using a dynamic Bayesian algorithm.
9. The intelligent rainfall estimation method according to claim 2, characterized in that, Historical rainfall data from each monitoring station were clustered to obtain rainfall data during the flood season and rainfall data outside the flood season, including: A data matrix was constructed by arranging the historical rainfall data of each monitoring station according to the time series. The historical rainfall data in the data matrix is divided into two clusters using a preset clustering algorithm; Based on the difference in the mean rainfall between the two clusters, the rainfall data at each time point is determined to be either flood season rainfall data or non-flood season rainfall data.
10. The intelligent rainfall estimation method according to claim 1, characterized in that, After obtaining historical rainfall data and atmospheric circulation factors, the following is also included: The historical rainfall data were subjected to stationarity testing and noise reduction. The atmospheric circulation factors were normalized.