A method and system for classifying spatially heterogeneous extreme precipitation events
By calculating a standardized precipitation anomaly index and using a self-organizing map network model for unsupervised training, the problem of identifying and classifying spatial distribution patterns of extreme precipitation events was solved, achieving stable and objective classification of extreme precipitation events and improving the accuracy of climate risk assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies struggle to objectively and automatically identify and classify the complex spatial distribution patterns of extreme precipitation events, leading to inaccurate climate risk assessments.
By calculating the standardized precipitation anomaly index, spatially heterogeneous extreme events are screened out, and a self-organizing map network model is used for unsupervised training to generate typical spatial distribution modes.
It has achieved a stable and objective classification of extreme precipitation events, providing a scientific basis for refined risk assessment.
Smart Images

Figure CN121211042B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of atmospheric science and climate analysis technology, specifically to a classification method and system for spatially heterogeneous extreme precipitation events. Background Technology
[0002] Against the backdrop of global climate change, the frequency and intensity of extreme precipitation events have increased significantly, exhibiting high spatial heterogeneity. This means that during the same period, a patchy distribution pattern of simultaneous drought and flooding may occur within a region. Accurately identifying and classifying these extreme events with different spatial distribution patterns is crucial for refined disaster risk assessment, water resource management, and the development of climate adaptation strategies.
[0003] Current technologies for analyzing extreme precipitation largely rely on single-site observation indicators or regional averages. When assessing large-scale events, this approach tends to homogenize complex spatial differences, ignoring the internal heterogeneity of events and potentially underestimating the true risk in local areas. For example, while some existing methods can process spatial precipitation data, their focus is primarily on predicting future events, rather than conducting in-depth classification studies of the complex spatial distribution patterns of historical events themselves.
[0004] Furthermore, while some studies have attempted to introduce clustering algorithms to classify meteorological data, these methods are typically sensitive to initial conditions, exhibit poor stability, and struggle to effectively handle high-dimensional climate spatial field data. More importantly, existing technologies generally lack a crucial preliminary step: how to objectively define and select truly spatially heterogeneous extreme events from massive datasets as the subjects of analysis. Therefore, existing methods struggle to systematically and objectively identify and classify the spatial heterogeneity characteristics of extreme precipitation events, failing to comprehensively characterize their diverse spatial distribution patterns. Summary of the Invention
[0005] To address this, the present invention provides a classification method and system for spatially heterogeneous extreme precipitation events, aiming to solve the technical problem that the existing technology lacks a method to objectively and automatically identify and classify the complex spatial distribution patterns of large-scale extreme precipitation events, resulting in an incomplete characterization of the spatial heterogeneity of extreme events and thus affecting the accuracy of climate risk assessment.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] According to a first aspect of the present invention, the present invention provides a method for classifying spatially heterogeneous extreme precipitation events, the method comprising:
[0008] Obtain raw precipitation data within a preset time scale and spatial range, and calculate a standardized precipitation anomaly index that characterizes the degree of precipitation anomaly based on the raw precipitation data;
[0009] For the standardized precipitation anomaly index spatial field corresponding to each time point, a statistical index is calculated to characterize the overall spatial heterogeneity of the spatial field.
[0010] Based on the time series of the aforementioned statistical indicators, spatially heterogeneous extreme events are selected according to a preset high quantile threshold.
[0011] The spatial distribution of the standardized precipitation anomaly index corresponding to the spatially heterogeneous extreme events is used as the input sample. The typical spatial distribution modes of the spatially heterogeneous extreme events are generated by clustering through unsupervised training of a network model based on a preset topology.
[0012] Furthermore, before performing unsupervised training using a network model based on a preset topology, the method further includes:
[0013] Multiple candidate topologies are obtained, and the spatial distribution of the standardized precipitation anomaly index corresponding to the spatial heterogeneous extreme events is used as the input sample. The network model based on each candidate topology is pre-trained, and the preset performance index of each candidate topology is calculated.
[0014] The optimal topology is selected based on the preset performance indicators, and is used as the preset topology.
[0015] The preset performance indicators include quantization error and / or topology error.
[0016] Furthermore, the raw precipitation data includes a cumulative precipitation sequence;
[0017] The raw precipitation data are used to calculate a standardized precipitation anomaly index characterizing the degree of precipitation anomaly, including:
[0018] The cumulative precipitation sequence is fitted with a Gamma distribution to calculate the target cumulative probability of precipitation.
[0019] The standardized precipitation anomaly index is obtained by converting the cumulative probability of the target to the quantiles of the standard normal distribution.
[0020] Further, the step of fitting the cumulative precipitation sequence to a Gamma distribution and calculating the target cumulative probability of precipitation includes:
[0021] The probability density function fitted to the Gamma distribution is calculated using the following formula:
[0022]
[0023] in, Let p represent the probability distribution function; p represents the precipitation. Represents the gamma function; Indicates shape parameters; Indicates the scale parameter;
[0024] The target cumulative probability of precipitation is calculated using the following formula:
[0025]
[0026] Where H(p) represents the cumulative probability of the target; q represents the probability of zero precipitation; and G(p) represents the cumulative probability.
[0027] Furthermore, for the standardized precipitation anomaly index spatial field corresponding to each time point, the statistical index used to characterize the overall spatial heterogeneity of the spatial field is calculated, including:
[0028] The area-weighted spatial variance of the standardized precipitation anomaly index at each time point is calculated using the following mathematical expression:
[0029]
[0030] in, The area-weighted spatial variance of the standardized precipitation anomaly index at each time point t represents the standardization precipitation anomaly index; x represents longitude; y represents latitude; t represents time; SPI (x,y,t) S represents the standardized precipitation index corresponding to time t, longitude x, and latitude y; (x,y) Represents the area of a grid point at longitude x and latitude y; This represents the area-weighted spatial average of all grid points at time point t.
[0031] Furthermore, the candidate topology is a self-organizing map network structure;
[0032] The calculation of preset performance metrics for each of the candidate topologies includes:
[0033] The quantization error, which characterizes the fitting accuracy of the candidate topology to the input sample, is calculated using the following formula:
[0034]
[0035] in, Indicates quantization error; The number of extreme events in spatial heterogeneity; For the first Spatial distribution of standardized precipitation anomaly indices for spatially heterogeneous extreme events; express The weight vector of the best matching unit; Represents the Euclidean geometric distance;
[0036] And / or,
[0037] The topology error, which characterizes the ability of a candidate topology to preserve the topology of the input space during the mapping process, is calculated using the following formula:
[0038]
[0039] in, Indicates topological error; This represents a judgment function. If the best matching unit and the second best matching unit of each input sample are adjacent, the value is 0; otherwise, the value is 1.
[0040] Further, the step of selecting the optimal topology based on the preset performance index as the preset topology includes:
[0041] The comprehensive index score is calculated based on the quantization error and the topology error, and the mathematical expression is as follows:
[0042]
[0043] Among them, S j This indicates a comprehensive indicator score; QE j TE represents the quantization error of the j-th candidate topology. j This represents the topological error of the j-th candidate topology. ( () represents the ascending sorting function; Indicates the weighting coefficient;
[0044] The candidate topology with the lowest comprehensive index score is selected as the preset topology.
[0045] Further, the step of using the spatial distribution of the standardized precipitation anomaly index corresponding to the spatially heterogeneous extreme events as input samples, and performing unsupervised training through a network model based on a preset topology to cluster and generate typical spatial distribution modes of the spatially heterogeneous extreme events includes:
[0046] The spatial distribution matrix of the standardized precipitation anomaly index corresponding to the spatially heterogeneous extreme events is converted into a one-dimensional matrix and standardized to generate an input feature vector.
[0047] Calculate the principal components of the input feature vector, and use the principal components to initialize the weight vector of the preset topology;
[0048] An initial neighborhood radius is selected, and unsupervised training is performed based on the input feature vector and the weight vector until the model converges and the iteration terminates, obtaining the weight vector of each neuron as the typical spatial distribution mode of the spatial heterogeneous extreme event.
[0049] Furthermore, the method also includes:
[0050] Based on the average distance between adjacent neurons, the typical spatial distribution patterns of the aforementioned spatially heterogeneous extreme events are interpreted, specifically including:
[0051] The average distance between adjacent neurons is calculated using the following mathematical expression:
[0052]
[0053] in, Indicates the coordinate position of the neuron. Indicates the neuron row number. Indicates the neuron column number; This represents the average distance matrix between adjacent neurons corresponding to each neuron; Indicates position The set of all neighbors of the neuron at that location; Indicates the number of neighbors; Indicates position The weight vector of the neuron at point w; k This represents the weight vector of the k-th neighboring neuron;
[0054] The average distance between adjacent neurons was normalized and visualized using a heatmap to assess the differences in the typical spatial distribution modes.
[0055] And / or,
[0056] Internal and external consistency assessments were performed on the typical spatial distribution modes.
[0057] The internal consistency assessment includes determining whether the typical spatial distribution mode conforms to the actual precipitation distribution to ensure the mathematical rationality of the clustering results; the external consistency assessment includes using associated data to compare and verify the typical spatial distribution mode to ensure the interpretability of the clustering results; the associated data includes at least one of meteorological knowledge, observational data, and physical precipitation mechanisms.
[0058] According to a second aspect of the present invention, the present invention provides a classification system for spatially heterogeneous extreme precipitation events, the system comprising:
[0059] The data acquisition module is used to acquire raw precipitation data within a preset time scale and spatial range, and calculate a standardized precipitation anomaly index that characterizes the degree of precipitation anomaly based on the raw precipitation data.
[0060] The index analysis module is used to calculate statistical indicators to characterize the overall spatial heterogeneity of the standardized precipitation anomaly index spatial field corresponding to each time point.
[0061] The event extraction module is used to filter out spatially heterogeneous extreme events based on the time series of the statistical indicators and according to a preset high quantile threshold.
[0062] The clustering and classification module is used to take the spatial distribution of the standardized precipitation anomaly index corresponding to the spatially heterogeneous extreme events as input samples, and perform unsupervised training through a network model based on a preset topology to cluster and generate typical spatial distribution modes of the spatially heterogeneous extreme events.
[0063] The present invention, by adopting the above technical solution, has at least the following beneficial effects:
[0064] This invention acquires raw precipitation data within a preset time scale and spatial range, and calculates a standardized precipitation anomaly index to characterize the degree of precipitation anomaly based on the raw precipitation data. For the spatial field of the standardized precipitation anomaly index corresponding to each time point, a statistical index is calculated to characterize the overall spatial heterogeneity of the spatial field. Based on the time series of the statistical index, spatially heterogeneous extreme events are screened according to a preset high quantile threshold. The spatial distribution of the standardized precipitation anomaly index corresponding to the spatially heterogeneous extreme events is used as input samples, and unsupervised training is performed through a network model based on a preset topology to cluster and generate typical spatial distribution modes of the spatially heterogeneous extreme events. Therefore, it can innovatively identify extreme events with significant spatial differences, and the classification results are objective and stable, providing a scientific basis for refined risk assessment.
[0065] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0066] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0067] Figure 1A flowchart illustrating a classification method for spatially heterogeneous extreme precipitation events provided in an embodiment of the present invention is shown.
[0068] Figure 2 An error ranking comparison chart of candidate topologies provided in an embodiment of the present invention is shown;
[0069] Figure 3 This diagram illustrates the iterative principle in an unsupervised training process according to an embodiment of the present invention.
[0070] Figure 4a A schematic diagram of the average distance between adjacent neurons provided in an embodiment of the present invention is shown;
[0071] Figure 4b A schematic diagram of the topological structure between neurons provided in an embodiment of the present invention is shown;
[0072] Figure 5 This diagram illustrates the spatial distribution types of spatially heterogeneous extreme precipitation events according to an embodiment of the present invention.
[0073] Figure 6 A schematic diagram of the structure of a classification system for spatially heterogeneous extreme precipitation events provided in another embodiment of the present invention is shown. Detailed Implementation
[0074] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0075] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0076] This invention provides a classification method for spatially heterogeneous extreme precipitation events. The method aims to first identify extreme precipitation events with significant spatial non-uniform distribution characteristics through a data-driven, objective process. Then, it uses an unsupervised neural network that preserves the data topology to cluster these events, thereby obtaining several stable and physically meaningful typical spatial distribution modes. For example... Figure 1 As shown, it may include at least the following steps S101~S104:
[0077] Step S101: Obtain raw precipitation data within a preset time scale and spatial range, and calculate a standardized precipitation anomaly index that characterizes the degree of precipitation anomaly based on the raw precipitation data.
[0078] Before acquiring raw precipitation data, it is necessary to determine the scale of precipitation events to be quantified for the target spatial area in order to calculate the standardized precipitation anomaly index (SPI) at the corresponding scale. Specifically, the scale of precipitation events refers to the length of time used to assess drought or heavy precipitation (flood) events, and is usually determined based on the climate characteristics of the target spatial area (region). Common scales include: less than 1 month, used to detect short-term drought or heavy precipitation events, suitable for meteorological drought diagnosis; about 3 months, which can reflect seasonal drought and is suitable for agricultural drought assessment; and 12 months or more, suitable for hydrological drought or long-term water resource management.
[0079] Raw precipitation data refers to the cumulative precipitation sequence. The core of calculating the standardized precipitation anomaly index is to fit the skewed cumulative precipitation sequence to a Gamma distribution and then standardize it into a standard normal distribution variable. First, based on the determined scale of the precipitation event, the precipitation is cumulatively calculated. For example, if a 3-month scale is chosen, the cumulative precipitation in March 2020 is the total precipitation from January to March 2020.
[0080] Furthermore, a Gamma distribution can be fitted to the cumulative precipitation series to calculate the target cumulative probability of precipitation. In practical applications, a Gamma distribution can be fitted to the cumulative precipitation series for each month, with the probability density function as follows:
[0081]
[0082] in, Let p represent the probability distribution function; p represents the precipitation. Represents the gamma function; Indicates shape parameters; The scaling parameter can be solved using the maximum likelihood estimation method.
[0083]
[0084] in, denoted by z, representing the sample mean of precipitation; z represents the sample size.
[0085] Based on the probability density function, calculate the cumulative probability G(p) of precipitation:
[0086]
[0087] It is understandable that, since the precipitation sequence may contain zero values, the cumulative probability needs to be adjusted to obtain the target cumulative probability:
[0088]
[0089] Where H(p) represents the cumulative probability of the target; q represents the probability of zero precipitation, which is usually estimated as the proportion of zero values in the precipitation sequence.
[0090] Finally, the cumulative probability of the target can be converted into the quantile of the standard normal distribution to obtain the standardized precipitation anomaly index (SPI). The range of SPI values can usually be limited to [-3.1, 3.1].
[0091] Step S102: For the standardized precipitation anomaly index spatial field corresponding to each time point, calculate the statistical index used to characterize the overall spatial heterogeneity of the spatial field.
[0092] In this embodiment of the invention, the statistical indicator preferably adopts the area-weighted spatial variance (SV). Specifically, the area-weighted spatial variance of the standardized precipitation anomaly index at each time point can be calculated, and the mathematical expression is as follows:
[0093]
[0094] in, The area-weighted spatial variance of the standardized precipitation anomaly index at each time point t represents the standardization precipitation anomaly index; x represents longitude; y represents latitude; t represents time; SPI (x,y,t) Indicates time t and longitude Standardized precipitation index corresponding to latitude y; S (x,y) Represents the area of a grid point at longitude x and latitude y; The area-weighted spatial average of all grid points at time point t is expressed mathematically as follows:
[0095]
[0096] Step S103: Based on the time series of statistical indicators, select spatially heterogeneous extreme events according to the preset high quantile threshold.
[0097] For example, the preset high quantile threshold can be set to 95%. That is, times exceeding the preset high quantile threshold (i.e., the times corresponding to the highest 5% of SV values) are identified as spatial heterogeneity extreme events, the event IDs are arranged in the order of occurrence, and the SPI spatial distribution of each spatial heterogeneity extreme event is extracted to perform subsequent data processing steps.
[0098] Step S104: The spatial distribution of the standardized precipitation anomaly index corresponding to the spatially heterogeneous extreme events is used as the input sample. Unsupervised training is performed through a network model based on a preset topology to generate typical spatial distribution modes of spatially heterogeneous extreme events by clustering.
[0099] It should be noted that before performing unsupervised training using a network model based on a preset topology, the network structure and parameters used for clustering analysis need to be determined first. Preferably, this embodiment of the invention uses a self-organizing map (SOM)-based network structure as a candidate topology, and then determines the grid parameters based on the SOM topology. Self-organizing maps, as an artificial neural network method, can achieve dimensionality reduction and classification of high-dimensional data while maintaining the data topology, offering advantages such as intuitive results, good visualization, and suitability for processing large-scale climate and environmental data.
[0100] Specifically, multiple candidate topologies with different grid parameters can be obtained. The spatial distribution of the standardized precipitation anomaly index corresponding to spatially heterogeneous extreme events is used as the input sample. The network model based on each candidate topology is pre-trained, and the preset performance index of each candidate topology is calculated. The optimal topology is selected as the preset topology based on the preset performance index.
[0101] In practical applications, the grid structure can be rectangular or hexagonal. The grid size is generally set to a reasonable upper limit based on the size of the grid data. All possible grid combinations within the upper limit are selected. For example, if the upper limit is set to 4×4 grid, then 2×1, 2×2, 3×1, 3×2, 3×3, 4×1, 4×2, 4×3, and 4×4 grids can be selected as candidate parameters.
[0102] Then, each candidate topology with different grid parameters is traversed, and the spatial distribution of the standardized precipitation anomaly index corresponding to each spatial heterogeneous extreme event is used as the input sample for SOM model pre-training. The corresponding quantization error and topology error are calculated.
[0103] Quantization error (QE) measures the average distance between all input samples and their best-fitting unit, reflecting the accuracy of the SOM in fitting the input samples. Its mathematical expression is as follows:
[0104]
[0105] in, Indicates quantization error; The number of extreme events in spatial heterogeneity; For the first Spatial distribution of standardized precipitation anomaly indices for spatially heterogeneous extreme events; express The weight vector of the best matching unit; This represents the Euclidean geometric distance.
[0106] Topology error (TE) measures the ability of a SOM to preserve the topology of the input space during mapping. It is evaluated by checking whether the best-matching cell and the second-best-matching cell for each input sample are adjacent on the output grid. The mathematical expression is as follows:
[0107]
[0108] in, Indicates topological error; This represents a judgment function. If the best matching unit and the second best matching unit of each input sample are adjacent, the value is 0; otherwise, the value is 1.
[0109] Understandably, a lower quantization error indicates that the data sample is closer to its best matching unit, and the intraclass consistency is high; a lower topology error indicates that the SOM network maintains the data topology well.
[0110] Furthermore, a comprehensive index score is calculated based on quantization error and topology error. This comprehensive index score is then used to evaluate the quality of each candidate topology. The mathematical expression is as follows:
[0111]
[0112] Among them, S j This indicates a comprehensive indicator score; QE j TE represents the quantization error of the j-th candidate topology. j This represents the topological error of the j-th candidate topology. ( () represents the ascending sorting function; This represents the weighting coefficient, reflecting the emphasis on SOM quantization accuracy and topology preservation, and its value ranges from [0, 1]. If... A value of 0.5 indicates that both are equally important. A value greater than 0.5 indicates partial vectorization precision. A value less than 0.5 indicates a bias towards topology.
[0113] An ideal SOM network structure should have low quantization error and low topology error. Therefore, the candidate topology structure with the lowest comprehensive index score is selected as the preset topology structure. Figure 2 The figure shows a comparison of error rankings for different candidate topologies in actual experiments of this invention. The horizontal axis represents the candidate SOM topologies, including different grid structures such as 2x1, 2x2, 3x1, 3x2, 3x3, 4x1, 4x2, 4x3, and 4x4; the vertical axis represents the error ranking, ranging from 0 to 10; the three broken lines represent the quantization error ranking, topographic error ranking, and average error ranking, respectively. As can be seen from the figure, the 3×3 network structure has the lowest average error ranking, therefore this topology is preferred as the optimal topology.
[0114] After determining the optimal topology of the SOM, it can be used as a network model with the preset topology for unsupervised training, mapping the three-dimensional precipitation distribution data to the two-dimensional topology. Specifically, this can include the following steps S104-1~S104-3:
[0115] Step S104-1: Convert the spatial distribution matrix of the standardized precipitation anomaly index corresponding to the spatial heterogeneous extreme event into a one-dimensional matrix and perform standardization processing to generate the input feature vector.
[0116] In this step, the training samples are the SPI spatial distributions of all spatially heterogeneous extreme events. Each sample is a spatial distribution matrix of a spatially heterogeneous extreme event (e.g., a 100×200 grid SPI). It needs to be converted into a one-dimensional vector (a total of 100×200=20000 effective grid points) and standardized as the input feature vector.
[0117] Step S104-2: Calculate the principal components of the input feature vector and use the principal components to initialize the weight vector of the preset topology.
[0118] In practical applications, the first two principal components of the input feature vector can be calculated, and then the preset topological structure can be projected onto the plane spanned by the principal components to obtain the initialized weight vector.
[0119] Step S104-3: Select the initial neighborhood radius, perform unsupervised training based on the input feature vector and weight vector until the model converges and the iteration terminates, and obtain the weight vector of each neuron as the typical spatial distribution mode of spatial heterogeneous extreme events.
[0120] In this embodiment of the invention, the specific steps for each iteration during unsupervised training are as follows: a~f:
[0121] a. Randomly select one sample from all input samples (with replacement);
[0122] b. Calculate the spatial weighted distance (Euclidean distance or Mahalanobis distance) between the current sample and the weight vectors of all neurons.
[0123] c. Find the neuron with the smallest distance as the best matching unit (i.e. the most similar topological node in the sample).
[0124] d. Based on the location of the best matching unit, update the weights of neurons in its neighborhood (similar extreme precipitation spatial distribution samples are mapped to adjacent regions of the two-dimensional grid).
[0125] e. Update the neighborhood radius (the initial radius gradually decays to 1 with each iteration, updating only the current cell);
[0126] f. The iteration termination condition is set as follows: the number of iterations equals the sample size multiplied by 10, or the change in quantization error value is less than 0.1% after 10 consecutive iterations.
[0127] like Figure 3 The diagram illustrates the principle of the iterative process in unsupervised training. In the diagram, the grid represents a two-dimensional SOM network, the purple shading represents high-dimensional data samples, and the yellow circles represent the neighborhood. It can be seen that during training, the weights of neurons in the neighborhood move towards the target direction. After multiple iterations, the grid region approximates the data distribution, forming a mapping from high-dimensional (samples) to low-dimensional (samples). Similar samples are mapped closer to each other, while samples with large differences are separated. After unsupervised training, the weight vector of each neuron represents the typical spatial distribution mode of spatially heterogeneous extreme events.
[0128] As an optional embodiment of the present invention, after obtaining the typical spatial distribution modes of spatially heterogeneous extreme events, the clustering results can also be analyzed and interpreted.
[0129] Specifically, the typical spatial distribution patterns of extreme events with spatial heterogeneity can be interpreted based on the average distance between adjacent neurons (U-matrix). The U-matrix reflects the similarity of neurons in the two-dimensional topology of the SOM; a larger value indicates a more significant difference in the extreme precipitation patterns corresponding to adjacent neurons, and vice versa. First, the average distance between adjacent neurons is calculated, as shown in the following mathematical expression:
[0130]
[0131] in, Indicates the coordinate position of the neuron. Indicates the neuron row number. Indicates the neuron column number; This represents the average distance matrix between adjacent neurons corresponding to each neuron; Indicates position ; Indicates the number of neighbors; Indicates position ;w k This represents the weight vector of the k-th neighboring neuron.
[0132] Furthermore, the average distance between adjacent neurons was normalized and visualized using a heatmap to assess the differences in different typical spatial distribution modes.
[0133] This invention utilizes self-organizing mapping to effectively distinguish the spatial distribution of precipitation in spatially heterogeneous extreme precipitation events within China. Taking this as an example, a 3×3 topological self-organizing mapping network is used to classify the spatial distribution of these events into nine categories. Based on this, a heatmap of the average distance between adjacent neurons is provided, such as... Figure 4a As shown. Figure 4a In the diagram, U-matrix values are distinguished by a gradient of colors from light to dark. Darker colors indicate greater distance between the neuron and other neurons, signifying stronger independence and a more significant difference in precipitation patterns compared to other categories. Each grid point is labeled with the total number of event days and a list of event IDs corresponding to that category. The topological structure between these nine neurons, i.e., the relationships between classes, is illustrated using a hexagonal structure resembling a honeycomb, as shown below. Figure 4b As shown.
[0134] In addition, this embodiment of the invention also provides another interpretation method, namely, to perform internal consistency assessment and external consistency assessment on typical spatial distribution modes.
[0135] Internal consistency assessment can include determining whether typical spatial distribution modes conform to the actual precipitation distribution to ensure the mathematical validity of the clustering results. In other words, internal consistency assessment focuses on the statistical characteristics and mathematical validity of the clustering results themselves, without relying on external labels or prior knowledge. For example, it can check whether each typical spatial distribution mode conforms to the actual precipitation distribution, excluding "spurious clusters" (distributions with scattered precipitation centers and no clear spatial pattern) that are meaningless or have large intra-cluster differences.
[0136] External consistency assessment can include using correlated data to compare and validate typical spatial distribution modes to ensure the interpretability of clustering results. For example, clustering results can be compared and validated with external meteorological knowledge, observational data, and physical mechanisms. Specifically, the typical spatial distribution modes of each type of event can be associated with the season of occurrence, weather systems, and circulation background. Depending on the research needs, types with significantly different generation mechanisms (such as frontal precipitation, local convective precipitation, typhoon precipitation, etc.) can be further subdivided within known categories, or types with similar occurrence backgrounds (such as spring precipitation, summer precipitation, autumn precipitation, winter precipitation, etc.) can be merged to ensure the interpretability of clustering results in meteorology or specific research.
[0137] In practical applications, the cluster centers (weight fields) of various spatially heterogeneous extreme precipitation events within the target area can be displayed in a spatial distribution format, along with the probability density function distributions of each category. For example... Figure 5 As shown, the spatial distribution patterns of cluster centers can be divided into nine categories. Category 1 is the north-south type, where precipitation is distributed in opposite directions (e.g., "northwest-southeast dipole type" or "north-south dipole type"). Category 9 is the east-west type, where precipitation is distributed diagonally in opposite directions (e.g., "east-west dipole type"). Category 3 is the wet type, where precipitation is generally humid with regional differences (e.g., "central wet type", "western wet type", "northern wet type", or "multipolar wet type"). Category 7 is the dry type, where precipitation is generally dry with regional differences (e.g., "multipolar dry type" or "western dry type"). It can be understood that the closer to a diagonal point, the more pronounced the corresponding characteristic.
[0138] This invention provides a classification method for spatially heterogeneous extreme precipitation events, which, compared with existing technologies, has at least the following advantages:
[0139] 1) By calculating statistical indicators that characterize the degree of spatial heterogeneity and defining and screening extreme events based on high quantile thresholds, it is possible to innovatively capture and identify extreme precipitation events with significant spatial differences that are easily overlooked by traditional methods, thereby improving the ability to identify complex events with high accuracy.
[0140] 2) The entire process is data-driven. The network structure is optimized by objectively evaluating performance indicators, and unsupervised neural networks that can maintain the data topology are used for clustering. This avoids the dependence of human-preset classification standards and clustering algorithms on initial conditions, and ensures the objectivity, stability and interpretability of the classification results.
[0141] 3) It can objectively and automatically classify the spatial distribution patterns of extreme precipitation events, providing a refined quantitative reference for risk assessment of different spatial distribution types. It can be widely used in infrastructure planning, agricultural disaster prevention and mitigation, insurance and financial risk assessment, and climate model verification, and has high application value.
[0142] Furthermore, as Figure 1 In its specific implementation, this invention provides a classification system for spatially heterogeneous extreme precipitation events, such as... Figure 6 As shown, the system may include: a data acquisition module 610, an indicator analysis module 620, an event extraction module 630, and a clustering and classification module 640.
[0143] The data acquisition module 610 can be used to acquire raw precipitation data within a preset time scale and spatial range, and calculate a standardized precipitation anomaly index that characterizes the degree of precipitation anomaly based on the raw precipitation data.
[0144] The indicator analysis module 620 can be used to calculate statistical indicators to characterize the overall spatial heterogeneity of the standardized precipitation anomaly index spatial field corresponding to each time point.
[0145] Event extraction module 630 can be used to filter out spatially heterogeneous extreme events based on time series data with statistical indicators and a preset high quantile threshold.
[0146] The clustering and classification module 640 can be used to take the spatial distribution of standardized precipitation anomaly indices corresponding to spatially heterogeneous extreme events as input samples, and perform unsupervised training through a network model based on a preset topology to cluster and generate typical spatial distribution modes of spatially heterogeneous extreme events.
[0147] It should be noted that other corresponding descriptions of the functional modules involved in the classification system for spatially heterogeneous extreme precipitation events provided in this embodiment of the invention can be found in the following references. Figure 1 The corresponding description of the method shown will not be repeated here.
[0148] Those skilled in the art will clearly understand that the specific working process of the systems, devices, modules and units described above can be referred to the corresponding process in the foregoing method embodiments. For the sake of brevity, it will not be repeated here.
[0149] Furthermore, the functional units in the various embodiments of the present invention can be physically independent of each other, or two or more functional units can be integrated together, or all functional units can be integrated into one processing unit. The integrated functional units described above can be implemented in hardware, or in software or firmware.
[0150] Those skilled in the art will understand that if the integrated functional unit is implemented in software and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or all or part of it, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computing device (e.g., a personal computer, server, or network device) to execute all or part of the steps of the methods described in the embodiments of the present invention when running the instructions. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0151] Alternatively, all or part of the steps of the foregoing method embodiments can be implemented by hardware (such as a computing device, personal computer, server, or network device) related to program instructions. The program instructions can be stored in a computer-readable storage medium. When the program instructions are executed by the processor of the computing device, the computing device executes all or part of the steps of the methods described in the various embodiments of the present invention.
[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that within the spirit and principles of the present invention, modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the corresponding technical solutions to depart from the protection scope of the present invention.
Claims
1. A method of typing spatially heterogeneous extreme precipitation events, characterized in that, The method comprises: obtaining original precipitation data in a preset time scale and spatial range, and calculating a standardized precipitation anomaly index representing the degree of precipitation anomaly based on the original precipitation data; for the standardized precipitation anomaly index spatial field corresponding to each time point, calculating a statistical index representing the overall spatial heterogeneity degree of the spatial field; based on the time series of the statistical index, an extreme spatial heterogeneity event is screened out according to a preset high quantile threshold; the spatial distribution of the standardized precipitation anomaly index corresponding to the extreme spatial heterogeneity event is taken as an input sample, and an unsupervised training is performed based on a preset topological structure network model to cluster and generate a typical spatial distribution mode of the extreme spatial heterogeneity event, comprising: the spatial distribution matrix of the standardized precipitation anomaly index corresponding to the extreme spatial heterogeneity event is converted into a one-dimensional matrix and standardized to generate an input feature vector; the principal component of the input feature vector is calculated, and the principal component is used to initialize the weight vector of the preset topological structure; an initial neighborhood radius is selected, and an unsupervised training is performed based on the input feature vector and the weight vector until the model converges to terminate the iteration, and the weight vector of each neuron is obtained as the typical spatial distribution mode of the extreme spatial heterogeneity event; Before the unsupervised training based on the preset topological structure network model, the method further comprises: obtaining a plurality of candidate topological structures, taking the spatial distribution of the standardized precipitation anomaly index corresponding to the extreme spatial heterogeneity event as an input sample, and pre-training the network model based on each candidate topological structure, and calculating a preset performance index of each candidate topological structure, comprising: the quantitative error for representing the fitting accuracy of the candidate topological structure to the input sample is calculated by the following formula: wherein, denotes the quantization error; is the number of spatially heterogeneous extreme events; is the standardized precipitation anomaly index of the spatial distribution of the denotes the weight vector of the best matching cell of ; and, denotes the Euclidean distance; and, the topological error for representing the ability of the candidate topological structure to maintain the topological structure of the input space in the mapping process is calculated by the following formula: wherein, denotes a topological error; denotes a decision function, which takes value 0 if the best matching cell and the second best matching cell of each input sample are adjacent, otherwise takes value 1; the preset performance index comprises the quantization error and the topological error; the optimal topological structure is selected according to the preset performance index as the preset topological structure, comprising: the comprehensive index score is calculated according to the quantitative error and the topological error, and the mathematical expression is as follows: wherein S j denotes the comprehensive index score; QE j denotes the quantization error of the jth candidate topology; TE j denotes the topology error of the jth candidate topology; denotes the ascending order sorting function; denotes the weight coefficient; the candidate topological structure with the minimum comprehensive index score is selected as the preset topological structure.
2. The method of claim 1, wherein, The original precipitation data comprises a cumulative precipitation sequence; the standardized precipitation anomaly index representing the degree of precipitation anomaly is calculated based on the original precipitation data, comprising: the cumulative precipitation sequence is fitted with Gamma distribution to calculate the target cumulative probability of precipitation; the target cumulative probability is converted into a quantile of standard normal distribution to obtain the standardized precipitation anomaly index.
3. The method of claim 1, wherein, for the standardized precipitation anomaly index spatial field corresponding to each time point, a statistical index representing the overall spatial heterogeneity degree of the spatial field is calculated, comprising: the area-weighted spatial variance of the standardized precipitation anomaly index at each time point is calculated, and the mathematical expression is as follows: wherein, represents the area-weighted spatial variance of the standardized precipitation anomaly index corresponding to each time point t; x represents longitude; y represents latitude; t represents time; SPI (x,y,t) represents the standardized precipitation index corresponding to time t, longitude x, and latitude y; S (x,y) represents the grid area of longitude x and latitude y. represents the area-weighted spatial mean value of all grids at time point t.
4. The method according to any one of claims 1 to 3, characterized in that, The method further comprises: the typical spatial distribution mode of the extreme spatial heterogeneity event is interpreted according to the average distance between adjacent neurons, specifically comprising: The average distance between the adjacent neurons is calculated, and the mathematical expression is as follows: wherein, represents the coordinate position of a neuron, represents the row index of a neuron, represents the column index of a neuron; represents the average distance matrix between neighboring neurons corresponding to each neuron; represents the position of all neighbors of a neuron at position represents the number of neighbors; represents the weight vector of a neuron at position ; w k represents the weight vector of the kth neighboring neuron; The average distance between the adjacent neurons is normalized and visualized by a heat map to evaluate the difference between different typical spatial distribution modes; And / or, The internal consistency and external consistency of the typical spatial distribution mode are evaluated; The internal consistency evaluation includes judging whether the typical spatial distribution mode conforms to the actual distribution of precipitation, to ensure the mathematical rationality of the clustering result; the external consistency evaluation includes comparing and verifying the typical spatial distribution mode by using correlation data, to ensure the interpretability of the clustering result; the correlation data includes at least one of meteorological knowledge, observation data and physical precipitation mechanism.
5. A system for typing spatially heterogeneous extreme precipitation events, characterized in that, The system comprises: A data acquisition module is configured to acquire original precipitation data within a preset time scale and spatial range, and calculate a standardized precipitation anomaly index representing the degree of precipitation anomaly based on the original precipitation data; An index analysis module is configured to calculate a statistical index representing the overall spatial heterogeneity of a spatial field of the standardized precipitation anomaly index corresponding to each time point; An event extraction module is configured to filter out spatial heterogeneity extreme events according to a preset high quantile threshold based on a time series of the statistical index; A clustering and typing module is configured to take the spatial distribution of the standardized precipitation anomaly index corresponding to the spatial heterogeneity extreme events as an input sample, and perform unsupervised training based on a network model with a preset topological structure to cluster and generate typical spatial distribution modes of the spatial heterogeneity extreme events.
Citation Information
Patent Citations
Meteorological waterlogging monitoring method based on Transform and GWRK
CN118501985A
Method for defining and identifying time-space continuous heavy rainfall event
CN118733996A