A verification method for extreme rain intensity forecast
By fusing station observations and radar data to generate a probabilistic fusion analysis field, extreme precipitation entities are identified and clustered, and a reasonable spatial search domain is set. This solves the problem of inaccurate verification benchmarks caused by sparse station data in traditional methods, and enables a more scientific evaluation of extreme rainfall intensity forecasts.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINESE ACAD OF METEOROLOGICAL SCI
- Filing Date
- 2026-03-13
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional methods rely on sparsely distributed ground station observation data in extreme rainfall intensity forecasting, resulting in inaccurate verification benchmarks and an inability to fairly assess the model's ability to capture weather entities. Furthermore, fuzzy matching methods are limited by insufficient reliability of the verification benchmarks, leading to inadequate assessment stability and credibility.
By fusing station observation data with radar quantitative precipitation estimation fields, a probabilistic fusion analysis field is generated. Spatially adjacent grid clusters are identified and extracted as observed extreme precipitation entities. An elliptical spatial search domain is set based on the prevailing wind direction to determine the hit and generate a comprehensive verification report.
This enables a more scientific, stable, and impartial assessment of the extreme precipitation forecasting capabilities of numerical models, solves the problem of inaccurate verification benchmarks caused by sparse station data in traditional methods, and improves the accuracy and reliability of the assessment.
Smart Images

Figure CN122131428A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of meteorological forecasting, and in particular to a verification method for forecasting extreme rainfall intensity. Background Technology
[0002] In the field of meteorological forecasting, the scientific verification of the ability of high-resolution numerical models to forecast extreme precipitation is crucial. Traditional point-to-point exact matching verification methods suffer from a double penalty effect due to their excessive sensitivity to unavoidable micro-spatial displacements in forecasts, making it impossible to fairly assess the model's ability to capture weather entities. To address this, a fuzzy matching method that allows for spatiotemporal deviations has been developed. This method tolerates reasonable errors by pre-setting spatial radii and time windows, and uses the hit rate as the core indicator, shifting the evaluation focus to whether the model successfully forecasts the corresponding weather system. This approach overcomes the limitations of traditional methods to some extent.
[0003] The effectiveness of fuzzy matching methods is limited by the reliability of their test benchmarks. They rely on sparsely distributed ground station observations to define extreme events and delineate search benchmarks. Heavy precipitation has strong spatial heterogeneity, and station data is difficult to accurately characterize the true spatial structure and intensity distribution of precipitation. There are inherent representativeness errors and measurement uncertainties. Using discrete and sparse station point data as deterministic ground truth to test continuous and high-resolution model field data results in a fundamental scale mismatch problem. This leads to the test results being overly dependent on the random distribution of stations and individual data errors, resulting in insufficient stability and reliability of the assessment. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a verification method for extreme rainfall intensity forecasts, which solves the problem of inaccurate extreme rainfall intensity forecast evaluation caused by insufficient representativeness of observation data and unreasonable matching rules.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a verification method for extreme rainfall intensity forecasts, which includes acquiring model forecast fields, station observation data and radar quantitative precipitation estimation fields; By fusing site observation data and radar quantitative precipitation estimation fields, a probabilistic fusion analysis field is generated. Based on the probabilistic fusion analysis of grid precipitation intensity, all grid points that reach the predetermined extreme precipitation intensity threshold are identified and extracted. Spatially adjacent grid points are clustered into discrete observation extreme precipitation entities, and each observation extreme precipitation entity is output. For each observed extreme precipitation entity, the atmospheric lower-level environmental flow field information corresponding to the time of occurrence is extracted to determine the prevailing wind direction. Based on the prevailing wind direction, elliptical spatial search domain parameters with the major axis direction parallel to the wind direction are set for the observed extreme precipitation entity. For each observed extreme precipitation entity, a hit determination is made based on the search parameters of the corresponding elliptical space search domain and the search mode forecast field within a preset time window. The hit determination results are statistically analyzed to generate a comprehensive test report.
[0007] As a preferred embodiment of the verification method for extreme rainfall intensity forecasting described in this invention, the acquisition of model forecast field, station observation data, and radar quantitative precipitation estimation field includes the following steps: Acquire gridded precipitation forecast fields, precipitation observation data from ground automatic weather stations, and gridded data for quantitative precipitation estimation from weather radar, and unify the timestamps; The time-uniform gridded precipitation forecast field, precipitation observation data from automatic weather stations, and quantitative precipitation estimation gridded data from weather radar are aligned to a common spatial grid through resampling and location recording. The gridded precipitation forecast field based on spatiotemporal alignment, precipitation observation data from automatic weather stations, and gridded data for quantitative precipitation estimation from weather radar are used as the model forecast field, the station observation dataset, and the radar quantitative precipitation estimation field, respectively.
[0008] As a preferred embodiment of the verification method for extreme rainfall intensity forecasting described in this invention, the method involves fusing station observation data and radar quantitative precipitation estimation fields to generate a probabilistic fusion analysis field, including the following steps: Based on the site observation dataset and radar quantitative precipitation estimation, the optimal interpolation algorithm is applied to fuse the site observation dataset and radar quantitative precipitation estimation field to generate a preliminary fusion analysis field. Based on the difference between the site observation dataset and the preliminary fusion analysis field, the uncertainty of the precipitation intensity estimate at each grid point in the preliminary fusion analysis field is calculated. By combining the grid point precipitation intensity estimates in the preliminary fusion analysis field with the corresponding uncertainty information, a probabilistic fusion analysis field is output.
[0009] As a preferred embodiment of the verification method for extreme rainfall intensity forecasting described in this invention, the method involves identifying and extracting all grid points that reach a predetermined extreme rainfall intensity threshold based on the grid point precipitation intensity in the probabilistic fusion analysis field, including the following steps: The gridded precipitation intensity field is extracted from the probabilistic fusion analysis field. The gridded precipitation intensity field contains the precipitation intensity estimate for each grid point in the probabilistic fusion analysis field. The estimated precipitation intensity at each grid point in the gridded precipitation intensity field is compared with a predetermined extreme rainfall intensity threshold. Mark all grid points in the grid precipitation intensity field whose estimated precipitation intensity is greater than or equal to a predetermined extreme precipitation intensity threshold, and output a list of grid point locations marked as candidate extreme precipitation grid points.
[0010] As a preferred embodiment of the verification method for extreme rainfall intensity forecasting described in this invention, the method involves: clustering spatially adjacent grid points into discrete observed extreme precipitation entities, and outputting each observed extreme precipitation entity, including the following steps: Based on the grid location list labeled as candidate extreme precipitation grid points, a spatial clustering algorithm is applied to group the grid points in the grid location list labeled as candidate extreme precipitation grid points to obtain each independent connected group; Each independent connected group is assigned an identifier for an observed extreme precipitation entity. Based on the position coordinates of all grid points within the independent connected group, the spatial center position of the observed extreme precipitation entity is determined. Based on the temporal attributes of all grid points within an independent connected group, the occurrence time of observed extreme precipitation entities is determined; Based on the precipitation intensity values of all grid points within the independent connected group, the maximum precipitation intensity of the observed extreme precipitation entity is determined; Based on the spatial distribution of all grid points within an independent connected group, the coverage area of the observed extreme precipitation entity is determined, and the observed extreme precipitation entity and its corresponding entity feature attributes are output.
[0011] As a preferred embodiment of the verification method for extreme rainfall intensity forecasting described in this invention, the method includes the following steps: for each observed extreme precipitation entity, extracting the lower atmospheric environmental flow field information corresponding to the occurrence time to determine the prevailing wind direction, and setting elliptical spatial search domain parameters with the major axis direction parallel to the wind direction for the observed extreme precipitation entity based on the prevailing wind direction. For each observed extreme precipitation entity, based on the occurrence time of the observed extreme precipitation entity, the corresponding time and the lower atmospheric environmental flow field information of the core area with the center location of the observed extreme precipitation entity are extracted from the external atmospheric reanalysis data; By performing regional vector averaging of the lower atmospheric environmental flow field information, the prevailing wind direction of the environment in which the observed extreme precipitation entity is located can be obtained; For each observed extreme precipitation entity, the core geometric parameters of the elliptical spatial search domain are defined using the prevailing wind direction.
[0012] As a preferred embodiment of the verification method for extreme rainfall intensity forecasting described in this invention, the method includes the following steps: For each observed extreme precipitation entity, a hit determination is performed based on the search parameters in the corresponding elliptical spatial domain and the retrieved model forecast field within a preset time window: For each observed extreme precipitation entity, the time of occurrence of the observed extreme precipitation entity and a preset time window are used to determine the time search range; For each observed extreme precipitation entity, a four-dimensional spatiotemporal search domain is defined by combining the temporal search range and the parameters of the elliptical spatial search domain. Within the four-dimensional spatiotemporal search domain, all grid point precipitation forecast values in the model forecast field are retrieved. All retrieved grid point precipitation forecast values are checked to determine whether at least one grid point precipitation forecast value reaches or exceeds the predetermined extreme rainfall intensity threshold. If a grid point precipitation forecast value that meets the conditions exists, the forecast for the observed extreme precipitation entity is determined as a hit event; otherwise, it is determined as a missed event.
[0013] As a preferred embodiment of the verification method for extreme rainfall intensity forecasting described in this invention, the method involves: statistically analyzing the hit determination results to generate a comprehensive verification report, including the following steps: The hit detection results of all observed extreme precipitation entities are statistically analyzed to obtain the total number of hit events and the total number of missed events; The hit rate score is calculated based on the total number of hit events and the total number of observed extreme precipitation entities. Using the list of hit events and the list of missed events, combined with the corresponding elliptical spatial search domain parameters, a geospatial distribution map is drawn, showing the distribution of hit events, missed events, and spatial search domains. The hit rate score and the geospatial distribution map are integrated to generate a comprehensive inspection report.
[0014] In a second aspect, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the verification method for extreme rainfall intensity forecasting as described in the first aspect of the present invention.
[0015] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the verification method for extreme rainfall intensity forecasting as described in the first aspect of the present invention.
[0016] The beneficial effects of this invention are as follows: by fusing site observations and radar data to generate a high-resolution probabilistic fusion analysis field, a more reliable verification benchmark is constructed. Based on the environmental flow field, an elliptical spatial search domain with its major axis parallel to the wind direction is set for each identified precipitation entity, thereby establishing a physically more reasonable fuzzy matching rule. This solves the problems of inaccurate verification benchmarks caused by the reliance on sparse site data in traditional methods, and the inability to reasonably tolerate reasonable forecast displacements due to the use of a fixed circular search domain. This enables a more scientific, stable, and fair evaluation of the extreme precipitation forecasting capability of numerical models. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of a verification method for extreme rainfall intensity forecasting. Detailed Implementation
[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0020] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0021] Secondly, the term "one embodiment" or "example" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the invention. The appearance of an embodiment in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that mutually excludes other embodiments.
[0022] Reference Figure 1 This is one embodiment of the present invention, which provides a verification method for extreme rainfall intensity forecasts, including the following steps: S1. Acquire model forecast fields, station observation data, and radar quantitative precipitation estimation fields.
[0023] S1.1 Acquire gridded precipitation forecast fields, precipitation observation data from ground automatic weather stations, and gridded data for quantitative precipitation estimation from weather radar, and unify the timestamps.
[0024] Furthermore, three sets of data are acquired in parallel from the operational archive or real-time data stream. The first set is a gridded precipitation forecast field covering the target area, output by a high-resolution numerical weather prediction model during the target verification period. The second set is a collection of precipitation observation data records from all automatic weather stations within the same time period and area. The third set is a quantitative precipitation estimation gridded data product generated by a weather radar network within the same spatiotemporal range. The acquisition operation must ensure the integrity and availability of the data. Subsequently, the time reference of these three independently acquired sets of data is unified. This involves identifying and converting the time stamps carried by each set of data and synchronizing them all to a common time reference frame. For example, the timestamps of all data are uniformly adjusted and standardized to the hour in Beijing time. This process eliminates potential time phase deviations caused by different data sources, different release cycles, or inconsistent timing methods. The final output is a gridded precipitation forecast field, automatic weather station precipitation observation data, and weather radar quantitative precipitation estimation gridded data that are synchronized in the time dimension.
[0025] Specifically, a preprocessing channel for synchronous acquisition and time alignment of multi-source, heterogeneous data was constructed. Its direct purpose is to lay a foundation of absolute temporal consistency for subsequent in-depth data fusion and comparative analysis. Traditional verification methods typically only handle model forecast fields and station observation data. This approach incorporates gridded data for quantitative precipitation estimation from weather radar as a third, independent, and crucial data source. This is to directly address and resolve the fundamental challenge of the inherent uncertainty and representativeness error of the observation data itself, as pointed out in the document. Relying solely on sparse station observation data as the true value has inherent limitations, while radar data provides high-resolution spatial structure information but may contain quantitative biases. By placing all three under the same time reference, this essentially creates the prerequisite for generating a superior probabilistic fusion analysis field using data fusion techniques in the next step. For example, without this step, directly fusing radar and station data from different time bases would produce incorrect spatiotemporal correlations, leading to a distorted fused field. However, by mandating timestamp unification, it is ensured that the spatial morphology of precipitation echoes from radar and the precise precipitation intensity values recorded by the stations are correctly matched in time. This allows subsequent fusion algorithms to correct the systematic errors in radar data to the greatest extent possible and use station data as anchor points. As a result, a more reliable and representative benchmark truth field is generated for the entire verification process, which fundamentally improves the accuracy and fairness of subsequent fuzzy matching verification.
[0026] S1.2 Align the time-uniform gridded precipitation forecast field, precipitation observation data from automatic weather stations, and quantitative precipitation estimation gridded data from weather radar to a common spatial grid through resampling and location recording.
[0027] Furthermore, a common spatial reference frame is defined, namely a target spatial grid with a clearly defined geographical extent, projection method, and grid resolution. Next, resampling techniques are used to transform the temporally unified gridded precipitation forecast field and the temporally unified weather radar quantitative precipitation estimation gridded data from their respective original grid systems to this predefined target spatial grid through interpolation calculations, ensuring that the two sets of gridded data have completely consistent spatial location definitions. For temporally unified automatic weather station precipitation observation data, the precise geographical coordinates of each station are recorded and mapped to the target spatial grid, clearly defining the target grid cell to which each station belongs. This achieves precise registration of all data in the spatial domain, outputting a gridded precipitation forecast field with defined target grids, automatic weather station precipitation observation data with defined target grids, and weather radar quantitative precipitation estimation gridded data with defined target grids.
[0028] Specifically, this contradiction is cleverly resolved through two differentiated but consistent techniques: resampling and location recording. Resampling standardizes continuous gridded field data, while location recording assigns discrete station data to a grid. For example, after this processing, observations from a station within a target grid can be rigorously compared with model forecasts and radar estimates obtained through resampling within that grid, making it possible to generate a unified grid for fusion analysis from multi-source data. This breaks down the barriers of differing spatial benchmarks between different data sources, ensuring that subsequent comparisons and calculations at each grid point or location are based on strict spatial consistency. This forms the geometric foundation for high-precision probabilistic fusion analysis and subsequent scientific spatial matching, ensuring that verification is not conducted in fuzzy spatial relationships but rather within a precisely quantified spatial grid system.
[0029] S1.3, the gridded precipitation forecast field based on spatiotemporal alignment processing, the precipitation observation data from ground automatic weather stations, and the gridded data for quantitative precipitation estimation by weather radar are used as the model forecast field, the station observation dataset, and the radar quantitative precipitation estimation field, respectively.
[0030] Furthermore, the gridded precipitation forecast field, the automatic weather station precipitation observation data, and the weather radar quantitative precipitation estimation gridded data, all defined for the target grid, undergo final format standardization and logical integration. This ensures that the data storage structure, variable naming, missing value identification, and physical units remain completely consistent throughout the entire process and facilitate indexing and retrieval in subsequent steps. Having completed these standardization processes, the gridded precipitation forecast field for the target grid is explicitly designated as the model forecast field used throughout the verification process. The automatic weather station precipitation observation data for the target grid is integrated into the station observation dataset. The weather radar quantitative precipitation estimation gridded data for the target grid is designated as the radar quantitative precipitation estimation field. The output consists of a well-formatted model forecast field, station observation dataset, and radar quantitative precipitation estimation field with completely unified spatiotemporal references.
[0031] S2. Integrate site observation data and radar quantitative precipitation estimation field to generate a probabilistic fusion analysis field.
[0032] S2.1 Based on the station observation dataset and radar quantitative precipitation estimation, the optimal interpolation algorithm is applied to fuse the station observation dataset and radar quantitative precipitation estimation field to generate a preliminary fusion analysis field; Furthermore, using the station observation dataset and the radar quantitative precipitation estimation field as inputs, an optimal interpolation algorithm is applied for data fusion. The optimal interpolation algorithm uses the radar quantitative precipitation estimation field as the background field, providing the prior spatial distribution of precipitation; and the station observation dataset as the observed values, providing precise point-based precipitation information. At each target grid point, the algorithm calculates a set of optimal weighting coefficients by calculating the background error covariance between the grid point and the locations of all station observation data within a certain influence range, as well as the observation error covariance of the station observation data itself. This set of optimal weighting coefficients is used to perform a weighted average of the deviation between the station observation data and the background field at the station location (i.e., the observation increment), and the weighted average correction is superimposed on the background field to generate a new precipitation intensity estimate at each grid point. After traversing all target grid points and completing the above calculations, a preliminary fused analysis field is obtained that is spatially continuous and integrates the spatial structure information of the radar data and the precise calibration information of the station data. This achieves the goal of correcting the radar quantitative precipitation estimation field using the station observation dataset and generating a more accurate spatially distributed analysis field.
[0033] Specifically, the optimal interpolation algorithm in the field of data assimilation is applied to construct a benchmark field for extreme precipitation verification, realizing a paradigm shift from single-source ground truth to multi-source fusion optimal estimation. Traditional verification methods directly treat sparse station observation data as deterministic ground truth, acknowledging that both the radar quantitative precipitation estimation field and the station observation dataset have errors, but their error characteristics are complementary. Radar fields have high spatial resolution but may have systematic biases; station data are accurate but have limited spatial representativeness. Through a rigorous error statistical model (background error covariance and observation error covariance), the optimal linear unbiased estimate for these two imperfect data sources is mathematically sought. For example, in sparsely populated areas, the algorithm relies more on the spatial gradient information provided by the radar field; while in densely populated areas, it primarily uses accurate station observations to strongly correct the radar field. By automatically and objectively weighing the contributions of different data sources, the algorithm generates a preliminary fusion analysis field that is closer to the actual precipitation situation at any location than a single data source. S2.2. Based on the difference between the site observation dataset and the preliminary fusion analysis field, calculate the uncertainty of the precipitation intensity estimate for each grid point in the preliminary fusion analysis field.
[0034] Furthermore, after generating the preliminary fusion analysis field, based on the same optimal interpolation theoretical framework, the uncertainty of this field is calculated using the already solved optimal weighting coefficients, the preset background error covariance matrix, and the observation error covariance matrix. According to the provided uncertainty expression, the analysis error variance is calculated for each target grid point in the preliminary fusion analysis field. The first part of the expression... This represents the prior uncertainty of the radar quantitative precipitation estimation field at this grid point. The double summation part of the expression physically represents the reduction in uncertainty after fusing the site observation datasets. The summation term involves a combination operation of the optimal weights of all site pairs and the background error covariance. By calculating this formula grid-by-grid, the final result is the analysis error variance value corresponding to the precipitation intensity estimate at each grid point in the preliminary fused analysis field. This value quantitatively characterizes the unreliability of the estimate, completing a quantitative assessment of the uncertainty of the estimate at each grid point in the preliminary fused analysis field.
[0035] Specifically, it quantifies the information gain and residual error resulting from the entire fusion process. For example, at a grid point surrounded by multiple high-quality sites, the summation term in the formula will be larger, resulting in a larger amount to be subtracted from Bii and a smaller calculated analysis error variance, making the fusion analysis value highly reliable. Conversely, at grid points far from all sites, the error reduction from fusion is limited, and the analysis error variance is close to the background error variance, indicating higher uncertainty in the estimated value. This capability is completely absent in traditional methods. It provides valuable confidence information for testing, improving the precision and reliability of the test conclusions.
[0036] The uncertainty expression for the grid point precipitation intensity estimate is: ; in, To initially integrate and analyze the field at the target grid point Uncertainty in the estimated precipitation intensity at the location The total number of observation data from the stations. For the index variable of the first observation station, For the index variable of the second observation station, For the first When assigning the analysis value of the nth target grid point, the first... The observation increment of observation data for each site For the first The and the first The covariance between the errors of the radar quantitative precipitation estimation field at the location of the observation data from each station. For the first When assigning the analysis value of the nth target grid point, the first... The observation increment of observation data for each site For the index of the target grid point, For radar quantitative precipitation estimation field at target grid points Prior uncertainty at the location, For the initial fusion analysis of the field; S2.3 Combine the grid point precipitation intensity estimates in the preliminary fusion analysis field with the corresponding uncertainty information to output the probabilistic fusion analysis field.
[0037] Furthermore, the preliminary fusion analysis field provides the optimal precipitation intensity estimate for each grid point, while the analysis error variance field provides a measure of the dispersion of the estimate for each corresponding grid point. These two fields are paired using the same spatial grid structure, ensuring that each spatial grid point simultaneously possesses a precipitation intensity estimate and a variance value representing the uncertainty of that value. The resulting paired and encapsulated dataset constitutes the final probabilistic fusion analysis field. This probabilistic fusion analysis field, as a complete data output, includes the optimal estimate of the spatial distribution of precipitation and the spatial distribution of the confidence level of that estimate. The goal is to generate a probabilistic fusion analysis field that contains both intensity and uncertainty information.
[0038] Specifically, for example, when identifying extreme precipitation entities, we can consider not only whether the estimated precipitation intensity exceeds a threshold, but also its uncertainty. For a grid point whose intensity just exceeds the threshold but has high uncertainty, and a grid point that also exceeds the threshold but has low uncertainty, different weights can be assigned or different processing strategies can be adopted. This provides a data foundation for the entire method to shift from deterministic testing to probabilistic testing or risk decision-making. The design of linking values with confidence levels in the output ensures that the generated test report not only tells where the forecast is correct and where it is incorrect, but also indicates the degree of confidence in these judgments, enhancing the depth and application value of the test method.
[0039] S3. Based on the probabilistic fusion analysis of the grid precipitation intensity in the field, identify and extract all grid points that reach the predetermined extreme rainfall intensity threshold.
[0040] S3.1 Extract the grid precipitation intensity field from the probabilistic fusion analysis field. The grid precipitation intensity field contains the precipitation intensity estimate of each grid point in the probabilistic fusion analysis field.
[0041] Furthermore, the precipitation intensity estimates for all grid points are read and separated from the probabilistic fusion analysis field. These precipitation intensity estimates, arranged according to the same spatial grid, are organized into an independent two-dimensional scalar field with a clear geographic coordinate reference; this data is the gridded precipitation intensity field. The gridded precipitation intensity field faithfully preserves the optimal estimation information about the spatial distribution and intensity of precipitation from the probabilistic fusion analysis field, but does not contain uncertainty components, thus achieving the goal of extracting a pure precipitation intensity estimation field from the composite data product.
[0042] Specifically, the basis for identifying extreme events is shifted from discrete, sparse station observation data to high-resolution, continuously spatially covered gridded precipitation intensity fields. Traditional methods begin by checking if a single station's observation exceeds a threshold. Limited by the spatial distribution of stations, this approach is prone to missing strong precipitation events occurring between stations or misjudging an extreme event at a single station as a large-scale event. The gridded precipitation intensity field used in this approach is the best estimate of the actual spatial structure of precipitation, obtained through multi-source fusion. For example, in mountainous or oceanic areas with few stations, traditional methods may be completely unable to detect a strong precipitation event. However, based on the gridded precipitation intensity field, every grid point in the entire region can be scanned without omission. This breaks down the traditional technical barrier of using points to represent areas for event identification, enabling true spatial completeness in the detection of extreme precipitation events. This represents a crucial first step in the transition from station event verification to weather system entity verification.
[0043] S3.2. Compare the estimated precipitation intensity of each grid point in the gridded precipitation intensity field with a predetermined extreme rainfall intensity threshold.
[0044] Furthermore, a fixed intensity criterion for defining extreme precipitation is established, namely a predetermined extreme rainfall intensity threshold. Then, the gridded precipitation intensity field is traversed grid by grid, sequentially reading the estimated precipitation intensity value for each grid point. Each grid point's estimated precipitation intensity value is compared numerically with the predetermined extreme rainfall intensity threshold. This comparison process is a simple scalar judgment operation, determining whether the current grid point's estimated precipitation intensity value reaches or exceeds the predetermined extreme rainfall intensity threshold. The traversal and comparison operations cover all grid points in the gridded precipitation intensity field, completing the preliminary judgment of whether all grid points meet the extreme precipitation intensity standard.
[0045] Specifically, the criteria for extreme events are applied to a physically consistent continuous field, rather than discrete, potentially physically inconsistent observation points. Traditional methods apply thresholds at each station, ignoring the conditions between stations, and station observations may be affected by local micro-topography or instrument errors, leading to poor physical consistency in extreme event identification. However, threshold comparison based on gridded precipitation intensity fields ensures the application of a unified extreme standard within the same physical spatial framework. For example, a mesoscale convective system might generate a region of heavy precipitation reaching the extreme threshold. Traditional methods might identify it as a few isolated points simply because one or two stations are located within it, while this method can identify all grid points covering the entire heavy precipitation region. This ensures that the subsequently extracted candidate grid points are spatially physically coherent, truly reflecting the spatial extent of extreme precipitation weather systems. This breaks the limitation of traditional methods, which can only identify scattered and partial extreme events due to data source constraints, making the identification results closer to real meteorological entities.
[0046] S3.3 Mark all grid points in the grid precipitation intensity field whose estimated precipitation intensity is greater than or equal to the predetermined extreme rainfall intensity threshold, and output a list of grid point locations marked as candidate extreme precipitation grid points.
[0047] Furthermore, based on the comparison results, all grid points in the gridded precipitation intensity field that meet the conditions are identified. A Boolean-type label field with the same spatial dimension as the gridded precipitation intensity field is created, or a list of location indexes is directly generated. For each location in the gridded precipitation intensity field, if its estimated gridded precipitation intensity is greater than or equal to a predetermined extreme rainfall intensity threshold, the corresponding location is marked as true or its geographic coordinates are recorded in the list; otherwise, it is marked as false or not recorded. Finally, the spatial location information (such as latitude and longitude coordinates or grid indexes) of all grid points marked as true is summarized and output to form an ordered set containing the spatial location information of all grid points that meet the extreme intensity conditions, i.e., a grid point location list marked as candidate extreme precipitation grid points, thus achieving the goal of extracting the location information of all discrete grid points that meet the extreme conditions from a continuous intensity field.
[0048] Specifically, this is reflected in the close connection between the form of the output and its subsequent use. The output is a list of candidate grid locations, rather than a binary image or a simple count. This list serves as a direct, structured input for subsequent spatial clustering analysis to identify complete precipitation entities. Traditional methods output a list of stations that have reached a threshold; these stations may be spatially discrete and isolated, failing to directly reflect the two-dimensional structure of the system. The output grid location list, however, has precise spatial coordinates for each element, and these points naturally possess spatial continuity (due to their origin from a continuous grid field). For example, a zonal convective system can identify a large number of spatially adjacent candidate grid points and record them in the list. Clustering algorithms can then aggregate these points into a complete observation of extreme precipitation entities, providing a perfect data foundation. This breaks away from the traditional mindset of treating each extreme record as an independent event. By outputting a structured location list, it proactively prepares for higher-level spatial analysis (entity identification), representing a crucial data transformation step from extreme point detection to extreme system identification.
[0049] S4. Cluster spatially adjacent grid points into discrete observation extreme precipitation entities and output each observation extreme precipitation entity.
[0050] S4.1 Based on the grid point location list labeled as candidate extreme precipitation grid points, a spatial clustering algorithm is applied to group the grid points in the grid point location list labeled as candidate extreme precipitation grid points to obtain each independent connected group.
[0051] Furthermore, taking as input a list of grid locations labeled as candidate extreme precipitation grid points, containing the spatial coordinates of all discrete grid points reaching the intensity threshold, a spatial clustering algorithm based on connectivity rules, such as using an eight-neighbor or four-neighbor criterion, is applied to traverse and group the grid points in the list. Starting from an unvisited grid point, the algorithm recursively searches for and merges all other grid points that are spatially directly adjacent to it and also exist in the list of grid locations labeled as candidate extreme precipitation grid points, until no new neighboring candidate grid points can be added, forming an independent connected group. This process is repeated until all grid points in the list of grid locations labeled as candidate extreme precipitation grid points have been visited and classified, ultimately outputting a set of spatially separated independent connected groups. This achieves the goal of aggregating discrete extreme precipitation grid points into potential weather system objects with spatial continuity.
[0052] Specifically, the concept of connected component analysis from image processing and pattern recognition is transplanted to the field of meteorological verification, enabling the object-oriented identification of extreme precipitation events. Traditional methods treat each observation station reaching a threshold as an independent event, completely ignoring the spatial continuity and organization of heavy precipitation, and failing to reflect the essence of mesoscale convective systems as two-dimensional entities. The application of spatial clustering algorithms completely breaks down this barrier. It does not presuppose the number, shape, or size of entities, but objectively groups them based solely on the spatial adjacency relationships between grid points. For example, an elliptical squall line system may generate hundreds or thousands of spatially adjacent candidate grid points. Traditional methods would treat these as a large number of isolated points, but this approach can automatically identify them as a complete independent connected group, corresponding to a complete squall line entity. This elevates the basic unit of verification from stations to weather system entities, laying the foundation for subsequent matching and evaluation based on entity features. This is a key step in realizing the paradigm shift from point-to-point verification to object-to-object verification.
[0053] S4.2 Assign an identifier to each independent connected group of extreme precipitation observation entities, and determine the spatial center position of the extreme precipitation observation entities based on the position coordinates of all grid points within the independent connected group.
[0054] Furthermore, each independent connected group is assigned a unique and traceable identifier as an identifier for the observed extreme precipitation entity. For each identified independent connected group, the location coordinates of all grid points within the group are read. Based on these location coordinates, the spatial center location of the observed extreme precipitation entity is determined by calculating the arithmetic mean of the latitude and longitude of all grid points, or by calculating the geometric centroid of the polygonal region formed by the independent connected group. This spatial center location is a single, representative geographic coordinate point used to characterize the spatial average location of the entire observed extreme precipitation entity, thus achieving the goal of assigning an identifier to each identified precipitation entity and determining its representative central location.
[0055] Specifically, the identified weather system is endowed with two core attributes: identity and location. This abstracts a complex spatial aggregate composed of numerous grid points into an operable object with a clearly identified identifier and a single location vector. In traditional methods, the location of an event is simply the location of a single station, failing to represent the entire spatial system. However, calculating the spatial center location based on all grid points within an independent connected group yields a representative point that comprehensively reflects the spatial distribution of the entity. For example, for a curved bow-shaped echo, its geometric centroid might be located in the middle to rear of the echo, which is more representative of the system's overall location than using the location of any single station included within it. This abstraction allows for unified and reasonable operations on the entire entity (e.g., by defining a search domain), breaking away from the fragmented nature of each station event in traditional methods and achieving a convergence of the verification target from scattered points to a holistic whole.
[0056] S4.3. Based on the temporal attributes of all grid points within an independent connected group, determine the occurrence time of observed extreme precipitation entities.
[0057] Furthermore, the observed extreme precipitation entities originate from the analysis of the probabilistic fusion analysis field at a single moment. Therefore, all grid points constituting each independent connected group share a common temporal attribute, namely the effective time corresponding to the field. This effective time is assigned to the corresponding observed extreme precipitation entity as its occurrence time. If there are theoretically minor differences in the temporal attributes among the grid points within an independent connected group, a unified representative time can be determined by taking the mode or average. Each observed extreme precipitation entity is associated with a specific occurrence time, thus achieving the goal of determining the occurrence time of each observed extreme precipitation entity.
[0058] Specifically, the processing of time attributes shifts from point time to entity time, assigning a unified time label to the entire two-dimensional weather system entity. In traditional methods, the time of an event at each station is simply the observation time of that station. When multiple stations are identified as part of the same system, their times may not be strictly synchronized, which can lead to confusion in subsequent time matching. This approach treats an independent connected group as a whole and assigns it a single occurrence time. For example, for a mesoscale convective complex analyzed at a specific time, regardless of its coverage area, it is considered a single entity occurring at that time. This processing simplifies the logic of subsequent time matching, ensuring consistent forecast retrieval for the entire entity within the same time window. It avoids the matching ambiguities that may arise from the inconsistency of times at multiple related stations in traditional methods, enhancing the clarity and consistency of matching in the time dimension.
[0059] S4.4. Based on the precipitation intensity values of all grid points within the independent connected group, determine the maximum precipitation intensity of the observed extreme precipitation entity.
[0060] Furthermore, after identifying the independent connected groups corresponding to the observed extreme precipitation entities, it is necessary to extract the precipitation intensity estimate of each grid point within that independent connected group from the gridded precipitation intensity field of the probabilistic fusion analysis field. By traversing and comparing all these grid point precipitation intensity estimates, the grid point with the largest value is identified. This maximum value is then formally recorded as the maximum precipitation intensity of the current observed extreme precipitation entity. This attribute characterizes the peak precipitation intensity reached within the spatial range of the observed extreme precipitation entity, thus achieving the goal of quantifying the strongest precipitation intensity within each observed extreme precipitation entity.
[0061] Specifically, the intensity characteristics of an entity are represented by the maximum intensity of all grid points within the entity, rather than relying solely on records from a single station. Traditional methods often define the intensity of an event solely by the station's observations, failing to reflect potentially stronger precipitation within its affected area. However, maximum extraction based on a high-resolution grid field more accurately captures the core intensity of weather systems. For example, the strongest precipitation core of a strong thunderstorm cell might lie between two stations; traditional methods underestimate its intensity, while this method captures this maximum value through grid analysis. This provides a more accurate true benchmark for assessing a model's ability to predict the core intensity of extreme precipitation, overcoming the limitations of traditional tests that generalize from a single point in intensity assessment, and making the evaluation of forecast intensity performance more precise and meteorologically significant.
[0062] S4.5. Based on the spatial distribution of all grid points within the independent connected group, determine the coverage area of the observed extreme precipitation entity, and output the observed extreme precipitation entity and its corresponding entity feature attributes for each observed extreme precipitation entity.
[0063] Furthermore, for each observed extreme precipitation entity and its corresponding independent connected group, the spatial coordinates of all grid points within the group are first obtained. Based on these discrete but spatially adjacent grid point coordinates, the outer contour of the continuous spatial region occupied by the independent connected group can be delineated. The coverage area of the observed extreme precipitation entity is determined by calculating the area enclosed by this outer contour, or by counting the total number of grid points contained in the independent connected group and multiplying it by the physical area represented by a single grid point. This coverage area is a scalar value characterizing the spatial influence range of the observed extreme precipitation entity. After completing all feature calculations, the entity feature attributes such as the identifier, spatial center location, occurrence time, maximum precipitation intensity, and coverage area of each observed extreme precipitation entity are packaged and associated, ultimately outputting structured information on the observed extreme precipitation entities.
[0064] Specifically, an entity feature vector was constructed for object-oriented verification, abstracting a space weather system into a complete set of quantifiable and comparable feature attributes, with coverage area being one of the key dimensions. Traditional verification methods provide no information about the spatial extent of extreme precipitation, while coverage area is a crucial physical quantity describing the importance and potential impact of mesoscale systems. For example, a large stratiform cloud precipitation area and a small but intense convective cell are drastically different in terms of meteorology and disaster prevention. Outputting entity features including coverage area allows subsequent evaluations to not only determine whether the model predicted extreme precipitation but also to further assess the reasonableness of the predicted extent. This opens up possibilities for multi-dimensional and more refined model performance diagnosis (such as evaluating system-scale forecasting capabilities), greatly enriching the information content and diagnostic value of the verification output.
[0065] S5. For each observed extreme precipitation entity, extract the atmospheric lower-level environmental flow field information corresponding to the time of occurrence to determine the prevailing wind direction, and set the parameters of the elliptical spatial search domain with the major axis direction parallel to the wind direction for the observed extreme precipitation entity based on the prevailing wind direction.
[0066] S5.1 For each observed extreme precipitation entity, based on the occurrence time of the observed extreme precipitation entity, extract the corresponding time and the lower atmospheric environmental flow field information of the core area with the center location of the observed extreme precipitation entity from the external atmospheric reanalysis data.
[0067] Furthermore, for each observed extreme precipitation entity, a specific time point is determined based on its occurrence time attribute. Based on this time point, an external atmospheric reanalysis database is accessed and retrieved. Grid data describing the lower atmospheric circulation state that perfectly matches the occurrence time of the observed extreme precipitation entity is extracted from the database. The extracted lower atmospheric environmental flow field information typically corresponds to the horizontal wind vector field on the 850 hPa or 925 hPa isobaric surface. The spatial range of the extraction operation is a pre-defined rectangular or circular region, geometrically centered on the spatial center of the observed extreme precipitation entity, extending outwards to cover the environmental field of the weather system containing the entity. This yields a three-dimensional data block of lower atmospheric environmental flow field information containing complete horizontal wind vectors, with the center of the observed extreme precipitation entity as the core region. This achieves the goal of obtaining the surrounding environmental wind field data at the time of occurrence for each observed extreme precipitation entity.
[0068] Specifically, this approach actively and dynamically links weather system entities with their large-scale environmental context, upgrading the matching rules from static geometric definitions to dynamic definitions bound to current weather dynamics. Traditional methods using fixed circular search domains implicitly assume that forecast errors have the same statistical characteristics in all directions and do not change with specific weather processes. By extracting specific, real-time environmental flow field information for each observed extreme precipitation entity, a physical basis is provided for the subsequent customized search domain shape. For example, a squall line system generated under a strong southwest low-level jet stream is primarily guided by the jet stream in its overall movement and propagation direction, and the predicted location error is also mainly along the jet stream direction. Obtaining the low-level flow field at this specific moment provides key physical information that dominates the system's movement, allowing the parameters of the matching rules to be derived from the data in real time, rather than being pre-coded. This enables the testing method to perceive and adapt to different weather conditions, providing a more refined and physical solution to the problem of spatial double penalties.
[0069] S5.2. By performing regional vector averaging of the low-level atmospheric environmental flow field information, the prevailing wind direction of the environment in which the observed extreme precipitation entity is located can be obtained.
[0070] Furthermore, the extracted lower atmospheric environmental flow field information, with the observed extreme precipitation entity's center location as the core area, is used as input. This input data is a two-dimensional vector field containing U (east-west) and V (north-south) components. The arithmetic mean of the U and V components at all grid points within the core area is calculated to obtain the regional average U and V components. Using these two average components, through vector synthesis and azimuth calculation, a single wind direction angle value representing the overall airflow direction of the region is obtained. This wind direction angle value is defined as the dominant wind direction of the environment where the observed extreme precipitation entity is located. The dominant wind direction is a scalar used to characterize the main environmental airflow direction affecting the movement or propagation of the entity, thus achieving the goal of extracting a representative dominant wind direction from local wind field data.
[0071] Specifically, a representative dominant wind direction is used to summarize the main forcing effect of the entire environmental field on the movement of precipitation systems, thus providing a concise and crucial directional parameter for defining the asymmetric search domain. Traditional circular search domains do not require directional parameters, but this also sacrifices the opportunity to optimize and validate rules using environmental field information. The dominant wind direction obtained by regional vector averaging is physically approximated by the geostrophic wind or steering airflow direction in that region, and is the dominant factor determining the translation of mesoscale convective systems. For example, for multiple convective cells influenced by a unified southerly airflow, even if they are spatially dispersed, the calculated dominant wind direction is consistent. This indicates that their overall predicted displacement error direction has a commonality. Introducing fundamental principles of fluid mechanics into the validation rules allows the major axis of the subsequently defined elliptical search domain to align with the dominant direction of atmospheric motion, thus allowing for a more lenient error tolerance in the airflow direction and maintaining relatively strict standards in the vertical airflow direction. This is much more physically reasonable than the traditional circular search domain, significantly enhancing the scientific content of fuzzy matching.
[0072] S5.3 For each observed extreme precipitation entity, the core geometric parameters of the elliptical spatial search domain are defined using the prevailing wind direction.
[0073] Furthermore, for each observed extreme precipitation entity, the core geometric parameters of the elliptical spatial search domain are defined using the prevailing wind direction. The definition process begins by setting the center point of the elliptical spatial search domain as the spatial center of the observed extreme precipitation entity. The major axis of the elliptical spatial search domain is then set to be parallel to the prevailing wind direction. A major axis radius and a minor axis radius are defined for the elliptical spatial search domain, with the major axis radius being greater than the minor axis radius. The specific values of these two lengths are predetermined based on statistical analysis of historical forecast errors of the model or typical scales of the convective system. A complete set of parameters is generated for each observed extreme precipitation entity, including the center point coordinates, major axis radius, minor axis radius, and major axis direction angle. This set of parameters collectively defines the parameters of the elliptical spatial search domain corresponding to that entity, with the major axis direction parallel to the wind direction. This achieves the goal of customizing the geometric parameters of the elliptical spatial search domain for each observed extreme precipitation entity.
[0074] Specifically, the permissible deviation is no longer uniformly distributed spatially, but rather aligned with the dominant direction of atmospheric motion. Traditional circular search domains assign the same tolerance along the system's direction of movement and vertically, which is unreasonable. Elliptical search domains, with their major axis along the dominant wind direction, mean setting a larger tolerance radius in the direction where the system is most likely to shift (which is also where forecast errors are typically larger), and a smaller tolerance radius in the vertical direction. For example, for convective systems guided by westerly winds, forecast positions are more likely to shift east-west rather than suddenly jump north-south. Elliptical search domains allow for this reasonable, directional shift while constraining unreasonable vertical displacements. This design breaks down the barrier between spatial tolerance rules and atmospheric dynamics, making hit determination not only based on geometric inclusion relationships but also including a judgment on the physical reasonableness of forecast errors. This allows for a fairer distinction between acceptable forecast displacements and complete forecast failures, greatly improving the accuracy and fairness of verification results in identifying the model's true forecasting capabilities.
[0075] S6. For each observed extreme precipitation entity, a hit determination is made based on the search domain parameters in the corresponding elliptical space and the preset time window to retrieve the forecast field of the mode.
[0076] S6.1 For each observed extreme precipitation entity, the time of occurrence of the observed extreme precipitation entity and a preset time window are used to determine the time search range.
[0077] Furthermore, for each observed extreme precipitation entity, the occurrence time of the observed extreme precipitation entity is used as the time reference point. A predefined time window value, representing the allowable time deviation duration, is added before and after the time reference point. The start time point is obtained by subtracting the time window value from the occurrence time of the observed extreme precipitation entity, and the end time point is obtained by adding the time window value to the occurrence time of the observed extreme precipitation entity. The start and end times together constitute a closed time interval, which is the time search range for that observed extreme precipitation entity. The time search range is a continuous time period, covering multiple whole-hour moments extending forward and backward from the occurrence time of the observed extreme precipitation entity, thus achieving the goal of defining a search time period with an allowable forecast time deviation for each observed extreme precipitation entity.
[0078] Specifically, a time-tolerance dimension is introduced for fuzzy matching, acknowledging and allowing reasonable phase deviations in precipitation triggering and peak times in numerical models, rather than requiring absolute time synchronization. Traditional rigorous hourly verification incorrectly decomposes these phase deviations into multiple false alarms and missed alarms, failing to identify their essence as time shifts within the same weather system. By defining a symmetrical time search range, the focus of verification shifts from whether a prediction is made at the precise moment to whether it is made within a reasonable timeframe. For example, a strong convective system that matures at 3 PM might be predicted by the model to reach its peak at 2 PM or 4 PM. The time search range allows for searching forecast signals within several hours before and after this point; if the model correctly predicts extreme precipitation at any of these times, it is considered a successful capture of the system with a time deviation. This breaks down the verification barrier of absolute time lock and solves the problem of traditional methods misjudging reasonable time phase errors caused by model spin-up and differences in development speed as complete failures, enabling assessments to more fairly reflect the model's forecasting ability for the life cycle of weather systems.
[0079] S6.2 For each observed extreme precipitation entity, a four-dimensional spatiotemporal search domain is defined by combining the time search range and the parameters of the elliptical spatial search domain.
[0080] Furthermore, the temporal search range is combined with the parameters of the elliptical spatial search domain determined by the observed extreme precipitation entity. This combined operation logically constructs a four-dimensional search volume. In the temporal dimension, this search volume is defined by the start and end times of the temporal search range. In the spatial dimension, it is defined by the parameters of the elliptical spatial search domain, which define a planar elliptical region centered on the spatial center of the observed extreme precipitation entity, possessing specific major and minor axes and orientations. This four-dimensional spatiotemporal search domain is a continuous spatiotemporal region that clearly defines the range of data to be retrieved in the model forecast field data. Specifically, it requires examining all grid points within the elliptical spatial region at every time point within the temporal search range. This achieves the goal of integrating the allowable spatiotemporal biases for each observed extreme precipitation entity and defining a clear four-dimensional search space.
[0081] Specifically, the matching rules for the test are expanded from a simple spatial circle to a physically meaningful spatiotemporal conduit. While traditional methods may also set separate time and spatial windows, they do not conceptually and operationally integrate them into a unified search entity. The defined spatiotemporal search domain is an anisotropic ellipse in space and a sliding window in time; the combined search conduit more closely reflects the spatiotemporal evolution trajectory of weather systems. For example, an eastward-moving convective system's trajectory in space and time is a diagonal line. The defined search domain allows forecasts to be advanced or delayed in time, while also allowing for spatial offsets (primarily along the direction of movement). As long as the strong forecast signal falls within this spatiotemporal conduit, it is considered reasonable. This design breaks down the barriers of independent handling and separate penalties for time and spatial errors in traditional tests, achieving true spatiotemporal collaborative fault tolerance. This allows the matching rules to more realistically simulate and allow for reasonable combinations of spatiotemporal errors in forecasts, greatly enhancing the physical consistency of the test and the comprehensiveness of forecast performance evaluation.
[0082] S6.3 Within the four-dimensional spatiotemporal search domain, retrieve all grid point precipitation forecast values in the model forecast field, check all retrieved grid point precipitation forecast values, and determine whether at least one grid point precipitation forecast value reaches or exceeds the predetermined extreme rainfall intensity threshold. If there is a grid point precipitation forecast value that meets the conditions, then the forecast for the observed extreme precipitation entity is determined as a hit event; otherwise, it is determined as a missed event.
[0083] Furthermore, for each observed extreme precipitation entity, data retrieval is performed on the model forecast field within its corresponding four-dimensional spatiotemporal search domain. The retrieval operation iterates through each forecast time within the time search range, and for each time, precipitation forecast values for all grid points falling within the elliptical spatial region are extracted. All these retrieved grid point precipitation forecast values are collected and organized to form a candidate forecast value list. Subsequently, each grid point precipitation forecast value in the candidate forecast value list is compared with a predetermined extreme rainfall intensity threshold. The check process determines whether at least one grid point precipitation forecast value is greater than or equal to the predetermined extreme rainfall intensity threshold. If the check result is yes, the forecast for this observed extreme precipitation entity is determined as a hit event and recorded; if the check result is no grid point precipitation forecast value reaching the threshold, the forecast is determined as a missed event and recorded, thus completing the final forecast signal retrieval and binary classification hit determination within a personalized spatiotemporal tolerance range.
[0084] Specifically, the determination of whether a forecast signal exists is conducted within a spatiotemporal channel tailored to each weather system entity, taking into account the direction of environmental steering airflow and reasonable life-history biases. Traditional determinations involve a simple presence check within a fixed-size cylindrical spatiotemporal region centered on a weather station. This region is dynamic, anisotropic, and tied to the physical characteristics of the weather process. For example, for a rapidly moving frontal precipitation band, its spatiotemporal search channel is long and slender, tilted along the airflow direction. This requires that the predicted heavy precipitation area also be organized into a similar shape and appear in a similar spatiotemporal location to be considered a hit. This effectively filters out random false alarms that, while meeting the intensity criteria, have no correlation between their location and timing. This integrated judgment rule breaks down the barriers of rigid rules and disconnection from physical processes in traditional methods. It means that hitting the result not only represents a forecast of heavy precipitation, but also implies that the corresponding weather system was forecast under the correct atmospheric environment and within a reasonable time and space range. This significantly improves the meteorological significance of the hit judgment result and the accuracy and depth of the diagnosis of the model's true forecasting ability.
[0085] S7. Statistically analyze the hit determination results and generate a comprehensive test report.
[0086] S7.1 Statistically analyze the hit determination results of all observed extreme precipitation entities to obtain the total number of hit events and the total number of missed events.
[0087] Furthermore, the input is the recorded results obtained after performing hit determination on all observed extreme precipitation entities. These records clearly indicate whether each observed extreme precipitation entity was determined as a hit event or a missed event. These records are iterated and counted. The total number of hit events is the sum of the counts for all observed extreme precipitation entities marked as hit events. Simultaneously, the total number of missed events is the sum of the counts for all observed extreme precipitation entities marked as missed events. The total number of hit events and the total number of missed events are two fundamental summary statistics, completing the goal of summarizing and counting the determination results of all tested cases to obtain the core statistics.
[0088] Specifically, the statistical objects are observed extreme precipitation entities. These entities originate from high-resolution probabilistic fusion analysis fields and are objectively identified through spatial clustering, representing more complete and realistic weather system objects. Therefore, the statistics of total hit events and total missed events are built upon a more robust and reliable event sample library. For example, a mesoscale convective system containing multiple strong cores might be identified as an uncertain number of station events due to uneven station distribution using traditional methods, leading to unstable counts. However, this method identifies it as one or a few entities, resulting in more stable and meteorologically significant counts. This allows subsequent scoring calculations to avoid noise caused by station randomness, and the count results better reflect the model's forecasting performance for real organized weather systems, improving the reliability of the statistical basis.
[0089] S7.2 Calculate the hit rate score based on the total number of hit events and the total number of observed extreme precipitation entities.
[0090] Furthermore, the total number of hit events is used as the number of predicted extreme rainfall intensities hit in the formula, while the total number of observed extreme precipitation entities (i.e., the sum of the total number of hit events and the total number of missed events) is used as the number of observed extreme rainfall intensities hit in the formula. Dividing the number of predicted extreme rainfall intensities hit (f) by the number of observed extreme rainfall intensities yields a small value between 0 and 1. Multiplying this small value by 100 and converting it to a percentage value, this percentage value is the hit rate score. The hit rate score is a single scalar value that directly reflects the proportion of observed extreme precipitation entities that the model forecast successfully captured in the entire test sample, thus achieving the goal of calculating the hit rate score, a core evaluation indicator, using basic statistics. Specifically, the hit rate calculated based on the total number of spatially significant observed extreme precipitation entities, with a more explicit physical meaning, directly answers the question of what proportion of these meteorologically significant heavy precipitation systems the model reported. For example, even if two regions have different station densities, as long as the entities identified by the method are complete weather systems, the calculated hit rates are comparable at the same conceptual level, both reflecting the forecasting capability for the system. This design frees the hit rate score from excessive reliance on observation network density, providing a purer performance indicator that better reflects the model's physical forecasting capability, and enhancing the universality and interpretability of the evaluation results.
[0091] The accuracy rating expression is: ; in, Rate the hit percentage. To predict the number of times extreme rainfall intensity will be hit, This refers to the number of times extreme rainfall intensities were observed.
[0092] S7.3 Using the list of hit events and the list of missed events, combined with the corresponding elliptical spatial search domain parameters, draw a geospatial distribution map showing the distribution of hit events, missed events, and spatial search domains. Integrate the hit rate score and the geospatial distribution map to generate a comprehensive inspection report.
[0093] Furthermore, the lists of all observed extreme precipitation entities identified as hit events and all those identified as missed events generated during the verification process are aggregated and compiled. Simultaneously, the corresponding defined elliptical spatial search domain parameters for these entities are obtained. Using geographic information mapping tools, the spatial center positions of all observed extreme precipitation entities for hit events and all missed events are marked on a base map using different graphic symbols (such as stars and circles). For typical cases requiring special emphasis, a semi-transparent elliptical region outline can be drawn centered on its spatial center position using the corresponding elliptical spatial search domain parameters, visually demonstrating the personalized spatial search range set for that entity. By overlaying the aforementioned spatial elements with base map information such as geographical boundaries and topography, a comprehensive geospatial distribution map is generated. This geospatial distribution map, along with the hit rate score and basic verification information (such as region, time period, and threshold), is then compiled and integrated to form a comprehensive verification report that combines text and graphics, including quantitative scoring and qualitative spatial distribution analysis. This achieves the goal of combining numerical statistical results with spatial distribution information to generate intuitive and comprehensive final verification results.
[0094] Specifically, the generated geospatial distribution map, especially with the overlay of elliptical spatial search domains, provides powerful diagnostic capabilities. For example, the map visually shows whether missed events are concentrated downwind of specific terrain features or scattered; it allows observation of whether the elliptical search domains of hit events generally align their longer axes with the direction of weather system movement, verifying the effectiveness of physical guidance rules; if a large number of missed events are found to be close to the edge of their elliptical search domains, it may indicate that the current search domain radius setting needs adjustment. This visualized report, combined with matching rule information, breaks down the barriers of traditional black-box verification results and the inability to deeply diagnose specific model shortcomings. It transforms evaluation results into spatially oriented diagnostic tools that forecasters and model developers can directly use, greatly enhancing the practical application value and in-depth analytical capabilities of verification methods.
[0095] This embodiment also provides a computer device suitable for verifying methods for extreme rainfall intensity forecasts, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the verification method for extreme rainfall intensity forecasts as proposed in the above embodiment.
[0096] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0097] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the verification method for extreme rainfall intensity forecasting as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0098] In summary, this invention generates a high-resolution probabilistic fusion analysis field by fusing site observations and radar data to construct a more reliable verification benchmark. Based on the environmental flow field, it sets an elliptical spatial search domain with its major axis parallel to the wind direction for each identified precipitation entity, thereby establishing a physically more reasonable fuzzy matching rule. This solves the problems of inaccurate verification benchmarks caused by the reliance on sparse site data in traditional methods, and the inability to reasonably tolerate reasonable forecast displacements due to the use of a fixed circular search domain. This enables a more scientific, stable, and fair evaluation of the extreme precipitation forecasting capability of numerical models.
[0099] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A verification method for extreme rainfall intensity forecasts, characterized in that: This includes acquiring model forecast fields, station observation data, and radar quantitative precipitation estimation fields; By fusing site observation data and radar quantitative precipitation estimation fields, a probabilistic fusion analysis field is generated. Based on the probabilistic fusion analysis of grid precipitation intensity, all grid points that reach the predetermined extreme precipitation intensity threshold are identified and extracted. Spatially adjacent grid points are clustered into discrete observation extreme precipitation entities, and each observation extreme precipitation entity is output. For each observed extreme precipitation entity, the atmospheric lower-level environmental flow field information corresponding to the time of occurrence is extracted to determine the prevailing wind direction. Based on the prevailing wind direction, elliptical spatial search domain parameters with the major axis direction parallel to the wind direction are set for the observed extreme precipitation entity. For each observed extreme precipitation entity, a hit determination is made based on the search parameters of the corresponding elliptical space search domain and the search mode forecast field within a preset time window. The hit determination results are statistically analyzed to generate a comprehensive test report.
2. The verification method for extreme rainfall intensity forecasting as described in claim 1, characterized in that: Acquiring model forecast fields, station observation data, and radar quantitative precipitation estimation fields includes the following steps: Acquire gridded precipitation forecast fields, precipitation observation data from ground automatic weather stations, and gridded data for quantitative precipitation estimation from weather radar, and unify the timestamps; The time-uniform gridded precipitation forecast field, precipitation observation data from automatic weather stations, and quantitative precipitation estimation gridded data from weather radar are aligned to a common spatial grid through resampling and location recording. The gridded precipitation forecast field based on spatiotemporal alignment, precipitation observation data from automatic weather stations, and gridded data for quantitative precipitation estimation from weather radar are used as the model forecast field, the station observation dataset, and the radar quantitative precipitation estimation field, respectively.
3. The verification method for extreme rainfall intensity forecasting as described in claim 2, characterized in that: By fusing site observation data and radar quantitative precipitation estimation fields, a probabilistic fusion analysis field is generated, including the following steps: Based on the site observation dataset and radar quantitative precipitation estimation, the optimal interpolation algorithm is applied to fuse the site observation dataset and radar quantitative precipitation estimation field to generate a preliminary fusion analysis field. Based on the difference between the site observation dataset and the preliminary fusion analysis field, the uncertainty of the precipitation intensity estimate at each grid point in the preliminary fusion analysis field is calculated. By combining the grid point precipitation intensity estimates in the preliminary fusion analysis field with the corresponding uncertainty information, a probabilistic fusion analysis field is output.
4. The verification method for extreme rainfall intensity forecasting as described in claim 3, characterized in that: Based on the gridded precipitation intensity in the probabilistic fusion analysis field, all gridded points that reach the predetermined extreme rainfall intensity threshold are identified and extracted, including the following steps: The gridded precipitation intensity field is extracted from the probabilistic fusion analysis field. The gridded precipitation intensity field contains the precipitation intensity estimate for each grid point in the probabilistic fusion analysis field. The estimated precipitation intensity at each grid point in the gridded precipitation intensity field is compared with a predetermined extreme rainfall intensity threshold. Mark all grid points in the grid precipitation intensity field whose estimated precipitation intensity is greater than or equal to a predetermined extreme precipitation intensity threshold, and output a list of grid point locations marked as candidate extreme precipitation grid points.
5. The verification method for extreme rainfall intensity forecasting as described in claim 4, characterized in that: Clustering spatially adjacent grid points into discrete observed extreme precipitation entities and outputting each observed extreme precipitation entity includes the following steps: Based on the grid location list labeled as candidate extreme precipitation grid points, a spatial clustering algorithm is applied to group the grid points in the grid location list labeled as candidate extreme precipitation grid points to obtain each independent connected group; Each independent connected group is assigned an identifier for an observed extreme precipitation entity. Based on the position coordinates of all grid points within the independent connected group, the spatial center position of the observed extreme precipitation entity is determined. Based on the temporal attributes of all grid points within an independent connected group, the occurrence time of observed extreme precipitation entities is determined; Based on the precipitation intensity values of all grid points within the independent connected group, the maximum precipitation intensity of the observed extreme precipitation entity is determined; Based on the spatial distribution of all grid points within an independent connected group, the coverage area of the observed extreme precipitation entity is determined, and the observed extreme precipitation entity and its corresponding entity feature attributes are output.
6. The verification method for extreme rainfall intensity forecasting as described in claim 5, characterized in that: For each observed extreme precipitation entity, the prevailing wind direction is determined by extracting the lower atmospheric environmental flow field information at the time of occurrence. Based on the prevailing wind direction, an elliptical spatial search domain parameter is set for the observed extreme precipitation entity with its major axis parallel to the wind direction. This includes the following steps: For each observed extreme precipitation entity, based on the occurrence time of the observed extreme precipitation entity, the corresponding time and the lower atmospheric environmental flow field information of the core area with the center location of the observed extreme precipitation entity are extracted from the external atmospheric reanalysis data; By performing regional vector averaging of the lower atmospheric environmental flow field information, the prevailing wind direction of the environment in which the observed extreme precipitation entity is located can be obtained; For each observed extreme precipitation entity, the core geometric parameters of the elliptical spatial search domain are defined using the prevailing wind direction.
7. The verification method for extreme rainfall intensity forecasting as described in claim 6, characterized in that: For each observed extreme precipitation entity, a hit determination is made based on the search parameters in the corresponding elliptical spatial domain and the retrieved model forecast field within a preset time window, including the following steps: For each observed extreme precipitation entity, the time of occurrence of the observed extreme precipitation entity and a preset time window are used to determine the time search range; For each observed extreme precipitation entity, a four-dimensional spatiotemporal search domain is defined by combining the temporal search range and the parameters of the elliptical spatial search domain. Within the four-dimensional spatiotemporal search domain, all grid point precipitation forecast values in the model forecast field are retrieved. All retrieved grid point precipitation forecast values are checked to determine whether at least one grid point precipitation forecast value reaches or exceeds the predetermined extreme rainfall intensity threshold. If a grid point precipitation forecast value that meets the conditions exists, the forecast for the observed extreme precipitation entity is determined as a hit event; otherwise, it is determined as a missed event.
8. The verification method for extreme rainfall intensity forecasting as described in claim 7, characterized in that, The hit determination results are statistically analyzed to generate a comprehensive test report, including the following steps: The hit detection results of all observed extreme precipitation entities are statistically analyzed to obtain the total number of hit events and the total number of missed events; The hit rate score is calculated based on the total number of hit events and the total number of observed extreme precipitation entities. Using the list of hit events and the list of missed events, combined with the corresponding elliptical spatial search domain parameters, a geospatial distribution map is drawn, showing the distribution of hit events, missed events, and spatial search domains. The hit rate score and the geospatial distribution map are integrated to generate a comprehensive inspection report.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the verification method for extreme rainfall intensity forecasting as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the verification method for extreme rainfall intensity forecasting as described in any one of claims 1 to 8.