A method of identifying precursory signatures of the start and end of an areal rainfall event

CN122734473APending Publication Date: 2026-09-11UNIV OF JINAN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610716116.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

现有雨/不雨判别多依赖降雨观测阈值、雷达回波反演或数值天气预报输出,难以在降雨存在系统偏差(如气候模式、未来情景外推)或仅有非降雨气象要素(温度、湿度、风、气压等)时稳定判定降雨事件边界,难以表达降雨过程的空间覆盖性与区域内部异质性;同时,深度学习端到端降雨分类虽可提升精度,但普遍缺乏可解释性,且在空间异质性强的区域容易出现训练样本代表性不足与泛化不稳的问题,导致在实际业务中稳定性与可信度受限

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122734473A_ABST
    Figure CN122734473A_ABST
Patent Text Reader

Abstract

This invention discloses a method for identifying the precursory features of the start and end of regional rainfall events. First, rainfall events are segmented in historical periods, and start / end labels are generated. Then, a grid is divided into several rain phenological layers using K-Means, and balanced sampling is performed within each layer. On the sampled grid, a window period of W days prior to the start / end date of the target rainfall event is set. By introducing a multivariate collaborative time series shape interpreter and traditional rate of change and trend statistics methods, multivariate feature engineering is constructed within the window period. Finally, APLR is used to learn the start and end discrimination rules, outputting daily start / end probabilities. This invention can stably identify rainfall event boundaries and provide interpretable discrimination criteria based on rainfall-related meteorological elements, possessing generalization, reproducibility, and engineering feasibility. It is applicable to operational scenarios such as daily deviation correction and data fusion for satellite rainfall products and multi-model rainfall products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent hydrological and meteorological identification, and in particular relates to a method for identifying the precursory features of the start and end of regional rainfall events. Background Technology

[0002] Accurate identification of regional rainfall events is a fundamental step in rainfall bias correction, data fusion, flood risk assessment, and hydrological forecasting. Current methods for determining rainfall / no-rain events largely rely on rainfall observation thresholds, radar echo inversion, or numerical weather prediction outputs. These methods struggle to reliably determine rainfall event boundaries when there are systematic biases in rainfall (such as climate models or future scenario extrapolation) or when only non-rainfall meteorological elements (temperature, humidity, wind, air pressure, etc.) are available. Furthermore, they fail to accurately represent the spatial coverage and regional heterogeneity of rainfall processes. While deep learning-based end-to-end rainfall classification can improve accuracy, it generally lacks interpretability and is prone to insufficient representativeness of training samples and unstable generalization in areas with strong spatial heterogeneity, thus limiting stability and reliability in practical applications. Summary of the Invention

[0003] To address the aforementioned problems, embodiments of the present invention propose a method for identifying the precursory features of the start and end of regional rainfall events.

[0004] The method for identifying the precursory features of the start and end of regional rainfall events according to the present invention includes the following steps:

[0005] S1. Acquire multivariate daily raster meteorological data and align other variables with time based on the rainfall time axis;

[0006] S2. Obtain or generate a spatial partition mask;

[0007] S3. Construct regional rainfall event labels;

[0008] S4. Construct a time window sequence;

[0009] S5. Parallel extraction of interpretable feature engineering representations;

[0010] S6. Parallel extraction of SVP-T depth characterization;

[0011] S7. Feature alignment and fusion;

[0012] Training and evaluation of the S8 APLR model;

[0013] S9. Application of model extrapolation.

[0014] The multivariate daily raster meteorological data in S1 includes rainfall data and one or more non-rainfall meteorological element data such as temperature, humidity, wind speed, and air pressure; the time alignment of other variables based on the rainfall time axis refers to the spatial consistency processing of each variable based on a unified spatial grid and the time alignment of each variable based on a unified date sequence, forming a multivariate daily raster data set under the same calendar index.

[0015] In step S2, obtaining or generating a spatial partitioning mask refers to dividing the study area into several weather layers or spatial sub-regions and generating spatial partitioning results. The spatial partitioning includes the following steps:

[0016] S201. During the training period, a spatial partition feature vector is constructed for each grid cell, and the spatial partition feature vector is used to characterize the rainfall process characteristics or rainfall climatology characteristics of the grid cell.

[0017] S202. Standardize the spatial partition feature vectors to obtain spatial representations;

[0018] S203. In the low-dimensional spatial representation, a clustering algorithm is used to divide the raster into a preset number of K spatial partitions. K-Means clustering algorithm is preferred to obtain K spaces, and each raster is assigned a corresponding partition number.

[0019] S204. The partition numbers are backfilled according to the grid spatial position to generate a spatial partition mask, wherein the spatial partition mask records the partition number corresponding to each grid in the form of a two-dimensional array.

[0020] The construction of regional rainfall event labels in S3 refers to generating rainfall start date labels and end date labels and constructing sequence samples aligned with W days prior to the target date.

[0021] The construction of the time window sequence in S4 refers to constructing a time window sequence X(r,t)=[x(tW),…,x(t-1)] for each (r,t) sample, which is W days before the target date. Here, x is a multivariate vector to ensure that the features only use historical information, r is a sub-region, and t is the date.

[0022] The parallel extraction of interpretable feature engineering representations in S5, used to construct interpretable statistical features from a W-day window sequence before the target date, includes the following steps:

[0023] S501. For each non-rainfall meteorological element variable in the window sequence, calculate the last value feature of the window, the slope feature of the trend within the window or the most recent k days, and the rise feature of the increase of the recent k-day mean relative to the previous m-day mean, and further calculate the difference feature and the relative rate of change feature.

[0024] S502. Can perform moving average or exponential smoothing on window sequences to suppress noise disturbances;

[0025] S503. Use the features in S501 as an interpretable input feature set for determining the start / end date of rainfall, so as to reflect the rate of change, trend direction and magnitude of abrupt changes of multiple variables before the target date.

[0026] The parallel extraction of SVP-T deep representation in S6 is used to learn temporal embedding features from a W-day window sequence before the target date via SVP-T, including the following steps:

[0027] S601. Extract or construct several sequence segments from different variables and different time positions from the multivariable window sequence as shape input units, and attach the corresponding variable identifier and time position to each shape input unit;

[0028] S602. Based on the attention mechanism, learn the relationship between shape input units to simultaneously capture temporal and variable-dimensional dependencies, and enhance the association weight of shape input units from different variables that have overlapping or adjacent relationships in time to highlight multivariate collaborative relationships;

[0029] S603. Pool the encoded output to obtain a low-dimensional embedding vector as a temporal embedding feature, and use the embedding vector as the input feature for the determination of the start / end date of rainfall, and concatenate and fuse it with interpretable statistical features to improve the stability and generalization ability of event boundary determination.

[0030] The feature alignment and fusion described in S7 refers to matching and aligning interpretable statistical features with temporal embedded features learned from the W-day window sequence before the target date using SVP-T, one by one according to the spatial index, spatial partition number, and time index of the sample, and then splicing or weighting the two types of features after alignment to form a fused feature vector.

[0031] The training and evaluation of the APLR model in S8 involves constructing two APLR models—a start discrimination model and an end discrimination model—with the start date label and end date label as supervision signals within the training period obtained by strict time segmentation, and then training and validating the two APLR models respectively.

[0032] The model extrapolation application of S9 refers to performing daily probability inference on the test set and subsequent dates after the precursor discrimination model has been trained and its parameters determined on the training and validation sets; inputting the fused feature vector into the start discrimination model and the end discrimination model respectively, and outputting the probability of the start date and the probability of the end date of rainfall on that date, thereby obtaining the daily probability sequence of subsequent dates; and converting the daily probability sequence into the discrimination results of the start date and end date of rainfall events in subsequent periods, or using the probability sequence as event-level prior information for subsequent deviation correction, data fusion or risk assessment operations.

[0033] SVP-T stands for Shape-Level Variable-Position Transformer, a Transformer method for multivariate time series classification. It changes the input from the traditional "timestamp level" (all variables at each time step) to "shape level" (time subsequences of each variable), and preserves the variable source and temporal location information of each subsequence through variable-position encoding (VP-layer). Finally, it uses a VP-based self-attention mechanism to specifically enhance the interaction between subsequences that "overlap temporally and come from different variables," thus simultaneously modeling temporal dependencies and inter-variable dependencies.

[0034] The beneficial effects of this invention are that it balances interpretability and high-order representation capabilities through dual-branch feature fusion; it improves cross-time stability through full-time embedding extraction and strict temporal segmentation; it can stably identify the boundaries of rainfall events and provide interpretable discrimination criteria based on meteorological elements related to rainfall, and it has the advantages of generalization, reproducibility and engineering feasibility. It is applicable to business scenarios such as daily deviation correction and data fusion of satellite rainfall products and multi-mode rainfall products. Attached Figure Description

[0035] Figure 1 This is a flowchart of a method for identifying the precursory features of the start and end of regional rainfall events according to an embodiment of the present invention. Detailed Implementation

[0036] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0037] The method for identifying the precursory features of the start and end of regional rainfall events according to the present invention includes the following steps:

[0038] S1. Obtain multivariate daily raster meteorological data for the study area and align other variables with time based on the rainfall time axis.

[0039] Multivariate daily raster meteorological data includes rainfall data as well as one or more non-rainfall meteorological elements such as temperature, humidity, wind speed, and air pressure.

[0040] Time alignment of other variables based on the rainfall timeline refers to spatially unifying each variable based on a unified spatial grid and aligning each variable based on a unified date series, forming a multivariate daily raster data set under the same calendar index.

[0041] The data for the entire period is divided into training and testing periods according to a preset time segmentation strategy, or further into training, validation and testing periods. The time segmentation strategy is any one or a combination of segmentation by year list, segmentation by time sequence ratio or time series cross-validation segmentation, and ensures that any statistics used for spatial stratification, clustering or model parameter estimation are calculated only from the training period data to avoid leakage of time information.

[0042] S2. Obtain or generate a spatial partition mask.

[0043] Spatial partitioning is used to divide the study area into several weather layers or spatial sub-regions and generate spatial partitioning results. Spatial partitioning includes the following steps:

[0044] S201. During the training period, construct a spatial partition feature vector for each grid cell. The spatial partition feature vector is used to characterize the rainfall process characteristics or rainfall climatology characteristics of the grid cell, and includes at least one or a combination of the following: effective rainfall daily frequency, rainfall event frequency, average event duration, average event interval, event intensity statistics, rainfall quantile statistics, and seasonality percentage statistics during the training period.

[0045] S202. Standardize the spatial partition feature vectors to obtain spatial representations;

[0046] S203. In the low-dimensional spatial representation, a clustering algorithm is used to divide the raster into a predetermined number of K spatial partitions. K-Means clustering algorithm is preferred to obtain K spatial partitions, and each raster is assigned a corresponding partition number r. id ;

[0047] S204. Fill in the partition number according to the grid spatial position to generate a spatial partition mask. The spatial partition mask records the partition number corresponding to each grid in the form of a two-dimensional array, where invalid or missing grids are marked as invalid values. The spatial partition mask is used as the spatial constraint basis for subsequent stratified sampling, sequence sample construction and start day discrimination and end day discrimination modeling.

[0048] S3. Construct regional rainfall event labels to generate rainfall start date and end date labels and construct sequence samples aligned W days prior to the target date, including the following steps:

[0049] S301. During the training period and in the daily rainfall sequence throughout the entire time period, the raster is first screened for effective rainfall. Raster-daily samples with daily rainfall not less than the preset effective rainfall threshold T are selected as effective rainfall samples, where T is preferably set to 1 mm. The selection method is as follows:

[0050] , where |·| represents the number of elements in the set (i.e. the number of grid points in the region), Pr(g,t) is the rainfall of grid point g on date t, and T is the effective rainfall threshold;

[0051] S302. Rainfall with a daily rainfall amount less than the effective rainfall threshold T is considered invalid rainfall and removed from the event statistics;

[0052] S303. On the timeline of the effective rainfall samples, rainfall events are divided into sessions based on the minimum rainfall interval (MIT):

[0053] When the interval between two consecutive effective rainfall events is less than MIT, they are merged into the same event; when the interval between two consecutive effective rainfall events is not less than MIT, they are divided into different events.

[0054] S304. The first effective rainfall day of each event is determined as the start date of the event, and the last effective rainfall day is determined as the end date of the event, thereby generating daily start and daily end labels;

[0055] S305. Subsequently, for each candidate target day, samples are constructed using the start label and end label as supervision signals, and multivariate sequences of W days prior to the target day are extracted to form a start window sequence for start day discrimination and an end window sequence for end day discrimination. The window sequence is a multivariate sequence fragment of W days prior to the target day and does not contain data of the target day itself. The corresponding spatial partition number, raster spatial index and time index are recorded for each window sequence sample for subsequent modeling and alignment fusion.

[0056] S4. Construct a time window sequence.

[0057] Regional rainfall event labels are constructed based on spatial partitioning masks: The study area is divided into multiple sub-regions. For any sub-region r and date t, the proportion of grid points in the sub-region that meet the rainfall threshold T is calculated. When the proportion is not lower than the coverage ratio threshold P, y(r,t)=1 is marked, otherwise y(r,t)=0 is marked.

[0058] For each sub-region r and date t, i.e. for each (r,t) sample, construct a time window sequence X(r,t)=[x(tW),…,x(t-1)] for W days before the target date, where x is a multivariate vector to ensure that the features use only historical information.

[0059] S5. Parallel extraction of interpretable feature engineering representations, used to construct interpretable statistical features from the W-day window sequence before the target date, including the following steps:

[0060] S501. For each non-rainfall meteorological element variable in the window sequence, calculate the last value feature of the window, the slope feature of the trend within the window or the most recent k days, and the rise feature of the increase of the mean of the most recent k days relative to the mean of the previous m days, and further calculate the difference feature and the relative rate of change feature.

[0061] S502. Can perform moving average or exponential smoothing on window sequences to suppress noise disturbances;

[0062] S503. Use the features obtained in S501 as an interpretable input feature set for determining the start and end dates of rainfall, so as to reflect the rate of change, trend direction and magnitude of abrupt changes of multiple variables before the target date.

[0063] S6. Parallel extraction of Shape-Level Variable-PositionTransformer (SVP-T) deep representations, used to learn temporal embedding features from a W-day window sequence before the target date via SVP-T, including the following steps:

[0064] S601. Extract or construct several sequence segments from different variables and different time positions from the multivariable window sequence as shape input units, and attach the corresponding variable identifier and time position to each shape input unit;

[0065] S602. Based on the attention mechanism, learn the relationship between shape input units to simultaneously capture temporal and variable-dimensional dependencies, and enhance the association weight of shape input units from different variables that have overlapping or adjacent relationships in time to highlight multivariate collaborative relationships;

[0066] S603. Pool the encoded output to obtain a low-dimensional embedding vector as a temporal embedding feature, and use the embedding vector as the input feature for distinguishing the start and end dates of rainfall, and concatenate and fuse it with interpretable statistical features to improve the stability and generalization ability of event boundary discrimination.

[0067] The shape input unit is processed by the SVP-T coding network to obtain temporal embedding features, which includes the following steps:

[0068] 1) During the training period, for the start date / end date discrimination samples (i.e., "the start window sequence used for start date discrimination and the end window sequence used for end date discrimination"), construct a multivariate window sequence for each candidate target date that is W days prior to the target date. , where W is the window length and V is the number of variables;

[0069] 2) Perform shape-level preprocessing on the multivariable window sequence to generate the shape input unit of SVP-T:

[0070] A large number of candidate subsequences are generated from the window sequence for each variable dimension, and a clustering algorithm is applied to each variable to obtain U cluster centers. The U subsequences closest to each cluster center are used as representative shape input units for that variable, thereby transforming a window sequence into a set of shapes. ,in And record the variable index for each shape input variable. and the start and end time positions in the original window ( );

[0071] 3) When inputting shape units into the SVP-T encoding network, a variable-position encoding layer (VP-layer) is introduced: a variable-position (VP) information vector is constructed for each shape input unit. The VP information is projected onto the model dimension through linear mapping and added to the linear projection result of the shape input unit to form the input representation of the Transformer encoder, enabling the model to simultaneously perceive the shape source variable and its relative position in the window.

[0072] 4) A variable-location-based self-attention mechanism (VP-based self-attention) is employed in the Transformer encoder of SVP-T:

[0073] When calculating attention weights, a shape overlap enhancement matrix is ​​introduced to explicitly enhance the attention weights between shape input unit pairs that are “from different variables and overlap in time”. The degree of overlap between the two shapes is calculated by the overlap length of their time intervals, and the attention weights are amplified by a parameterized enhancement function, thereby strengthening the collaborative change relationship of multiple variables in the same time period.

[0074] S7. Feature alignment and fusion refers to matching and aligning interpretable statistical features with time-series embedded features learned from the W-day window sequence of the target date through SVP-T, one by one according to the spatial index, spatial partition number and time index of the sample, and then splicing or weighting the two types of features after alignment to form a fused feature vector.

[0075] The training and evaluation of the S8 APLR model involved constructing two APLR models—a start-discrimination model and an end-discrimination model—with the start date label and end date label as supervision signals within a training period obtained through strict time segmentation. The two APLR models were then trained and validated separately.

[0076] The APLR model training includes:

[0077] S801. Perform standardization processing on the fusion features;

[0078] S802. Set several segmented nodes for each continuous feature and perform a spline or piecewise linear expansion to map the original feature into a piecewise linear basis function;

[0079] S803. Regularized logistic regression is used to learn classification weights in the basis function space to output the daily probabilities of the start and end days.

[0080] S804. The number of segment nodes, node positions, and regularization strength are determined through validation period evaluation or time series cross-validation, and class imbalance caused by sample scarcity on the start and end days can be mitigated through class weighting, stratified sampling quotas, or oversampling / undersampling strategies.

[0081] During the model evaluation phase, relevant evaluation metrics are calculated based on the probability output from the validation or testing phase.

[0082] 1) Youden exponential maximization strategy based on ROC curve;

[0083] 2) A strategy based on maximizing the Critical Success Index (CSI);

[0084] 3) A strategy based on the frequency bias index (FBI) being close to 1;

[0085] 4) Output at least one of the following metrics: POD, FAR, POFD, CSI, FBI, Precision, Accuracy, and AUC. POD (Probability of Detection) represents the hit rate, indicating the proportion of samples correctly identified as having a start / end date. FAR (False Alarm Ratio) represents the false alarm rate, indicating the proportion of samples predicted as having a start / end date that are not actually having one. POFD (Probability of False Detection) represents the false alarm rate, indicating the proportion of non-rainy days misidentified as rainy days. CSI (Critical Success Index) is a composite score of hit and false alarms, with a value closer to 1 being better. FBI (Frequency Bias Index) represents the ratio of predicted frequency to actual frequency, with a value of 1 indicating no bias. Precision represents the proportion of samples predicted as positive that are actually positive. Accuracy represents the proportion of all samples correctly identified. AUC (Area of ​​Detection) represents the accuracy of a prediction. The area under the ROC curve (AUC) is a comprehensive measure of the model's discriminative ability; the closer to 1, the better. Each metric is calculated based on the true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN) in the confusion matrix.

[0086]

[0087] 5) Combine the evaluation indicators to subsequently convert the probabilities into start-day and end-day discrimination results;

[0088] 6) Save the evaluation results, optimal threshold configuration, and model parameters to support the reproduction of experiments and engineering deployment.

[0089] S9. Model extrapolation application refers to the daily probability inference of the test set and subsequent dates after the training and parameter determination of the precursor discrimination model are completed during the training and validation periods.

[0090] After completing the selection of hyperparameters, determination of thresholds, and model finalization during the validation period, the model parameters, standardized parameters, and segmented node configurations of the starting and ending discrimination models are frozen.

[0091] For any subsequent date t after the validation period, the same sample construction rules as those in the training and validation phases are used to generate corresponding samples from a multivariate window sequence of W days before the target date. The interpretable statistical features and SVP-T embedded features are extracted and fused to form a fused feature vector in the manner described in step S7.

[0092] The fused feature vectors are input into the start and end discrimination models respectively, and the probability of the start and end of rainfall on the given date is output, thus obtaining the daily probability sequence for subsequent dates. The daily probability sequence is then converted into the start and end date discrimination results of rainfall events in subsequent periods, or the probability sequence is used as event-level prior information for subsequent deviation correction, data fusion, or risk assessment.

[0093] Example

[0094] like Figure 1 As shown, a method for identifying the precursor features of the start and end of regional rainfall events based on spatial partitioning and bi-branch feature representation is provided, including the following steps:

[0095] S1. Acquire multivariate daily raster meteorological data. The data includes at least a precipitation variable Pr and one or more precursor variables, each with time, latitude, and longitude. Align the other variables with the time axis of the precipitation variable as a reference, and interpolate or fill in missing dates if necessary.

[0096] S2. Obtain or generate a spatial partitioning mask. If an external spatial partitioning mask file exists, it is read directly; otherwise, a multivariate statistical feature vector is constructed for each valid grid point within the study area. After standardization and principal component analysis for dimensionality reduction, a clustering algorithm is used in the dimensionality-reduced space to divide the grid points into K sub-regions, and the spatial partitioning mask is output, where invalid grid points are marked as invalid values.

[0097] S3. Construct regional rainfall event labels. For any date t and sub-region r, count the percentage of grid points in the region that satisfy Pr(g,t)≥T. If the percentage is ≥P, set y(r,t)=1; otherwise, set y(r,t)=0.

[0098] S4. Construct a time window sequence. For each (r,t) sample, take a window of W days before the target date [tW,…,t-1] to form a window sequence X(r,t) to ensure that the features only use historical information.

[0099] S5. Parallel Extraction of Interpretable Feature Engineering Representations. For each variable in the window sequence, construct the last value (last), trend slope (slope), and recent increase (rise), and further calculate difference or relative rate of change features. Construct a central pattern in the positive examples of the training samples: when using a single center, the center is the mean of the positive example window; when using multiple prototypes, cluster the positive example windows to obtain multiple prototype centers. Calculate the similarity (sim) feature between any window and the central pattern. If the features are first calculated at the grid scale, then aggregate the grid features within the same (r,t) region at the regional scale to obtain the regional-daily feature vector.

[0100] S6. Parallel extraction of SVP-T deep representations. Using the region mean window sequence as sample input, construct an SVP sample library and record samples containing r. id With t id Meta-information, where r id For spatial partition numbering, used for sub-region identification, t id The target date index is used for time identification. The window sequence is tokenized by time-variable, and the time position embedding and variable embedding are superimposed, projected onto the numerical values, and then input into the Transformer encoder. Pooling yields the embedding vector z(r,t). After fitting the model during the training period, the SVP model extracts the embeddings from samples across the entire time period, outputting the SVP representation and its meta-information. This supports the reuse of deep representations under any strict segmentation scheme and avoids missing matches.

[0101] S7. Feature alignment and fusion. (r) id ,t id The region-day features obtained in step S5 are matched and concatenated with the embedding vector obtained in step S6 to form the fused feature F(r,t).

[0102] S8. APLR modeling outputs probabilities. After standardizing F(r,t), piecewise linear spline expansion is performed, and a logistic regression model is trained to output the probability p(r,t) of rainfall events for sub-region r and date t. The number of spline nodes and the regularization intensity are determined through grid search or cross-validation. Training and evaluation employ a strict time-segmentation strategy and can be trained separately for different seasons and regional groups.

[0103] S9. Threshold Selection, Evaluation, and Interpretation Output. A threshold scan is performed on p(r,t), and a threshold is selected based on criteria such as Youden, maximum CSI, or FBI close to 1. Indicators such as AUC, POD, FAR, POFD, CSI, and FBI are calculated, and the results are output as curves and tables. Simultaneously, rule mining is performed based on interpretable features to obtain a rule set, and a central pattern visualization curve and regional partitioning map are output to enhance the model's credibility and analyzability.

[0104] Although the above embodiments have been shown and described, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Any changes, modifications, substitutions and variations made to the above embodiments by those skilled in the art are within the protection scope of the present invention.

Claims

1. A method for identifying precursory features of the onset and end of regional rainfall events, characterized in that, Includes the following steps: S1. Acquire multivariate daily raster meteorological data and align other variables with time based on the rainfall time axis; S2. Obtain or generate a spatial partition mask; S3. Construct regional rainfall event labels; S4. Construct a time window sequence; S5. Parallel extraction of interpretable feature engineering representations; S6. Parallel extraction of SVP-T depth characterization; S7. Feature alignment and fusion; Training and evaluation of the S8 APLR model; S9. Application of model extrapolation.

2. The method for identifying the precursory features of the start and end of regional rainfall events according to claim 1, characterized in that, The multivariate daily raster meteorological data in S1 includes rainfall data and one or more non-rainfall meteorological element data such as temperature, humidity, wind speed, and air pressure; the time alignment of other variables based on the rainfall time axis refers to the spatial consistency processing of each variable based on a unified spatial grid and the time alignment of each variable based on a unified date sequence, forming a multivariate daily raster data set under the same calendar index.

3. The method for identifying the precursory features of the start and end of regional rainfall events according to claim 1, characterized in that, In step S2, obtaining or generating a spatial partitioning mask refers to dividing the study area into several weather layers or spatial sub-regions and generating spatial partitioning results. The spatial partitioning includes the following steps: S201. During the training period, a spatial partition feature vector is constructed for each grid cell, and the spatial partition feature vector is used to characterize the rainfall process characteristics or rainfall climatology characteristics of the grid cell. S202. Standardize the spatial partition feature vectors to obtain spatial representations; S203. In the low-dimensional spatial representation, a clustering algorithm is used to divide the raster into a preset number of K spatial partitions. K-Means clustering algorithm is preferred to obtain K spaces, and each raster is assigned a corresponding partition number. S204. The partition numbers are backfilled according to the grid spatial position to generate a spatial partition mask, wherein the spatial partition mask records the partition number corresponding to each grid in the form of a two-dimensional array.

4. The method for identifying the precursory features of the start and end of regional rainfall events according to claim 1, characterized in that, The construction of regional rainfall event labels in S3 refers to generating rainfall start date labels and end date labels and constructing sequence samples aligned with W days prior to the target date.

5. The method for identifying the precursory features of the start and end of regional rainfall events according to claim 4, characterized in that, The construction of the time window sequence in S4 refers to constructing a time window sequence X(r,t)=[x(tW),…,x(t-1)] for each (r,t) sample, which is W days before the target date. Here, x is a multivariate vector to ensure that the features only use historical information, r is a sub-region, and t is the date.

6. The method for identifying the precursory features of the start and end of regional rainfall events according to claim 5, characterized in that, The parallel extraction of interpretable feature engineering representations in S5, used to construct interpretable statistical features from a W-day window sequence before the target date, includes the following steps: S501. For each non-rainfall meteorological element variable in the window sequence, calculate the last value feature of the window, the slope feature of the trend within the window or the most recent k days, and the rise feature of the increase of the recent k-day mean relative to the previous m-day mean, and further calculate the difference feature and the relative rate of change feature. S502. Can perform moving average or exponential smoothing on window sequences to suppress noise disturbances; S503. Use the features in S501 as an interpretable input feature set for determining the start / end date of rainfall, so as to reflect the rate of change, trend direction and magnitude of abrupt changes of multiple variables before the target date.

7. The method for identifying the precursory features of the start and end of regional rainfall events according to claim 6, characterized in that, The parallel extraction of SVP-T deep representation in S6 is used to learn temporal embedding features from a W-day window sequence before the target date via SVP-T, including the following steps: S601. Extract or construct several sequence segments from different variables and different time positions from the multivariable window sequence as shape input units, and attach the corresponding variable identifier and time position to each shape input unit; S602. Based on the attention mechanism, learn the relationship between shape input units to simultaneously capture temporal and variable-dimensional dependencies, and enhance the association weight of shape input units from different variables that have overlapping or adjacent relationships in time to highlight multivariate collaborative relationships; S603. Pool the encoded output to obtain a low-dimensional embedding vector as a temporal embedding feature, and use the embedding vector as the input feature for the determination of the start / end date of rainfall, and concatenate and fuse it with interpretable statistical features to improve the stability and generalization ability of event boundary determination.

8. The method for identifying the precursory features of the start and end of regional rainfall events according to claim 7, characterized in that, The feature alignment and fusion described in S7 refers to matching and aligning interpretable statistical features with temporal embedded features learned from the W-day window sequence before the target date using SVP-T, one by one according to the spatial index, spatial partition number, and time index of the sample, and then splicing or weighting the two types of features after alignment to form a fused feature vector.

9. The method for identifying the precursory features of the start and end of regional rainfall events according to claim 8, characterized in that, The training and evaluation of the APLR model in S8 involves constructing two APLR models—a start discrimination model and an end discrimination model—with the start date label and end date label as supervision signals within the training period obtained by strict time segmentation, and then training and validating the two APLR models respectively.

10. The method for identifying the precursory features of the start and end of regional rainfall events according to claim 9, characterized in that, The model extrapolation application of S9 refers to performing daily probability inference on the test set and subsequent dates after the precursor discrimination model has been trained and its parameters determined on the training and validation sets; inputting the fused feature vector into the start discrimination model and the end discrimination model respectively, and outputting the probability of the start date and the probability of the end date of rainfall on that date, thereby obtaining the daily probability sequence of subsequent dates; and converting the daily probability sequence into the discrimination results of the start date and end date of rainfall events in subsequent periods, or using the probability sequence as event-level prior information for subsequent deviation correction, data fusion or risk assessment operations.