Multi-source data fusion disaster abnormal area identification and early warning method

Through multi-source data fusion and feature extraction, combined with SARIMA and GWR models, the problems of disaster warning inaccuracy and timeliness caused by a single data source are solved, and efficient and accurate identification and early warning of disaster abnormal areas are achieved.

CN120470276APending Publication Date: 2025-08-12ANHUI COMM IND SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510470101.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing disaster warning methods rely on a single data source, resulting in incomplete information, insufficient accuracy and limited timeliness, making it difficult to effectively integrate multi-source heterogeneous data for accurate disaster abnormal areas identification and early warning.

Method used

Multi-source data fusion method is adopted to extract and fusion of image data, GIS data and meteorological data, combine SARIMA and GWR models for abnormal area identification and early warning, and use local binary mode and terrain undulation calculation and other algorithms to extract features, and use KPCA algorithm to realize feature dimensionality reduction and nonlinear feature extraction.

Benefits of technology

It improves the accuracy and comprehensive information of disaster abnormal areas identification, reduces misjudgment and misjudgment, provides more reliable disaster warning decision support, and improves data processing efficiency and timeliness of early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470276A_ABST
    Figure CN120470276A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source data fusion disaster abnormal area identification and early warning method, and relates to the technical field of disaster early warning, and the method comprises the following steps: data collection and preprocessing are carried out in advance, and the data collection at least comprises image data, GIS data and meteorological data; performing feature extraction, and extracting texture and edge features from the image data based on a local binary pattern; a topographic relief, a topographic slope and a river network density are calculated from GIS data. According to the invention, image data, GIS data and meteorological data are fused, and information contained in different types of data can be comprehensively integrated. Through comprehensive analysis of multi-source data, multi-dimensional information about disaster occurrence environments, geographical conditions, meteorological factors and the like can be obtained, and compared with a traditional single-data-source early warning method, comprehensiveness and accuracy of the information are greatly improved, and a more reliable basis is provided for recognition and early warning of disaster abnormal areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of disaster early warning technology, and in particular to a method for identifying and warning disaster abnormal areas by fusing multi-source data. Background Art

[0002] In the field of disaster early warning technology, timely and accurate identification of disaster anomalies and issuance of early warnings are crucial for protecting people's lives and property and reducing disaster losses. Traditional disaster early warning methods often rely on a single data source, such as providing meteorological disaster warnings based solely on meteorological data or assessing geological disaster risks solely through topographic data. This single-data source early warning approach has many limitations, as follows:

[0003] 1. Incomplete information: A single type of data cannot fully reflect the complex mechanisms of disaster occurrence and development. For example, relying solely on meteorological data makes it difficult to accurately assess the risk of flood convergence in mountainous areas due to topographic factors, as it does not consider the impact of topography on water flow.

[0004] 2. Inaccuracy: Due to the lack of cross-correlation and complementarity between multi-dimensional information, early warning conclusions drawn from a single data source are less accurate. For example, when predicting landslide hazards, if only terrain data is used without incorporating data such as soil moisture and vegetation cover, potential risk points for landslides caused by excessive soil moisture may be missed.

[0005] 3. Limited Timeliness: The acquisition and analysis process of a single data source is relatively fixed, making it difficult to quickly adapt to the dynamic changes in the disaster situation. As a disaster develops, relying solely on a single data source may fail to capture emerging risk factors in a timely manner, resulting in delayed warnings.

[0006] With the development of information technology, the acquisition of multi-source data, such as satellite remote sensing imagery, geographic information system (GIS) data, and meteorological monitoring data, has become more convenient. However, how to effectively fuse these heterogeneous multi-source data and use the fused data to accurately identify and warn of disaster anomaly areas has become a difficult and hot topic in current research. Existing data fusion methods suffer from poor fusion results, high computational complexity, and an inability to fully exploit data features when processing multi-source data, making them difficult to meet the needs of actual disaster warning.

[0007] Currently, no effective solutions have been proposed for the problems in related technologies. Summary of the Invention

[0008] In response to the problems in the related technologies, the present invention proposes a disaster abnormal area identification and early warning method based on multi-source data fusion to overcome the above technical problems existing in the existing related technologies.

[0009] The technical solution of the present invention is achieved as follows:

[0010] A multi-source data fusion disaster abnormal area identification and early warning method includes the following steps:

[0011] Perform data collection and preprocessing in advance, wherein the data collection includes at least: image data, GIS data and meteorological data;

[0012] Perform feature extraction, extracting texture and edge features from image data based on local binary patterns; calculate terrain relief, terrain slope, and river network density from GIS data; and calculate water vapor flux divergence, precipitation intensity, and wind speed change rate from meteorological data;

[0013] Perform data fusion, calibrate the features extracted from each data source into a feature matrix and splice them by column, and standardize the spliced feature matrix; calibrate the radial basis kernel function as the kernel function to calculate the kernel matrix, center the kernel matrix, perform eigenvalue decomposition, and calculate the low-dimensional feature vector after fusion;

[0014] Identify abnormal areas, use the ADF test to stabilize time series data, determine SARIMA model parameters by observing autocorrelation and partial autocorrelation function graphs and combining them with the Akaike Information Criterion, build a model and perform a residual white noise test, and use the model to predict time series values; construct a geographic weight matrix and estimate the coefficients of the geographically weighted regression model using the weighted least squares method; calculate the residuals between the SARIMA model predictions and the actual observations, combine the geographic information into the GWR model, determine the anomaly threshold based on the residual mean and standard deviation, and identify disaster anomaly areas based on the GWR model results and the anomaly threshold;

[0015] The warning level is determined based on the preset disaster loss assessment model, and warning information is released through a multilingual warning release system.

[0016] The image data feature extraction includes: texture and edge feature extraction based on local binary patterns, including the following steps:

[0017] Calibrate the 3×3 neighborhood, the gray value of the central pixel is g c , the gray value of the neighborhood pixel is g i (i=0,1,…,7), obtain the LBP code, expressed as:

[0018]

[0019] Where P is the number of neighborhood pixels, R is the neighborhood radius, and s(x) is the sign function;

[0020] When x ≥ 0, s(x) = 1;

[0021] When x<0, s(x)=0;

[0022] By counting the histograms of different LBP codes, the texture and edge features of the image are obtained.

[0023] Among them, GIS data feature extraction includes the following steps:

[0024] Calculate the terrain relief, which reflects the degree of terrain change and is expressed as:

[0025]

[0026] Among them, z ij is the elevation value of grid cell ij, is the average elevation value in the study area, n and m are the number of grid rows and columns in the study area respectively;

[0027] Calculate the terrain slope. The terrain slope indicates the degree of terrain inclination. In raster data, a 3×3 window is used for calculation, which is expressed as:

[0028]

[0029] in, and are the slope components in the x and y directions, respectively, which are calculated by the central difference method and are expressed as:

[0030]

[0031] Among them, Δx and Δy are the grid resolutions;

[0032] Calculate the river network density, which is used to measure the density of river distribution in a region and is expressed as:

[0033]

[0034] Where L is the total length of rivers in the study area, and A is the area of the study area.

[0035] The feature extraction of meteorological data includes the following steps:

[0036] Calculate the water vapor flux divergence. The water vapor flux divergence reflects the convergence and divergence of water vapor in the atmosphere and is expressed as:

[0037]

[0038] Where q is the specific humidity, is the horizontal wind speed vector, u and v are the wind speed components in the x and y directions respectively;

[0039] Calculate the precipitation intensity. The precipitation intensity PI represents the amount of precipitation per unit time and is expressed as:

[0040]

[0041] Where ΔP is the amount of precipitation in the time period Δt;

[0042] Calculate the wind speed change rate. The wind speed change rate VCR is used to measure how fast the wind speed changes over time and is expressed as:

[0043]

[0044] in, and are the wind speeds at time t1 and t2 respectively.

[0045] The data fusion process includes the following steps:

[0046] The features extracted from image data, GIS data and meteorological data are calibrated as feature matrices respectively, and these feature matrices are spliced into a unified feature matrix X by column, X=[x1,x2,…,x n ], where x i is the feature vector of the i-th sample, n is the number of samples, and X is standardized so that the mean of each feature is 0 and the variance is 1, which is expressed as:

[0047]

[0048] Among them, x ij is the jth eigenvalue of the i-th sample in the original feature matrix X, is the mean of the jth feature, σ j is the standard deviation of the jth feature, x ij is the normalized eigenvalue,

[0049] The radial basis kernel function is calibrated as the kernel function, and the kernel matrix is calculated, which is expressed as:

[0050]

[0051] Among them, x i and x j is the feature vector of the two samples, ||x i -x j || is the Euclidean distance between them, σ is the bandwidth parameter of the kernel function,

[0052] According to the selected kernel function, calculate the kernel matrix K, whose element K ij =K(x i ,x j ), i, j = 1, 2, ..., n;

[0053] The central kernel matrix is constructed, including the following steps:

[0054] Calculate the row mean vector of the kernel matrix K Its elements i=1,…,n;

[0055] Calculate the centralized kernel matrix K, expressed as:

[0056]

[0057] Calculate the eigenvalues and eigenvectors, perform eigenvalue decomposition on the centralized kernel matrix K, that is, solve the characteristic equation Kα=λα, where λ is the eigenvalue and α is the corresponding eigenvector; the obtained eigenvalues (λ1≥λ2≥…≥λ n ), and the corresponding eigenvectors (α1,α2,…,α n );

[0058] The number of principal components selected is determined based on the cumulative contribution rate. The cumulative contribution rate calculation formula is:

[0059]

[0060] Calculate the fused feature vector, project the original data into the new feature space, and obtain the fused low-dimensional feature vector. For the i-th sample, the fused feature vector y i The calculation formula is:

[0061]

[0062] Among them, α ji is the i-th element of the j-th eigenvector.

[0063] The abnormal area identification includes: performing SARIMA model analysis, including the following steps:

[0064] Perform data stabilization processing and use ADF test to test the stability of time series data, including: calibrating the time series to y t , the sequence after d-order difference is Δ d y t , where Δy t =y t -y t-1 , Δ d y t =Δ(Δ d-1 y t );

[0065] Model order determination: by observing the autocorrelation function and partial autocorrelation function graphs, the autoregressive order p, moving average order q, and seasonal period D of the SARIMA model are determined. Then, the Akaike Information Criterion is used to search for the optimal model parameter combination (p, d, q) and (P, D, Q) within a certain range.

[0066] Use the selected parameters to build a SARIMA model, and use the maximum likelihood estimation method to estimate the model parameters. Use the fitted SARIMA model to predict the data at future time points to obtain the predicted value sequence

[0067] The method also includes: performing GWR model analysis, including the following steps:

[0068] Construct a geographic weight matrix and determine the Euclidean distance between each sample point based on the geographic location information of the data;

[0069] Construct a geographic weight matrix W based on the distance, whose element w ij Represents the weight between sample points i and j, which is calculated using the Gaussian kernel function and is expressed as:

[0070]

[0071] Among them, d ij is the distance between sample points i and j, h is the bandwidth parameter;

[0072] For model estimation, the expression of the geographically weighted regression model is:

[0073]

[0074] Among them, u i ,v i is the geographic coordinate of sample point i, y i is the dependent variable, x ik is the kth independent variable, β k (u i ,v i ) is at position u i ,v i The regression coefficient at ∈ i is a random error term, and the regression coefficient β of each position is estimated using weighted least squares method k (u i ,v i ), minimize.

[0075] The method further includes: identifying abnormal areas, including the following steps:

[0076] Calculate the residuals and convert the predicted values of the SARIMA model and the actual observed value y t For comparison, calculate the time series residual, which is expressed as:

[0077]

[0078] These residuals are combined with geographic information and input into the GWR model for spatial analysis;

[0079] Based on the residual distribution of historical data, use statistical methods to determine the abnormal threshold;

[0080] To identify abnormal areas, the spatial regression results obtained based on the GWR model are combined with the abnormal threshold to judge each spatial location; if the residuals of multiple adjacent locations in a certain area exceed the abnormal threshold, the area is identified as a disaster abnormal area.

[0081] Beneficial effects of the present invention:

[0082] 1. This invention integrates image data, GIS data, and meteorological data, fully integrating the information contained in these different data types. Through comprehensive analysis of multi-source data, it can obtain multi-dimensional information on the disaster environment, geographical conditions, and meteorological factors. Compared with traditional early warning methods based on a single data source, this greatly improves the comprehensiveness and accuracy of information, providing a more reliable basis for identifying and warning abnormal disaster areas.

[0083] 2. The present invention analyzes time series data through the SARIMA model and combines it with the GWR model for spatial analysis, which can accurately identify disaster abnormal areas from both time and space dimensions. By using residual analysis and abnormal threshold judgment, the accuracy of abnormal area identification is effectively improved, the occurrence of misjudgment and missed judgment is reduced, and accurate decision-making support is provided for disaster warning. At the same time, in the data processing process, a variety of algorithms such as local binary pattern (LBP), terrain relief calculation, and water vapor flux divergence calculation are used to extract features from different data sources respectively. These features can accurately reflect the intrinsic characteristics of the data. At the same time, the KPCA algorithm is used to achieve feature dimensionality reduction and nonlinear feature extraction, which retains key information while reducing the data dimension, improves data processing efficiency, and makes subsequent abnormal area identification and warning analysis more efficient and accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0085] Figure 1 The present invention is a flowchart of a method for identifying and warning abnormal disaster areas by fusion of multi-source data according to an embodiment of the present invention. DETAILED DESCRIPTION

[0086] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of the present invention.

[0087] According to an embodiment of the present invention, a method for identifying and warning abnormal disaster areas by fusing multi-source data is provided.

[0088] like Figure 1 As shown, the disaster abnormal area identification and early warning method based on multi-source data fusion according to an embodiment of the present invention includes the following steps:

[0089] Perform data collection and preprocessing in advance, wherein the data collection includes at least: image data, GIS data and meteorological data;

[0090] Image data processing involves acquiring image data from the mountainous region using a data acquisition device mounted on a drone. Due to the complex terrain, images are susceptible to noise and geometric distortion. Based on the image noise characteristics, the standard deviation of the adaptive Gaussian filter was set to 3, the contrast limit in the CLAHE algorithm was set to 4.0, and the tile size was set to 8×8. The position and attitude data of the image acquisition device were acquired using the drone's satellite positioning system and IMU sensor. Combined with the RFM model, geometric correction was performed on the image to ensure that the locations of features in the image accurately corresponded to their actual geographic coordinates, providing an accurate image foundation for subsequent analysis.

[0091] GIS data processing involves collecting various GIS data types, including the mountain area's transportation network, infrastructure distribution, and topographic contours. This data may come from different departments and be in various formats, such as GeoJSON and Shapefile. Using specialized data conversion tools, all data is converted to a unified GeoPackage format. The mountain area is located at the edge of the projection zone. Considering the angular and area accuracy requirements for subsequent terrain analysis, the Gauss-Krüger projection was selected based on a dynamic projection transformation algorithm, and the coordinate system was standardized to WGS84.

[0092] Among them, meteorological data processing collects multi-source meteorological data, including precipitation, wind speed, humidity, etc., through meteorological radar stations and satellite meteorological data receiving equipment around the mountainous area. An outlier detection algorithm based on the K-Means clustering algorithm (K value is set to 5) is used to perform quality control on the collected meteorological data. If the wind speed data of a monitoring station is found to deviate significantly from the cluster center at a certain moment, it is verified that it is caused by a sensor failure and the outlier is removed. Combined with the terrain data of the mountainous area, the terrain-adaptive Kriging interpolation method is used to perform spatial interpolation on the meteorological data to fill the blank areas between the monitoring stations, making the meteorological data more continuous and accurate in space. At the same time, the meteorological data is set to be updated every 3 hours to ensure the timeliness of the data.

[0093] Features are extracted from image data, GIS data and meteorological data respectively to provide a basis for subsequent analysis.

[0094] Among them, image data feature extraction includes: texture and edge feature extraction based on local binary pattern (LBP), including:

[0095] Calibrate the 3×3 neighborhood, the gray value of the central pixel is g c , the gray value of the neighborhood pixel is g i (i=0,1,…,7), obtain the LBP code, expressed as:

[0096]

[0097] Where P is the number of neighborhood pixels, R is the neighborhood radius, and s(x) is the sign function;

[0098] When x ≥ 0, s(x) = 1;

[0099] When x<0, s(x)=0;

[0100] By counting the histograms of different LBP codes, the texture and edge features of the image are obtained.

[0101] Among them, GIS data feature extraction includes the following steps:

[0102] Calculate the terrain relief. The terrain relief (RD) reflects the degree of terrain change and is expressed as:

[0103]

[0104] Among them, z ij is the elevation value of grid cell ij, is the average elevation value in the study area, and n and m are the number of grid rows and columns in the study area, respectively.

[0105] Calculate the terrain slope. The terrain slope (Slope) indicates the degree of terrain inclination. In raster data, a 3×3 window is used for calculation, which is expressed as:

[0106]

[0107] in, and are the slope components in the x and y directions, respectively, which are calculated by the central difference method and are expressed as:

[0108]

[0109] Where Δx and Δy are the grid resolutions.

[0110] Calculate the river network density. The river network density (RND) is used to measure the density of river distribution in a region and is expressed as:

[0111]

[0112] Where L is the total length of rivers in the study area, and A is the area of the study area. In GIS, these two parameters are obtained by calculating the length of river vector data and the area of regional polygons.

[0113] The feature extraction of meteorological data includes the following steps:

[0114] Calculate the water vapor flux divergence. The water vapor flux divergence reflects the convergence and divergence of water vapor in the atmosphere and is expressed as:

[0115]

[0116] Where q is the specific humidity, is the horizontal wind speed vector, u and v are the wind speed components in the x and y directions, respectively.

[0117] Calculate the precipitation intensity. The precipitation intensity (PI) represents the amount of precipitation per unit time and is expressed as:

[0118]

[0119] Here, ΔP is the amount of precipitation during the time period Δt. In practice, this can be calculated based on observations from meteorological monitoring stations. For example, if a monitoring station receives 5 mm of precipitation within an hour, the precipitation intensity for that period is 5 mm / h.

[0120] Calculate the wind speed change rate. The wind speed change rate (VCR) is used to measure how fast the wind speed changes over time and is expressed as:

[0121]

[0122] in, and are the wind speeds at time t1 and t2, respectively. By calculating the rate of change of wind speed at different times, we can analyze the dynamic trend of wind speed and provide a reference for disaster warning.

[0123] Data fusion includes the following steps:

[0124] The features extracted from image data, GIS data and meteorological data are calibrated as feature matrices respectively. These feature matrices are spliced into a unified feature matrix X by column, X=[x1,x2,…,x n ], where x i is the feature vector of the i-th sample, and n is the number of samples. X is standardized so that the mean of each feature is 0 and the variance is 1, which is expressed as:

[0125]

[0126] Among them, x ij is the jth eigenvalue of the i-th sample in the original feature matrix X, is the mean of the jth feature, σ j is the standard deviation of the jth feature, x ij is the normalized eigenvalue.

[0127] The radial basis kernel function is calibrated as the kernel function, and the kernel matrix is calculated, which is expressed as:

[0128]

[0129] Among them, x i and x j is the feature vector of the two samples, ||x i -x j || is the Euclidean distance between them, and σ is the bandwidth parameter of the kernel function.

[0130] According to the selected kernel function, calculate the kernel matrix K, whose element K ij =K(x i ,x j ), i, j = 1, 2,…, n.

[0131] To make the KPCA results consistent with the principal component analysis (PCA) in data processing, the kernel matrix K needs to be centered. This includes the following steps:

[0132] Calculate the row mean vector of the kernel matrix K Its elements i=1,…,n;

[0133] Calculate the centralized kernel matrix K, expressed as:

[0134]

[0135] Calculate the eigenvalues and eigenvectors, and perform eigenvalue decomposition on the centralized kernel matrix K, that is, solve the characteristic equation Kα=λα, where λ is the eigenvalue and α is the corresponding eigenvector. The obtained eigenvalues (λ1≥λ2≥…≥λ n ), and the corresponding eigenvectors (α1,α2,…,α n ).

[0136] The number of principal components selected is determined based on the cumulative contribution rate. The cumulative contribution rate calculation formula is:

[0137]

[0138] Among them, the first k principal components of 80%-95% can be selected to make CR(k) reach a certain threshold. The eigenvectors α1, α2, ..., αk corresponding to the k principal components here are the basis vectors of KPCA transformation.

[0139] Calculate the fused feature vector and project the original data into the new feature space to obtain the fused low-dimensional feature vector. For the i-th sample, its fused feature vector y i The calculation formula is:

[0140]

[0141] Among them, α ji is the i-th element of the j-th eigenvector.

[0142] This technical solution uses the KPCA algorithm to complete the fusion of features from different data sources, achieve feature dimensionality reduction and nonlinear feature extraction, and provide more effective data features for subsequent abnormal area identification and early warning analysis.

[0143] Identifying abnormal areas includes the following steps:

[0144] Performing SARIMA model analysis includes the following steps:

[0145] Perform data stabilization processing and use ADF test (Augmented Dickey-Fuller Test) to test the stability of time series data, including: calibrating the time series to y t , the sequence after d-order difference is Δ d y t , where Δy t =y t -y t-1 , Δ dy t =Δ(Δ d-1 y t ).

[0146] Model order determination: By observing the autocorrelation function (ACF) and partial autocorrelation function (PACF) plots, we preliminarily determine the autoregressive order p, moving average order q, and seasonal period D (if the data has seasonality) of the SARIMA model. We then use the Akaike Information Criterion (AIC) to search for the optimal model parameter combinations (p, d, q) (for the non-seasonal portion) and (P, D, Q) (for the seasonal portion) within a certain range.

[0147] Model estimation and testing: A SARIMA model is constructed using the selected parameters, and the model parameters are estimated using maximum likelihood estimation. After estimation, the model residuals are tested for white noise. If the residual sequence is white noise, the model fits the data well; otherwise, the model parameters need to be readjusted.

[0148] Predicting time series values: Using the fitted SARIMA model to predict data at future time points, we can obtain a series of predicted values. Among them, the prediction formula is determined according to the structure of the SARIMA model.

[0149] The GWR model analysis includes the following steps:

[0150] Construct a geographic weight matrix and determine the Euclidean distance between each sample point based on the geographic location information of the data;

[0151] Construct a geographic weight matrix W based on the distance, whose element w ij Represents the weight between sample points i and j. It is calculated using the Gaussian kernel function and is expressed as:

[0152]

[0153] Among them, d ij is the distance between sample points i and j, and h is the bandwidth parameter, which controls the decay rate of spatial weights. The choice of bandwidth h can be determined by methods such as cross-validation.

[0154] For model estimation, the expression of the geographically weighted regression model is:

[0155]

[0156] Among them, u i ,v i is the geographic coordinate of sample point i, y i is the dependent variable, x ik is the kth independent variable, β k (ui ,v i ) is at position u i ,v i The regression coefficient at ∈ i is a random error term. The regression coefficient β of each location is estimated using weighted least squares method. k (u i ,v i ), minimize.

[0157] Abnormal area identification includes the following steps:

[0158] Calculate the residuals and convert the predicted values of the SARIMA model and the actual observed value y t For comparison, calculate the time series residual, which is expressed as:

[0159]

[0160] These residuals are combined with geographic information and input into the GWR model for spatial analysis.

[0161] According to the residual distribution of historical data, the abnormal threshold is determined using statistical methods.

[0162] This technical solution calculates the mean of the residuals and standard deviation σ, set the abnormal threshold to (k is 2 or 3.) If the residual at a certain location exceeds the threshold, it is considered that there may be an anomaly at that location.

[0163] To identify abnormal areas, we use the spatial regression results obtained from the GWR model and the abnormal threshold to judge each spatial location. If the residuals of multiple adjacent locations in a region exceed the abnormal threshold, the region is identified as a disaster abnormal area.

[0164] This technical solution visually displays the distribution of abnormal areas by drawing maps and other means, providing a decision-making basis for disaster warning.

[0165] Issue an early warning, use the preset disaster loss assessment model, consider various factors such as the population distribution of the mountainous area, the value of infrastructure, the expected precipitation intensity and the scope of the abnormal area to determine the warning level and issue the early warning information.

[0166] This technical solution sets the warning level to red (the highest level) if a large number of residents are expected to live in the abnormal area and the rainfall intensity is high, which may cause serious flooding. New data is collected regularly and model parameters are updated to ensure the accuracy and timeliness of the warning level.

[0167] For the aforementioned warning information release, a multilingual warning release system utilizes SMS messaging to send text messages to local residents, including the warning level, potentially affected areas, and precautionary recommendations. Warning messages are also pushed to relevant apps, with pop-up alerts appearing. Detailed warning information is also released on social media platforms such as Weibo and WeChat official accounts. A multi-channel backup mechanism ensures the stability and delivery of information, ensuring that the public and relevant departments receive warning information in a timely manner and take effective disaster prevention and mitigation measures.

[0168] In summary, with the help of the above technical solution of the present invention, the following effects can be achieved:

[0169] 1. This invention integrates image data, GIS data, and meteorological data, fully integrating the information contained in these different data types. Through comprehensive analysis of multi-source data, it can obtain multi-dimensional information on the disaster environment, geographical conditions, and meteorological factors. Compared with traditional early warning methods based on a single data source, this greatly improves the comprehensiveness and accuracy of information, providing a more reliable basis for identifying and warning abnormal disaster areas.

[0170] 2. The present invention analyzes time series data through the SARIMA model and combines it with the GWR model for spatial analysis, which can accurately identify disaster abnormal areas from both time and space dimensions. By using residual analysis and abnormal threshold judgment, the accuracy of abnormal area identification is effectively improved, the occurrence of misjudgment and missed judgment is reduced, and accurate decision-making support is provided for disaster warning. At the same time, in the data processing process, a variety of algorithms such as local binary pattern (LBP), terrain relief calculation, and water vapor flux divergence calculation are used to extract features from different data sources respectively. These features can accurately reflect the intrinsic characteristics of the data. At the same time, the KPCA algorithm is used to achieve feature dimensionality reduction and nonlinear feature extraction, which retains key information while reducing the data dimension, improves data processing efficiency, and makes subsequent abnormal area identification and warning analysis more efficient and accurate.

[0171] The foregoing is merely a preferred embodiment of the present invention and is not intended to limit the present invention. A person skilled in the art will readily appreciate other embodiments of the present invention after considering the disclosure in the specification and examples. This application is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered merely exemplary, and the true scope and spirit of the present invention are indicated by the claims.

[0172] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A multi-source data fusion method for identifying and warning abnormal disaster areas, characterized by: The following steps are involved: Perform data collection and preprocessing in advance, wherein the data collection includes at least: image data, GIS data and meteorological data; Perform feature extraction, extracting texture and edge features from image data based on local binary patterns; calculate terrain relief, terrain slope, and river network density from GIS data; and calculate water vapor flux divergence, precipitation intensity, and wind speed change rate from meteorological data; Perform data fusion, calibrate the features extracted from each data source into a feature matrix and splice them by column, and standardize the spliced feature matrix; calibrate the radial basis kernel function as the kernel function to calculate the kernel matrix, center the kernel matrix, perform eigenvalue decomposition, and calculate the low-dimensional feature vector after fusion; Identify abnormal areas, use the ADF test to stabilize time series data, determine SARIMA model parameters by observing autocorrelation and partial autocorrelation function graphs and combining them with the Akaike Information Criterion, build a model and perform a residual white noise test, and use the model to predict time series values; construct a geographic weight matrix and estimate the coefficients of the geographically weighted regression model using the weighted least squares method; calculate the residuals between the SARIMA model predictions and the actual observations, combine the geographic information into the GWR model, determine the anomaly threshold based on the residual mean and standard deviation, and identify disaster anomaly areas based on the GWR model results and the anomaly threshold; The warning level is determined based on the preset disaster loss assessment model, and warning information is released through a multilingual warning release system.

2. The method for identifying and warning abnormal disaster areas based on multi-source data fusion according to claim 1 is characterized in that: Image data feature extraction, including: texture and edge feature extraction based on local binary patterns, including the following steps: Calibrate the 3×3 neighborhood, the gray value of the central pixel is g c , the gray value of the neighborhood pixel is g i (i=0,1,…,7), obtain the LBP code, expressed as: Where P is the number of neighborhood pixels, R is the neighborhood radius, and s(x) is the sign function; When x ≥ 0, s(x) = 1; When x<0, s(x)=0; By counting the histograms of different LBP codes, the texture and edge features of the image are obtained.

3. The method for identifying and warning abnormal disaster areas based on multi-source data fusion according to claim 2 is characterized in that: GIS data feature extraction includes the following steps: Calculate the terrain relief, which reflects the degree of terrain change and is expressed as: Among them, z ij is the elevation value of grid cell ij, is the average elevation value in the study area, n and m are the number of grid rows and columns in the study area respectively; Calculate the terrain slope. The terrain slope indicates the degree of terrain inclination. In raster data, a 3×3 window is used for calculation, which is expressed as: in, and are the slope components in the x and y directions, respectively, which are calculated by the central difference method and are expressed as: Among them, Δx and Δy are the grid resolutions; Calculate the river network density, which is used to measure the density of river distribution in a region and is expressed as: Where L is the total length of rivers in the study area, and A is the area of the study area.

4. The method for identifying and warning abnormal disaster areas based on multi-source data fusion according to claim 3 is characterized in that: Meteorological data feature extraction includes the following steps: Calculate the water vapor flux divergence. The water vapor flux divergence reflects the convergence and divergence of water vapor in the atmosphere and is expressed as: Where q is the specific humidity, is the horizontal wind speed vector, u and v are the wind speed components in the x and y directions respectively; Calculate the precipitation intensity. The precipitation intensity PI represents the amount of precipitation per unit time and is expressed as: Where ΔP is the amount of precipitation in the time period Δt; Calculate the wind speed change rate. The wind speed change rate VCR is used to measure how fast the wind speed changes over time and is expressed as: in, and are the wind speeds at time t1 and t2 respectively.

5. The method for identifying and warning abnormal disaster areas based on multi-source data fusion according to claim 1, characterized in that: Data fusion includes the following steps: The features extracted from image data, GIS data and meteorological data are calibrated as feature matrices respectively, and these feature matrices are spliced into a unified feature matrix X by column, X=[x1,x2,…,x n ], where x i is the feature vector of the i-th sample, n is the number of samples, and X is standardized so that the mean of each feature is 0 and the variance is 1, which is expressed as: Among them, x ij is the jth eigenvalue of the i-th sample in the original feature matrix X, is the mean of the jth feature, σ j is the standard deviation of the jth feature, x ij is the normalized eigenvalue, The radial basis kernel function is calibrated as the kernel function, and the kernel matrix is calculated, which is expressed as: Among them, x i and x j is the feature vector of the two samples, ||x i -x j || is the Euclidean distance between them, σ is the bandwidth parameter of the kernel function, According to the selected kernel function, calculate the kernel matrix K, whose element K ij =K(x i ,x j ), i, j = 1, 2, ..., n; The central kernel matrix is constructed, including the following steps: Calculate the row mean vector of the kernel matrix K Its elements Calculate the centralized kernel matrix K, expressed as: Calculate the eigenvalues and eigenvectors, perform eigenvalue decomposition on the centralized kernel matrix K, that is, solve the characteristic equation Kα=λα, where λ is the eigenvalue and α is the corresponding eigenvector; the obtained eigenvalues (λ1≥λ2≥…≥λ n ), and the corresponding eigenvectors (α1,α2,…,α n ); The number of principal components selected is determined based on the cumulative contribution rate. The cumulative contribution rate calculation formula is: Calculate the fused feature vector, project the original data into the new feature space, and obtain the fused low-dimensional feature vector. For the i-th sample, the fused feature vector y i The calculation formula is: Among them, α ji is the i-th element of the j-th eigenvector.

6. The method for identifying and warning abnormal disaster areas based on multi-source data fusion according to claim 1, characterized in that: Identify abnormal areas, including: Perform SARIMA model analysis, including the following steps: Perform data stabilization processing and use ADF test to test the stability of time series data, including: calibrating the time series to y t , the sequence after d-order difference is Δ d y t , where Δy t =y t -y t-1 , Δ d y t =Δ(Δ d-1 y t ); Model order determination: by observing the autocorrelation function and partial autocorrelation function graphs, the autoregressive order p, moving average order q, and seasonal period D of the SARIMA model are determined. Then, the Akaike Information Criterion is used to search for the optimal model parameter combination (p, d, q) and (P, D, Q) within a certain range. Use the selected parameters to build a SARIMA model, and use the maximum likelihood estimation method to estimate the model parameters. Use the fitted SARIMA model to predict the data at future time points to obtain the predicted value sequence 7. The method for identifying and warning abnormal disaster areas based on multi-source data fusion according to claim 6, characterized in that: Also includes: Performing GWR model analysis includes the following steps: Construct a geographic weight matrix and determine the Euclidean distance between each sample point based on the geographic location information of the data; Construct a geographic weight matrix W based on the distance, whose element w ij Represents the weight between sample points i and j, which is calculated using the Gaussian kernel function and is expressed as: Among them, d ij is the distance between sample points i and j, h is the bandwidth parameter; For model estimation, the expression of the geographically weighted regression model is: Among them, u i ,v i is the geographic coordinate of sample point i, y i is the dependent variable, x ik is the kth independent variable, β k (u i ,v i ) is at position u i ,v i The regression coefficient at ∈ i is a random error term, and the regression coefficient β of each position is estimated using weighted least squares method k (u i ,v i ), minimize.

8. The method for identifying and warning abnormal disaster areas based on multi-source data fusion according to claim 7, characterized in that: Also includes: Identifying abnormal areas includes the following steps: Calculate the residuals and convert the predicted values of the SARIMA model and the actual observed value y t For comparison, calculate the time series residual, which is expressed as: These residuals are combined with geographic information and input into the GWR model for spatial analysis; Based on the residual distribution of historical data, use statistical methods to determine the abnormal threshold; To identify abnormal areas, the spatial regression results obtained based on the GWR model are combined with the abnormal threshold to judge each spatial location; if the residuals of multiple adjacent locations in a certain area exceed the abnormal threshold, the area is identified as a disaster abnormal area.