A landslide hazard multi-modal identification fusion method and system
Patent Information
- Application Number
- CN202411142406.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2044-08-20
AI Technical Summary
[0094]本发明针对现有的滑坡隐患识别技术存在的局限性,提出了一种滑坡隐患多模态识别融合方法及系统,旨在实现高精度滑坡隐患识别。一方面,以光学、热红外、微波、地形以及InSAR多源遥感数据为基础,引入光谱特征、地形属性以及水热因子的长时序变化趋势变量构成多光谱变化特征参数集,以地形指数结合InSAR形变数据构成地形特征参数集,进行多模态数据融合,全面地表征滑坡灾害潜在的变化特征,由于额外引入多源长时序多光谱变化特征能够提供更全面的背景信息及演化趋势,耦合两者可互补不足,可克服因单一依赖InSAR形变结果而影响滑坡隐患识别精度的问题。另一方面,滑坡隐患识别模型完成训练后,只需要将多光谱变化特征集以及地形特征参数集分别输入模型,即可快速获取识别结果,显著提升滑坡识别效率,大幅降低人力和物力成本,与传统方法需要进行人工判识相比,滑坡预测将更加智能便捷。
Smart Images

Figure CN118965270B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and system, and more particularly to a multimodal identification and fusion method and system for landslide hazards. Background Technology
[0002] Landslide hazard identification involves monitoring and analyzing the geological structure and geomorphological features of the Earth's surface and subsurface to identify potential landslide risks and predict disaster events. With continuous technological advancements and deeper applications, landslide hazard identification technology has gradually become a crucial tool in disaster prevention and mitigation. Traditional landslide identification methods primarily rely on manual observation and ground surveys. These methods are limited by human resources and geographical scope, making comprehensive monitoring and identification of large-scale, complex terrain impossible. However, the development of modern remote sensing technology has provided a new solution for landslide identification. Remote sensing can acquire vast amounts of data on and beneath the Earth's surface, including image data across multiple bands such as optical, thermal infrared, and radar, as well as various geographic information data such as topography, geomorphology, and vegetation cover. This data provides a rich information foundation for landslide identification. In recent years, with the development of machine learning and artificial intelligence technologies, landslide identification has entered a new stage. Machine learning algorithms can extract features from massive amounts of remote sensing data, identify the potential characteristics and patterns of landslides, and achieve automatic landslide identification and prediction. Traditional machine learning algorithms have been widely used in the field of landslide identification and have achieved remarkable results. However, landslides have complex and diverse features. How to extract effective features from a large amount of data and select them is a challenge that may directly affect the performance and final accuracy of the identification model.
[0003] Existing landslide hazard identification technologies mainly suffer from the following technical problems, which are reflected in the following two aspects: First, they rely on a limited number of InSAR images to obtain surface deformation, which is easily affected by a combination of factors such as time interval, surface cover, sensor noise, and atmospheric conditions; Second, the current main methods require combining surface deformation obtained from InSAR images with expert judgment, and then manually identifying and delineating the hazard range area area by area. This requires a lot of time for visual identification and manual operation, resulting in low identification efficiency and the results being highly susceptible to the influence of human subjective experience. Summary of the Invention
[0004] To address the shortcomings of the aforementioned technologies, this invention provides a multimodal identification and fusion method and system for landslide hazard identification.
[0005] To solve the above technical problems, the technical solution adopted by the present invention is: a multimodal identification and fusion method for landslide hazards, comprising the following steps:
[0006] S1. Acquisition and preprocessing of long-time-series multi-source data;
[0007] S2. Construction of Multispectral Variation Feature Parameter Set: Based on the close relationship between long-term multi-source data and landslides, a multispectral feature parameter set is selected and constructed, and the long-term series variation characteristics of the corresponding parameters are estimated one by one to construct the multispectral variation feature parameter set.
[0008] S3. Construction of the terrain feature parameter set:
[0009] A set of terrain feature parameters is constructed by combining deformation rate and terrain index obtained from InSAR data;
[0010] S4. Selection of training samples and model training: Based on the characteristics of surface changes in the landslide hazard area, the surface state is divided and model training samples are selected for different surface state types. The model is trained on the multispectral change feature parameter set and the terrain feature parameter set respectively until the model based on the multispectral change feature parameter and the model based on the terrain feature parameter are obtained.
[0011] S5. Multimodal data fusion: Multimodal fusion is performed on the prediction results of models based on multispectral variation characteristic parameters and models based on terrain characteristic parameters.
[0012] S6. Surface condition determination: Based on the multimodal fusion model obtained in the above steps, the surface condition is determined.
[0013] Preferably, in S1, the acquired data includes multi-source optical, thermal infrared, microwave, topographic, and InSAR remote sensing data, specifically including:
[0014] Landsat and MODIS surface reflectance datasets at different scales over long periods; digital elevation models; Landsat and MODIS surface temperature products; microwave soil moisture products, such as SMAP, including global surface and root soil moisture data; and corresponding time-series InSAR data.
[0015] Preferably, in S2, the process of constructing the multispectral variation characteristic parameter set specifically includes:
[0016] 2.1) Construction of a multispectral feature parameter set;
[0017] 2.2) Construction of a set of multispectral variation characteristic parameters;
[0018] In 2.1), the construction of the multispectral feature parameter set includes two steps: first, constructing the spectral index parameter set; second, constructing the hydrothermal factor spatial gradient parameter set.
[0019] The spectral index parameter set is constructed by combining spectral indices closely related to landslides. The spectral indices include: Normalized Difference Vegetation Index (NDVI), FVC (Fresh Cover), and Soil Modified Vegetation Index (SAVI).
[0020] The constructed set of spatial gradient parameters for hydrothermal factors includes: the surface temperature gradient index NLSTG and the soil moisture gradient index NSMG;
[0021] In step 2.2), based on the spectral indices NDVI, FVC, SAVI and hydrothermal factor spatial gradient parameter indices NLSTG and NSMG obtained in the previous step, a set of surface state characteristics is constructed. The corresponding set of multispectral variation characteristic parameters is further estimated. The variation characteristic parameters include: principal slope m, slope change rate r, intercept b, parameter free correlation ρ, and coefficient of variation CV. The long-term variation dynamics reflect the parameter dynamics related to landslides.
[0022] Therefore, each of the five surface state parameters has five variation characteristic parameters, resulting in a set V consisting of 25 multispectral variation characteristic parameters. SUR .
[0023] Preferably, in 2.1), the spectral indices include the Normalized Difference Vegetation Index (NDVI), the FVC (Fresh Cover), and the Soil Modified Vegetation Index (SAVI). The specific calculation formulas are as follows:
[0024]
[0025]
[0026] Among them, R nir Reflectance in the near-infrared band; R red is the reflectance in the red light band; L is the soil brightness correction factor, with a value ranging from (0,1), which is adjusted according to the specific vegetation type and soil characteristics; NDVI M and NDVI m These are the maximum and minimum NDVI values, corresponding to the NDVI values of bare soil and dense vegetation, respectively.
[0027] Preferably, in 2.1), the process of constructing the spatial gradient parameter set of the hydrothermal factor includes:
[0028] The first step is to downscale to a spatial resolution consistent with the spectral index parameters in order to obtain high-resolution surface temperature and soil moisture data.
[0029] The second step involves combining the high-resolution surface temperature and soil moisture data obtained in the previous step to estimate the spatial gradient of hydrothermal parameters according to the following steps: First, select the eight neighboring pixels around each pixel to estimate the gradient; second, calculate the pixel value difference between each pixel and its surrounding neighboring pixels; then, calculate the distance between the center pixel and each neighboring pixel; finally, divide the difference by the distance to calculate the spatial gradient of each pixel.
[0030] For each pixel, the spatial gradients of land surface temperature and soil moisture are respectively calculated by the following formulas:
[0031]
[0032] wherein, ΔX q is the average value of the parameter difference between the central pixel and all adjacent pixels; Δd q is the distance between the central pixel and the q-th adjacent pixel; XG is the calculated parameter spatial gradient of each pixel;
[0033] Third step, normalization processing of the spatial gradients of water and heat parameters; a linear normalization method is selected to scale the spatial gradients of land surface temperature and soil moisture to [-1, 1], so that their values are within a unified scale. The specific normalization formula is as follows:
[0034]
[0035] wherein, XG is the calculated parameter spatial gradient of each pixel; XG M and XG m are respectively the maximum value and the minimum value corresponding to the parameter spatial gradient of the central pixel; according to formulas (4) and (5), the normalized land surface temperature gradient index NLSTG and soil moisture gradient index NSMG can be calculated pixel by pixel.
[0036] Preferably, in step 2.1) and step 2.2), the specific calculation method of the change characteristic parameter is as follows:
[0037] (2.21) The calculation formula of the main slope m is:
[0038]
[0039] wherein, m n represents the slope of the n-th point pair, n represents the n-th point pair, i and j respectively represent different period numbers, satisfying i<j, and i, j∈[1,k]; y i and y j are respectively the values of the multispectral characteristic parameter corresponding to period i and period j; x i and x j are respectively the time corresponding to period i and period j; k is the total number of long time series data;
[0040] according to formula (6), two periods i and j are sequentially selected from the long time series data, and N main slopes can be finally calculated, wherein N m values are sequentially sorted from small to large to obtain an ordered point pair slope sequence {m1,m2...m N}, if N is an odd number, the median slope is the middle value of the ordered sequence, that is, the The slope corresponding to each point As the principal slope; if N is even, then the median slope is the first slope of the ordered sequence. and The average of the slopes at each point As the principal slope;
[0041] (2.22) The formula for calculating the rate of change of slope r is:
[0042]
[0043] Where r represents the rate of change of the average slope, k represents the total number of long-term time series data, and m ij Δt represents the rate of change of the slope corresponding to period i and period j. ij This represents the time interval between period i and period j;
[0044] (2.23) The formula for calculating the principal intercept b is:
[0045]
[0046] Where p represents the p-th pixel, s represents the s-th point pair, m represents the principal slope, and N is the total number of point pairs, which can be represented as: k is the total number of long-series data, P p,s t represents the value of the multispectral feature parameter corresponding to the s-th point pair of the p-th pixel. s This represents the time corresponding to the s-th point pair;
[0047] Sort the intercept values of all point pairs in ascending order to obtain an ordered point pair intercept sequence {b1, b2, ..., bb...} N If N is odd, then the median slope is the median value of the ordered sequence, i.e., the th... The intercepts of each point As the principal intercept; if N is even, then the median slope is the first digit of the ordered sequence. and The average of the slopes at each point As the principal intercept;
[0048] (2.24) The formula for calculating the autocorrelation ρ of the parameter is:
[0049]
[0050] in, Let P be the average value of parameter P corresponding to a long time series, and ρ(z) represent the degree of correlation between two parameter values separated by z time units in the long time series. k represents the total number of long-series data;
[0051] (2.25) The formula for calculating the coefficient of variation (CV) is:
[0052]
[0053] Where σ represents the standard deviation of the surface state parameter, This shows the average value of parameter P corresponding to a long time series, where k represents the total number of long time series data, P i Let i represent the i-th sample value in the surface state parameter dataset, where i∈[1,k];
[0054] After the above steps, five variation characteristic parameters {m,r,b,ρ,CV} corresponding to the spectral indices NDVI, FVC, SAVI, and the hydrothermal parameter spatial gradient indices NLSTG and NSMG are estimated. Once the calculation of these five variation characteristic parameters corresponding to the five surface state parameters is completed, a set V consisting of 25 multispectral variation characteristic parameters will be obtained. SUR The specific expression is as follows:
[0055]
[0056] Variation feature parameter set V SUR This will subsequently serve as the independent variable for machine learning, thus completing the input variable for training the machine learning algorithm.
[0057] Preferably, in S3, the process of constructing the terrain feature parameter set includes:
[0058] (3.1) Construct a set of terrain feature parameters: Use the preprocessed digital elevation model to obtain landslide-related terrain indices, including: terrain curvature index TCI and terrain turbulence index TPI;
[0059] First, the second derivative of the Earth's surface is calculated based on the surface elevation data, as shown in the following formula:
[0060]
[0061] Where C represents the surface curvature and E represents the surface elevation. Representing the partial derivative, x and y are the horizontal coordinates of the Earth's surface, both of which can be obtained from the digital elevation model;
[0062] Then, the calculated curvature is normalized as follows:
[0063]
[0064] Among them, C M and C mare the minimum and maximum values of surface curvature respectively; TCI is used to judge the concave-convex degree of terrain, TCI>0.7: indicates that the terrain is very flat, 0.3<TCI≤0.7: indicates that the terrain has certain undulations, TCI≤0.3: indicates that the terrain is very rugged;
[0065] The Topographic Position Index TPI can be obtained through statistics of elevation values around pixels, and the formula is as follows:
[0066]
[0067] wherein, E0 represents the elevation value of the central pixel, represents the average elevation of the surrounding neighborhood, both of which can be obtained from the digital elevation model; a positive TPI indicates that the central pixel is located at a relatively high position, a negative TPI indicates that the central pixel is located at a relatively low position, and 0 indicates that the elevation value of the central pixel is equal to the average value of the surrounding neighborhood;
[0068] (3.2) InSAR-based deformation rate estimation:
[0069] An interferogram is generated from the registered multi-temporal InSAR images, and the deformation rate v of each pixel is extracted; the deformation rate v, the topographic curvature index TCI and the topographic position index TPI data are combined into a data set, wherein the deformation rate data v is subjected to downscaling processing; the combined data set is subjected to normalization processing, thereby obtaining the topographic feature parameter set V DEF , V DEF will subsequently be used as an independent variable for machine learning, thus completing the input variables for machine learning algorithm training.
[0070] Preferably, in S4, the selection method of model training samples is:
[0071] According to the characteristics of surface change in hidden landslide hazard areas, the surface is divided into three states, including stable, slight change and unstable; a combination of random sample selection and hierarchical artificial sample selection is used to select training samples for different surface state types;
[0072] In S4, the model training process includes:
[0073] the set V composed of multispectral change characteristic parameters SUR and the topographic feature parameter set V DEF are respectively used as independent variables of machine learning, and the surface state sample set S is used as the dependent variable of machine learning;
[0074] The input variables and surface state samples are divided into 70% and 30% proportions, wherein 70% of V SUR_train and V DEF_train data are used for training of machine learning on surface state changes, and the model mapping training process is represented by the following formula:
[0075] f SUR =f train (V SUR_train ,S train (15)
[0076] f DEF =f train (V DEF_train ,S train (16)
[0077] Among them, S strain It is the 70% surface state sample set used for training; V SUR_train V DEF_train These are the 70% multispectral variation feature parameter set and terrain feature parameters used for training, respectively. SUR f DEF These are nonlinear models trained by machine learning, obtained by training a set of 70% of the selected input parameters and surface state samples.
[0078] The remaining 30% of the dataset is used to evaluate the performance of the machine learning algorithms to determine whether they have met the predetermined performance standards. If the algorithm performance fails to meet the predetermined accuracy requirements, iterative training is performed until its recognition accuracy reaches or exceeds the expected threshold, ultimately resulting in the two types of models mentioned above: the model based on multispectral variation feature parameters and the model based on terrain feature parameters.
[0079] Preferably, in S5, a probability averaging method is used for multimodal fusion, outputting the probability prediction for each category. For the two-class model, the probability prediction for each category is first estimated:
[0080] P sur (Y|V SUR_pre )=f SUR (V SUR_pre (17)
[0081] P def (Y|V DEF_pre )=f DEF (V DEF_pre (18)
[0082] Among them, f SUR f DEF V represents the prediction model trained using a multispectral variation feature parameter set and a terrain feature parameter set, respectively; SUR_pre and V DEF_pre P represents the set of multispectral variation feature parameters and topographic feature parameters involved in the prediction, respectively; sur f SUR For V SUR_preThe predicted probability P for each category Y def f DEF For V DEF_pre Predict the probability for each category Y;
[0083] Then, a weighted average of the probabilities for each category is calculated to obtain the overall probability for each category. Finally, the category with the highest overall probability is selected as the final prediction.
[0084] P(Y)=ω sur ×P sur (Y|V SUR_pre )+ω sar ×P sar (Y|V SAR_pre (19)
[0085] Where P(Y) represents the final class probability obtained by weighted averaging the probability predictions of different models; ω sur and ω def Let represent the probability weight coefficients of each model, and satisfy ω. sur +ω def =1;
[0086] After obtaining the weighted average probability P(Y) for each category, the category with the highest overall probability is finally selected as the final prediction result.
[0087] A multimodal identification and fusion system for landslide hazards includes the following modules:
[0088] The long-time-series multi-source data preprocessing module is used for the preprocessing of acquired long-time-series multi-source data.
[0089] The multispectral variation characteristic parameter estimation module uses 30m optical data to estimate the Normalized Difference Vegetation Index (NDVI), FVC (Further Vegetation Cover), and SAVI (Soil-Adjusted Vegetation Index). It employs various machine learning algorithms to train downscaling models, obtaining high-resolution land surface temperature and soil moisture data. This allows for the estimation of spatial gradient indices for hydrothermal parameters: the Land Surface Temperature Gradient Index (NLSTG) and the Soil Moisture Gradient Index (NSMG). For NDVI, FVC, SAVI, NLSTG, and NSMG, the module calculates the principal slope, slope change rate, intercept, parameter autocorrelation, and coefficient of variation, respectively, to obtain the multispectral variation characteristic parameter set.
[0090] The terrain feature parameter estimation module uses a digital elevation model to obtain the corresponding elevation, slope, and aspect, and estimates the terrain curvature index (TCI) and terrain turbulence index (TPII). It obtains the deformation rate of each pixel through multi-temporal InSAR images. The deformation rate data v is downscaled to a spatial resolution of 30 meters using a statistical downscaling method, and merged with the TCI and TPI data into a dataset. After normalization, the terrain feature parameter set is obtained.
[0091] The training sample selection and model training module selects samples from stable, slightly changing, and unstable regions respectively. The multispectral variation feature parameter set and the terrain feature parameter are used as independent variables for machine learning, and the surface state sample set is used as the dependent variable for machine learning. The model is iteratively trained and evaluated until a model based on multispectral variation feature parameters and a model based on terrain feature parameters are obtained.
[0092] The multimodal fusion module fuses two types of models, and the category with the highest overall probability will be used as the final prediction model.
[0093] The surface condition determination module calls the final prediction model to determine the surface condition and outputs a spatial distribution map of the regional surface condition.
[0094] This invention addresses the limitations of existing landslide hazard identification technologies by proposing a multimodal identification and fusion method and system for landslide hazard identification, aiming to achieve high-precision landslide hazard identification. On one hand, based on multi-source remote sensing data including optical, thermal infrared, microwave, topographic, and InSAR data, a multispectral variation feature parameter set is constructed by introducing long-term trend variables of spectral features, topographic attributes, and hydrothermal factors. A topographic feature parameter set is constructed by combining topographic indices with InSAR deformation data. Multimodal data fusion comprehensively characterizes the potential changing features of landslide hazards. The additional introduction of multi-source long-term multispectral variation features provides more comprehensive background information and evolution trends; coupling the two complements each other, overcoming the problem of relying solely on InSAR deformation results affecting the accuracy of landslide hazard identification. On the other hand, after the landslide hazard identification model is trained, only the multispectral variation feature set and the topographic feature parameter set need to be input into the model to quickly obtain identification results, significantly improving landslide identification efficiency and greatly reducing manpower and material costs. Compared with traditional methods that require manual identification, landslide prediction will be more intelligent and convenient. Attached Figure Description
[0095] Figure 1 This is a flowchart illustrating the overall technical process of the present invention.
[0096] Figure 2 This is a system module connection diagram of the present invention. Detailed Implementation
[0097] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0098] This invention fully considers the limitations of existing landslide hazard identification technologies and proposes a multimodal identification and fusion method and system for landslide hazards. First, multi-source optical, thermal infrared, microwave, topographic, and InSAR remote sensing data are acquired. Based on their relevance to landslides, corresponding spectral features and hydrothermal factor multispectral feature parameter sets are selected and constructed. The long-term series variation characteristics of each parameter are estimated to construct a multispectral variation feature parameter set. Then, a topographic feature parameter set is constructed by combining InSAR data deformation rate and topographic index. Models are trained on both the multispectral variation feature parameter set and the topographic feature parameter set. Finally, multimodal fusion is performed based on the prediction results of the model based on multispectral variation features and the model based on topographic feature parameters to obtain a comprehensive judgment result of the surface state, thereby improving the accuracy of landslide hazard identification.
[0099] The overall technical flow chart of the present invention is as follows: Figure 1 As shown, the process mainly consists of the following six steps: S1, acquisition and preprocessing of long-term multi-source data; S2, construction of multispectral variation feature parameter set; S3, construction of terrain feature parameter set; S4, selection of training samples and model training; S5, fusion of multimodal data; and S6, determination of surface condition.
[0100] Combination Figure 1 The technical flowchart and specific implementation methods of the landslide hazard multimodal identification and fusion method proposed in this invention are as follows:
[0101] S1. Acquisition and preprocessing of long-term multi-source data
[0102] The acquired data includes multi-source optical, thermal infrared, microwave, topographic, and InSAR remote sensing data, specifically:
[0103] Acquire long-term Landsat and MODIS surface reflectance datasets (30m and 500m spatial resolutions) at different scales, with time series >= 10 years. Months with minimal cloud, rain, and snow disturbances, and cloud cover less than 70%, are selected based on the actual region for estimating spectral indices. Digital elevation models (such as SRTM and ASTER GDEM) with a spatial resolution of 30m, covering the entire globe, can be used to extract elevation and slope, and will be subsequently used to estimate topographic indices. Download Landsat and MODIS surface temperature products (90m and 1km spatial resolutions) with time series >= 10 years. Download microwave soil moisture products since 2015, such as SMAP, including global surface and root soil moisture data, with a time resolution of 3 days and a spatial resolution of 1-3 kilometers, in HDF5 format. Acquire corresponding time-series InSAR data (such as Sentinel 1 and ALOS-2) with time series >= 5 years, and perform preprocessing including denoising, topographic correction, and atmospheric correction.
[0104] The downloaded multi-source products were geometrically corrected and reprojected, and the land surface temperature products and SMAP data were resampled to a resolution consistent with other data sources. Outliers were removed to ensure data accuracy and reliability, and the products were uniformly cropped into small areas of 30×30km (including 1000×1000 pixels). This preprocessing ensured the high quality and consistency of the acquired long-term remote sensing data.
[0105] S2. Construction of Multispectral Variation Characteristic Parameter Set
[0106] Based on the close relationship between long-term multi-source data and landslides, a set of multispectral feature parameters was selected and constructed. The long-term series variation characteristics of each parameter were then estimated to construct the multispectral variation feature parameter set. The specific process includes:
[0107] 2.1) Construction of Multispectral Feature Parameter Set
[0108] The construction of a multispectral feature parameter set mainly involves two steps: first, constructing a spectral index parameter set; and second, constructing a hydrothermal factor spatial gradient parameter set.
[0109] (2.11) Constructing the spectral index parameter set
[0110] Spectral indices closely related to landslides will be combined to construct a set of spectral index parameters, including the Normalized Difference Vegetation Index (NDVI), Fractional Vegetation Cover (FVC), and Soil-Adjusted Vegetation Index (SAVI). Among them, NDVI is an important indicator for measuring vegetation status and can be used to identify sparse or damaged vegetation areas and provide early warning of landslide risks; FVC can reflect the impact of changes in land cover type on vegetation-related indices; SAVI takes into account the moderating effect of soil on vegetation reflectance and can be used to monitor areas of soil exposure.
[0111] The three spectral indices mentioned above are obtained by calculating specific spectral bands using pre-processed surface reflectance products. The specific calculation formulas are as follows:
[0112]
[0113] Wherein, NDVI is the Normalized Difference Vegetation Index; FVC is vegetation cover; SAVI is the Soil-Regulated Vegetation Index; R nir Reflectance in the near-infrared band; R red is the reflectance in the red light band; L is the soil brightness correction factor, with a value range of (0,1), usually set to 0.5, which can be adjusted according to the specific vegetation type and soil characteristics to obtain more accurate results; NDVI M and NDVI m These are the maximum and minimum NDVI values, corresponding to the NDVI values of bare soil and dense vegetation, respectively.
[0114] The above steps complete the construction of three types of spectral index parameter sets.
[0115] (2.12) Constructing the spatial gradient parameter set of hydrothermal factors
[0116] The first step is to downscale to a spatial resolution consistent with the spectral index parameters.
[0117] To maintain spatial resolution consistent with spectral and topographic indices for both surface temperature and soil moisture, 30-meter resolution spectral indices (such as NDVI) can be used as auxiliary data. Various machine learning algorithms are employed to construct regression relationships between surface temperature and auxiliary variables, and between soil moisture and auxiliary variables, such as random forest, gradient boosting regression tree, and neural network methods. Simultaneously, for each regression model, the optimal parameter combination is selected, and the best machine learning algorithm is chosen to train the downscaling model, thereby obtaining high-resolution surface temperature and soil moisture data.
[0118] The second step is to estimate the spatial gradient of hydrothermal parameters.
[0119] Based on the high-resolution surface temperature and soil moisture data obtained in the previous step, the spatial gradients are estimated according to the following steps: First, the gradient is estimated by selecting the eight neighboring pixels (i.e., a 3×3 neighborhood) around each pixel; second, the pixel value difference between the central pixel and its neighboring pixels is calculated pixel by pixel, where the average of the pixel differences between the central pixel and each neighboring pixel can be used as the difference for that pixel; then, the distance between the central pixel and each neighboring pixel is calculated pixel by pixel, where Euclidean distance between pixels can be used to calculate the distance; finally, the spatial gradient of each pixel is calculated by dividing the difference by the distance. That is, for each pixel, the spatial gradients of surface temperature and soil moisture can be calculated using the following formulas:
[0120]
[0121] Where, ΔX q It is the average of the parameter differences between the central pixel and all its neighboring pixels; Δd q XG is the distance between the center pixel and the q-th neighboring pixel; XG is the parameter spatial gradient calculated for each pixel.
[0122] The third step is to normalize the spatial gradient of hydrothermal parameters.
[0123] Here, a linear normalization method can be used to scale the spatial gradients of surface temperature and soil moisture to [-1, 1], so that their values are within a uniform scale. The specific normalization formula is as follows:
[0124]
[0125] Where XG is the calculated parameter space gradient for each pixel; XG M With XG m These are the maximum and minimum values corresponding to the spatial gradient of the parameters of the central pixel, respectively. According to formulas (4) and (5), the normalized surface temperature gradient index NLSTG and soil moisture gradient index NSMG can be calculated pixel by pixel.
[0126] The above completes the acquisition of the spatial gradient exponents of two types of hydrothermal parameters, NLSTG and NSMG.
[0127] 2.2) Construction of Multispectral Variation Characteristic Parameter Set
[0128] Based on the set of surface state features obtained in the previous step, consisting of spectral indices (NDVI, FVC, SAVI) and hydrothermal parameter spatial gradient indices (NLSTG, NSMG), the corresponding set of multispectral variation feature parameters is further estimated. The variation feature parameters will be used as input for model training.
[0129] For the five parameters in the multi-spectral feature parameter set of the long time series, the following five change characteristic parameters are calculated respectively, including: main slope, slope change rate, intercept, parameter free correlation, and coefficient of variation. The dynamics of landslide-related parameters are reflected through the long time-series change dynamics. Thus, each of the five surface state parameters has five change characteristic parameters, that is, each pixel corresponds to 25 different types of information.
[0130] The specific calculation methods of the five change characteristic parameters are as follows:
[0131] (2.21) Main slope m
[0132] For long time-series data of a given parameter, the main slope is calculated using the pair-wise slope, that is, for each data pair in the time series, the slope between two adjacent data points is calculated sequentially, and the median value of these slopes is taken as the main slope on the time series. Assuming there are k phases of data in total, k-1 slopes need to be calculated in total, and the specific calculation formula of the slope is as follows:
[0133]
[0134] where, m n represents the slope of the nth point pair, n represents the nth point pair, i and j respectively represent different phase numbers, satisfying i<j, and i, j∈[1,k]; y i and y j respectively represent the multi-spectral feature parameter values corresponding to phases i and j; x i and x j respectively represent the time corresponding to phases i and j; k is the total number of long time-series data;
[0135] According to formula (6), two phases i and j are sequentially selected from the long time-series data, and N main slopes can be finally calculated, wherein Sort N m values from small to large to obtain an ordered point pair slope sequence {m1,m2...m N}, if N is an odd number, the median slope is the middle value of the ordered sequence, that is, the slope corresponding to the th point pair is taken as the main slope; if N is an even number, the median slope is the average of the slopes of the and th point pairs in the ordered sequence taken as the main slope.
[0136] (2.22) Slope change rate r
[0137] The rate of change of slope can help identify whether the trend of parameter change is accelerating or slowing down. The overall average rate of change of slope is obtained by averaging the rates of change of slope across all adjacent data points. The formula is:
[0138]
[0139] Where r represents the rate of change of the average slope, k represents the total number of long-term time series data, and m ij Δt represents the rate of change of the slope corresponding to period i and period j. ij This represents the time interval between period i and period j.
[0140] (2.23) Principal intercept b
[0141] After obtaining the principal slope corresponding to the long-term parameters, the difference between the multispectral feature parameter value of each point pair and the fitted trend line (the product of the principal slope and time) is calculated. All differences are averaged to obtain the principal intercept b of the p-th pixel. p It can be expressed as follows:
[0142]
[0143] Where p represents the p-th pixel, s represents the s-th point pair, m represents the principal slope, and N is the total number of point pairs, which can be represented as: k is the total number of long-series data, P p,s t represents the value of the multispectral feature parameter corresponding to the s-th point pair of the p-th pixel. s This represents the time corresponding to the s-th point pair.
[0144] Sort the intercept values of all point pairs in ascending order to obtain an ordered point pair intercept sequence {b1, b2, ..., bb...} N If N is odd, then the median slope is the median value of the ordered sequence, i.e., the th... The intercepts of each point As the principal intercept; if N is even, then the median slope is the first digit of the ordered sequence. and The average of the slopes at each point As the principal intercept.
[0145] (2.24) Parameter autocorrelation ρ
[0146] The autocorrelation of land surface state parameters is the correlation between the parameter value at a certain moment in a long-term series and the parameter value at a previous moment. The autocorrelation function can be used to calculate the autocorrelation. Assume the long-term land surface state parameters are represented as P = {P1, P2...P...} k}, where P kThis represents the parameter value for the k-th period. The formula for calculating the autocorrelation of the parameter is as follows:
[0147]
[0148] in, Let P be the average value of parameter P corresponding to a long time series, and ρ(z) represent the degree of correlation between two parameter values separated by z time units in the long time series. k represents the total number of long-term data. Here, the time interval z will be set according to the natural period of different surface state parameters.
[0149] (2.25) Coefficient of variation (CV)
[0150] The coefficient of variation (CV) is used to measure the relative variability or dispersion of data, that is, the magnitude of data fluctuation relative to the mean. A high CV value indicates that the data fluctuates significantly compared to the mean, while a low CV value indicates that the data is relatively stable. It is calculated as the ratio of the standard deviation of the land surface state parameter to the mean, and the specific formula is as follows:
[0151]
[0152] Where σ represents the standard deviation of the surface state parameter, This shows the average value of parameter P corresponding to a long time series, where k represents the total number of long time series data, P i Let i represent the i-th sample value in the surface state parameter dataset, where i∈[1,k].
[0153] After the above steps, five variation characteristic parameters {m,r,b,ρ,CV} corresponding to the spectral indices (NDVI, FVC, SAVI) and the spatial gradient indices of hydrothermal parameters (NLSTG, NSMG) are estimated. Once the calculation of these five variation characteristic parameters corresponding to the five surface state parameters is completed, a set V consisting of 25 multispectral variation characteristic parameters will be obtained. SUR The specific expression is as follows:
[0154]
[0155] Here, V SUR This is a set of change characteristic parameters consisting of 25 surface parameter variables corresponding to a pixel; NDVI is the Normalized Difference Vegetation Index, FVC is vegetation cover, SAVI is the Soil-Regulated Vegetation Index, NLSTG is the Spatial Gradient Index of Surface Temperature, and NSMG is the Spatial Gradient Index of Soil Moisture. m represents the principal slope, r represents the rate of change of slope, b represents the intercept, ρ represents the parameter autocorrelation, and CV represents the coefficient of variation. Change characteristic parameter set V SUR This will subsequently serve as the independent variable for machine learning, thus completing the input variable for training the machine learning algorithm.
[0156] S3. Construction of terrain feature parameter set
[0157] A terrain feature parameter set is constructed by combining the deformation rate obtained from InSAR data and terrain indices (Terrain Curvature Index TCI and Topographic Position Index TPI). The specific process includes:
[0158] (3.1) Construction of terrain feature parameter set
[0159] The corresponding elevation, slope and aspect are obtained by using the preprocessed digital elevation model, and terrain indices related to landslides, namely Terrain Curvature Index (TCI, Terrain Curvature Index) and Topographic Position Index (TPI, Topographic Position Index), are obtained through corresponding calculation.
[0160] First, the second derivative of the surface is calculated according to the surface elevation data. The specific formula is as follows:
[0161]
[0162] Wherein, C represents the surface curvature, E represents the surface elevation, represents the partial derivative, x and y are the horizontal positions of the surface coordinates, both of which can be obtained from the digital elevation model.
[0163] Then, the calculated curvature is normalized to ensure that the values are within a uniform range. The normalization calculation is as follows:
[0164]
[0165] Wherein, C M and C m are the minimum and maximum values of surface curvature respectively; TCI is the Terrain Curvature Index, which can be used to judge the convex-concave degree of terrain: TCI>0.7 indicates that the terrain is very flat, which usually occurs in plains, gentle slopes and other areas; 0.3<TCI≤0.7 indicates that the terrain has certain undulations, which usually occurs in hills or relatively gentle mountains; TCI≤0.3 indicates that the terrain is very rugged, which usually occurs in ridges, steep slopes and other areas.
[0166] The Topographic Position Index TPI can be obtained through statistics of elevation values around the pixel. The formula is as follows:
[0167]
[0168] Here, the value of TPI represents the relative position of the elevation value of the central pixel relative to the elevation values of the surrounding neighborhood, which can be estimated by selecting 8 neighboring pixels around each pixel (that is, a 3×3 neighborhood). E0 represents the elevation value of the central pixel, The average elevation of the surrounding neighborhood can be obtained from the digital elevation model. A positive TPI value indicates that the center pixel is located at a relatively high position, a negative value indicates that the center pixel is located at a relatively low position, and 0 indicates that the elevation value of the center pixel is equal to the average value of the surrounding neighborhood.
[0169] (3.2) Deformation rate estimation based on InSAR
[0170] Interferograms are generated from registered multi-temporal InSAR images. Using DInSAR technology or other deformation monitoring techniques, the deformation phase data of the multi-temporal images are analyzed by phase unwrapping and time series analysis to extract the deformation rate of each pixel.
[0171] Deformation rate data (v), TCI, and TPI data are merged into a single dataset. Statistical downscaling methods, such as regression models or interpolation techniques, are applied to the deformation rate data (v) to reduce it to a spatial resolution of 30 meters. The merged dataset is then normalized to avoid bias caused by differences in numerical ranges during analysis, thereby obtaining the terrain feature parameter set V. DEF This will subsequently serve as the independent variable for machine learning, thus completing the input variable for training the machine learning algorithm.
[0172] S4. Training Sample Selection and Model Training
[0173] Based on the characteristics of surface changes in landslide-prone areas, surface states are classified, and model training samples are selected for different surface state types. Models are trained on both multispectral variation feature parameter sets and topographic feature parameter sets until models based on multispectral variation feature parameters and topographic feature parameters are obtained. The specific process includes:
[0174] The first step is to select training samples for the model.
[0175] Based on the characteristics of surface changes in landslide-prone areas, the surface is divided into three states: stable, slightly changing, and unstable. A stable state indicates that the surface shows no or only minor deformation, demonstrating long-term stability without significant abnormal changes. A slightly changing state indicates that the surface begins to show slight but identifiable changes, possibly due to seasonal factors, minor human disturbance, or initial natural changes. An unstable state indicates that the surface shows significant deformation or other significant changes in surface state parameters, potentially indicating a potential landslide risk.
[0176] For training samples of different state types, a combination of random sample selection and stratified manual sample selection can be used:
[0177] The "stable" state samples utilize historical geological records, selecting areas that have not experienced geological disasters (such as landslides and debris flows) for a long period. Through years of optical and SAR imagery, historical records of surface deformation and geological disasters are observed to identify areas where no significant changes have occurred. Using a random point generation tool in GIS software, sample points are randomly selected from these considered stable areas. These generated sample points are then verified through field surveys or expert evaluation to ensure they are indeed located within stable regions. Based on the verification results, the sample point locations are adjusted to ensure the accuracy and representativeness of the samples, as well as their random geographical distribution. Considering that the number of "stable" state samples may far exceed those of the other two states, oversampling or undersampling techniques can be used to balance the samples, ensuring the model is not biased towards either state.
[0178] The "slightly changed" and "unstable" state samples are intended to use existing historical landslide event records to label the surface state of known disaster areas. In addition, by leveraging the knowledge of geologists and remote sensing experts, areas that may not have experienced recorded disaster events but have potential risks will be labeled. These areas will be labeled as "slightly changed" and "unstable" respectively, based on the existing InSAR deformation results. Expert review and possible on-site verification can be adopted to ensure the accuracy of the sample labeling, thereby constructing a complete surface state sample set S.
[0179] The second step is model training.
[0180] The set V consists of multispectral variation characteristic parameters. SUR and terrain feature parameter set V DEF The independent and dependent variables are used as independent variables in machine learning, while the surface state sample set S is used as the dependent variable. Surface states include three types: stable, slightly changing, and unstable. Both the independent and dependent variables serve as the input and output of the machine learning algorithm. During this process, appropriate machine learning or deep learning models, such as random forests, support vector machines, and neural networks, can be selected to simulate nonlinear and complex relationships.
[0181] To facilitate the training of the machine learning algorithm and subsequent error estimation, the input variables and surface state samples are divided into 70% and 30% respectively. The 70% V... SUR_train and V DEF_train The data will be used to train machine learning on changes in land surface conditions. The model mapping training process can be represented by the following formula:
[0182] f SUR =f train (V SUR_train ,S train (15)
[0183] f DEF =f train(V DEF_train ,S train (16)
[0184] Among them, S strain It is a sample set of surface conditions; V SUR_train V DEF_train These are the 70% multispectral variation feature parameter set and terrain feature parameters used for training, respectively. SUR f DEF These are nonlinear models trained by machine learning, obtained from a set of selected 70% input parameters and surface state samples.
[0185] The remaining 30% of the dataset will be used to evaluate the performance of the machine learning algorithms to determine whether they have met the predetermined performance standards. If the algorithm's performance fails to meet the established accuracy requirements, it must be iteratively trained until its recognition accuracy reaches or exceeds the expected threshold, ultimately resulting in the two types of models mentioned above: models based on multispectral variation feature parameters and models based on terrain feature parameters.
[0186] S5, Multimodal Data Fusion
[0187] Here, for the prediction results of the model based on multispectral variation feature parameters and the model based on terrain feature parameters, a probability averaging method is used for multimodal fusion, which can output the probability prediction for each category. For both types of models, the probability prediction for each category is first estimated:
[0188] P sur (Y|V SUR_pre )=f SUR (V SUR_pre (17)
[0189] P def (Y|V DEF_pre )=f DEF (V DEF_pre (18)
[0190] Among them, f SUR f DEF V represents the prediction model trained using a multispectral variation feature parameter set and a terrain feature parameter set, respectively; SUR_pre and V DEF_pre P represents the set of multispectral variation feature parameters and topographic feature parameters involved in the prediction, respectively; sur and P def They represent f respectively SUR and f DEF Predict the conditional probability for each possible category Y.
[0191] Then, the probabilities of each category are weighted and averaged to obtain the overall probability of each category. Finally, the category with the highest overall probability is selected as the final prediction.
[0192] P(Y)=ω sur ×P sur (Y|V SUR_pre )+ω sar ×P sar (Y|V SAR_pre (19)
[0193] Where P(Y) represents the final class probability obtained by weighted averaging of the probability predictions from different models; V SUR_pre and V DEF_pre P represents the set of multispectral variation feature parameters and topographic feature parameters involved in the prediction, respectively; sur f SUR For V SUR_pre The predicted probability P for each category Y def f DEF For V DEF_pre The predicted probability for each category Y; ω sur and ω def Let represent the probability weight coefficients of each model, and satisfy ω. sur +ω def =1;
[0194] After obtaining the weighted average probability P(Y) for each category, the category with the highest overall probability is finally selected as the final prediction result.
[0195] S6. Surface Condition Determination
[0196] Based on the multimodal fusion model obtained through the above steps, the aforementioned multispectral variation feature parameter set and terrain feature parameters are used as independent variables of the identification model to obtain corresponding prediction probabilities. Weighted averaging allows for the acquisition of the final surface state pixel by pixel. The results are mainly divided into three states: a stable state, a slightly changing state, and an unstable state. Finally, combining the multimodal fusion prediction results, a spatial distribution map of the regional surface state can be drawn.
[0197] like Figure 2 As shown in the figure, this embodiment also provides a multimodal fusion landslide hazard identification system, which includes six modules: a long-time-series multi-source data preprocessing module, a multispectral variation feature parameter set estimation module, a terrain feature parameter set estimation module, a training sample selection and model training module, a multimodal fusion module, and a surface state determination module. The functions of each module are as follows:
[0198] (1) Long-term multi-source data preprocessing module. Performs preprocessing including denoising, topographic correction, and atmospheric correction of raw remote sensing images, as well as geometric correction, reprojection, resampling, outlier removal, and region cropping of remote sensing products.
[0199] (2) Multispectral variation characteristic parameter set estimation module. NDVI, FVC, and SAVI are estimated using 30m optical data. Multiple machine learning algorithms are used to train downscaling models to obtain high-resolution surface temperature and soil moisture, and the spatial gradient exponents of hydrothermal parameters (NLSTG, NSMG) are estimated. For NDVI, FVC, SAVI, NLSTG, and NSMG, the principal slope, slope change rate, intercept, parameter autocorrelation, and coefficient of variation are calculated respectively to obtain the multispectral variation characteristic parameter set.
[0200] (3) Estimation of terrain feature parameters. The corresponding elevation, slope and aspect are obtained using the digital elevation model, and the TCI and TPI are estimated. The deformation rate of each pixel is obtained through multi-temporal InSAR images. The deformation rate data v is downscaled to a spatial resolution of 30 meters using a statistical method, and merged with the TCI and TPI data into a dataset. After normalization, the terrain feature parameter set is obtained.
[0201] (3) Training Sample Selection and Model Training Module. Samples were selected for stable, slightly changing, and unstable regions, respectively. The multispectral variation feature parameter set and terrain feature parameters were used as independent variables for machine learning, and the surface state sample set was used as the dependent variable for machine learning. The input variables and surface state samples were divided into 70% and 30% respectively. 70% of the data was used for training the machine learning model on surface state changes, and the remaining 30% of the dataset was used to evaluate the performance of the machine learning algorithm.
[0202] (4) Multimodal fusion module. The probability prediction for each category is estimated for both models, and the probability of each category is weighted and averaged to obtain the comprehensive probability of each category. The category with the highest comprehensive probability will be used as the final prediction.
[0203] (5) Surface State Determination Module. The surface state of each pixel is mainly divided into three states according to the category with the highest comprehensive probability: one is a stable state; the second is a state with slight changes; and the third is an unstable state. A spatial distribution map of the regional surface state is drawn.
[0204] In summary, the multimodal fusion landslide hazard identification method and system proposed in this invention addresses the limitations of current landslide hazard identification technologies and the high labor costs and low efficiency of existing technologies. Based on multi-source remote sensing data, the method trains models using long-term multispectral variation feature sets of parameters potentially related to landslides and topographic feature parameters, respectively. The results are then fused using multimodal methods. This approach can more comprehensively characterize the multispectral variation features related to landslides, capture dynamic change details, and efficiently and quickly identify landslide hazards, effectively improving the accuracy of identification. Compared with existing technologies, it has the following technical advantages:
[0205] This invention simultaneously considers the characteristic differences of the land surface in optical, thermal infrared, microwave and InSAR signals, introduces spectral, topographic and hydrothermal characteristic parameters that are closely related to landslide hazards, and performs multimodal fusion of the results based on long-time multispectral variation trend variables and topographic characteristic parameters to achieve a comprehensive characterization of the potential change characteristics of the land surface.
[0206] This invention only requires applying the multispectral variation feature set and terrain feature parameters to the machine learning model to quickly obtain pixel-by-pixel recognition results, greatly reducing manpower and material costs, and solving the problem of low efficiency caused by traditional methods that require a lot of time and manual processes.
[0207] The above embodiments are not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the technical solution of the present invention are also within the protection scope of the present invention.
Claims
1. A multimodal identification and fusion method for landslide hazard, characterized in that: Includes the following steps: S1. Acquisition and preprocessing of long-time-series multi-source data; S2. Construction of Multispectral Variation Feature Parameter Set: Based on the close relationship between long-term multi-source data and landslides, a multispectral feature parameter set is selected and constructed, and the long-term series variation characteristics of the corresponding parameters are estimated one by one to construct the multispectral variation feature parameter set. S3. Construction of the terrain feature parameter set: A set of terrain feature parameters is constructed by combining deformation rate and terrain index obtained from InSAR data; S4. Selection of training samples and model training: Based on the characteristics of surface changes in the landslide hazard area, the surface state is divided and model training samples are selected for different surface state types. The model is trained on the multispectral change feature parameter set and the terrain feature parameter set respectively until the model based on the multispectral change feature parameter and the model based on the terrain feature parameter are obtained. S5. Multimodal data fusion: Multimodal fusion is performed on the prediction results of the model based on multispectral variation characteristic parameters and the model based on terrain characteristic parameters to obtain a multimodal fusion model; In S5, a probability averaging method is used for multimodal fusion, outputting the probability prediction for each class. For the two-class model, the probability prediction for each class is first estimated: (17) (18) in, , These represent the prediction models trained using a set of multispectral variation feature parameters and terrain feature parameters, respectively. as well as These represent the set of multispectral variation feature parameters and the terrain feature parameters involved in the prediction, respectively. express for The predicted probability for each category Y, express for Predict the probability for each category Y; Then, a weighted average of the probabilities for each category is calculated to obtain the overall probability for each category. Finally, the category with the highest overall probability is selected as the final prediction. (19) Where P(Y) represents the final class probability obtained by weighted averaging the probability predictions of different models; as well as Let represent the probability weight coefficients of each model, and satisfy . + =1; After obtaining the weighted average probability P(Y) for each category, the category with the highest overall probability is finally selected as the final prediction result. S6. Surface condition determination: Based on the multimodal fusion model obtained in step S5, the surface condition is determined.
2. The landslide hazard multimodal identification and fusion method according to claim 1, characterized in that: In S1, the acquired data includes multi-source optical, thermal infrared, microwave, topographic, and InSAR remote sensing data, specifically including: Landsat and MODIS surface reflectance datasets at different scales over long periods; digital elevation models; Landsat and MODIS surface temperature products; microwave soil moisture products, SMAP, including global surface and root soil moisture data; and corresponding time-series InSAR data.
3. The landslide hazard multimodal identification and fusion method according to claim 1, characterized in that: In S2, the specific process of constructing the multispectral variation characteristic parameter set includes: 2.1) Construction of a multispectral feature parameter set; 2.2) Construction of a set of multispectral variation characteristic parameters; In 2.1), the construction of the multispectral feature parameter set includes two steps: first, constructing the spectral index parameter set; second, constructing the hydrothermal factor spatial gradient parameter set. The spectral index parameter set is constructed by combining spectral indices closely related to landslides. The spectral indices include: Normalized Difference Vegetation Index (NDVI), FVC (Fresh Cover), and Soil Modified Vegetation Index (SAVI). The constructed set of spatial gradient parameters for hydrothermal factors includes: the surface temperature gradient index NLSTG and the soil moisture gradient index NSMG; In step 2.2), based on the spectral indices NDVI, FVC, SAVI and hydrothermal factor spatial gradient parameter indices NLSTG and NSMG obtained in the previous step, a set of surface state characteristics is constructed. The corresponding set of multispectral variation characteristic parameters is further estimated. The variation characteristic parameters include: principal slope m, slope change rate r, intercept b, parameter free correlation ρ, and coefficient of variation CV. The long-term variation dynamics reflect the parameter dynamics related to landslides. Therefore, each of the five surface state parameters has five variation characteristic parameters, resulting in a set of 25 multispectral variation characteristic parameters. .
4. The landslide hazard multimodal identification and fusion method according to claim 3, characterized in that: In section 2.1), the specific calculation formulas for the spectral indices, including the Normalized Difference Vegetation Index (NDVI), FVC (Fresh Cover), and SAVI (Soil-Adjusted Vegetation Index), are as follows: (1) (2) (3) in, The reflectivity is in the near-infrared band; is the reflectance in the red light band; L is the soil brightness correction factor, with a value range of (0,1), which is adjusted according to the specific vegetation type and soil characteristics; and These are the maximum and minimum NDVI values, corresponding to the NDVI values of bare soil and dense vegetation, respectively.
5. The landslide hazard multimodal identification and fusion method according to claim 4, characterized in that: In section 2.1), the process of constructing the spatial gradient parameter set of the hydrothermal factor includes: The first step is to downscale to a spatial resolution consistent with the spectral index parameters in order to obtain high-resolution surface temperature and soil moisture data. The second step involves combining the high-resolution surface temperature and soil moisture data obtained in the previous step to estimate the spatial gradient of hydrothermal parameters according to the following steps: First, select the eight neighboring pixels around each pixel to estimate the gradient; second, calculate the pixel value difference between each pixel and its surrounding neighboring pixels; then, calculate the distance between the center pixel and each neighboring pixel; finally, divide the difference by the distance to calculate the spatial gradient of each pixel. For each pixel, the spatial gradients of surface temperature and soil moisture are calculated using the following formulas: (4) in, It is the average of the parameter differences between the center pixel and all its neighboring pixels; XG is the distance between the center pixel and the q-th neighboring pixel; XG is the parameter spatial gradient calculated for each pixel. The third step is to normalize the spatial gradient of hydrothermal parameters. A linear normalization method is selected to scale the spatial gradients of surface temperature and soil moisture to [-1, 1], so that their values are within a uniform scale. The specific normalization formula is as follows: (5) Where XG is the calculated parameter space gradient for each pixel; and These are the maximum and minimum values corresponding to the spatial gradient of the parameters of the central pixel, respectively; the normalized surface temperature gradient index NLSTG and soil moisture gradient index NSMG are calculated pixel by pixel according to formulas (4) and (5).
6. The landslide hazard multimodal identification and fusion method according to claim 5, characterized in that: In sections 2.1) and 2.2), the specific calculation methods for the changing characteristic parameters are as follows: (2.21) The formula for calculating the principal slope m is: (6) wherein, represents the slope of the n-th point pair, n represents the n-th point pair, i and j respectively represent different period numbers, satisfying i<j, and ; and are respectively the multispectral feature parameter values corresponding to periods i and j; and are respectively the times corresponding to periods i and j; k is the total number of long time-series data; According to formula (6), two periods i and j are selected sequentially from the long-term time series data, and N main slopes are finally calculated, among which Sort the N m values in ascending order to obtain an ordered sequence of point-pair slopes. If N is odd, then the median slope is the median value of the ordered sequence, i.e., the th... The slope corresponding to each point As the principal slope; if N is even, then the median slope is the first slope of the ordered sequence. and The average of the slopes at each point As the principal slope; (2.22) The formula for calculating the rate of change of slope r is: (7) Where r represents the rate of change of the average slope, and k represents the total number of long-term time series data. This represents the rate of change of the slope corresponding to period i and period j. This represents the time interval between period i and period j; (2.23) The formula for calculating the principal intercept b is: (8) Where p represents the p-th pixel, s represents the s-th point pair, m represents the principal slope, and N is the total number of point pairs, denoted as k is the total number of long-series data. This represents the value of the multispectral feature parameter corresponding to the s-th point pair of the p-th pixel. This represents the time corresponding to the s-th point pair; Sort the intercept values of all point pairs in ascending order to obtain an ordered point pair intercept sequence. If N is odd, then the median slope is the median value of the ordered sequence, i.e., the th... The intercepts of each point As the principal intercept; if N is even, then the median slope is the first digit of the ordered sequence. and The average of the slopes at each point As the principal intercept; (2.24) The formula for calculating the autocorrelation ρ of the parameter is: (9) in, Let P be the average value of parameter P corresponding to a long time series, and ρ(z) represent the degree of correlation between two parameter values separated by z time units in the long time series. k represents the total number of long-series data; (2.25) The formula for calculating the coefficient of variation (CV) is: (10) Where σ represents the standard deviation of the surface state parameter, The figure shows the average value of parameter P corresponding to the long time series, where k represents the total number of long time series data. This represents the i-th sample value in the surface state parameter dataset. ; After the above steps, five characteristic parameters corresponding to the spectral indices NDVI, FVC, SAVI, and the hydrothermal parameter spatial gradient indices NLSTG and NSMG were estimated. Once the calculations for the five change characteristic parameters corresponding to the five surface state parameters are completed, a set of 25 multispectral change characteristic parameters will be obtained. The specific expression is as follows: (11) Change feature parameter set This will subsequently serve as the independent variable for machine learning, thus completing the input variable for training the machine learning algorithm.
7. The landslide hazard multimodal identification and fusion method according to claim 1, characterized in that: In S3, the process of constructing the terrain feature parameter set includes: (3.1) Construct a set of terrain feature parameters: Use the preprocessed digital elevation model to obtain landslide-related terrain indices, including: terrain curvature index TCI and terrain turbulence index TPI; First, the second derivative of the Earth's surface is calculated based on the surface elevation data, as shown in the following formula: (12) Where C represents the surface curvature and E represents the surface elevation. Representing the partial derivative, x and y are the horizontal coordinates of the Earth's surface, both of which can be obtained from the digital elevation model; Then, the calculated curvature is normalized as follows: (13) in, and These are the minimum and maximum values of the surface curvature, respectively; TCI is used to determine the degree of unevenness of the terrain. TCI > 0.7: indicates that the terrain is very flat; 0.3 < TCI ≤ 0.7: indicates that the terrain has some undulations; TCI ≤ 0.3: indicates that the terrain is very rugged. The Terrain Turbulence Index (TPI) can be obtained by statistically analyzing the elevation values around a pixel, as shown in the following formula: (14) in, Indicates the elevation value of the center pixel. The average elevation of the surrounding neighborhood can be obtained from the digital elevation model; a positive TPI value indicates that the center pixel is located at a relatively high position, a negative value indicates that the center pixel is located at a relatively low position, and 0 indicates that the elevation value of the center pixel is equal to the average value of the surrounding neighborhood. (3.2) Deformation rate estimation based on InSAR: Interferograms are generated from registered multi-temporal InSAR images, and the deformation rate v of each pixel is extracted. The deformation rate (v), topographic curvature index (TCI), and topographic turbulence index (TPI) data were merged into a single dataset, with the deformation rate data (v) being downscaled. The merged dataset was then normalized to obtain the set of topographic feature parameters. , This will subsequently serve as the independent variable for machine learning, thus completing the input variable for training the machine learning algorithm.
8. The landslide hazard multimodal identification and fusion method according to claim 1, characterized in that: In S4, the method for selecting model training samples is as follows: Based on the characteristics of surface changes in landslide hazard areas, the surface is divided into three states: stable, slightly changed, and unstable. Training samples are selected for different surface state types by combining random selection and stratified manual selection. In S4, the model training process includes: Multispectral variation characteristic parameter set and terrain feature parameter set The surface state sample set S is used as the independent variable in machine learning, and the surface state sample set S is used as the dependent variable in machine learning. The input variables and surface condition samples were divided into 70% and 30% respectively, with the 70% being... as well as The data will be used to train machine learning on changes in land surface condition. The model mapping training process is represented by the following formula: (15) (16) in, It is the 70% surface state sample set used for training; , These are the 70% multispectral variation feature parameter set and terrain feature parameters used for training, respectively. , These are nonlinear models trained by machine learning, obtained by training with 70% of the selected input parameters and a sample set of surface state data. The remaining 30% of the dataset is used to evaluate the performance of machine learning algorithms to determine whether they have met the predetermined performance standards. If the algorithm performance fails to meet the predetermined accuracy requirements, iterative training is performed until its recognition accuracy reaches or exceeds the expected threshold, ultimately resulting in the two types of models mentioned above: the model based on multispectral variation feature parameters and the model based on terrain feature parameters.
9. A multimodal identification and fusion system for landslide hazards, employing the multimodal identification and fusion method for landslide hazards as described in claim 1, characterized in that: The landslide hazard multimodal identification and fusion system includes the following modules: The long-time-series multi-source data preprocessing module is used for the preprocessing of acquired long-time-series multi-source data. The multispectral variation characteristic parameter estimation module uses 30m optical data to estimate the Normalized Difference Vegetation Index (NDVI), FVC (Further Vegetation Cover), and SAVI (Soil-Adjusted Vegetation Index). It employs various machine learning algorithms to train downscaling models, obtaining high-resolution land surface temperature and soil moisture data. This allows for the estimation of spatial gradient indices for hydrothermal parameters: the Land Surface Temperature Gradient Index (NLSTG) and the Soil Moisture Gradient Index (NSMG). For NDVI, FVC, SAVI, NLSTG, and NSMG, the module calculates the principal slope, slope change rate, intercept, parameter autocorrelation, and coefficient of variation, respectively, to obtain the multispectral variation characteristic parameter set. The terrain feature parameter estimation module uses a digital elevation model to obtain the corresponding elevation, slope, and aspect, and estimates the terrain curvature index (TCI) and terrain turbulence index (TPII). It obtains the deformation rate of each pixel through multi-temporal InSAR images. The deformation rate data v is downscaled to a spatial resolution of 30 meters using a statistical method, and merged with the TCI and TPI data into a dataset. After normalization, the terrain feature parameter set is obtained. The training sample selection and model training module selects samples from stable, slightly changing, and unstable regions respectively. The multispectral variation feature parameter set and the terrain feature parameter are used as independent variables for machine learning, and the surface state sample set is used as the dependent variable for machine learning. The model is iteratively trained and evaluated until a model based on multispectral variation feature parameters and a model based on terrain feature parameters are obtained. The multimodal fusion module fuses two types of models, and the category with the highest comprehensive probability will be used as the final prediction model. The surface condition determination module calls the final prediction model to determine the surface condition and outputs a spatial distribution map of the regional surface condition.
Citation Information
Patent Citations
Hot melt pond state identification method combining long time sequence data and machine learning algorithm
CN112801007A
Method, system, device and medium for landslide identification based on full polarimetric SAR
US11747498B1