Hidden danger identification method and system based on multi-source disaster factor and dynamic grid model
Patent Information
- Application Number
- CN202610873421.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-08
AI Technical Summary
[0014]鉴于传统地灾隐患识别方案中存在的技术缺陷,本发明的目的是提供一种基于多源地灾因子与动态格网模型的隐患识别方法,通过动态格网和集成机器学习模型进行多源遥感数据的地灾隐患识别,以解决目前地灾隐患识别方案中存在的采用单一数据源、需要人工布设工程等问题,实现非接触、大范围、自适应、高精度、可解释的地质灾害隐患识别
[0024]从上面的技术方案可知,本发明提供的基于多源地灾因子与动态格网模型的隐患识别方法及系统,以动态格网替代逐像素识别,能够克服逐像素识别忽略区域相关性、结果不稳定的缺陷,实现多源异构遥感与非遥感地灾数据的标准化融合;通过将多源地灾因子深度融合,全面反映致灾机理,且通过动态格网自适应划分适配不同规模、不同形态滑坡隐患;通过异构集成模型提升模型在样本不均衡、高维特征下的识别精度与泛化能力;决策路径清晰,保证识别模型具备可解释性,满足地质灾害监管与机理分析需求;并且非接触式空天地一体化识别方案,无需布设传感器,能够有效避免人员涉险与工程扰动。
Smart Images

Figure CN122715005A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of geological disaster monitoring and remote sensing intelligent identification technology, and more specifically, to a method and system for identifying hidden dangers based on multi-source geological disaster factors and a dynamic grid model. Background Technology
[0002] Frequent geological disasters, such as landslides, collapses, and debris flows, severely disrupt and threaten the daily operation of infrastructure in various fields, and also pose a serious threat to the lives and property of people living along the affected areas. Therefore, the monitoring and early warning of geological disasters has always been an important topic in the fields of geological engineering and earth sciences.
[0003] Traditional methods for monitoring geological hazards mainly include ground observation stations, GPS measurements, and InSAR (Interferometric Synthetic Aperture Radar). While these methods can provide accurate data to a certain extent, they suffer from limitations such as limited coverage, high costs, and insufficient real-time performance. With the rapid development of remote sensing technology, especially the widespread application of high-resolution satellite remote sensing and UAV remote sensing, new solutions have been provided for monitoring geological hazards.
[0004] The core of addressing this challenge lies in the accurate early identification and continuous monitoring of landslide hazards. For a long time, traditional methods based on geodesy have formed the cornerstone of identification and monitoring work. For example, by deploying a high-precision Global Navigation Satellite System (GNSS) monitoring network, the three-dimensional displacement changes of key points on the Earth's surface can be tracked in real time; inclinometers deployed deep underground can detect the activity of deep slip surfaces invisible to the naked eye, providing direct evidence for analyzing the overall extent of the landslide, its dynamic deformation process, its intrinsic deformation mechanism, and the inducing and disastrous factors. These methods are targeted, technologically mature, and provide direct data. They are crucial for key monitoring and disaster assessment of landslides that have already shown initial signs of deformation, and are an indispensable basis for taking targeted prevention and control measures.
[0005] In addition, Synthetic Aperture Radar Interferometry (InSAR) technology has gained widespread application due to its unique advantages. This technology utilizes radar sensors mounted on satellites or aircraft to repeatedly observe the same area. Through precise calculation of radar signal phase information, it can invert large-scale subtle surface deformations with centimeter- to millimeter-level accuracy. It boasts outstanding characteristics such as high spatial resolution, wide coverage (up to thousands of square kilometers in a single imaging session), immunity to adverse weather conditions like clouds, fog, rain, and snow, and the ability to operate around the clock. This makes it extremely valuable in monitoring long-term, slow deformations in areas such as seismic coseismic deformation, surface subsidence in mining areas, and glacial migration, and it is gradually becoming an important means of early identification and surveying of regional landslide hazards. By processing time-series SAR data on a large scale (such as multi-temporal InSAR technology), it is possible to extract signs of long-term, slow surface creep from massive amounts of data, pointing to potential unstable slopes.
[0006] The accurate identification of landslide hazards is now moving towards a comprehensive observation model that integrates multi-source data fusion and three-dimensional air-space-ground collaboration. In addition to SAR radar satellites, high-resolution optical satellite imagery and UAV remote sensing technology are also showing great potential. UAV platforms are flexible and maneuverable, flying near the ground to acquire centimeter-resolution three-dimensional terrain (digital surface model, DSM), detailed surface texture, and fracture information. By carrying multispectral / thermal infrared sensors, they can capture potential precursors of instability such as surface water content and temperature anomalies, effectively compensating for the shortcomings of spaceborne InSAR in capturing surface details, making them particularly suitable for detailed investigations of small areas or key locations.
[0007] Meanwhile, at the technical level, with the development of artificial intelligence (AI) and machine learning (ML) technologies, by training deep learning models, it is possible to automatically and intelligently identify landslide-related geomorphic markers (such as abnormal slope depressions, tensile crack groups at the top of the slope, and mounds), abnormal vegetation changes, and other indirect deformation indicators, as well as the correlation between various markers, from massive multi-source remote sensing images. This greatly improves the screening efficiency and data mining depth for potential hidden danger targets.
[0008] In addition, there are low-cost, low-power, high-density intelligent sensor networks (covering parameters such as micro-seismic activity, ground sound, soil moisture, and pore water pressure) and real-time data transmission combined with IoT technology. By constructing a ground-based perception network covering key hazard areas, it forms a three-dimensional synergy with aerospace remote sensing technology to achieve integrated "air-space-ground" perception of the entire life cycle of landslides.
[0009] Existing landslide hazard identification technologies also have obvious limitations. For example, GNSS monitoring networks are costly to deploy and maintain, and their coverage is relatively limited, only allowing for fixed-point monitoring of known or highly suspected areas. Especially for hidden landslides occurring in sparsely populated, steep, and dangerous areas, effective control is often difficult to implement in the early stages of disaster development, resulting in monitoring blind spots. This method still has some objective shortcomings in the large-scale application of landslide hazard identification. In addition, the deployment of monitoring equipment requires personnel to venture into dangerous areas, and manual deployment activities may further induce landslides and other hazards.
[0010] Furthermore, the dramatic topographical variations in landslide-prone areas pose significant challenges to InSAR technology. Firstly, the inherent geometric limitations of radar side-looking imaging make signal acquisition difficult on steep slopes (such as areas moving towards the radar or perpendicular to the line of sight), sometimes even resulting in "shadow" blind spots. Secondly, dense vegetation growth leads to severe decorrelation of radar echo signals between observations (spatiotemporal decoherence); and fluctuations in atmospheric water vapor content (atmospheric decoherence) significantly interfere with the reliability of phase information in SAR data. This decoherence results in large "no-data areas" or significantly reduced accuracy in radar images, thus limiting the effectiveness of InSAR technology in monitoring these complex terrain areas.
[0011] Currently, when using machine learning methods for geological disaster hazard identification, it is typically necessary to determine the region of interest (ROI) of potential geological disaster hazards pixel by pixel. This ROI is then used to identify relevant hazard characteristic indicators within the ROI region, thereby constructing training samples to analyze the risk level and probability of geological disaster hazards existing within that ROI. This pixel-by-pixel hazard identification method struggles to estimate the domain relevance and regional statistical characteristics of hazard pixels. This limits identification accuracy and, because the final identification result is pixel-by-pixel, it is prone to generating a large number of unstable identification pixel results.
[0012] Despite significant progress in remote sensing technology and multi-source data fusion methods for geological disaster monitoring, some limitations remain. For example, existing technologies often focus on the analysis of single or a few geological disaster factors, lacking comprehensive integration of multiple sources. Furthermore, due to the complexity and uncertainty of geological disasters, traditional models still need improvement in prediction accuracy and timeliness. In addition, with the advent of the big data era, how to efficiently process and analyze massive amounts of remote sensing data to extract useful information is also a current technical challenge.
[0013] In addition, although multi-source data fusion has certain advantages in identifying landslide hazards, there are still shortcomings in how to effectively integrate massive heterogeneous data with different spatiotemporal scales, different precisions, and different physical meanings, and transform them into reliable instability early warning thresholds that can be used for decision-making, as well as how to further optimize the generalization ability and interpretability of AI models. Summary of the Invention
[0014] In view of the technical defects in traditional geological hazard identification schemes, the purpose of this invention is to provide a hazard identification method based on multi-source geological hazard factors and dynamic grid models. By using dynamic grids and integrated machine learning models to identify geological hazard hazards from multi-source remote sensing data, this invention aims to solve the problems of using a single data source and requiring manual engineering deployment in current geological hazard identification schemes, thereby achieving non-contact, large-scale, adaptive, high-precision, and interpretable geological hazard identification.
[0015] On the one hand, this invention provides a hazard identification method based on multi-source geological disaster factors and a dynamic grid model, comprising: Acquire multi-source geological hazard factor data of the target area, and preprocess the multi-source geological hazard factor data to form disaster-causing hazard factors; wherein, the multi-source geological hazard factor data includes remote sensing data and non-remote sensing data; The scale, shape and data source resolution of geological disaster hazards are determined based on the disaster-causing hazard factors, and the size and shape of the dynamic grid are adaptively determined to divide the target area into multiple dynamic grid units. For each dynamic grid cell, regional statistical features of disaster-causing factors are extracted based on the pixel values of the grid area. These disaster-causing factors include terrain-related factors, symptom factors, and external triggering factors. The pre-constructed heterogeneous ensemble learning model is trained based on the regional statistical characteristics of the disaster-causing hazard factors; wherein, the heterogeneous ensemble learning model includes three base learners: decision tree, support vector machine, and feedforward neural network, as well as a linear meta-model for fusion output; The prediction results of each base learner on the test set are input into the linear meta-model for weighted fusion to obtain the geological disaster hazard identification result.
[0016] Alternatively, the remote sensing data may include optical remote sensing data, DEM data, and SAR time-series deformation data, while the non-remote sensing data may include meteorological precipitation data, geological lithology data, hydrological data, land use data, and human engineering activity data.
[0017] In addition, an optional approach is to preprocess the remote sensing data including radiometric correction, atmospheric correction, georegistration and geometric correction, orthorectification, and panchromatic sharpening; wherein, The radiometric correction is used to eliminate the interference of the sensor itself, atmospheric conditions and the angle of sunlight on the reflectivity or radiance of ground objects, so as to restore the true physical properties of ground objects. The atmospheric correction is used to estimate and remove the absorption and scattering effects of atmospheric molecules, aerosols and water vapor on electromagnetic waves through an atmospheric transmission model, thereby eliminating the interference and influence of the atmospheric environment on imaging. The georegistration and geometric correction are used to correct geometric distortions in remote sensing images, giving them accurate spatial location information and enabling them to be accurately overlaid with data from other data sources. The orthorectification is used to correct projection differences and displacements caused by terrain undulations using the DEM, generating a map with orthorectified projection effects to ensure the orthorectified perspective of the data. The panchromatic sharpening is used to fuse a low spatial resolution multispectral image with a high spatial resolution panchromatic image to generate an image that has both high spatial resolution and high spectral resolution.
[0018] In addition, alternative options include the following: the topographic factors include slope, slope height, slope position, slope aspect, topographic curvature, and lithology; the symptom factors include normalized vegetation index and deformation; and the external inducing factors include distance from river systems, rainfall, land use type, and intensity of engineering activities.
[0019] In addition, an alternative approach is that the regional statistical characteristics of the disaster-causing hazard factors include mean characteristics and standard deviation characteristics.
[0020] In addition, an optional approach is to further include, before training the pre-constructed heterogeneous ensemble learning model based on the regional statistical characteristics of the disaster-causing hazard factors: The regional statistical characteristics of the disaster-causing hazard factors are normalized using the following formula: in, Represents the normalized features. x This represents the characteristics of the original disaster-causing hazard factors to be normalized. Represents the characteristic of sample mean. It represents the standard deviation.
[0021] Alternatively, the decision tree can be constructed using the Gini index and employ a hybrid pruning and multi-granularity integration strategy that integrates business rules, thus providing an interpretable decision path. The support vector machine uses an improved loss function to handle sample imbalance and employs a mixture of linear and RBF kernels to adapt to nonlinear geological disaster factor relationships. The feedforward neural network is used to fuse static terrain factors and temporal InSAR deformation factors, and is trained through nonlinear activation functions and Adam optimization. The linear meta-model performs a weighted summation and fusion of the outputs of the three base learners, with the weights determined by model training, and outputs the final probability and level of potential hazards.
[0022] Alternatively, the input data undergoes a two-step process of linear computation and nonlinear transformation at each layer of the feedforward neural network to obtain the predicted output; and, after obtaining the predicted output, The error between the predicted output and the true value is propagated backward along the feedforward neural network, and the gradient of the loss function with respect to each weight and bias is calculated according to the chain rule. The optimizer uses the gradient to update all weights and biases in the feedforward neural network.
[0023] On the other hand, the present invention also provides a hazard identification system based on multi-source geological disaster factors and a dynamic grid model, used for identifying geological disaster hazards using the hazard identification method based on multi-source geological disaster factors and a dynamic grid model as described above. The system includes: The factor data acquisition unit is used to acquire multi-source geological disaster factor data of the target area and preprocess the multi-source geological disaster factor data to form disaster-causing hazard factors; wherein, the multi-source geological disaster factor data includes remote sensing data and non-remote sensing data; The dynamic grid division unit is used to determine the scale, shape and data source resolution of geological disaster hazards based on the disaster-causing hazard factors, adaptively determine the size and shape of the dynamic grid, and divide the target area into multiple dynamic grid units. The regional statistical feature extraction unit is used to extract regional statistical features of disaster-causing factors for each dynamic grid cell based on the pixel values of the grid area. The disaster-causing factors include terrain-related factors, symptom factors, and external inducing factors. The model training unit is used to train a pre-constructed heterogeneous ensemble learning model based on the regional statistical characteristics of the disaster-causing hazard factors; wherein, the heterogeneous ensemble learning model includes three base learners: decision tree, support vector machine, and feedforward neural network, as well as a linear meta-model for fusing outputs; The model application unit is used to input the prediction results of each base learner on the test set into the linear meta-model for weighted fusion to obtain the geological disaster hazard identification result.
[0024] As can be seen from the above technical solution, the hazard identification method and system based on multi-source geological disaster factors and dynamic grid models provided by this invention replaces pixel-by-pixel identification with dynamic grids, which can overcome the defects of pixel-by-pixel identification that ignores regional correlation and results are unstable, and realize the standardized fusion of multi-source heterogeneous remote sensing and non-remote sensing geological disaster data; by deeply integrating multi-source geological disaster factors, it comprehensively reflects the disaster-causing mechanism, and adapts to landslide hazards of different scales and forms through dynamic grid adaptive division; the heterogeneous integrated model improves the model's identification accuracy and generalization ability under imbalanced samples and high-dimensional features; the decision path is clear, ensuring that the identification model has interpretability and meets the needs of geological disaster supervision and mechanism analysis; and the non-contact air-space-ground integrated identification solution does not require the deployment of sensors, which can effectively avoid personnel risks and engineering disturbances.
[0025] To achieve the foregoing and related objectives, one or more aspects of the invention include the features that will be described in detail below. The following description and accompanying drawings illustrate certain exemplary aspects of the invention. However, these aspects indicate only a few of the various ways in which the principles of the invention can be used. Furthermore, the invention is intended to encompass all such aspects and their equivalents. Attached Figure Description
[0026] Other objects and results of the invention will become more apparent and readily understood with reference to the following description taken in conjunction with the accompanying drawings. In the drawings: Figure 1 This is a flowchart illustrating the hazard identification method based on multi-source geological disaster factors and a dynamic grid model according to an embodiment of the present invention. Figure 2 This is the overall technical approach of the hazard identification method based on multi-source geological disaster factors and dynamic grid model according to embodiments of the present invention; Figure 3 This is a schematic diagram of the grid classification of hidden danger factors according to an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the calculation of the water system distance factor according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the heterogeneous integrated learning model structure according to an embodiment of the present invention; Figure 6 This is a schematic diagram of a decision tree model according to an embodiment of the present invention; Figure 7 This is a schematic diagram illustrating the principle of a support vector machine according to an embodiment of the present invention; Figure 8 This is a schematic diagram of an MLP neural network structure according to an embodiment of the present invention; Figure 9 This is a schematic diagram of a two-layer integrated learning architecture according to an embodiment of the present invention.
[0027] In all the accompanying drawings, the same reference numerals indicate similar or corresponding features or functions. Detailed Implementation
[0028] In the following description, numerous specific details are set forth for illustrative purposes and to provide a thorough understanding of one or more embodiments. However, it will be apparent that these embodiments may also be implemented without these specific details. In other instances, well-known structures and devices are shown in block diagram form for ease of description of one or more embodiments.
[0029] To address the aforementioned issues with traditional methods for identifying geological disaster hazards, such as deploying GNSS base stations or other sensor equipment, which require personnel to venture into dangerous areas and whose manual deployment activities may further induce landslides and other hazards, as well as the low accuracy resulting from relying on a single data source, this invention provides a hazard identification method and system based on multi-source geological disaster factors and a dynamic grid model. This method uses non-contact remote sensing data as a foundation, supplemented by meteorological, hydrological, and lithological data, which are then processed and fused with the remote sensing data source to obtain necessary information such as deformation data and DEM data of the target area. Furthermore, it dynamically segments the multi-source data features into a grid and utilizes machine learning methods to comprehensively analyze and identify landslide hazards in the grid samples, thereby improving the accuracy of identifying landslide hazards of different scales.
[0030] To better illustrate the technical solution of the present invention, some of the technical terms involved in the present invention will be briefly explained below.
[0031] Multi-source data refers to data from different platforms, different sensors, and different types. The multi-source data in this invention includes optical remote sensing, SAR, DEM, meteorological, hydrological, lithological, and human engineering activity data.
[0032] SAR (Synthetic Aperture Radar) is an active microwave remote sensing technology that can acquire surface information around the clock and in all weather conditions, unaffected by clouds, fog, rain, or snow.
[0033] Temporal InSAR is used to utilize a large number of long-term SAR images (dozens or more) to invert the surface deformation process in continuous time dimensions through temporal modeling, denoising, and phase unwrapping, in order to obtain high-precision, long-term deformation rates, cumulative deformation, deformation trends, and anomalous accelerations.
[0034] Multi-temporal InSAR is used to perform differential interferometric processing on two or more SAR images of the same area acquired at different times, pairing them together to detect whether surface deformation has occurred and where the deformation has occurred. In this invention, multiple SAR observations of the same area are performed using temporal InSAR and multi-temporal InSAR, the phase difference is calculated, and slow surface deformation at the millimeter-centimeter level is inverted for early landslide identification.
[0035] Dynamic grid refers to grid units that are adaptively divided according to the scale, shape, and data resolution of geological disasters.
[0036] Regional statistical characteristics refer to the mean, standard deviation, and other features that reflect the overall and local variations of a region, calculated using grid units.
[0037] An ensemble learning model refers to a geological hazard identification model that uses decision trees, support vector machines, and neural networks as base learners and outputs the results through a weighted fusion of linear meta-models.
[0038] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0039] To illustrate the hazard identification method based on multi-source geological disaster factors and dynamic grid model provided by this invention, Figure 1 and Figure 2 The flowchart and overall technical roadmap of the hazard identification method based on multi-source geological disaster factors and dynamic grid model according to embodiments of the present invention are shown respectively.
[0040] like Figure 1 and Figure 2 As shown in the figure, the hazard identification method based on multi-source geological disaster factors and dynamic grid model provided by the present invention includes: S1: Obtain multi-source geological hazard factor data of the target area, and preprocess the multi-source geological hazard factor data to form disaster-causing hazard factors.
[0041] This step corresponds to Figure 2 The overall technical approach outlined in this invention comprises two stages: data acquisition and preprocessing, and disaster-causing factor extraction. The disaster hazard identification and prediction method designed in this invention is based on multi-source remote sensing data and other non-remote sensing disaster-related data. Therefore, the multi-source disaster factor data in this invention includes both remote sensing and non-remote sensing data. The remote sensing data includes high-resolution optical remote sensing data, DEM data, and SAR synthetic aperture radar imagery data. The non-remote sensing data includes regional meteorological precipitation, hydrology, geological lithology, and human engineering activities—data closely related to the formation of disaster hazards. This invention first requires acquiring the aforementioned relevant data for the target study area and then performing necessary preprocessing on both the remote sensing and non-remote sensing data to ensure the quality of subsequent data.
[0042] Specifically, during the imaging process of remote sensing images, factors such as solar elevation angle, atmospheric absorption and scattering, sensor parameters, and topography often vary over time, resulting in inconsistent radiometric values for the same ground feature in images from different time phases. To construct high-quality, standardized raw data, these images need to undergo standardized preprocessing. As an example, preprocessing of remote sensing data includes radiometric correction, atmospheric correction, georegistration and geometric correction, orthorectification, and panchromatic sharpening.
[0043] Radiometric correction is used to eliminate interference from the sensor itself, atmospheric conditions, and the angle of sunlight on the reflectance or radiance values of ground objects, in order to restore the true physical properties (reflectance, radiance) of the ground objects and make image values comparable across different times and different sensors. The main method of radiometric correction is radiometric calibration, which uses the calibration parameters built into the sensor to convert the raw digital quantization values (DN values) recorded by the sensor into physically meaningful radiance values or apparent reflectance. The specific formula is shown below: Formula 1 in, L These are the pixel values after radiometric calibration. DN It is the raw digital quantized value recorded by the sensor. G This is the gain coefficient. Bias Using the offset value, Formula 1 above can convert the raw digital quantization values recorded by the sensor into physically meaningful pixel values. The pixel values of the radiometrically corrected image can reflect the true physical properties of the ground features.
[0044] Atmospheric correction is used to estimate and remove the absorption and scattering effects of atmospheric molecules, aerosols, and water vapor on electromagnetic waves through an atmospheric transmission model, thereby eliminating the interference and influence of the atmospheric environment on imaging. In a specific embodiment of this invention, the atmospheric transmission model used is as follows: Formula 2 As shown in Formula 2, where, L_surface It is the clean reflectance value after removing atmospheric errors. L It is the value obtained from the radiation calibration in Formula 1. p_path The path radiation reflectivity is the radiation portion caused by atmospheric scattering that enters the sensor directly without being reflected from the ground surface. T It is atmospheric transmittance, with a value between 0 and 1, representing the attenuation of radiation as it travels from the sun to the Earth's surface and then to the sensor.
[0045] The above formula can be used to perform radiometric correction preprocessing on the original remote sensing data to obtain the true surface reflectance after atmospheric correction, which can be used in the subsequent analysis of geological disaster risks.
[0046] Georegistration and geometric correction are used to correct geometric distortions in remote sensing images, giving them accurate spatial location information and enabling accurate overlay with data from other data sources.
[0047] Geometric correction is divided into two parts: fine geometric correction and coarse geometric correction. Coarse geometric correction uses the satellite's orbital parameters, sensor attitude parameters (roll, pitch, yaw), and an Earth model to initially correct systematic distortions (such as Earth's rotation, curvature, and sensor scanning distortion). Fine geometric correction, to further eliminate residual random distortions, requires establishing a mathematical model using ground control points (GCPs) for fine correction, generally including the following steps: (1) Selection of ground control points: On the image to be corrected and on the reference data with accurate geographic coordinates, select a series of feature points at the same location. These control points should be evenly distributed on the image and have a certain number; (2) Transformation model establishment: Using the coordinates of control points, solve for the parameters of the geometric distortion model. Commonly used models include polynomial transformations, rational function models, etc. (3) Pixel resampling: After determining the transformation relationship, the original image pixels need to be repositioned into the new regular geographic grid. Commonly used resampling methods include: nearest neighbor method, bilinear interpolation method, etc.
[0048] Orthorectification is used to correct projection discrepancies and displacements caused by topographic relief using a DEM (Digital Elevation Model), generating a map with orthophoto projection to ensure the orthophoto perspective of the data. For mountainous areas with undulating terrain, this step is crucial for ensuring image quality.
[0049] Panchromatic sharpening is used to fuse a low spatial resolution multispectral (MS) image with a high spatial resolution panchromatic (PAN) image, generating an image with both high spatial and spectral resolution. Typically, the multispectral image is transformed from the RGB color space to another color space (such as IHS), then the luminance component (I) is replaced with a high-resolution panchromatic band, and finally, it is inversely transformed back to the RGB space to enhance both spectral and spatial resolution.
[0050] The core value of the above preprocessing lies in transforming the raw data containing various errors into standard data with consistent physical meaning and precise spatial location through systematic correction steps, so as to facilitate the subsequent identification of geological hazard risks using multi-source remote sensing data.
[0051] After preprocessing the multi-source geological hazard factor data, hazard-causing factors can be extracted based on the preprocessed multi-source geological hazard factor data. The hazard-causing factors in this invention include topographic-related factors, symptom factors, and external inducing factors. Among them, topographic-related factors include slope, slope height, slope position, slope aspect, topographic curvature, and lithology; symptom factors include normalized vegetation index and deformation; and external inducing factors include distance from river systems, rainfall, land use type, and intensity of engineering activities.
[0052] The occurrence of geological hazards such as landslides is a complex and multifaceted process driven by multiple factors. Specific geological and topographical conditions provide the intrinsic basis for landslides. These include slippery soil and rock types, a certain slope gradient, and the softening effect of water activity on the soil and rock mass. These intrinsic conditions together constitute a potentially unstable slope system. Subsequently, certain external triggering factors can disrupt the original equilibrium of this system. Rainfall, especially heavy or prolonged rainfall, is the most active and primary natural triggering factor. Rainwater infiltration significantly increases the weight of the soil and rock mass and softens the potential sliding surface, reducing its shear strength and increasing the risk of landslides. In addition, human engineering activities, such as excavation of the slope toe, loading of the upper part of the slope, blasting, and destruction of vegetation, directly alter the stress state and stability of the slope, and are also important external triggering factors.
[0053] The formation and occurrence of landslide hazards is a complex geophysical process involving multiple factors, and a single factor is often insufficient to comprehensively assess its risk. Based on the above considerations, in a specific embodiment of this invention, the following multi-source geological hazard factors are processed and obtained for subsequent analysis, as shown in the table below: Table 1. Risk Factors for Disasters This invention uses the disaster-causing hazard factors described in the table above to conduct landslide hazard identification and analysis. Among them, topographic slope, topographic slope height, slope position, slope aspect, topographic curvature, and lithological factors are related to topographic height and describe the basic topographic conditions of the area, providing a basis for determining whether the area has the basic geological conditions for landslides.
[0054] Furthermore, normalized difference in vegetation index (NDVI) and deformation can serve as indicators of the impending occurrence of potential landslides. For instance, before a landslide forms, it is generally accompanied by certain deformation and abnormal vegetation growth on the landslide surface. Therefore, providing data on deformation and NDVI can help measure the characteristics preceding a landslide.
[0055] In addition, factors such as engineering activities and rainfall, as important triggers for landslides, are also taken into consideration. Further analysis will extract the characteristics of these factors and combine them with machine learning models to classify and identify potential geological hazards.
[0056] After determining the disaster-causing hazard factors, step S2 can be entered: determine the scale, shape and data source resolution of the geological disaster hazard based on the disaster-causing hazard factors, adaptively determine the size and shape of the dynamic grid, and divide the target area into multiple dynamic grid units.
[0057] Classic machine learning models, such as random forests and decision trees, often use individual pixels in images as basic units to identify and classify disaster hazards one by one. This approach often ignores the regional correlation characteristics of disaster hazards, using isolated pixel units as the basis for hazard identification, resulting in insufficient identification accuracy and identification results that are mostly isolated single pixel areas. Therefore, this invention proposes a dynamic grid-based hazard factor feature extraction scheme. Figure 3 An example of a grid classification of hazard factors according to an embodiment of the present invention is shown.
[0058] like Figure 3 As shown, most of the potential hazard factors in this embodiment are raster data. Traditional machine learning models, when processing this type of data, often extract the corresponding factor features (pixel values) based on individual pixels, such as... Figure 3 As shown in region 1 (each grid represents a pixel). However, this method cannot consider the correlation between pixels, and only uses single pixel features for classification, which easily leads to limited final classification accuracy. To address this problem, this invention proposes a region feature extraction scheme based on dynamic grids. This scheme does not use a single pixel as the base unit, but rather uses a gridded region as the basic unit to obtain its more robust regional statistical features. As shown in regions 2, 3, and 4 in the above figure, specifically, the gridded region, combined with the geometric characteristics of the disaster hazard, can be further divided into elongated types (such as...). Figure 3 (as shown in area 4) or ordinary type (such as...) Figure 3 The grid includes regions 2 and 3, as well as various other trends and shapes (adjusted based on actual geological disaster samples). Furthermore, the grid size can be dynamically and adaptively determined based on the resolution of the hazard factor data source and the scale of the disaster. The basic principle of the division is: given the resolution of the factor data, the size of the divided grid needs to be able to encompass the corresponding geological disaster hazard sample range. The grid size is dynamically set according to the sample situation and is not limited or fixed to a specific size, ensuring good adaptability to geological disaster samples at different scales.
[0059] With the core principle of unifying and matching the actual spatial range, geometric shape, and remote sensing data resolution of geological hazard risks, the grid size and shape are adaptively determined. This ensures that a grid can completely contain a typical hazard, that the grid does not disrupt the spatial continuity of the hazard, and that the grid scale matches the data accuracy, avoiding the problems of large grids containing small hazards or small grids exceeding resolution.
[0060] Specifically, as an example, the step of adaptively determining the size and shape of a dynamic grid based on the scale, geometry, and remote sensing data resolution of potential geological hazards further includes: S21: Determine the grid size based on the equivalent diameter of the hazard, so that a single grid can completely accommodate a single hazard. S22: Determine the grid shape based on the aspect ratio and extension direction of the hidden danger, including squares, rectangles and long strips; S23: Determine the minimum grid size based on the resolution of the remote sensing data, ensuring that the grid contains no less than 4 effective pixels to support the extraction of regional statistical features.
[0061] In one specific embodiment of the present invention, the scheme for determining the grid size (side length or area) based on the scale of the hidden danger is as follows: First, determine the scale and classification of the potential hazard. For example, based on historical landslide samples and remote sensing interpretation results, the scale of the potential hazard can be divided into three levels: Minor hazards: < 500㎡ (shallow surface collapse, small-scale deformation); Medium-sized potential hazard: 500㎡~5000㎡ (typical slope landslide); Major potential hazards: > 5000㎡ (ancient landslides, large creep lesions); Secondly, the grid side length is dynamically determined according to the scale. In a specific embodiment of the invention, the grid side length is set to = average equivalent diameter of the hazard × 1.0 to 1.2, which ensures that one grid can accommodate exactly one hazard, without breakage or overlap. The grid side length defined according to the hazard scale is as follows: The grid side length for small-scale hazards is 5m to 15m; The grid side length for medium-sized hazards is 15m to 50m; The grid side length for large-scale hidden dangers is 50m to 200m; In this embodiment, the key logic of the hazard classification system is as follows: The smaller the potential hazard, the smaller the grid size, ensuring that no detail is lost; The greater the potential hazard, the larger the grid should be, to avoid computational redundancy and excessive noise.
[0062] In determining the grid shape (geometric type) based on the hazard morphology, the grid shape can be dynamically selected according to the actual development morphology of the landslide / hazard, rather than a fixed square. For example, in hazard morphologies with concentrated slopes and approximately circular or square areas (such as typical landslide bodies and collapse zones), square or rectangular grids can be used; in hazard morphologies with a large aspect ratio (e.g., aspect ratio > 3:1) extending along the slope length, slender rectangular grids can be used, ensuring that the grid wind direction is consistent with the slope aspect or the main sliding direction, to avoid the grid cutting through the landslide body and to ensure the integrity of regional characteristics; in hazard morphologies with varied undulations and discrete distribution, irregular adaptive grids can be used, automatically fitting according to the terrain curvature and the boundaries of abnormal deformation areas to preserve the local abrupt change characteristics of the hazard (standard deviation characteristics are more sensitive).
[0063] In determining the minimum grid size based on the data source resolution, the grid size cannot be less than twice the resolution of the remote sensing image, ensuring that each grid contains at least four effective pixels to calculate statistical characteristics such as mean and standard deviation. The lower the resolution, the larger the grid must be; the higher the resolution, the smaller the grid can be, resulting in higher accuracy. Specifically, as an example, for UAV remote sensing and high-precision DEMs with a ground spatial resolution of 0.5m to 2m, the minimum grid size is limited to 5m × 5m; for high-resolution satellite data with a ground spatial resolution of 2m to 5m, the minimum grid size is limited to 10m × 10m; and for medium-resolution satellite and SAR data with a ground spatial resolution of 10m to 30m, the minimum grid size is limited to 20m × 20m or 30m × 30m.
[0064] The present invention, through the above-mentioned grid division scheme, can take into account the different geometric shapes of the hidden danger area on the one hand, and on the other hand, it can also take into account the relationship between the diverse scale of the hidden danger and the data resolution.
[0065] After completing the dynamic grid division, proceed to step S3: for each dynamic grid unit, extract the regional statistical features of disaster-causing factors based on the pixel values of the grid area.
[0066] This invention statistically analyzes the regional mean and standard deviation of topographic factors such as slope and curvature, as well as NDVI-normalized vegetation factors. The mean generally reflects the basic numerical value of the corresponding factor in the region, while the standard deviation reflects the spatial variation of the factor's regional value. It is worth noting that this embodiment uses both mean and standard deviation, but this does not mean that the invention only uses these two features. This embodiment primarily introduces a concept, using mean and standard deviation as examples. The use of other factors such as gradient and median should also be considered within the scope of this disclosure.
[0067] The factors and features used in this embodiment are shown in Table 2 below. The different features of different factors describe the probability of disaster hazards from different aspects.
[0068] Table 2 Factor Feature Extraction - 1 As shown in the table above, the mean slope reflects the average inclination of the area and is directly related to the basic potential energy for gravity-driven disasters (such as landslides and collapses). The larger the mean, the worse the overall stability. Additionally, the standard deviation of slope reflects the degree of variation in surface slope. A large standard deviation indicates significant topographic relief within the area, potentially indicating steep slopes, landslide-prone areas, or uneven slope deformation due to erosion.
[0069] Other metrics, such as the mean slope position, reflect the average condition of the landform location of a pixel within a region (e.g., slope top, middle, and toe). Different slope positions experience different stress conditions, affecting stability. This reflects the complexity and diversity of landform locations. A large standard deviation indicates that the region is located in an area with multiple different types of slope positions.
[0070] The mean aspect reflects the dominant direction of sunlight in a region, influencing rock weathering rate, soil moisture content, and vegetation growth, indirectly affecting the mechanical properties of the soil and rock mass. The standard deviation of aspect reflects spatial differences in sunlight and microclimate conditions. A large standard deviation indicates a significant effect of sunny and shady slopes, which may lead to uneven weathering of slopes and inconsistent vegetation protection effects, resulting in varying distribution of potential hazards.
[0071] The mean curvature reflects the average bending morphology (convexity and concavity) of the Earth's surface. The mean indicates whether the region as a whole is an erosion zone (convex, prone to stretching) or a deposition zone (concave, prone to water saturation). Its standard deviation reflects the complexity and fragmentation of the surface morphology. A large standard deviation indicates drastic changes in surface unevenness, potentially revealing numerous gullies, fissures, or unstable blocks, signifying active deformation.
[0072] The mean slope height reflects the relative elevation difference of the terrain and is related to gravitational potential energy and the potential transport distance of the landslide. The larger the mean, the greater the potential scale and energy of the disaster. The standard deviation of slope height reflects the steepness and undulation of the local terrain. A large standard deviation often indicates the presence of cliffs, steep slopes, or other free-faced surfaces, which are favorable terrain for landslides.
[0073] The mean vegetation index (NDVI) reflects the average vegetation cover of a region. A high mean usually indicates strong soil-fixing effect of vegetation roots, which is beneficial to stability; a low mean indicates exposed soil and rock, making it susceptible to erosion. The standard deviation reflects the uniformity of vegetation cover. A large standard deviation may indicate localized vegetation damage (such as landslides or excavation), and these abnormal patches are important clues for identifying potential hazards.
[0074] This invention establishes a core system for geological hazard identification using the aforementioned factors. The mean characteristic provides a macroscopic understanding of the overall stability pattern of the study area. The standard deviation characteristic effectively reveals the heterogeneity and local variations within the region, which are often key areas for the incubation and occurrence of geological hazards. These factors enable a preliminary assessment of the regional topography and geomorphology. Further comprehensive analysis, combining triggering factors and other related factors, is needed, as shown in Table 3 below. Table 3 Factor Feature Extraction - 2 In addition to the topographic factors mentioned in Table 2, the data in Table 3, such as river system distribution, lithology distribution, land use distribution, human engineering activities, time-series InSAR deformation data, and precipitation distribution, are also of great significance in inducing and identifying potential geological hazards.
[0075] The hardness and structure of different rock strata directly control slope stability. Rock formations with alternating layers of hard and soft rock, or those prone to weathering (such as mudstone and siltstone), are more susceptible to landslides and collapses. River erosion continuously alters slope morphology, reducing its stability. Areas closer to rivers and with higher river network density are more strongly affected by lateral erosion and downcutting, and typically have a higher concentration of potential geological hazards. Rainfall is the most direct factor triggering geological hazards. Heavy rainfall infiltrates the soil and rock mass, increasing its weight and reducing the strength of structural surfaces, further increasing the likelihood of geological hazards. Therefore, processing these factors and extracting useful features is of significant value for identifying potential geological hazards.
[0076] Figure 4 This is a schematic diagram illustrating the calculation of river system distance factors according to an embodiment of the present invention. Figure 4 As shown, the river system distribution is calculated based on the spatial horizontal distance from the center of the sample. The minimum horizontal distance D from the center of the dynamically divided grid area sample to the river system distribution is calculated as a characteristic factor for the development and formation of the disaster sample. This factor reflects the distribution distance between the potential hazard area and the river system; the closer the distance, the stronger the erosion and cutting effect of the river, and the higher the risk level of the corresponding potential hazard.
[0077] For the lithology distribution factor, the main lithological types in the gridded samples are statistically analyzed and used as the lithological characteristics of the samples. Similarly, for the land use distribution factor, the main land use types in the gridded samples are statistically analyzed and used as the land use type characteristics of the samples.
[0078] Finally, for the InSAR deformation factor and precipitation factor, the average deformation rate and average precipitation in the corresponding grid samples are calculated to measure the deformation and precipitation factors of the grid samples.
[0079] After processing each factor as described above, the corresponding features are further obtained for subsequent training and analysis of machine learning models.
[0080] S4: Train the pre-constructed heterogeneous ensemble learning model based on the regional statistical characteristics of the disaster-causing hazard factors; the heterogeneous ensemble learning model includes three base learners: decision tree, support vector machine, and feedforward neural network, as well as a linear meta-model for fusion output.
[0081] Considering the lack of stability of single-model methods in decision classification, this invention specifically adopts a multi-method integrated architecture to construct a hazard identification model, thereby improving the reliability and accuracy of identification and prediction. Figure 5 The structure of a multi-method integrated hazard identification model according to an embodiment of the present invention is shown. For example... Figure 5 As shown, the multi-method integrated hazard identification model provided in this embodiment first divides the input multi-factor hazard feature data into a training set and a test set in an 8:2 ratio. Then, the training set is fed in parallel into three different base learners (decision tree model, neural network model, and support vector machine model) for training, and each learner generates prediction results (result-1, result-2, and result-3) on the test set. Finally, the outputs of these three base learners are aggregated into an integrated decision module (linear meta-model) for integration, thereby generating a final, more robust prediction result.
[0082] The model ensemble learning model designed in this embodiment improves the generalization performance and stability of the overall model by combining the advantages of multiple heterogeneous models, ensuring the accuracy and reliability of geological disaster hazard prediction. A detailed introduction follows.
[0083] Input data: As mentioned above, to accurately identify potential geological hazards, this disclosure uses factors such as Normalized Difference Vegetation Index (NDVI), topographic slope, topographic slope height, slope position, slope aspect, topographic curvature, lithology, rainfall, deformation, land use type, engineering activities, and distance from river systems as input data. After processing such as dynamic grid division, the final data types and basic information used for model input are shown in Table 4 below: Table 4. Schematic diagram of input data for model training As shown in Table 4 above, each specific sample area is obtained from the study area through dynamic grid division. Each sample should then have the disaster-causing factor characteristics mentioned above, such as normalized vegetation index (NDVI), topographic slope, topographic slope height, slope position, slope aspect, etc., as well as a label (hazardous or non-hazardous) to distinguish the type of the sample.
[0084] The samples obtained from the study area are organized into the above basic form, and the features of each sample are normalized before being used for subsequent model training. The normalization formula is as follows: Formula 3 This embodiment uses Z-Score normalization as shown in the formula above. This method is insensitive to outliers in the features. These are the normalized features. It is the sample mean. This is the sample standard deviation. The main purpose of normalization is to eliminate the adverse effects on the model caused by excessive differences in the dimensions and ranges of different features.
[0085] Finally, after normalizing the features of all samples, the data is used as sample data for subsequent model training.
[0086] The training set was then fed into three different types of base learners (decision tree, neural network, and support vector machine) for independent training. Each trained model then made predictions on the test set, producing result-1, result-2, and result-3 respectively. Finally, the predictions from all models were aggregated into an ensemble decision module. This module used a weighted average to fuse the three results into a unified and more reliable final prediction.
[0087] The geological disaster hazard identification module of this invention has high requirements for the accuracy and robustness of the model method; therefore, it employs methods such as... Figure 5 The overall architecture is designed by integrating multiple models, each with its own advantages, forming a complementary system. The following will provide illustrative examples of three different types of base learners.
[0088] Decision tree model Considering that interpretability is crucial in hazard identification, and given that the decision-making process of the decision tree model is very intuitive and easy to understand and interpret, this invention uses it as a sub-model of the ensemble model.
[0089] A decision tree is a tree-structured predictive model that simulates the human thought process in decision-making. It progressively divides complex data through a series of rules, ultimately forming a tree structure composed of nodes and directed edges.
[0090] Figure 6 The structure of a decision tree model according to an embodiment of the present invention is shown. Figure 6 As shown, the process of constructing the decision tree model in this embodiment is as follows: Feature selection determines how to find the most discriminative features from all available features. This disclosure uses the Gini index (CART algorithm) to measure the impurity of the data. The smaller the Gini index, the higher the data purity, i.e., the better the classification effect. The formula for calculating the Gini index is as follows: Formula 4 As shown in Formula 4, Gini It is the calculated Gini index. P_k For the first in the dataset k The proportion of samples of the same class. The closer the calculated index is to 0, the higher the purity of the samples in that node. This index will be used to control whether the decision tree node splits further.
[0091] The tree model is built starting from the root node, which contains the entire training dataset. Then, the next node is partitioned from the root node. Before partitioning the current node, it is determined whether certain conditions are met. If they are met, the node is marked as a leaf node, and the partitioning of that branch stops. Key stopping conditions include: all samples in the current node belong to the same category; no remaining features are available; the number of samples is less than a predetermined threshold; or the maximum depth of the tree is reached.
[0092] If the stopping condition is not met, the "purity" improvement resulting from splitting the dataset using all currently available features is calculated according to a predetermined criterion (the Gini index mentioned above is used in this disclosure), and the feature with the largest improvement is selected as the splitting criterion for the current node. Based on the selected optimal feature and its value, the dataset under the current node is divided into several mutually exclusive subsets. Each subset corresponds to a value branch of that feature.
[0093] The above steps are repeated to build sub-branches of the tree structure. When the recursive process returns due to the satisfaction of any stopping condition, a leaf node is created, which represents the final decision result.
[0094] In constructing the decision tree model, to enhance its practicality, interpretability, and dynamic adaptability in real-world scenarios, this embodiment first employs a hybrid pruning strategy that integrates business rules. This strategy combines statistical indicators with domain knowledge (such as known hidden danger patterns). During the pruning process, business rule weights and expert evaluation mechanisms are introduced to ensure that the decision tree remains concise without losing key business logic, thereby improving the model's professionalism and interpretability.
[0095] Secondly, a multi-granularity decision tree integration framework is adopted to construct specialized decision subtrees for different levels or types of hidden danger characteristics (such as equipment level, system level, and environment level). The judgments of each subtree are integrated through interpretable integration rules (such as weight allocation based on attention mechanism), thereby improving the overall identification performance while maintaining the transparency and traceability of the final decision.
[0096] Finally, a dynamic incremental learning mechanism is introduced. By embedding a concept drift detection module and a local update strategy, the decision tree can adapt to the environment where the data distribution changes over time. Only the subtrees most affected are updated, enabling the model to continuously learn from changes in hazard patterns while ensuring the consistency of decision logic.
[0097] The above design ensures that the constructed decision tree model meets the stringent requirements of accuracy and transparency in hazard identification tasks. Subsequently, this decision tree model will be used as a sub-model within an integrated model framework for hazard identification decisions.
[0098] Support Vector Machine Considering that Support Vector Machine (SVM) can still maintain good generalization ability in small sample and high-dimensional data, this invention also regards it as a sub-model in the multi-method ensemble model.
[0099] The core idea of SVM is to find an optimal decision boundary (a "hyperplane") that can separate samples of different classes to the greatest extent possible, and whose distance (i.e., "margin") from the nearest data point (i.e., "support vector") is as large as possible.
[0100] Figure 7 The principle of a support vector machine according to an embodiment of the present invention is illustrated. For example... Figure 7 As shown, in the identification of geological disaster hazards, this "hyperplane" is the optimal decision boundary line that divides "high-risk areas" and "safe areas" in a high-dimensional feature space composed of various disaster-causing factors (such as the slope, lithology, precipitation, etc. mentioned above).
[0101] The aforementioned normalized hazard factors are used as inputs for training the SVM model, and an SVM model is constructed. First, considering that the number of hazard point samples is usually far fewer than non-hazard points, resulting in a severe sample imbalance problem, the loss function of the traditional SVM is improved by assigning a higher penalty weight to misclassification of minority class (hazard point) samples, forcing the model to focus more on the correct identification of hazard points. Furthermore, this invention also designs a hybrid kernel function to address the complex nonlinear relationships between various geological hazard factors. The linear kernel, which excels at capturing linear relationships, is combined with the RBF kernel, which excels at handling complex nonlinear relationships, leveraging their respective strengths to better fit the true structure of geographic space and improve the model's expressive power and prediction accuracy.
[0102] Feedforward Neural Network Feedforward Neural Networks (NNNs) are among the most fundamental and crucial models in the field of deep learning. By mimicking the connection patterns of neurons in the human brain, they possess the ability to learn complex patterns from data. Figure 8The general structure of an MLP (Multi-Layer Perceptron) feedforward neural network according to an embodiment of the present invention is shown. Figure 8 As shown, the MLP feedforward neural network in this embodiment includes an input layer, multiple hidden layers, and an output layer. The input layer receives the raw input data, and the number of neurons in this layer equals the number of input features. The hidden layers are the core computational layers, responsible for feature extraction and nonlinear transformation. There can be at least one hidden layer, or multiple layers; the number of layers and the number of neurons per layer are adjustable. The specific number of layers is related to the task complexity and the amount of data, and a balance between fitting ability and generalization ability must be ensured. The output layer produces the final prediction result; the number of neurons and the activation function are set according to the task. For classification tasks, the Softmax activation function is commonly used.
[0103] Input data enters from the first layer (input layer) of the MLP network, passes through one or more hidden layers, and finally reaches the output layer. In each layer, the data undergoes two computational steps: (1) Linear calculation: The input of each neuron in this layer is the weighted sum of the outputs of all neurons in the previous layer, plus a bias term. Equation 5 is expressed as follows: Z = W·a + b (Formula 5) Where w is the weight matrix, a is the output vector of the previous layer, and b is the bias vector.
[0104] (2) Nonlinear transformation: The result z of the linear calculation is input into a nonlinear activation function (such as ReLU) to obtain the final output a=f(z) of the layer, and it is used as the input of the next layer. In this way, MLP can abstract and extract features layer by layer.
[0105] After obtaining the predicted output, the loss function calculates the error between the predicted and the true values. The task of the backpropagation algorithm is to propagate this error back along the network from the output layer, and calculate the gradient (i.e., derivative) of the loss function with respect to each weight and bias according to the chain rule. Subsequently, the optimizer (such as Adam) uses these gradients to update all weights and biases, so that the predictive ability of the model is gradually improved.
[0106] Considering that traditional methods rely heavily on static geological environmental factors, this embodiment also deeply integrates time-series remote sensing monitoring data (such as the surface deformation rate obtained by SBAS-InSAR technology) with static environmental factors and uses MLP for collaborative analysis to improve prediction accuracy.
[0107] Step S5: Input the prediction results of each base learner on the test set into the linear meta-model for weighted fusion to obtain the geological disaster hazard identification result.
[0108] Considering that the reliability of geological hazard identification results is crucial in the geological hazard identification task, the ensemble learning in this invention improves the overall performance and reliability by constructing and combining multiple machine learning models. Therefore, the above three types of base models are trained, and their prediction results are then used as new features to input into a meta-model. The meta-model learns how to optimally combine these predictions to obtain the final result, thereby integrating the advantages of different models.
[0109] Figure 9 An integrated learning structure according to an embodiment of the present invention is shown. Figure 9 As shown, in this embodiment of the ensemble learning framework, the first layer is a base learner (heterogeneous model). It takes the same multi-factor feature data of geological disaster hazards as input and inputs it in parallel into three different types of basic machine learning models (decision tree, neural network, and support vector machine). These heterogeneous models learn and predict independently.
[0110] Decision tree models are highly interpretable and align with the traditional thinking patterns of geologists. After a heterogeneous model makes a prediction, the entire decision-making process can be clearly traced, revealing the decision-making process that identified the area as high-risk. Neural network models are adept at handling complex nonlinear relationships. Geological disasters are often the result of complex interactions between multiple factors (topography, geology, hydrology, human activities, etc.). Neural networks excel at learning and expressing such complex, nonlinear interactions, helping to find correlations between multiple factors. Support vector machines (SVMs) demonstrate stable generalization ability with small sample / high-dimensional data. Samples of major geological disasters (positive samples) are usually scarce and valuable. SVMs are highly effective for learning from small samples and handling imbalanced sample problems, and geological disaster features typically have high dimensionality (dozens of factors). SVMs can work effectively in high-dimensional spaces using kernel function techniques to find complex nonlinear classification boundaries.
[0111] The second layer is the meta-model: the predictions (probability of hazard occurrence or hazard level) from the three base learners in the first layer are used as new input features and fed into the second-layer meta-model. The meta-model further learns how to optimally combine these three predictions to obtain a more stable and accurate final output.
[0112] In this embodiment, a linear model with strong interpretability is selected as the basic structure of the meta-model, and its general expression is shown below: Formula 6 in, This is the output of the meta-model. , , These are the weights of the input part of the corresponding model. , , These represent the outputs of different base learners. After training and determining the corresponding weight parameters, the meta-model can be used to integrate the outputs of different base models. This disclosure employs a linear meta-model, which boasts high computational efficiency, fast training and prediction speeds, and low resource consumption. Furthermore, it exhibits strong resistance to overfitting and interpretability, intuitively reflecting the relative importance of each base model in the final decision.
[0113] Ultimately, this invention, through its integrated model architecture, enables the maintenance of high prediction accuracy, better stability, and a certain degree of interpretability when facing complex and highly uncertain tasks such as geological disaster identification, thereby providing more reliable technical support for risk management and early warning of geological disasters.
[0114] Corresponding to the above-mentioned hazard identification method based on multi-source geological hazard factors and dynamic grid models, this invention also provides a hazard identification system based on multi-source geological hazard factors and dynamic grid models, used for identifying geological hazard hazards using the above method. The system includes: The factor data acquisition unit is used to acquire multi-source geological disaster factor data of the target area and preprocess the multi-source geological disaster factor data to form disaster-causing hazard factors; wherein, the multi-source geological disaster factor data includes remote sensing data and non-remote sensing data; The dynamic grid division unit is used to determine the scale, shape and data source resolution of geological disaster hazards based on the disaster-causing hazard factors, adaptively determine the size and shape of the dynamic grid, and divide the target area into multiple dynamic grid units. The regional statistical feature extraction unit is used to extract regional statistical features of disaster-causing factors for each dynamic grid cell based on the pixel values of the grid area. The disaster-causing factors include terrain-related factors, symptom factors, and external inducing factors. The model training unit is used to train a pre-constructed heterogeneous ensemble learning model based on the regional statistical characteristics of the disaster-causing hazard factors; wherein, the heterogeneous ensemble learning model includes three base learners: decision tree, support vector machine, and feedforward neural network, as well as a linear meta-model for fusing outputs; The model application unit is used to input the prediction results of each base learner on the test set into the linear meta-model for weighted fusion to obtain the geological disaster hazard identification result.
[0115] For the embodiment of the system for identifying potential geological hazards assisted by the multi-source time-series remote sensing soil water content prediction technology provided by the present invention, since it is basically similar to the embodiment of the hazard identification method based on multi-source geological hazard factors and dynamic grid model, the relevant parts can be referred to the description of the method embodiment, and will not be repeated here.
[0116] As can be seen from the specific embodiments provided by the present invention above, the hazard identification method and system based on multi-source geological disaster factors and dynamic grid model provided by the present invention has the following advantages compared with the prior art: 1. Geological hazard identification is based on multi-factor joint analysis, and all factors are divided into static basic factors related to terrain and dynamic risk factors that induce hazards. Through comprehensive processing and analysis of different types of factors, geological hazard identification can be achieved. 2. To better adapt to the development shape and scale of geological disaster hazards, a dynamic grid is used to divide the risk area of hazard samples. Compared with pixel-based machine learning classification methods, this method can better take into account the regional characteristics of geological disaster hazards and enhance the accuracy of machine learning methods in the identification of geological disaster hazards. 3. Considering the characteristics of the geological disaster hazard identification task, an ensemble learning model architecture was designed and adopted. Multiple base learners were selected and trained separately, and the results were finally integrated through a linear learner. Compared with the prediction results of a single model, this method further ensures the reliability and stability of the identification results. 4. Considering the high requirements for interpretability and traceability in the task of identifying geological disaster hazards, the design and selection of all methods in this invention tend to favor models with strong interpretability, while avoiding models with opaque mechanisms as much as possible. The design concept of this invention ensures that the prediction results of the designed model have strong interpretability and can be used for the study of the development mechanism of hazard points.
[0117] The hazard identification method and system based on multi-source geological disaster factors and dynamic grid models according to the present invention have been described above by way of example with reference to the accompanying drawings. However, those skilled in the art should understand that various modifications can be made to the hazard identification method and system based on multi-source geological disaster factors and dynamic grid models proposed in the present invention without departing from the scope of the present invention. Therefore, the scope of protection of the present invention should be determined by the content of the appended claims.
Claims
1. A method for identifying potential hazards based on multi-source geological disaster factors and a dynamic grid model, characterized in that, include: Acquire multi-source geological hazard factor data of the target area, and preprocess the multi-source geological hazard factor data to form disaster-causing hazard factors; wherein, the multi-source geological hazard factor data includes remote sensing data and non-remote sensing data; The scale, shape and data source resolution of geological disaster hazards are determined based on the disaster-causing hazard factors, and the size and shape of the dynamic grid are adaptively determined to divide the target area into multiple dynamic grid units. For each dynamic grid cell, regional statistical features of disaster-causing factors are extracted based on the pixel values of the grid area. These disaster-causing factors include terrain-related factors, symptom factors, and external triggering factors. The pre-constructed heterogeneous ensemble learning model is trained based on the regional statistical characteristics of the disaster-causing hazard factors; wherein, the heterogeneous ensemble learning model includes three base learners: decision tree, support vector machine, and feedforward neural network, as well as a linear meta-model for fusion output; The prediction results of each base learner on the test set are input into the linear meta-model for weighted fusion to obtain the geological disaster hazard identification result.
2. The hazard identification method based on multi-source geological disaster factors and dynamic grid model as described in claim 1, characterized in that, The remote sensing data includes optical remote sensing data, DEM data, and SAR time-series deformation data, while the non-remote sensing data includes meteorological precipitation data, geological lithology data, hydrological data, land use data, and human engineering activity data.
3. The hazard identification method based on multi-source geological disaster factors and dynamic grid model as described in claim 2, characterized in that, The preprocessing of the remote sensing data includes radiometric correction, atmospheric correction, georegistration and geometric correction, orthorectification, and panchromatic sharpening; among which, The radiometric correction is used to eliminate the interference of the sensor itself, atmospheric conditions and the angle of sunlight on the reflectivity or radiance of ground objects, so as to restore the true physical properties of ground objects. The atmospheric correction is used to estimate and remove the absorption and scattering effects of atmospheric molecules, aerosols and water vapor on electromagnetic waves through an atmospheric transmission model, thereby eliminating the interference and influence of the atmospheric environment on imaging. The georegistration and geometric correction are used to correct geometric distortions in remote sensing images, giving them accurate spatial location information and enabling them to be accurately overlaid with data from other data sources. The orthorectification is used to correct projection differences and displacements caused by terrain undulations using the DEM, generating a map with orthorectified projection effects to ensure the orthorectified perspective of the data. The panchromatic sharpening is used to fuse a low spatial resolution multispectral image with a high spatial resolution panchromatic image to generate an image that has both high spatial resolution and high spectral resolution.
4. The hazard identification method based on multi-source geological disaster factors and dynamic grid model as described in claim 1, characterized in that, The topographic factors include slope, slope height, slope position, slope aspect, topographic curvature, and lithology; The symptom factors include the normalized vegetation index and deformation. The external inducing factors include distance from the river system, rainfall, land use type, and intensity of engineering activities.
5. The hazard identification method based on multi-source geological disaster factors and dynamic grid model as described in claim 1, characterized in that, The regional statistical characteristics of the disaster-causing factors include mean characteristics and standard deviation characteristics.
6. The hazard identification method based on multi-source geological disaster factors and dynamic grid model as described in claim 5, characterized in that, Before training the pre-constructed heterogeneous ensemble learning model based on the regional statistical characteristics of the disaster-causing hazard factors, the following steps are also included: The regional statistical characteristics of the disaster-causing hazard factors are normalized using the following formula: in, Represents the normalized features. x This represents the characteristics of the original disaster-causing hazard factors to be normalized. Represents the characteristic of sample mean. This indicates the characteristic of standard deviation.
7. The hazard identification method based on multi-source geological disaster factors and dynamic grid model as described in claim 6, characterized in that, The decision tree is constructed using the Gini index and employs a hybrid pruning and multi-granularity integration strategy that integrates business rules, thus providing an interpretable decision path. The support vector machine uses an improved loss function to handle sample imbalance and employs a mixture of linear and RBF kernels to adapt to nonlinear geological disaster factor relationships. The feedforward neural network is used to fuse static terrain factors and temporal InSAR deformation factors, and is trained through nonlinear activation functions and Adam optimization. The linear meta-model performs a weighted summation and fusion of the outputs of the three base learners, with the weights determined by model training, and outputs the final probability and level of potential hazards.
8. The hazard identification method based on multi-source geological disaster factors and dynamic grid model as described in claim 7, characterized in that, The input data undergoes two steps of processing—linear computation and nonlinear transformation—at each layer of the feedforward neural network to obtain the predicted output; and, after obtaining the predicted output, The error between the predicted output and the true value is propagated backward along the feedforward neural network, and the gradient of the loss function with respect to each weight and bias is calculated according to the chain rule. The optimizer uses the gradient to update all weights and biases in the feedforward neural network.
9. A hazard identification system based on multi-source geological disaster factors and a dynamic grid model, characterized in that, The system is used for identifying geological hazard risks using the hazard identification method based on multi-source geological hazard factors and dynamic grid model as described in any one of claims 1 to 8, and includes: The factor data acquisition unit is used to acquire multi-source geological disaster factor data of the target area and preprocess the multi-source geological disaster factor data to form disaster-causing hazard factors; wherein, the multi-source geological disaster factor data includes remote sensing data and non-remote sensing data; The dynamic grid division unit is used to determine the scale, shape and data source resolution of geological disaster hazards based on the disaster-causing hazard factors, adaptively determine the size and shape of the dynamic grid, and divide the target area into multiple dynamic grid units. The regional statistical feature extraction unit is used to extract regional statistical features of disaster-causing factors for each dynamic grid cell based on the pixel values of the grid area. The disaster-causing factors include terrain-related factors, symptom factors, and external inducing factors. The model training unit is used to train a pre-constructed heterogeneous ensemble learning model based on the regional statistical characteristics of the disaster-causing hazard factors; wherein, the heterogeneous ensemble learning model includes three base learners: decision tree, support vector machine, and feedforward neural network, as well as a linear meta-model for fusing outputs; The model application unit is used to input the prediction results of each base learner on the test set into the linear meta-model for weighted fusion to obtain the geological disaster hazard identification result.
10. The hazard identification system based on multi-source geological disaster factors and a dynamic grid model as described in claim 9, characterized in that, The factor data acquisition unit includes a radiation correction module, an atmospheric correction module, a georegistration and geometric correction module, an orthorectification module, and a panchromatic sharpening module; wherein... The radiation correction module is used to eliminate the interference of the sensor itself, atmospheric conditions and the angle of sunlight on the reflectivity or radiance of ground objects, so as to restore the true physical properties of ground objects. The atmospheric correction module is used to estimate and remove the absorption and scattering effects of atmospheric molecules, aerosols and water vapor on electromagnetic waves through an atmospheric transmission model, thereby eliminating the interference and influence of the atmospheric environment on imaging. The georegistration and geometric correction module is used to correct geometric distortions in remote sensing images, giving them accurate spatial location information and enabling them to be accurately overlaid with data from other data sources. The orthorectification module is used to use the DEM to correct the projection difference and displacement caused by terrain undulation, and generate a map with orthorectification effect to ensure the orthorectification perspective of the data. The panchromatic sharpening module is used to fuse a low spatial resolution multispectral image with a high spatial resolution panchromatic image to generate an image that has both high spatial resolution and high spectral resolution.