Geological disaster occurrence trend prediction system based on historical data

By combining multi-level anomaly data analysis and clustering techniques with support vector machine classification, the shortcomings of traditional geological disaster prediction systems in handling nonlinear and dynamic changes are addressed, achieving higher accuracy and real-time geological disaster trend prediction.

CN121765418APending Publication Date: 2026-03-31GUANGDONG HUIZHOU GEOLOGICAL ENG SURVEY INST +2
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-21
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional geological hazard occurrence trend prediction systems fail to fully consider complex nonlinear relationships and dynamic changes in data, resulting in low prediction accuracy and an inability to effectively cope with geological hazard occurrence trends under multivariable environments. In particular, they lack real-time performance and flexibility when processing multidimensional data such as rainfall, pore water pressure, and surface displacement.

Method used

By employing an anomaly identification module, clustering labeling module, stress zoning module, and trend classification module, and through multi-level anomaly data analysis and clustering techniques, potential abnormal changes are identified, the differences in multi-dimensional data are calculated, and trend prediction is performed using the support vector machine classification method to generate geological disaster trend prediction results.

Benefits of technology

It significantly improves the accuracy and timeliness of disaster prediction, and can adjust prediction results based on real-time data, providing more timely and reliable support for decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765418A_ABST
    Figure CN121765418A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a geological disaster occurrence trend prediction system based on historical data, which comprises an anomaly recognition module, a clustering labeling module, a stress partitioning module, a trend classification module and a disaster prediction module. According to the method, by introducing a multi-level abnormal data analysis and clustering technology, potential abnormal changes can be recognized more accurately when multi-dimensional data are processed, and therefore the problem that non-linear and dynamic changes of data are neglected in a traditional method is effectively solved. By calculating the difference of the multi-dimensional data, different stress partition areas can be distinguished, trend prediction is performed based on the characteristics of the areas, and the accuracy and timeliness of disaster prediction are remarkably improved. A support vector machine classification method is utilized to integrate focusing features and stress features, not only is the adaptability of a prediction model optimized, but also a prediction result can be adjusted according to real-time data, so that support with higher timeliness and reliability is provided for decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a geological disaster occurrence trend prediction system based on historical data. Background Technology

[0002] The field of data processing technology encompasses the application of techniques in data acquisition, storage, management, analysis, computation, and visualization. Its core aspects include data preprocessing, data mining, pattern recognition, machine learning, data analysis, and prediction, and are widely used in fields such as finance, healthcare, transportation, and energy. Data processing technologies typically rely on big data analytics and artificial intelligence algorithms to extract valuable information and patterns from large amounts of historical data through in-depth analysis, thereby supporting decision-making processes and predicting future trends.

[0003] Traditional geological hazard occurrence trend prediction systems refer to systems that predict the occurrence trend of geological hazards by collecting and analyzing historical geological hazard data and applying statistical methods and data mining techniques. These systems typically rely on data from various aspects, including historical hazard records, climate change, and geological conditions, and use mathematical models to quantitatively analyze and predict potential hazard risks. Traditional trend prediction methods are often based on fixed statistical models or simple regression analysis, failing to fully consider complex nonlinear relationships and dynamic changes in data. Therefore, existing systems have certain limitations in accuracy and real-time performance.

[0004] Traditional geological hazard prediction systems rely on fixed statistical models or simple regression analysis in practice. These methods fail to fully consider the complex nonlinear relationships and dynamic changes between data, resulting in low prediction accuracy and an inability to effectively address geological hazard trends under multivariate environments. Especially when processing multidimensional data such as rainfall, pore water pressure, and surface displacement, existing systems often fail to capture anomalous changes, lacking real-time performance and flexibility. Disaster prediction depends on historical data, but historical records may not cover sudden extreme events; therefore, traditional systems struggle to provide accurate trend predictions when facing constantly changing natural conditions. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and propose a geological disaster occurrence trend prediction system based on historical data.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a geological disaster occurrence trend prediction system based on historical data includes: The anomaly identification module collects historical rainfall, pore water pressure, and surface displacement, divides the data into time windows, standardizes the data, judges the deviation of the data, marks historical data that exceeds the anomaly identification threshold, generates anomaly structure data, and transmits it to the clustering and labeling module. The clustering and labeling module clusters the abnormal structural data, sorts the abnormal data in the same cluster according to the cluster center, generates structured focused data, and transmits it to the stress partitioning module. The stress zoning module, based on the structured focused data, calculates the rainfall increment, pore water pressure gradient, and displacement acceleration for the corresponding time window and makes difference judgments, distinguishes the change range of multiple parameters, generates stress zoning data, and transmits it to the trend classification module. The trend classification module extracts focusing features and stress features based on the stress zoning data and concatenates them into a trend prediction feature vector. This vector is then input into a support vector machine for classification, generating a trend classification result which is then transmitted to the disaster prediction module. The disaster prediction module extracts historical data for the corresponding time period based on the trend classification results, calculates the probability of disaster occurrence trends in multiple time periods, filters time periods, and generates geological disaster trend prediction results.

[0007] As a further aspect of the present invention, the abnormal structure data includes rainfall deviation, pore water pressure deviation, and displacement offset rate; the structured focusing data includes cluster center value, cluster sequence number, and cluster ranking value; the stress zoning data includes rainfall increment interval, pore water pressure gradient interval, and displacement acceleration interval; the trend classification result includes trend category quantity, feature vector magnitude, and classification hyperplane distance; and the geological disaster trend prediction result includes time period probability value, threshold exceedance marker quantity, and risk change segment.

[0008] As a further aspect of the present invention, the acquisition step of the anomaly identification module specifically includes: The data analysis submodule acquires historical rainfall, pore water pressure, and surface displacement sequences, divides the three types of sequences into time windows, calls the segmented values ​​of rainfall, pore water pressure, and surface displacement within the time window, calculates the window difference value, and generates the window difference result. The standardized calculation submodule, based on the window difference results, calls the segmented values ​​of rainfall, pore water pressure, and surface displacement within the corresponding window, calculates the ratios based on the window difference values, adjusts the ratio intervals, calculates the interval ratios, and obtains the interval ratio sequence. The deviation screening submodule, based on the interval ratio sequence, calls the multi-window ratio, and screens the deviation amount according to the difference between the ratio value and the preset anomaly identification threshold. It obtains the deviation amount of all windows and marks the windows whose deviation amount exceeds the threshold, generating abnormal structure data.

[0009] As a further aspect of the present invention, the acquisition step of the clustering annotation module is specifically as follows: The data clustering submodule, based on the abnormal structure data, groups the data using the K-means clustering algorithm, spatially partitions the data according to the attributes of each data point, calculates the cluster center of each group, assigns the cluster label to each data point, and obtains a cluster label sequence. The sorting processing submodule, based on the cluster label sequence, sorts the data points according to the time position of each cluster center by calling the timestamp corresponding to each cluster center, and obtains the time sorting sequence; The structured output submodule integrates the data based on the time-sorted sequence and the clustering label of each data point, performs structured processing on the data according to the clustering label and time order, and generates structured focused data.

[0010] As a further aspect of the present invention, the step of obtaining the stress partitioning module specifically includes: The parameter calculation submodule extracts rainfall, pore water pressure and surface displacement within the corresponding time window based on the structured focused data, calculates rainfall increment, pore water pressure gradient and surface displacement acceleration, performs calculations based on the numerical gradient, and generates gradient change. The difference interval submodule calls the gradient change amount and performs a difference judgment action on the rainfall increment, pore water pressure gradient and surface displacement acceleration based on the gradient change amount. The three types of values ​​are divided into change intervals according to the preset interval division benchmark value to obtain the interval range value sequence. The label segmentation submodule, based on the interval range value sequence, performs label numbering for each group of interval range value sequences, matches the interval range value sequences to the label numbers according to the position of the numerical intervals and generates corresponding labels, thereby generating stress zoning data.

[0011] As a further aspect of the present invention, the process of generating gradient change based on numerical gradient calculation is specifically to perform differential calculation on adjacent time series values ​​of rainfall, pore water pressure and surface displacement based on a uniform time window length, so that the rainfall increment, pore water pressure gradient and surface displacement acceleration corresponding to each time window respectively form gradient change with a fixed time interval.

[0012] As a further aspect of the present invention, the acquisition step of the trend classification module specifically includes: The feature splicing submodule, based on the stress partitioning data and the structured focusing data, extracts focusing feature quantities and stress feature quantities within the time window and performs splicing action at the same time sequence position, arranges the focusing feature quantities and stress feature quantities in sequence order, calculates the splicing vector value, and generates a splicing vector set; The vector construction submodule calls the concatenated vector set, arranges the concatenated vector values ​​according to the feature dimension to form a trend prediction input vector and forms an input vector magnitude, and performs vector normalization according to the input vector magnitude to obtain a normalized vector; The trend classification submodule, based on the normalized vector, performs classification on the support vector machine model by inputting the normalized vector, calculates the category output and uses it as the trend classification value, and obtains the trend classification result.

[0013] As a further aspect of the present invention, the calculation of the splicing vector value is performed using the following formula: ; in: Representing the A concatenated vector value, Represents the first time window Normalized focus features at time points. Represents the first time window Normalized stress characteristic at time t, Represents the number of moments in the time window. The weighting coefficients represent the focused feature quantities; Weighting coefficients representing stress characteristic quantities The impact index representing the focus characteristics. The influence index representing stress characteristics The balance index represents the composite characteristic quantity. Representing the The time correlation coefficient at any given moment Representing the Temperature correction factor at any given time.

[0014] As a further aspect of the present invention, the acquisition step of the disaster prediction module specifically includes: The historical data extraction submodule extracts historical data of rainfall, pore water pressure, and surface displacement for the corresponding time period based on the trend prediction feature vector of the trend classification result. It then synchronizes and aligns the three types of data in chronological order to generate a historical data sequence. The disaster trend submodule calls the historical data sequence and calculates the probability of disaster occurrence trends for multiple time periods based on rainfall, pore water pressure, and surface displacement data for each time period, generating disaster trend probability calculation results. The disaster screening submodule, based on the disaster trend probability calculation results, filters the probability of disaster occurrence in multiple time periods according to a preset probability threshold, marks the time periods that exceed the threshold, and generates disaster trend prediction results.

[0015] As a further aspect of the present invention, the process of filtering the probability of disaster occurrence over multiple time periods based on a preset probability threshold specifically involves setting the probability threshold to a fixed value between zero and one, comparing the probability of disaster occurrence trends over time periods based on quantification parameters, and marking the time periods in which the probability of disaster occurrence trends exceeds the probability threshold as disaster trend prediction results.

[0016] Compared with the prior art, the advantages and positive effects of the present invention are as follows: This invention, by introducing multi-level anomaly data analysis and clustering techniques, can more accurately identify potential anomalies when processing multi-dimensional data, effectively solving the problem of traditional methods neglecting data nonlinearity and dynamic changes. By calculating the differences in multi-dimensional data, different stress zones can be distinguished, and trend predictions can be made based on the characteristics of these zones, significantly improving the accuracy and timeliness of disaster prediction. Utilizing the support vector machine classification method to integrate focusing features and stress features not only optimizes the adaptability of the prediction model but also allows for adjustments to the prediction results based on real-time data, thus providing more timely and reliable support for decision-making. Attached Figure Description

[0017] Figure 1 This is a system flowchart of the present invention; Figure 2 This is a flowchart of the anomaly identification module of the present invention; Figure 3 This is a flowchart of the clustering annotation module of the present invention; Figure 4 This is a flowchart of the stress partitioning module of the present invention; Figure 5 This is a flowchart of the trend classification module of the present invention; Figure 6 This is a flowchart of the disaster prediction module of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0019] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0020] Please see Figure 1 A geological disaster occurrence trend prediction system based on historical data includes: The anomaly identification module collects historical rainfall, pore water pressure, and surface displacement, divides the data into time windows, standardizes the data, judges the deviation of the data, marks historical data that exceeds the anomaly identification threshold, generates anomaly structure data, and transmits it to the clustering and labeling module. The clustering and labeling module clusters abnormal structural data, sorts abnormal data in the same cluster according to the cluster center, generates structured focused data, and transmits it to the stress partitioning module. The stress zoning module, based on structured focused data, calculates the rainfall increment, pore water pressure gradient, and displacement acceleration for the corresponding time window and makes differences, distinguishes the variation range of multiple parameters, generates stress zoning data, and transmits it to the trend classification module. The trend classification module extracts focusing features and stress features based on stress zoning data and concatenates them into a trend prediction feature vector. This vector is then input into a support vector machine for classification, generating trend classification results which are then passed to the disaster prediction module. The disaster prediction module extracts historical data for the corresponding time period based on the trend classification results, calculates the probability of disaster occurrence trends in multiple time periods, filters the time periods, and generates geological disaster trend prediction results.

[0021] The abnormal structure data includes rainfall deviation, pore water pressure deviation, and displacement offset rate; the structured focused data includes cluster center value, cluster sequence number, and cluster ranking value; the stress zoning data includes rainfall increment interval, pore water pressure gradient interval, and displacement acceleration interval; the trend classification results include trend category quantity, feature vector magnitude, and classification hyperplane distance; and the geological disaster trend prediction results include time period probability value, threshold exceedance label quantity, and risk change segment.

[0022] Please see Figure 2 The specific steps for obtaining the anomaly detection module are as follows: The data analysis submodule acquires historical rainfall, pore water pressure, and surface displacement sequences, divides the three types of sequences into time windows, calls the segmented values ​​of rainfall, pore water pressure, and surface displacement within the time window, calculates the window difference value, and generates the window difference result. First, a communication connection with the geological monitoring database is established, and data acquisition commands are set. Historical monitoring sequences spanning the most recent 30 days and sampled hourly are extracted from the rainfall sensor, pore water pressure gauge, and surface displacement gauge at monitoring point A. The collected raw data streams are converted into floating-point numerical arrays, with the rainfall, pore water pressure, and surface displacement sequences serving as independent data stream inputs. Then, the time window length is set to 24 hours, and the sliding step size to 1 hour. A sliding window algorithm is used to simultaneously extract data segments from the three historical sequences. For example, data from hour t to hour t plus 24 hours constitutes the i-th time window, and so on, generating a total of N time windows. Next, the data within each independent time window is segmented. The 24-hour window is divided into four sub-segments of 6 hours each. The cumulative rainfall, average pore water pressure, and surface displacement increment are calculated for each sub-segment, resulting in four sub-segments within that window. The feature values, as shown in Table 1, list the raw values ​​of some monitoring data before entering the segmentation process. Taking the 10th window as an example, the cumulative values ​​of the four rainfall segments it contains are 10.5 mm, 15.2 mm, 8.0 mm, and 2.1 mm, respectively. Then, the dispersion of these four segment values ​​is calculated. Specifically, the arithmetic mean of the four segment values ​​is first calculated as (10.5 + 15.2 + 8.0 + 2.1) / 4 = 8.95. Then, the arithmetic mean of each segment value is calculated separately. The absolute value of the difference is calculated, and these four absolute values ​​are summed. The calculation process is 1.55 + 6.25 + 0.95 + 6.85 = 15.6. This summation result is defined as the rainfall window difference value under this window. Similarly, the pressure window difference value is calculated based on the segmented average value sequence of pore water pressure, and the displacement window difference value is calculated based on the segmented incremental value of surface displacement. Segmented statistics and absolute deviation summation are performed on all N windows. The three types of difference values ​​corresponding to each window are associated and stored to generate the window difference result.

[0023] Table 1. Historical sequence data fragments of monitoring points As shown in Table 1, the table displays the original data segments of the monitoring points from the 10th to the 14th hour. These data are then processed by summation, averaging, or difference to form the segmented feature values ​​required for subsequent steps.

[0024] The standardized calculation submodule, based on the window difference results, calls the segmented values ​​of rainfall, pore water pressure, and surface displacement within the corresponding window, calculates the ratio based on the window difference values, adjusts the ratio interval, calculates the interval ratio, and obtains the interval ratio sequence. The system reads the window difference results containing N sets of data. First, it sets standardized reference values, which are selected as the maximum fluctuation extreme value of historical monitoring data for the same period. The reference values ​​are set as follows: rainfall difference 50.0, pore water pressure difference 20.0, and surface displacement difference 5.0. For any i-th window, the corresponding rainfall window difference value is divided by the reference value to obtain the rainfall ratio. Similarly, the pressure ratio and displacement ratio are calculated. For example, if the displacement window difference value of a certain window is 1.2, the calculated displacement ratio is 1.2 / 5.0 = 0.24. Next, the effective judgment interval of the ratio is defined, setting the lower limit of the ratio to 0.1 and the upper limit to 0.9. The calculated three... The analogy values ​​undergo interval mapping checks. If a ratio is less than 0.1, it is forcibly corrected to 0.1; if it is greater than 0.9, it is corrected to 0.9; and if it is between 0.1 and 0.9, it remains unchanged. After correction, a weighted calculation rule is introduced to synthesize the interval ratios. The weights are set as follows: rainfall (0.3), pore water pressure (0.3), and surface displacement (0.4). The three corrected ratios are multiplied by their respective weights, i.e., a weighted summation is performed: rainfall ratio multiplied by 0.3, pressure ratio multiplied by 0.3, and displacement ratio multiplied by 0.4. This weighted sum is defined as the final interval ratio for that window. For example, if the three corrected ratios for a window are 0.4, 0.5, and 0.24, the calculated ratio is 0.4. 0.3 + 0.5 0.3 + 0.24 0.4 = 0.12 + 0.15 + 0.096 = 0.366. Following this process, all window data are processed one by one, and the calculated weighted values ​​are arranged according to time to obtain the interval ratio sequence.

[0025] The deviation screening submodule, based on the interval ratio sequence, calls the multi-window ratio, and screens the deviation amount according to the difference between the ratio value and the preset anomaly identification threshold. It obtains the deviation amount of all windows and marks the windows whose deviation amount exceeds the threshold, generating abnormal structure data. The system receives an interval ratio sequence containing N values. First, a preset threshold for identifying anomalies is set. This threshold is based on the arithmetic mean of the ratio sequence under normal geological conditions plus twice the standard deviation. Based on calculations using historical stable period data, the anomaly identification threshold is set to 0.65. Then, each value in the interval ratio sequence is iterated over, and a value comparison operation is performed. The preset anomaly identification threshold is subtracted from the ratio value of the current window to obtain the ratio difference. For example, if the interval ratio of a certain window is 0.82, the difference is 0.82 - 0.65 = 0.17; if the ratio of another window is 0.50, the difference is 0.50 - 0.65 = -0.15. Finally, the direction of deviation is determined based on the sign of the difference; a positive value indicates that the data in that window exhibits [anomaly / abnormality]. A deviation exceeding the norm is considered high-risk, while a negative value represents a low-risk deviation. The absolute value of the difference is calculated as the degree of deviation. A minimum sensitivity threshold of 0.05 is set to filter out minor fluctuations and noise. The calculated degree of deviation for each window is compared with the sensitivity threshold. Only when the degree of deviation is greater than 0.05 and the deviation direction is positive is the window considered to have a substantial anomaly. For windows that meet the above conditions, their window number, start timestamp, corresponding degree of deviation, and original interval ratio value are packaged and recorded. For example, if the 45th window is determined to be an anomaly, the record includes the window ID, occurrence time, deviation value of 0.17, and original ratio of 0.82. All the information of the marked windows is summarized and integrated to generate anomaly structure data.

[0026] Please see Figure 3 The specific steps for obtaining the clustering annotation module are as follows: The data clustering submodule, based on abnormal structure data, uses the K-means clustering algorithm to group the data, spatially partitions the data according to the attributes of each data point, calculates the cluster center of each group, assigns a cluster label to each data point, and obtains a cluster label sequence. Based on the generated abnormal structure data, retrieve the window number and deviation level contained within it. and interval ratio values As a core clustering attribute, a feature vector space is first constructed to describe the abnormal state, mapping the data points of each abnormal window to two-dimensional coordinates, and setting the parameter for the number of cluster groups. The coordinates of the three cluster centers are set to 3, corresponding to the three geological state categories of "minor disturbance," "moderate risk," and "severe mutation," respectively. Based on historical monitoring experience, the initial centers are set to... , , Then, Euclidean distance calculation is performed for each outlier data point. Taking the data in window number 45 listed in Table 2 as an example, its characteristic coordinates are... Calculate the distance from this point to The distance is the square root of the sum of the squares of 0.17 minus 0.08 and 0.82 minus 0.70, which is the square root of the sum of the squares of 0.09 and 0.12. The result is 0.15. The same logic applies to the distance... The distance is the square root of 0.17 minus 0.15 squared plus the square root of 0.82 minus 0.78 squared, resulting in 0.0447. (Calculation up to...) The distance is calculated as the square root of the sum of the squares of 0.17 minus 0.25 and 0.82 minus 0.85, which gives 0.0854. Through numerical comparison, this data point is... Since the distance is the smallest, window number 45 is grouped into category 2 and given a temporary label. After the first round of allocation of all data points, the geometric center of each group is recalculated, and all data points belonging to the second group are included. Sum the values ​​and divide by the number of points to obtain the new center x-coordinate. The summation of the values ​​divided by the number of points yields the new center ordinate. If the newly calculated center coordinates... for If the old center is replaced with the new coordinates, the distance calculation and center update steps are repeated until the change in the center coordinates is less than the convergence threshold of 0.001. The fixed category to which each data point belongs is confirmed, and the cluster label sequence is obtained.

[0027] Table 2. Abnormal Structure Data Sampling Table As shown in Table 2, the table lists some of the abnormal window data that entered the cluster analysis stage. The deviation degree and interval ratio are used as two-dimensional features to input the clustering algorithm to determine the risk category label to which they belong.

[0028] The sorting submodule, based on the cluster label sequence, sorts the data points according to the time position of each cluster center by calling the timestamp corresponding to each cluster center, and obtains the time sorting sequence. Based on the cluster label sequence, the timestamp attribute of all data points within each cluster group is first identified and extracted. The timestamps are then converted to Unix time values ​​for easier arithmetic operations. For each cluster, the centroid of its time dimension is calculated. Specifically, the timestamps of all data points within that cluster are summed and divided by the total number of data points in that cluster to obtain the time position of the cluster center. For example, if the first cluster contains three data points with timestamps of 1764921600, 1764925200, and 1764928800, the time position of the cluster center is the sum of these three timestamps divided by 3, resulting in 1764925200. The same calculation is performed for the second and third clusters. The time position of the center is determined, and then the three cluster categories are macroscopically sorted according to the calculated time position values. If the center time of the first category is earlier than that of the second category, and the second category is earlier than that of the third category, then the processing order of the categories is established as category 1, category 2, and category 3. After determining the category order, the data points are sorted microscopically by calling the specific collection timestamp of each data point. For the data in the first category, the timestamp size is compared, and the data points are rearranged according to the ascending order from earliest to latest, placing the data point with time 08:00 before the data point with time 09:00. The data points of all categories are sorted in turn, and the sorted data streams of each category are spliced ​​together end to end according to the macroscopic category order to obtain the time sorted sequence.

[0029] The structured output submodule integrates the data based on the time-sorted sequence and the clustering labels of each data point, performs structured processing on the data according to the clustering labels and time order, and generates structured focused data. Based on the time-sorted sequence, a structured data container with multiple fields is created. This container pre-sets standard fields including a unique identifier, risk level label, occurrence time, core indicator value, and related feature parameters. The sorted data sequence is traversed, and the attribute information of each data point is read and populated into the corresponding fields. For the data in window 45, cluster labels are defined. The data is mapped to the text description "Medium Risk" and filled into the Risk Level field. The timestamp "2025-12-05 08:00" is extracted and filled into the Occurrence Time field. The Deviation Level (0.17) and Interval Ratio (0.82) are filled into the Core Indicator field. At the same time, according to the preset data encapsulation protocol, a hash checksum for indexing is generated for this record. All fields are combined into a standard JSON format object, for example, the generated content is "{ID: 45, Level: Medium, Time: 2025-12-05, 08:00, Metrics: {D: 0.17, R: 0.82}}". The above encapsulation operation is performed on all data points in the sequence, and the encapsulated objects are stored in the result array according to the predetermined time sorting order to generate structured focused data.

[0030] Please see Figure 4 The specific steps for obtaining the stress partitioning module are as follows: The parameter calculation submodule extracts rainfall, pore water pressure and surface displacement within the corresponding time window based on structured focused data, calculates rainfall increment, pore water pressure gradient and surface displacement acceleration, performs calculations based on numerical gradients, and generates gradient change. Based on the generated structured focused data, the JSON format objects contained within are parsed to extract the original monitoring sequences within the target time window (taking the anomaly window ID-45 as an example). Rainfall arrays, pore water pressure readings at different depths, and surface displacement time-series data for that time period are obtained. Weighted incremental calculations are performed on the rainfall data using the following formula: In the formula This represents an increase in rainfall, used to quantify the overall impact of rainfall events on geological bodies. (Superscript) The total number of sub-time periods within the time window is set to 3 hourly units. For sub-time period index, Representing the The absolute increase in precipitation within each sub-period (unit: mm). The dimensionless coefficient representing the area influence of the rain zone within this time window. The instantaneous rainfall intensity is represented by the derivative of precipitation with respect to time (unit: mm / hour). Representative at the The minute variation in precipitation within a time window, denoted by [symbol]. The differential time is taken as 1 hour, and the multiplication operations in the formula... The aim is to construct a "rainfall intensity-rainfall amount" coupling term similar to the definition of kinetic energy. By multiplying the total precipitation by the precipitation rate, the signal characteristics of short-duration heavy rainfall are nonlinearly amplified, while the coefficients... The introduction of this method aims to correct point-based monitoring data into areal influence factors through topographic projection relationships, and the final summation sign... This is used to accumulate all instantaneous impact effects within the entire window period, thereby accurately distinguishing the different degrees of stress activation in the soil caused by gradual and continuous rainfall versus sudden torrential rain. The specific calculation process is as follows: First, the influence coefficient of the rain area is set. This coefficient is calculated by measuring the actual surface area of ​​the monitored area. With horizontal projected area The ratio is obtained if It is 1100 square meters. If it is 1000 square meters, then Next, we extracted rainfall data for three consecutive hours, as shown in Table 3. The rainfall in the first hour... It is 12.5 mm, and the rate of change is... That is, 12.5 mm / hour, in the second hour. The value was 5.0 mm, with a rate of change of 5.0 mm / hour, in the 3rd hour. The value is 22.0 mm, and the rate of change is 22.0 mm / hour; substituting the value into the formula... The first item is calculated as The second item is calculated as follows: The third item is calculated as follows: Finally, the sum of these three results yields the rainfall increment. This result indicates that a high-impact rainfall event occurred during this period. Simultaneously, the pore water pressure gradient was calculated using readings from two sensors at depths of 5.0 meters and 7.5 meters. The upper reading was 30.5 kPa, and the lower reading was 38.9 kPa. The difference was calculated. kPa, divided by the distance The pressure gradient was calculated to be 3.36 kPa / m. The surface displacement acceleration was calculated by extracting the cumulative displacement values ​​of 100.10 mm, 100.35 mm, and 100.75 mm from three consecutive time points. The velocity in the first stage was then calculated. mm / h, the second stage speed is mm / hour, difference in calculation speed The acceleration is calculated as 0.15 mm / sqm / h by dividing the time interval by 1 hour. The calculated rainfall increment of 731.775, pore water pressure gradient of 3.36, and surface displacement acceleration of 0.15 are then combined to generate the gradient change.

[0031] Table 3 Detailed Parameters for Calculating Rainfall Increment As shown in Table 3, the table details the parameter values ​​and sub-item calculation results for each sub-period when calculating the rainfall increment for the ID-45 window. The final accumulated value constitutes the rainfall impact index for this window.

[0032] The difference interval submodule calls the gradient change amount and performs difference judgment on the rainfall increment, pore water pressure gradient and surface displacement acceleration based on the gradient change amount. It divides the three types of values ​​into change intervals according to the preset interval division benchmark value and obtains the interval range value sequence. The generated gradient changes—namely, rainfall increment 731.775, pore water pressure gradient 3.36, and surface displacement acceleration 0.15—are used. Differentiated interval classification benchmarks are set for these three physical quantities. These benchmarks are derived from inversion analysis of historical geological disaster cases. For rainfall increment, the upper limit benchmark value for the low-intensity interval is set at 500.0, and the lower limit benchmark value for the high-intensity interval is set at 1000.0. Intervals less than 500.0 are defined as "Interval I," 500.0 to 1000.0 as "Interval II," and greater than 1000.0 as "Interval III." For pore water pressure gradient, benchmark values ​​are set at 2.0 kPa / m and 5.0 kPa / m. For surface displacement acceleration, benchmark values ​​are set at 0.10 mm / h² and 0.30 mm / h². A difference judgment is then performed, comparing the calculated rainfall increment 731.775 with the benchmark values. The baseline value was compared, and the value was determined to be between 500.0 and 1000.0, belonging to "Interval II", indicating a moderate rainfall-induced state. The pore water pressure gradient of 3.36 was compared with the baseline value, and the value was determined to be between 2.0 and 5.0, belonging to "Interval II", indicating a moderate seepage pressure state. The surface displacement acceleration of 0.15 was compared with the baseline value, and the value was determined to be between 0.10 and 0.30, belonging to "Interval II", indicating an accelerated deformation state. These three determination results were combined in the order of "rainfall-pressure-displacement" to form an interval range text describing the physical state of the window, such as "Rain: [Interval-II], Press: [Interval-II], Disp: [Interval-II]". The same baseline comparison and interval mapping operation was performed on all abnormal windows to obtain the interval range value sequence.

[0033] The label segmentation submodule, based on the interval range value sequence, performs label numbering for each group of interval range value sequences, matches the interval range value sequence to the label number according to the position of the numerical interval and generates the corresponding label, thus generating stress zoning data; Based on the generated interval range value sequence, a multi-dimensional stress zoning mapping table is pre-constructed. This table defines the geological stress state labels corresponding to different interval combinations, and the mapping rules are set as follows: if all three parameters are in interval I, the corresponding label number is "L-01", representing "stable bearing zone"; if all three parameters are in interval II, the corresponding label number is "L-02", representing "transitional creep zone"; if any parameter enters interval III and the other parameters are not lower than interval II, the corresponding label number is "L-03", representing "critical slip zone"; for the interval sequence "interval II-interval" generated by the aforementioned ID-45 window, ... If the sequence of another window is "Interval II-Interval II-Interval II", the mapping rule base is traversed, and the combination is found to fully meet the definition conditions of "L-02". Therefore, the "L-02" label is assigned to this window, and the semantic description of "transitional creep zone" is attached. If the sequence of another window is "Interval III-Interval II-Interval II", then the "L-03" label is matched. Each combination item in the entire sequence is scanned one by one, and the abstract interval range value is transformed into a specific stress state code according to the position matching logic. Finally, the classification information containing the label number (such as L-02) and its physical meaning (such as transitional creep zone) is bound to the original window ID to generate stress partition data.

[0034] Please see Figure 5 The specific steps for obtaining the trend classification module are as follows: The feature splicing submodule, based on stress partitioning data and structured focusing data, extracts focusing feature quantities and stress feature quantities within the time window and performs splicing action at the same time sequence position. It arranges the focusing feature quantities and stress feature quantities in sequence order, calculates the splicing vector value, and generates a splicing vector set. Based on the generated stress zoning data and structured focusing data, an association index for the two types of data is first established using a unique window number (ID-45), and continuous data within that time window is extracted. At any time (set) (Focusing feature quantity) With stress characteristic quantity Data alignment and splicing are performed at the same time series positions, where normalization focuses on feature quantities. The value is a measure of the degree of deviation in structured data. Ratio of intervals The arithmetic mean, normalized focused feature quantity This is obtained by numerically mapping the stress zoning labels. The mapping standard is set as follows: label "L-01" is mapped to 0.2, "L-02" to 0.5, and "L-03" to 0.8. Soil moisture content is also introduced as an additional characteristic quantity. and the historical average as a relative reference characteristic. Then, the splicing calculation formula is called: In the formula Representing the A concatenated vector value, Represents the first time window Normalized focus features at time points. Represents the first time window Normalized stress characteristic at time t, Represents the number of moments in the time window. The weighting coefficients represent the focused feature quantities; Weighting coefficients representing stress characteristic quantities The impact index representing the focus characteristics. The influence index representing stress characteristics The balance index represents the composite characteristic quantity. Representing the The time correlation coefficient at any given moment Representing the Temperature correction factor at time, and The weighting coefficients for the focusing characteristic quantity and the stress characteristic quantity are set respectively. , This is used to balance the contributions of statistical and physical characteristics. Set to 2.0. Set to 2.0, these are all influence indices on feature differences, used to enhance the sensitivity to biases between features. It is set to 2.0 as a balancing index for the synthesized features, serving to smooth the numerical magnitude. The time correlation coefficient represents the coefficient used to correct for the time lag effect in data acquisition. It is calculated as follows: If there is no lag, then take 1.0. This represents the temperature correction factor, with a value of [value missing]. ,in The current surface temperature; a specific example is: for window ID-45, Take 1.0, It is 30 degrees Celsius, therefore As shown in Table 4, at time 1 (Depend on (Calculate the mean) (Corresponding to L-02), Time 2 , , third moment , Substitute into the formula Calculate: The term at time 1 is The term at time 2 is The term at time 3 is The sum of the three terms is Take it The square root of the exponent is obtained by taking the square root. Finally, multiply by a coefficient. Calculation The result indicates that the feature synthesis intensity within this window is at a medium-to-high level, generating a spliced ​​vector set.

[0035] Table 4. Detailed Table of Concatenation Vector Calculation Parameters As shown in Table 4, the table lists the intermediate calculation results of various instantaneous characteristic parameters and their differences required to calculate the ID-45 window splicing vector value, reflecting the dispersion of the monitoring data and the baseline state.

[0036] The vector construction submodule calls the concatenated vector set, arranges the concatenated vector values ​​according to the feature dimension to form the trend prediction input vector and forms the input vector magnitude, and performs vector normalization according to the input vector magnitude to obtain the normalized vector; The program calls a concatenated vector set containing the calculation results of multiple time windows, and extracts the concatenated vector values ​​corresponding to each window in chronological order. For example, it extracts the values ​​of three consecutive windows from ID-43 to ID-45. These three values ​​are arranged sequentially to construct a three-dimensional trend prediction input vector. Then, vector normalization is performed based on the magnitude of the input vector to eliminate the influence of the absolute magnitude of the values ​​on the classification model. The upper bound of normalization is set to 1.5, and the lower bound is set to 0.0 (these boundary values ​​are determined based on the theoretical extreme value range of the formula under normalized input). Min-max normalization logic is used, and for each element in the vector, the first element is calculated... ; calculate for the second element ; calculate the third element After processing, a normalized vector is obtained. If the calculation result exceeds the range of 0 to 1, it is forcibly truncated to the boundary value. For example, if it is greater than 1, it is taken as 1. This process ensures that the data input into the model has a uniform scale standard and obtains a normalized vector.

[0037] The trend classification submodule, based on normalized vectors, performs classification on the support vector machine model by inputting normalized vectors, calculates the category output and uses it as the trend classification metric, and obtains the trend classification result. Based on the generated normalized vector The input vector is fed as a feature input into a pre-trained Support Vector Machine (SVM) classification model. This model uses a radial basis function (RBF) as the kernel function to handle nonlinear classification problems. Internally, the model maps the input vector to a high-dimensional feature space and calculates the distance between the input vector and each support vector. It then performs a classification decision based on the decision boundary function. The preset category labels are "C1 - stationary trend", "C2 - gradual trend", and "C3 - sudden trend". The model outputs the probability confidence of each category. For example, the probability of belonging to C1 is calculated to be 0.15, the probability of belonging to C2 is 0.25, and the probability of belonging to C3 is 0.60. According to the principle of maximum probability, "C3 - sudden trend" with the highest probability value is selected as the final classification output. This result is associated with the original window ID and stored to form a trend classification result. This result indicates that the current geological state is on the trajectory of evolving towards a high-risk sudden change, thus obtaining the trend classification result.

[0038] Please see Figure 6 The specific steps to obtain the disaster prediction module are as follows: The historical data extraction submodule extracts historical data of rainfall, pore water pressure, and surface displacement for the corresponding time period based on the trend prediction feature vector of the trend classification results. It then synchronizes and aligns the three types of data in chronological order to generate a historical data sequence. Based on the output trend classification results, the abnormal window number (ID-45) marked as "C3-mutation trend" was identified. The time span covered by this window was analyzed, and the start time point was determined to be 08:00 on December 5, 2025, and the end time point was determined to be 10:00 on December 5, 2025. Subsequently, a targeted search was launched on the underlying multidimensional database to retrieve three types of core monitoring data within this time period. The first type was hourly rainfall data, with search results of 12.5 mm at 08:00, 5.0 mm at 09:00, and 22.0 mm at 10:00. The second type was pore water pressure sensor data at a depth of 5 meters. The readings retrieved were 30.5 kPa at 08:00 and 38.9 kPa at 10:00. The third category consisted of cumulative displacement values ​​from surface GNSS monitoring points, with results of 100.10 mm at 08:00, 100.35 mm at 09:00, and 100.75 mm at 10:00. A time synchronization and alignment operation was performed, constructing a data matrix with timestamps as row indices and physical quantities as column attributes. Data integrity was checked; if pore water pressure data for a certain time (e.g., 09:00) was missing, linear interpolation was performed using values ​​from adjacent times (08:00 and 10:00) to fill the missing data. Using kilopascals as a substitute value ensures that the data sequence is continuous and uninterrupted. The cleaned and aligned 3D dataset is encapsulated in chronological order to generate a historical data sequence.

[0039] The disaster trend submodule calls historical data sequences and calculates the probability of disaster occurrence trends for multiple time periods based on rainfall, pore water pressure, and surface displacement data for each time period, generating disaster trend probability calculation results. The generated historical data sequence is used to perform a quantitative calculation of the probability of disaster occurrence for each time slice. This calculation process adopts a multi-factor weighted normalization scoring logic. First, the critical values ​​of each parameter are defined: the critical value for hourly rainfall is 50.0 mm, the critical value for pore water pressure is 50.0 kPa, and the critical value for surface displacement rate is 1.0 mm / hour. Weighting coefficients are also assigned: the weight for rainfall is set to 0.3, the weight for pore water pressure is set to 0.3, and the weight for displacement rate is set to 0.4. For the data at 10:00, the calculation shows a rainfall of 22.0 mm, and its normalized risk value is... After weighting, we get The pore water pressure is 38.9 kPa, and the normalized risk value is [value missing]. After weighting, we get For displacement data, first calculate the displacement rate at this moment relative to the previous moment (09:00), i.e. mm / hour, its normalized risk value is After weighting, we get Finally, the three weighted values ​​are added together. The score is converted into a final probability value through a preset Sigmoid mapping function. The probability of disaster occurrence at that moment is set to 0.82 after mapping. Similarly, the probability at 08:00 is calculated to be 0.35 and the probability at 09:00 is calculated to be 0.45. The system binds the calculated probability values ​​with the corresponding timestamps one by one. As shown in Table 5, the normalization results of the parameters at each moment and the final calculated trend probability are listed in detail, generating the disaster trend probability calculation results.

[0040] Table 5. Detailed Table for Calculating Disaster Trend Probability As shown in Table 5, the table displays the intermediate values ​​of various monitoring indicators at three consecutive times within the ID-45 window after normalization and weighting, as well as the final derived probability of disaster occurrence, clearly reflecting the evolution of risk accumulation over time.

[0041] The disaster screening submodule, based on the disaster trend probability calculation results, filters the probability of disaster occurrence in multiple time periods according to the preset probability threshold, marks the time periods that exceed the threshold, and generates disaster trend prediction results. Based on the disaster trend probability calculation results, i.e., the sequence A stringent disaster warning screening threshold was set, which was based on the average probability characteristics of the pre-slide phase in historical landslide events, and was set to 0.80. This means that only when the calculated disaster probability reaches or exceeds 80% is the period considered to have a substantial risk of disaster. A step-by-step screening process was then performed, first comparing the probability at 08:00 to 0.35. This was determined to be a safe fluctuation and was ignored; then the probability at 09:00 was compared to 0.45, because... The risk accumulation period has been determined, and no alarm will be triggered for the time being; the probability at 10:00 is 0.82, therefore... If the screening criteria are met, the time period is immediately locked and marked as a "high-risk time of landslide". All feature parameters and probability values ​​associated with this time are extracted and packaged into an independent early warning object. If multiple consecutive times exceed the threshold, they are merged into a continuous disaster early warning interval. Finally, a screening report containing information on high-risk time periods and corresponding probability values ​​is output, generating disaster trend prediction results.

[0042] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1.A geological disaster occurrence trend prediction system based on historical data, characterized by, The system comprises: An anomaly identification module collects historical rainfall, pore water pressure and ground surface displacement, divides time windows, normalizes data and judges data deviation, marks historical data exceeding an anomaly identification threshold, generates anomaly structure data and passes it to a clustering labeling module; The clustering labeling module clusters the anomaly structure data, sorts anomaly data of the same cluster according to the cluster center, generates structured focused data and passes it to a stress partition module; The stress partition module calculates rainfall increment, pore water pressure gradient and displacement acceleration of the corresponding time window based on the structured focused data, judges the differences, distinguishes the change intervals of multiple parameters, generates stress partition data and passes it to a trend classification module; The trend classification module extracts focused features and stress features based on the stress partition data, splices them into a trend prediction feature vector, inputs a support vector machine classification, generates a trend classification result and passes it to a disaster prediction module; The disaster prediction module extracts historical data of the corresponding time based on the trend classification result, calculates disaster occurrence trend probabilities of multiple time periods and filters time periods, generates a geological disaster trend prediction result. 2.The geological disaster occurrence trend prediction system based on historical data according to claim 1, wherein, The anomaly structure data includes rainfall deviation, pore water pressure deviation and displacement deviation rate, the structured focused data includes cluster center value, cluster sequence number and cluster sorting amount, the stress partition data includes rainfall increment interval, pore water pressure gradient interval and displacement acceleration interval, the trend classification result includes trend category amount, feature vector magnitude and classification hyperplane distance, and the geological disaster trend prediction result includes time period probability value, threshold exceeding label amount and risk change section. 3.The geological disaster occurrence trend prediction system based on historical data according to claim 2, wherein, The acquisition step of the anomaly identification module is specifically as follows: A data analysis submodule acquires historical rainfall, pore water pressure and ground surface displacement sequences, divides time windows for the three types of sequences, calls rainfall segment values, pore water pressure segment values and ground surface displacement segment values in the time windows, calculates window difference values and generates window difference results; A normalization calculation submodule calls rainfall segment values, pore water pressure segment values and ground surface displacement segment values in the corresponding windows based on the window difference results, calculates ratios according to the window difference values, adjusts the ratio interval and calculates interval ratios, and acquires interval ratio sequences; A deviation screening submodule calls multi-window ratios based on the interval ratio sequences, screens deviation degree amounts according to the difference between the ratio values and a preset anomaly identification threshold, acquires all window deviation degree amounts and marks windows with deviation degree amounts exceeding the threshold, and generates anomaly structure data. 4.The geological disaster occurrence trend prediction system based on historical data according to claim 3, wherein, The acquisition step of the clustering labeling module is specifically as follows: A data clustering submodule groups data based on the anomaly structure data based on a K-means clustering algorithm, divides space according to the attributes of each data point, calculates the cluster center of each group, divides the cluster label to which each data point belongs, and obtains a cluster label sequence; An ordering processing submodule sorts each cluster center according to the time position, calls the timestamp corresponding to each cluster center to sort data points, and acquires a time ordering sequence. The structured output submodule integrates the data based on the time sequence and in combination with the cluster labels of each data point, and generates structured focused data by structuring the data according to the cluster labels and the time sequence. 5.The geological disaster occurrence trend prediction system based on historical data according to claim 4, wherein, The stress partition module comprises: The parameter calculation submodule extracts the rainfall, pore water pressure and surface displacement in the corresponding time window based on the structured focused data, calculates the rainfall increment, pore water pressure gradient and surface displacement acceleration, performs a calculation action according to the numerical gradient, and generates a gradient change amount. The difference interval submodule calls the gradient change amount, performs a difference judgment action on the rainfall increment, pore water pressure gradient and surface displacement acceleration according to the gradient change amount, divides the three types of values into change intervals according to a preset interval division reference value, and obtains an interval range value sequence. The label division submodule generates stress partition data by performing label numbering on each interval range value sequence based on the interval range value sequence, matching the interval range value sequence to a label number according to the numerical interval position, and generating a corresponding label. 6.The geological disaster occurrence trend prediction system based on historical data according to claim 5, wherein, The trend classification module comprises: 7.The geological disaster occurrence trend prediction system based on historical data according to claim 6, wherein, The feature splicing submodule extracts the focused feature quantity and the stress feature quantity in the time window based on the stress partition data and the structured focused data, performs a splicing action at the same time sequence position, arranges the focused feature quantity and the stress feature quantity in sequence, calculates a splicing vector value, and generates a splicing vector set. The vector construction submodule calls the splicing vector set, arranges the splicing vector value into a trend prediction input vector according to the feature dimension, forms an input vector magnitude, performs vector normalization according to the input vector magnitude, and obtains a normalized vector. The trend classification submodule performs classification on the normalized vector by inputting the normalized vector into a support vector machine model, calculates a class output quantity as a trend classification quantity, and obtains a trend classification result. The formula for calculating the splicing vector value is: 8.The geological disaster occurrence trend prediction system based on historical data according to claim 7, wherein, The disaster prediction module comprises: ; wherein: represents the nth spliced vector value, represents the normalized focusing feature quantity at the nth moment within the time window, represents the normalized stress feature quantity at the nth moment within the time window, represents the number of moments within the time window, represents the weight coefficient of the focusing feature quantity; represents the weight coefficient of the stress feature quantity, represents the influence index of the focusing feature, represents the influence index of the stress feature, represents the balance index of the synthesized feature quantity, represents the time correlation coefficient at the nth moment, represents the temperature correction coefficient at the nth moment. 9.The geological disaster occurrence trend prediction system based on historical data according to claim 1, wherein, The historical data extraction submodule extracts the rainfall, pore water pressure and surface displacement historical data of the corresponding time period based on the trend prediction feature vector of the trend classification result, synchronously aligns the three types of data according to the time sequence, and generates a historical data sequence. The disaster trend submodule calls the historical data sequence, calculates the disaster occurrence trend probability of multiple time periods based on the rainfall, pore water pressure and surface displacement data of each time period, and generates a disaster trend probability calculation result. The disaster screening submodule screens the disaster occurrence probability of multiple time periods according to a preset probability threshold based on the disaster trend probability calculation result, marks the time periods that exceed the threshold, and generates a disaster trend prediction result. ​ 10.The geological disaster occurrence trend prediction system based on historical data according to claim 9, wherein, The process of screening the disaster occurrence probabilities of the multiple time periods according to the preset probability threshold value is specifically that the probability threshold value is set as a fixed value between zero and one, and the disaster occurrence trend probabilities are compared time period by time period based on the quantification parameters, and the time periods with the disaster occurrence trend probabilities exceeding the probability threshold value are marked as the disaster trend prediction results.

Citation Information

Cited By

  • A cold region ice-containing joint rock mass instability early warning method based on multi-source data fusion

    CN122392282A