Hydrological data prediction method based on data analysis
Through the hydrological data prediction method based on data analysis, historical precipitation, soil moisture content and soil permeability data are obtained and analyzed, precipitation intensity and soil harshness are calculated, clustering and disaster possibility assessment are carried out, and disaster prediction models are constructed, which solves the problem that existing methods cannot accurately predict mudslide disasters, and achieves more accurate disaster prediction and prevention measures.
Patent Information
- Application Number
- CN202510526005.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The existing methods cannot accurately predict the occurrence of mudslide disasters, mainly because there are large differences in hydrogeological data in the same area at different times.
A hydrological data prediction method based on data analysis is proposed. By obtaining the historical precipitation, soil moisture content and soil permeability data of the area to be measured, the days where the mudslides occur are used as disaster days, and the days where the mudslides have not occurred are used as normal days, the precipitation intensity and soil harshness are calculated, the precipitation characteristic similarity and soil characteristic similarity are analyzed, clustering and disaster possibility assessment are carried out, and the disaster prediction model is finally constructed.
Through detailed data analysis and model construction, the occurrence of mudslide disasters can be more accurately predicted and the effectiveness of prevention and remedial measures can be improved.
Smart Images

Figure CN120046386A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of disaster prediction, and particularly to a hydrological data prediction method based on data analysis. Background Art
[0002] Debris flow is a special flood that occurs due to precipitation in valleys or on slopes, carrying a large amount of solid substances such as sediment, stones, and boulders. It has the characteristics of suddenness, fast flow velocity, large flow rate, large material capacity, and strong destructive power. The occurrence of debris flow is usually accompanied by huge casualties and property losses. Therefore, accurate prediction of the occurrence of debris flow disasters can help relevant personnel take preventive and remedial measures in a timely manner.
[0003] In related technologies, hydrological and geological data related to debris flow disasters, such as precipitation, soil water content, soil looseness, etc., are usually analyzed to achieve the prediction of debris flow disasters. However, due to the large differences in hydrological and geological data in the same area at different times, the possibility of debris flow disasters occurring at different times is different, and thus the occurrence of debris flow disasters cannot be accurately predicted by existing methods. Summary of the Invention
[0004] In order to solve the technical problem that the occurrence of debris flow disasters cannot be accurately predicted by existing methods, the purpose of the present invention is to provide a hydrological data prediction method based on data analysis, and the specific technical solution adopted is as follows: The present invention provides a hydrological data prediction method based on data analysis, and the method includes: Obtain the precipitation, soil water content, and soil infiltration rate of each day in a preset historical time period in the area to be measured. Take the days corresponding to the occurrence of debris flow in the preset historical time period as disaster days, and take other historical days except the disaster days as normal days; Obtain the precipitation intensity of each day according to the precipitation of each day and the precipitation of each day in the preset time domain; obtain the soil severity of each day according to the soil water content, the soil infiltration rate, and the precipitation intensity of each day; Take any normal day as the target normal day, and obtain the precipitation feature similarity of each disaster day with respect to the target normal day according to the difference in the precipitation intensity between each disaster day and the target normal day; obtain the soil feature similarity of each disaster day with respect to the target normal day according to the difference in the soil severity between each disaster day and the target normal day; cluster all the disaster days according to the differences in the precipitation feature similarity and the soil feature similarity of each disaster day with respect to the target normal day, and obtain multiple clustering clusters of the target normal day; obtain the disaster likelihood of the target normal day according to the precipitation feature similarity and the soil feature similarity of each disaster day in each clustering cluster and the number of disaster days in each clustering cluster; based on the disaster likelihood, screen out similar normal days from all normal days; Based on the precipitation amount, the soil water content, and the soil infiltration rate of each similar normal day and disaster day, construct a disaster prediction model and predict the debris flow disaster of the current day.
[0005] Further, the obtaining of the precipitation intensity of each day includes: Take any day as the target day, use the precipitation amount of the target day as the numerator, use the sum of the precipitation amounts of all days in the preset time domain of the target day as the denominator, and use the ratio as the precipitation proportion of the target day; Use the number of days with precipitation amount greater than 0 in the preset time domain of the target day as the numerator, use the total number of days in the preset time domain of the target day as the denominator, and use the ratio as the precipitation day proportion of the target day; Use the average value of the precipitation amounts of all days in the preset time domain of the target day as the overall precipitation of the preset time domain of the target day, perform normalization processing on the difference between the precipitation amount of the target day and the overall precipitation amount, and obtain the precipitation relative difference value of the target day; After comprehensively processing the precipitation proportion, the precipitation day proportion, and the precipitation relative difference value of the target day and performing normalization processing, obtain the precipitation intensity of the target day.
[0006] Further, the obtaining of the soil severity of each day includes: After comprehensively processing the soil water content, the soil infiltration rate, and the precipitation intensity of the target day and performing normalization processing, obtain the soil severity of the target day.
[0007] Further, the obtaining of the precipitation feature similarity of each disaster day with respect to the target normal day includes: Perform negative-correlated normalization processing on the absolute value of the difference in precipitation intensity between each disaster day and the target normal day, and obtain the precipitation feature similarity of each disaster day with respect to the target normal day.
[0008] Further, the obtaining of the soil feature similarity of each disaster day with respect to the target normal day includes: Performing a negative correlation normalization on the absolute value of the difference in the soil severity between each disaster day and the target normal day to obtain the soil feature similarity of each disaster day with respect to the target normal day.
[0009] Further, the obtaining of multiple clustering clusters of the target normal day includes: Based on the precipitation feature similarity and the soil feature similarity of each disaster day with respect to the target normal day, calculating the Euclidean distance between any two disaster days as the distance metric between any two disaster days; Using the K-means clustering algorithm and based on the distance metric between any two disaster days, clustering all disaster days to obtain multiple clustering clusters of the target normal day.
[0010] Further, the obtaining of the disaster possibility of the target normal day includes: Among all the clustering clusters of the target normal day, selecting the reference clustering cluster of the target normal day, where the number of disaster days in the reference clustering cluster is the largest; Taking the average value of the precipitation feature similarities of all disaster days in the reference clustering cluster of the target normal day with respect to the target normal day as the first overall similarity of the target normal day, taking the average value of the soil feature similarities of all disaster days in the reference clustering cluster of the target normal day with respect to the target normal day as the second overall similarity of the target normal day, and comprehensively considering the first overall similarity, the second overall similarity of the target normal day, and the number of disaster days in the reference clustering cluster of the target normal day to obtain the disaster environment similarity of the target normal day; According to the difference between the precipitation intensity of the target normal day and the overall level of the precipitation intensities of all normal days, and the difference between the soil severity of the target normal day and the overall level of the soil severities of all normal days, obtaining the disaster risk coefficient of the target normal day; After comprehensively considering the disaster environment similarity and the disaster risk coefficient of the target normal day and performing normalization processing, obtaining the disaster possibility of the target normal day.
[0011] Further, the obtaining of the disaster risk coefficient of the target normal day includes: Taking the precipitation intensity of the target normal day as the numerator and the average value of the precipitation intensities of all normal days as the denominator, and taking the ratio as the first risk coefficient of the target normal day; Taking the soil severity of the target normal day as the numerator and the average value of the soil severities of all normal days as the denominator, and taking the ratio as the second risk coefficient of the target normal day; Integrate the first risk coefficient and the second risk coefficient of the target normal day to obtain the disaster risk coefficient of the target normal day.
[0012] Further, the screening of similar normal days from all normal days based on the disaster possibility includes: Take the normal days with a disaster possibility greater than the preset possibility threshold as similar normal days.
[0013] Further, the prediction of the debris flow disaster on the current day includes: Standardize the precipitation, soil water content, and soil infiltration rate of each similar normal day and each disaster day respectively to obtain the standard precipitation, standard soil water content, and standard soil infiltration rate of each similar normal day and each disaster day; Take any one of the similar normal days or any one of the disaster days as the test day, and take the vector composed of the standard precipitation, standard soil water content, and standard soil infiltration rate of the test day as the feature vector of the test day. If the test day is a similar normal day, set the class label of the feature vector of the test day to the value 0. If the test day is a disaster day, set the class label of the feature vector of the test day to the value 1, and form a set of feature vectors with the class labels of all similar normal days and all disaster days as the training set; Input the data in the training set into a support vector machine for training, and use the trained support vector machine as a disaster prediction model; Input the precipitation, soil water content, and soil infiltration rate of the current day into the disaster prediction model, and the disaster prediction model outputs the debris flow disaster prediction result of the current day.
[0014] The present invention has the following beneficial effects: Considering that the existing methods cannot accurately predict the occurrence of debris flow disasters, this invention first obtains the daily precipitation, soil water content, and soil infiltration rate in the area to be measured during a preset historical time period. The days corresponding to the occurrence of debris flow are regarded as disaster days, and the days without debris flow are regarded as normal days. Considering that different precipitation amounts and precipitation conditions within a local time period have different degrees of difficulty in triggering debris flow disasters, the precipitation intensity obtained can be used to reflect the possibility of triggering debris flow disasters based on the daily precipitation status. At the same time, considering that the state characteristics of the soil are also key factors in triggering debris flow disasters, the degree of soil deterioration obtained can be used to reflect the possibility of triggering debris flow disasters based on the daily soil status in the area to be measured, providing a data basis for subsequent analysis of the similarity of environmental characteristics between normal days and disaster days. Considering the interference of external factors, some normal days are close to the conditions for debris flow occurrence but no debris flow occurs. These normal days with similar soil and precipitation status to disaster days also need to be emphasized. Therefore, the similarity of precipitation characteristics between the target normal day and each disaster day is first reflected by the obtained precipitation feature similarity, and the similarity of soil status between the target normal day and each disaster day is reflected by the soil feature similarity. Then, the disaster days with similar precipitation feature similarity and soil feature similarity are grouped into the same clustering cluster. Further, the possibility of debris flow disasters occurring on the target normal day is reflected by the disaster possibility, and similar normal days close to the conditions for debris flow occurrence are selected. Using various data of each similar normal day and disaster day, an effective disaster prediction model is constructed to accurately predict the debris flow disaster on the current day. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a flowchart of a hydrological data prediction method based on data analysis provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details a method for predicting hydrological data based on data analysis proposed according to the present invention, including its specific implementation manner, structure, features, and effects, as follows. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0019] The following specifically describes the specific solution of a method for predicting hydrological data based on data analysis provided by the present invention in conjunction with the accompanying drawings.
[0020] Please refer to Figure 1 , which shows a flowchart of a method for predicting hydrological data based on data analysis provided by an embodiment of the present invention. The method includes: Step S1: Obtain the daily precipitation, soil moisture content, and soil infiltration rate in a preset historical time period for the area to be measured. The days corresponding to debris flow occurrences within the preset historical time period are regarded as disaster days, and the other historical days except the disaster days are regarded as normal days.
[0021] For areas with frequent debris flow occurrences, relevant surveyors usually measure hydrological and geological data related to debris flow disasters in the area, such as precipitation, soil moisture content, and soil infiltration rate, etc., and then record and store these data in the hydrological and geological database for subsequent data analysis.
[0022] Therefore, in the embodiment of the present invention, the daily precipitation, soil moisture content, and soil infiltration rate in a preset historical time period are first extracted from the hydrological and geological database of the area to be measured. Among them, the preset historical time period is set to 1 year, and the specific value of the preset historical time period can also be set by the implementer according to the specific real-time scenario and is not limited herein.
[0023] It should be noted that multiple debris flow disasters usually occur within the preset historical time period. Therefore, in the embodiment of the present invention, the days corresponding to debris flow occurrences within the preset historical time period are regarded as disaster days, and the other historical days except the disaster days are regarded as normal days. That is to say, no debris flow disasters occur on normal days. Subsequently, the similarity of environmental characteristics between normal days and disaster days can be analyzed, and various hydrological data of normal days and disaster days can be combined to accurately predict debris flow disasters.
[0024] Step S2: Obtain the precipitation intensity for each day based on the daily precipitation and the precipitation of each day in the preset time domain of each day; obtain the soil severity for each day based on the daily soil water content, soil infiltration rate, and precipitation intensity.
[0025] Different precipitation amounts have different possibilities of triggering debris flow disasters. Before each debris flow disaster occurs, the more precipitation and the longer the duration, the easier it is to trigger a debris flow. Or if the precipitation increases within a short period of time, it is more likely to trigger a debris flow. Therefore, in the embodiments of the present invention, the daily precipitation and the precipitation of each day in the preset time domain of each day are analyzed. By obtaining the precipitation intensity, the possibility of the daily precipitation state triggering a debris flow disaster is reflected. The greater the precipitation intensity of a certain day, the more likely the precipitation characteristics of that day are to trigger a debris flow disaster, providing a data basis for subsequent analysis of the similarity of environmental characteristics between normal days and disaster days. Among them, the length of the preset time domain of a certain day is set to 7, that is, the preset time domain of that day includes the six days closest to that day and that day itself.
[0026] Preferably, in an embodiment of the present invention, the method for obtaining the precipitation intensity for each day specifically includes: Take any day as the target day, use the precipitation of the target day as the numerator, and use the sum of the precipitation of all days in the preset time domain of the target day as the denominator. Take the ratio as the precipitation proportion of the target day. The greater the precipitation proportion, the greater the precipitation of the target day relative to the precipitation of other days in the preset time domain, and thus the more likely the precipitation of the target day is to trigger a debris flow disaster.
[0027] Take the number of days with precipitation greater than the value 0 in the preset time domain of the target day as the numerator, and take the total number of days in the preset time domain of the target day as the denominator. Take the ratio as the precipitation day proportion of the target day. Among them, if the precipitation of a certain day is greater than the value 0, it means that day is a rainy day. The greater the precipitation day proportion, the more frequent the precipitation in the preset time domain of the target day, and thus the more likely it is to trigger a debris flow disaster.
[0028] Take the average value of the precipitation of all days in the preset time domain of the target day as the overall precipitation of the preset time domain of the target day. Normalize the difference between the precipitation of the target day and the overall precipitation, and limit the calculation result within the range, so as to obtain the relative precipitation difference value of the target day. The greater the relative precipitation difference value, the greater the precipitation of the target day relative to the overall precipitation in the preset time domain, and thus the more likely it is to trigger a debris flow disaster.
[0029] After comprehensively processing the precipitation proportion, precipitation day proportion, and relative precipitation difference value of the target day and performing normalization processing, limit the calculation result within the range, so as to obtain the precipitation intensity of the target day.
[0030] In an embodiment of the present invention, the sum or product value of the precipitation proportion, precipitation days proportion, and precipitation relative difference value of the target day can be calculated to achieve the integration of the three, and no limitation is made here.
[0031] In an embodiment of the present invention, the normalization process can be specifically, for example, the maximum-minimum normalization process, and the normalization in subsequent steps can all adopt the maximum-minimum normalization process. In other embodiments of the present invention, other normalization methods can be selected according to the specific range of values, and this will not be elaborated here.
[0032] As an example, in an embodiment of the present invention, the expression of the precipitation intensity of the target day can be specifically, for example: Among them, represents the precipitation intensity of the target day; represents the precipitation of the target day; represents the sum value of the precipitation of all days in the preset time domain of the target day; represents the precipitation proportion of the target day. Since the embodiments of the present invention are for predicting areas with frequent debris flows, the precipitation in such areas is relatively frequent, so ; represents the number of days when the precipitation in the preset time domain of the target day is greater than the value 0; represents all days in the preset time domain of the target day, that is, the length of the preset time domain. In an embodiment of the present invention ; represents the overall precipitation of the preset time domain of the target day; represents the activation function, which is used for the normalization process; represents the normalization function.
[0033] The daily precipitation intensity can be obtained by the same method as above. During the precipitation process, a large amount of precipitation can quickly make the soil reach a saturated state. Excessive water content leads to soil saturation, reduces the soil's water resistance, accelerates the fluidity and erosion of the soil, and increases the risk of debris flow disasters. The soil infiltration rate determines the penetration rate and distribution uniformity of precipitation water in the soil, and affects the change rate of soil water content and soil saturation. When the soil reaches a saturated state, runoff will begin to form on the surface. When the rainfall continues to increase, it is more likely to cause debris flow disasters. Therefore, the embodiment of the present invention analyzes the daily soil water content and soil infiltration rate. At the same time, the greater the precipitation intensity on a certain day, it means that the precipitation on that day will cause the soil state to be more likely to cause debris flow disasters. Therefore, the soil severity obtained in combination with the daily precipitation intensity can reflect the possibility of the soil state of the tested area on each day causing debris flow disasters. The greater the soil severity on a certain day, the more likely the soil state on that day is to cause debris flow disasters, providing a data basis for subsequent analysis of the similarity of environmental characteristics between normal days and disaster days.
[0034] Preferably, in one embodiment of the present invention, the method for obtaining the soil severity every day specifically includes: The soil moisture content, soil infiltration rate and precipitation intensity of the target day are integrated and normalized, and the calculation results are limited to range, thereby obtaining the soil severity on the target day.
[0035] In the embodiment of the present invention, the integration of soil moisture content, soil infiltration rate and precipitation intensity on the target day can be achieved by calculating the sum or product of the three, which is not limited here.
[0036] As an example, in one embodiment of the present invention, the expression of soil severity on the target day may be specifically, for example, as follows: in, Indicates the severity of the soil on the target day; Indicates the soil moisture content on the target day; represents the soil infiltration rate on the target day; Indicates the precipitation intensity of the target day; Represents the normalization function.
[0037] The daily soil severity can be obtained by the same method as above. At this point, the daily precipitation intensity and soil severity are obtained.
[0038] Step S3: Take any normal day as the target normal day. According to the difference in precipitation intensity between each disaster day and the target normal day, obtain the precipitation feature similarity of each disaster day with respect to the target normal day; according to the difference in soil severity between each disaster day and the target normal day, obtain the soil feature similarity of each disaster day with respect to the target normal day; according to the differences in precipitation feature similarity and soil feature similarity of each disaster day with respect to the target normal day among all disaster days, cluster all disaster days to obtain multiple clustering clusters of the target normal day; according to the precipitation feature similarity and soil feature similarity of each disaster day in each clustering cluster with respect to the target normal day, and the number of disaster days in each clustering cluster, obtain the disaster likelihood of the target normal day; based on the disaster likelihood, screen out similar normal days from all normal days.
[0039] Since the occurrence of debris flow disasters is affected by various factors, under the interference of other external factors not considered, some normal days are close to the conditions for the occurrence of debris flow but no debris flow occurs, but are on the verge of debris flow disasters, with a great possibility of debris flow disasters. These normal days with similar soil and precipitation states to disaster days also need to be taken seriously. Therefore, in the embodiment of the present invention, first, any normal day is analyzed, any normal day is taken as the target normal day, and according to the difference in precipitation intensity between each disaster day and the target normal day, the precipitation feature similarity of each disaster day with respect to the target normal day is obtained. The precipitation feature similarity reflects the similarity of precipitation features between the target normal day and each disaster day, providing a data basis for subsequent clustering analysis.
[0040] Preferably, in an embodiment of the present invention, the method for obtaining the precipitation feature similarity of each disaster day with respect to the target normal day specifically includes: The smaller the difference in precipitation intensity between the disaster day and the target normal day, the more similar the precipitation features of the disaster day and the target normal day. Therefore, the absolute value of the difference in precipitation intensity between each disaster day and the target normal day can be subjected to negative-correlated normalization processing to obtain the precipitation feature similarity of each disaster day with respect to the target normal day.
[0041] In an embodiment of the present invention, a negative exponential function with the natural constant as the base or a function form of can be used to achieve negative-correlated normalization processing, where represents the normalization function.
[0042] As an example, in an embodiment of the present invention, the expression of the precipitation feature similarity of each disaster day with respect to the target normal day can be specifically, for example: where represents the The similarity of precipitation characteristics of each disaster day to those of the target normal day; Denote the precipitation intensity of the n-th disaster day; Denote the precipitation intensity of the target normal day;
[0043] Similarly, it is also necessary to analyze the differences in soil severity between each disaster day and the target normal day, and reflect the similarity of soil conditions between the target normal day and each disaster day through the obtained soil feature similarity, providing a data basis for subsequent clustering analysis.
[0044] Preferably, in an embodiment of the present invention, the method for obtaining the soil feature similarity of each disaster day to the target normal day specifically includes: The smaller the difference in soil severity between a disaster day and the target normal day, the more similar the soil conditions of the disaster day and the target normal day. Therefore, the absolute value of the difference in soil severity between each disaster day and the target normal day can be subjected to a negatively correlated normalization process to obtain the soil feature similarity of each disaster day to the target normal day.
[0045] As an example, in an embodiment of the present invention, the expression of the soil feature similarity of each disaster day to the target normal day can be specifically, for example: Wherein, Denote the soil feature similarity of the n-th disaster day to the target normal day; Denote the soil severity of the n-th disaster day; Denote the
[0046] The above-obtained precipitation feature similarity and soil feature similarity reflect the similarity of precipitation and soil characteristics between each disaster day and the target day. Therefore, all disaster days can be further clustered according to the differences in precipitation feature similarity and soil feature similarity of each disaster day to the target normal day, so as to divide the disaster days with similar precipitation feature similarity and soil feature similarity into the same clustering cluster. Subsequently, based on the differences in precipitation feature similarity and soil feature similarity of each disaster day in the clustering cluster, the possibility of debris flow disasters occurring on the target normal day can be accurately calculated and analyzed.
[0047] Preferably, in one embodiment of the present invention, the method for obtaining multiple clustering clusters of the target normal day specifically includes: First, based on the precipitation feature similarity and soil feature similarity of each disaster day with respect to the target normal day, calculate the Euclidean distance between any two disaster days as the distance metric between any two disaster days.
[0048] As an example, in one embodiment of the present invention, the expression of the distance metric between any two disaster days can be specifically, for example: Wherein, represents the distance metric between the th disaster day and the th disaster day; represents the precipitation feature similarity of the th disaster day; represents the precipitation feature similarity of the th disaster day; represents the soil feature similarity of the th disaster day; represents the soil feature similarity of the th disaster day.
[0049] Then, use the K-means clustering algorithm and, based on the distance metric between any two disaster days, cluster all disaster days to obtain multiple clustering clusters of the target normal day. In other embodiments of the present invention, other clustering algorithms such as the DBSCAN clustering algorithm can also be used for clustering processing, which is not limited herein.
[0050] After obtaining multiple clustering clusters of the target normal day, the degree of similarity of the precipitation state and soil state between each disaster day in the same clustering cluster and the target normal day is relatively close, and the degree of this similarity feature of different clustering clusters is different. The clustering cluster containing more disaster days can more accurately reflect the relationship between the target normal day and the disaster days. Therefore, the disaster possibility of the target normal day can be obtained according to the precipitation feature similarity and soil feature similarity of each disaster day in each clustering cluster with respect to the target normal day and the number of disaster days in each clustering cluster. Subsequently, normal days with precipitation and soil environments similar to those of the disaster days can be screened based on the disaster possibility, improving the accuracy of the subsequent construction of the disaster prediction model.
[0051] Preferably, in one embodiment of the present invention, the method for obtaining the disaster possibility of the target normal day specifically includes: First, among all the clustering clusters of the target normal day, select the reference clustering cluster of the target normal day. Among them, the number of disaster days in the reference clustering cluster is the largest. Since the number of disaster days in the reference clustering cluster is the largest, its representativeness is the strongest, which improves the accuracy of subsequent environmental similarity analysis between the target normal day and disaster days.
[0052] Take the average value of the precipitation feature similarities of all disaster days in the reference clustering cluster of the target normal day with respect to the target normal day as the first overall similarity of the target normal day. The larger the first overall similarity, the more similar the precipitation features between the target normal day and each disaster day. Take the average value of the soil feature similarities of all disaster days in the reference clustering cluster of the target normal day with respect to the target normal day as the second overall similarity of the target normal day. The larger the second overall similarity, the more similar the soil features between the target normal day and each disaster day. Synthesize the first overall similarity, the second overall similarity of the target normal day, and the number of disaster days in the reference clustering cluster of the target normal day to obtain the disaster environment similarity of the target normal day. The larger the disaster environment similarity, the more similar the precipitation and soil features of the target normal day are to those of disaster days, and further indicates that the target normal day is more likely to occur a debris flow disaster.
[0053] In an embodiment of the present invention, the sum value or product value of the first overall similarity, the second overall similarity of the target normal day, and the number of disaster days in the reference clustering cluster of the target normal day can be used as the disaster environment similarity of the target normal day to achieve the synthesis of the three, and no limitation is made here.
[0054] As an example, in an embodiment of the present invention, the expression formula of the disaster environment similarity of the target normal day can be specifically, for example: Among them, represents the disaster environment similarity of the target normal day; represents the first overall similarity of the target normal day; represents the second overall similarity of the target normal day; represents the number of disaster days in the reference clustering cluster of the target normal day.
[0055] Then, the greater the precipitation intensity and soil severity of the target normal day relative to those of other normal days, the more likely the target normal day is to occur a debris flow disaster. Therefore, the disaster risk coefficient of the target normal day can be obtained based on the difference between the precipitation intensity of the target normal day and the overall level of precipitation intensity of all normal days, and the difference between the soil severity of the target normal day and the overall level of soil severity of all normal days. The larger the disaster risk coefficient, the more likely the target normal day is to occur a debris flow disaster.
[0056] Preferably, in an embodiment of the present invention, the method for obtaining the disaster risk coefficient of the target normal day specifically includes: Taking the precipitation intensity of the target normal day as the numerator, taking the average value of the precipitation intensities of all normal days as the denominator, and taking the ratio as the first risk coefficient of the target normal day. The larger the first risk coefficient, the greater the precipitation intensity of the target normal day relative to the precipitation intensities of each normal day, and thus the more likely the target normal day is to have a debris flow disaster risk.
[0057] Taking the soil deterioration degree of the target normal day as the numerator, taking the average value of the soil deterioration degrees of all normal days as the denominator, and taking the ratio as the second risk coefficient of the target normal day. The larger the second risk coefficient, the greater the soil deterioration degree of the target normal day relative to the soil deterioration degrees of each normal day, and thus the more likely the target normal day is to have a debris flow disaster risk.
[0058] Therefore, the first risk coefficient and the second risk coefficient of the target normal day can be combined to obtain the disaster risk coefficient of the target normal day.
[0059] In an embodiment of the present invention, the sum value or product value of the first risk coefficient and the second risk coefficient of the target normal day can be used as the disaster risk coefficient of the target normal day to achieve the combination of the two, and no limitation is made here.
[0060] As an example, in an embodiment of the present invention, the expression of the disaster risk coefficient of the target normal day can be specifically, for example: Wherein, represents the disaster risk coefficient of the target normal day; represents the precipitation intensity of the target normal day; represents the average value of the precipitation intensities of all normal days; represents the first risk coefficient of the target normal day; represents the soil deterioration degree of the target normal day; represents the average value of the soil deterioration degrees of all normal days; represents the second risk coefficient of the target normal day.
[0061] Furthermore, after combining and normalizing the disaster environment similarity and the disaster risk coefficient of the target normal day, the calculation result is limited within the range to obtain the disaster possibility of the target normal day.
[0062] In an embodiment of the present invention, the combination of the two can be achieved by calculating the sum value or product value of the disaster environment similarity and the disaster risk coefficient of the target normal day, and no limitation is made here.
[0063] As an example, in an embodiment of the present invention, the expression of the disaster possibility of the target normal day can be specifically, for example: Wherein, represents the disaster possibility of the target normal day; represents the similarity of the disaster environment of the target normal day; represents the disaster risk coefficient of the target normal day; represents a normalization function.
[0064] By the same method as above, the disaster possibility of each normal day can be obtained. The greater the disaster possibility, the more likely it is that a debris flow disaster will occur on the normal day, and the closer it is to the edge of the debris flow disaster risk. Therefore, based on the disaster possibility, normal days that are similar to the precipitation and soil environment of the disaster day and are extremely likely to occur debris flow disasters can be screened out from all normal days, that is, similar normal days. Subsequently, based on various hydrogeological data of the similar normal days and the disaster day, a more accurate disaster prediction model can be constructed to improve the accuracy of debris flow disaster prediction.
[0065] Preferably, in an embodiment of the present invention, normal days with a disaster possibility greater than a preset possibility threshold are used as similar normal days, wherein the preset possibility threshold is set to 0.8, and the specific value of the preset possibility threshold can also be set by the implementer according to the specific implementation scenario and is not limited herein.
[0066] Thus, the similar normal days are screened out.
[0067] Step S4: Based on the precipitation, soil water content, and soil infiltration rate of each similar normal day and the disaster day, construct a disaster prediction model and predict the debris flow disaster on the current day.
[0068] Since there is a great possibility that a debris flow disaster will occur on the similar normal days, the hydrogeological data of the similar normal days should also receive more attention to ensure the accuracy of debris flow disaster prediction. Therefore, based on the precipitation, soil water content, and soil infiltration rate of each similar normal day and the disaster day, a disaster prediction model is constructed and the debris flow disaster on the current day is predicted.
[0069] Preferably, in an embodiment of the present invention, the method for predicting the debris flow disaster on the current day specifically includes: First, in order to train the model on the same scale and improve the final model accuracy, it is necessary to standardize the precipitation, soil water content, and soil infiltration rate of each similar normal day and each disaster day respectively, eliminate the influence of dimensions, and obtain the standard precipitation, standard soil water content, and standard soil infiltration rate of each similar normal day and each disaster day. Among them, data standardization is a technical means well-known to those skilled in the art and will not be elaborated here.
[0070] Then, take any similar normal day or any disaster day as the day to be measured, and take the vector composed of the standard precipitation, standard soil water content, and standard soil infiltration rate of the day to be measured as the feature vector of the day to be measured. If the day to be measured is a similar normal day, set the class label of the feature vector of the day to be measured to the value 0. If the day to be measured is a disaster day, set the class label of the feature vector of the day to be measured to the value 1. Through the above same method, the feature vectors and corresponding class labels of each similar normal day and each disaster day can be obtained. Furthermore, the set composed of the feature vectors with class labels of all similar normal days and all disaster days can be used as the training set.
[0071] Finally, input the data in the training set into the support vector machine for training, and use the trained support vector machine as the disaster prediction model. Then, input the precipitation, soil water content, and soil infiltration rate of the current day into the disaster prediction model, and the disaster prediction model outputs the debris flow disaster prediction result of the current day. Among them, the support vector machine is a technical means well-known to those skilled in the art and will not be elaborated here.
[0072] It should be noted that the above sequence of embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0073] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.
Claims
1. A hydrological data prediction method based on data analysis, characterized in that: The method comprises: Obtain the daily precipitation, soil moisture content and soil infiltration rate of the area to be tested within a preset historical time period, and take the days corresponding to the occurrence of debris flow within the preset historical time period as disaster days, and take the other historical days except disaster days as normal days; According to the daily precipitation and the precipitation of each day in the preset time domain, the daily precipitation intensity is obtained; according to the daily soil moisture content, the soil infiltration rate and the precipitation intensity, the daily soil severity is obtained; Take any normal day as the target normal day, and obtain the precipitation feature similarity of each disaster day with respect to the target normal day according to the difference in precipitation intensity between each disaster day and the target normal day; obtain the soil feature similarity of each disaster day with respect to the target normal day according to the difference in soil severity between each disaster day and the target normal day; cluster all disaster days according to the difference in precipitation feature similarity and soil feature similarity between each disaster day with respect to the target normal day, and obtain multiple clusters of the target normal day; obtain the disaster possibility of the target normal day according to the precipitation feature similarity and soil feature similarity of each disaster day in each cluster and the number of disaster days in each cluster; and screen out similar normal days from all normal days based on the disaster possibility; Based on the precipitation, soil moisture content and soil infiltration rate of similar normal days and disaster days, a disaster prediction model is constructed, and the debris flow disaster of the current day is predicted.
2. A hydrological data prediction method based on data analysis according to claim 1, characterized in that: The daily precipitation intensity is obtained by: Take any day as the target day, take the precipitation of the target day as the numerator, take the sum of the precipitation of all days in the preset time domain of the target day as the denominator, and take the ratio as the precipitation ratio of the target day; The number of days in the preset time domain of the target day with precipitation greater than the value 0 is used as the numerator, all days in the preset time domain of the target day is used as the denominator, and the ratio is used as the ratio of precipitation days on the target day; The average of the precipitation of all days in the preset time domain of the target day is taken as the overall precipitation of the preset time domain of the target day, and the difference between the precipitation of the target day and the overall precipitation is normalized to obtain a relative difference value of precipitation of the target day; The precipitation ratio of the target day, the ratio of the number of precipitation days and the relative difference value of precipitation are integrated and normalized to obtain the precipitation intensity of the target day.
3. A hydrological data prediction method based on data analysis according to claim 2, characterized in that: The soil severity obtained each day includes: The soil moisture content, the soil infiltration rate and the precipitation intensity of the target day are integrated and normalized to obtain the soil severity of the target day.
4. The hydrological data prediction method based on data analysis according to claim 1 is characterized in that: The obtaining of the similarity of precipitation characteristics of each disaster day with respect to the target normal day includes: The absolute value of the difference in precipitation intensity between each disaster day and the target normal day is subjected to negative correlation normalization processing to obtain the precipitation feature similarity of each disaster day with respect to the target normal day.
5. The hydrological data prediction method based on data analysis according to claim 1 is characterized in that: The obtaining of the soil characteristic similarity of each disaster day with respect to the target normal day comprises: The absolute value of the difference in soil severity between each disaster day and the target normal day is normalized by negative correlation to obtain the soil characteristic similarity of each disaster day with respect to the target normal day.
6. The hydrological data prediction method based on data analysis according to claim 1 is characterized in that: The method of obtaining a plurality of clusters of target normal days comprises: Based on the precipitation feature similarity and the soil feature similarity of each disaster day with respect to the target normal day, calculating the Euclidean distance between any two disaster days as a distance measure between any two disaster days; Using the K-means clustering algorithm and based on the distance metric between any two disaster days, all disaster days are clustered to obtain multiple clusters of target normal days.
7. The hydrological data prediction method based on data analysis according to claim 1 is characterized in that: The disaster probability of obtaining a target normal day includes: Among all clusters of target normal days, a reference cluster of target normal days is selected, wherein the number of disaster days in the reference cluster is the largest; The average value of the precipitation feature similarities of all disaster days in the reference cluster of the target normal day with respect to the target normal day is taken as the first overall similarity of the target normal day, the average value of the soil feature similarities of all disaster days in the reference cluster of the target normal day with respect to the target normal day is taken as the second overall similarity of the target normal day, and the first overall similarity of the target normal day, the second overall similarity and the number of disaster days in the reference cluster of the target normal day are integrated to obtain the disaster environment similarity of the target normal day; Obtaining a disaster risk coefficient for a target normal day according to a difference between the precipitation intensity of the target normal day and the overall level of the precipitation intensity of all normal days, and a difference between the soil severity of the target normal day and the overall level of the soil severity of all normal days; The disaster environment similarity and the disaster risk coefficient of the target normal day are integrated and normalized to obtain the disaster possibility of the target normal day.
8. The hydrological data prediction method based on data analysis according to claim 7 is characterized in that: The disaster risk coefficient of the target normal day is obtained by: The precipitation intensity of the target normal day is used as the numerator, the average of the precipitation intensities of all normal days is used as the denominator, and the ratio is used as the first risk coefficient of the target normal day; The soil severity of the target normal day is used as the numerator, the average value of the soil severity of all normal days is used as the denominator, and the ratio is used as the second risk coefficient of the target normal day; The first risk coefficient and the second risk coefficient of a target normal day are integrated to obtain a disaster risk coefficient of the target normal day.
9. The hydrological data prediction method based on data analysis according to claim 1 is characterized in that: Based on the disaster possibility, similar normal days are selected from all normal days, including: Normal days with disaster probability greater than the preset probability threshold are regarded as similar normal days.
10. The hydrological data prediction method based on data analysis according to claim 1, characterized in that: The prediction of the debris flow disaster on the current day includes: Standardizing the precipitation, soil moisture content, and soil infiltration rate of each similar normal day and each disaster day, respectively, to obtain the standard precipitation, standard soil moisture content, and standard soil infiltration rate of each similar normal day and each disaster day; Any similar normal day or any disaster day is taken as the day to be tested, and the vector composed of the standard precipitation, the standard soil moisture content and the standard soil infiltration rate of the day to be tested is taken as the feature vector of the day to be tested; if the day to be tested is a similar normal day, the class label of the feature vector of the day to be tested is set to a value of 0; if the day to be tested is a disaster day, the class label of the feature vector of the day to be tested is set to a value of 1; and a set composed of the feature vectors with the class labels of all similar normal days and all disaster days is taken as a training set; Inputting the data in the training set into a support vector machine for training, and using the trained support vector machine as a disaster prediction model; The current day's precipitation, soil moisture content and soil infiltration rate are input into the disaster prediction model, and the disaster prediction model outputs the debris flow disaster prediction result for the current day.
Citation Information
Patent Citations
Water-soil coupling type disaster fluid occurrence probability prediction model construction method and debris flow forecasting method
CN114357777A
Artificial intelligence and Bayesian theory coupled flood forecasting method and system
CN116663404A
Regional geological disaster early warning method and device, computer equipment and storage medium
CN117079425A
Data acquisition method and system for laboratory informatization system
CN118965025A
Ammeter fault prediction method, system and equipment based on deep learning
CN119720028A
Cited By
Loess slope deformation monitoring data processing method based on MEMS inertial sensor data and GNSS data
CN120849978A