Mountain fire risk distribution space-time evaluation method based on data full-process processing
Through resampling and decision tree testing, the wildfire factor weight is given, the factor importance and correlation confidence are analyzed, and the factor selection is optimized, which solves the redundant noise and non-balance problems of the existing wildfire risk assessment model, and improves the evaluation accuracy and performance.
Patent Information
- Application Number
- CN202510321199.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The existing wildfire risk assessment model has redundant feature noise interference in feature selection and data processing, ignoring the unbalance and correlation of factor values, resulting in a degradation of evaluation performance.
The data set is divided by resampling technology, the wildfire factor category weight is assigned, the decision tree is established and the out-of-bag error rate is tested, the factor importance and correlation confidence are analyzed, the factor selection is optimized, and the naive Bayesian network model is trained.
It improves the accuracy and performance of the wildfire risk assessment model, reduces the noise interference of redundant factors, and enhances the ability to evaluate extreme environmental factors.
Smart Images

Figure CN120278510A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of forest fire prevention and control, and specifically to a spatio-temporal assessment method for wildfire risk distribution based on the full-process data processing. Background Art
[0002] The transmission line is an important part of the power system. The transmission network is laid across regions, and the lines and other equipment are exposed in the wild for a long time. Affected by factors such as the continuous natural environment and material aging, the risk of wildfires caused by transmission line failures is relatively high. Fires caused by power lines often result in large-scale safety risks and ecological damage. In addition, under the influence of wildfires, the insulation strength of the air gap of the transmission line decreases, which is extremely likely to cause the line to short-circuit and break down, induce the tripping of the transmission line, and the success rate of reclosing is low, seriously endangering the safe and stable operation of the power grid. There are differences in the distribution of wildfire risk factors such as geographical factors, vegetation types, and climate conditions in the areas where different lines are located, and the risks of wildfires caused by transmission line failures are different. Establishing a wildfire risk assessment model and guiding the maintenance personnel to carry out differentiated wildfire prevention work are important means to prevent wildfire hazards.
[0003] There are many wildfire risk factors in the original wildfire sample data set. The risks of wildfires caused by different factors are different, and they will show certain differences with the different regions where they are located. Too many redundant factors will bring unnecessary noise, increase the complexity of the model, and affect the assessment performance of the model for wildfire risks. At present, the method of selecting wildfire sample data risk factors ignores the imbalance of factor values in the training set and the correlation between different factors, resulting in the mis-elimination of important wildfire risk factors and the decline of the wildfire risk assessment performance of the model.
[0004] Patent CN109447331A discloses a wildfire risk prediction method based on the stacking algorithm, which realizes the secondary processing and generation of features through the stacking method, and improves the overall effect of wildfire risk prediction. However, it ignores the noise interference caused by redundant features in the feature set to model training, resulting in the decline of the prediction accuracy of the wildfire risk prediction method. Patent CN109829583A discloses a wildfire risk prediction method based on probabilistic programming technology, which uses deep learning and probabilistic programming technology to predict the risk of wildfires. Due to its insufficient analysis of the correlation between wildfire risks caused by different types of data, the wildfire risk prediction model established by this method has poor performance in predicting the risks of wildfires caused by multiple factors. Summary of the Invention
[0005] In view of the above, it is necessary to provide a spatio-temporal assessment method for wildfire risk distribution based on the full-process data processing to solve the above problems.
[0006] An embodiment of the present application provides a spatio-temporal assessment method for wildfire risk distribution based on full-process data processing. The method includes:
[0007] Obtain all wildfire factors and fire point sequences of each sample point to form a wildfire sample data set;
[0008] Divide the wildfire sample data set into a resampled data set and an out-of-bag data set through resampling technology; for a resampled data set, determine all subsets of each wildfire factor; analyze the length characteristics of the fire point sequence of each wildfire factor, and combine the proportion of the number of values of the corresponding subset in the resampled data set to obtain the class weight of each wildfire factor in each resampled data set;
[0009] Establish decision trees based on a set number of wildfire factors in each resampled data set, test the decision trees based on the out-of-bag data set corresponding to the resampled data set, and obtain the out-of-bag error rate of each decision tree; analyze the distribution difference characteristics of the out-of-bag error rates of each wildfire factor and the remaining wildfire factors on all decision trees, and combine the corresponding class weights to obtain the factor importance of each wildfire factor;
[0010] Preset a training set; according to the numerical difference characteristics and distance characteristics between pairwise combined sample points of any two wildfire factors in the training set, combine the factor importance of each wildfire factor and the conditional probability obtained by the Bayesian network model to obtain the association confidence of the any two wildfire factors;
[0011] Train a wildfire risk assessment model according to the numerical characteristics of the association confidence between each wildfire factor and the remaining wildfire factors to obtain the wildfire risk probability of the sample point.
[0012] Among them, all subsets of each wildfire factor are determined based on the 3σ principle.
[0013] Among them, the formula for the class weight of each wildfire factor in each resampled data set is specifically: Denote the class weight of the kth wildfire factor in the nth resampled data set as w k,n , and its formula form is: Among them, T represents the number of subsets of the wildfire factor in the resampled data set; t represents the subset serial number; represents the mean length of the fire point sequences of all sample points in the tth subset of the kth wildfire factor in the nth resampled data set; r t represents the proportion of the number of values in the tth subset of the kth wildfire factor in the nth resampled data set; ε is a preset value.
[0014] Among them, when establishing a decision tree, a set number of wildfire factors are used as decision tree features.
[0015] Among them, the specific formula for obtaining the factor importance of each wildfire factor is as follows: Among them, s k represents the factor importance of the k-th wildfire factor; w k,n represents the class weight of the k-th wildfire factor in the n-th resampled dataset; N k represents the number of resampled datasets with the k-th wildfire factor as the decision tree feature; exp() represents the exponential function with the natural constant as the base; μ k,n represents the mean of the out-of-bag error rates of the remaining wildfire factors except the k-th wildfire factor in the n-th decision tree; er k,n represents the out-of-bag error rate of the k-th wildfire factor in the n-th decision tree.
[0016] Among them, the out-of-bag error rate is specifically the mean of the out-of-bag error rates of all decision trees with each wildfire factor as the decision tree feature.
[0017] Among them, the process of the association confidence of any two wildfire factors is specifically as follows:
[0018] Compare the difference features between the vectors composed of different sample points of any two wildfire factors, and combine the distance features between the sample points to obtain the association feature value between the any two wildfire factors at the sample points;
[0019] Denote the association confidence of the i-th and j-th wildfire factors as CI ij , and its formula form is: Among them, represents the average value of the association feature values between all sample points of the i-th and j-th wildfire factors in the training set; s i and s j respectively represent the factor importance of the i-th and j-th wildfire factors; P i and P j respectively represent the conditional probabilities of the i-th and j-th wildfire factors; norm[] represents the normalization function.
[0020] Among them, the process of obtaining the association feature value between any two wildfire factors at the sample points is specifically as follows:
[0021] Denote the vector composed of any two wildfire factors at a sample point as the factor vector;
[0022] Obtain the positive correlation mapping result of the distance metric between two sample points, and combine it with the distance metric between the factor vectors of the any two wildfire factors among the two sample points to obtain the association feature value between the any two wildfire factors at the sample points.
[0023] Among them, the training wildfire risk assessment model includes:
[0024] Calculate the sum of the cumulative correlation confidence levels between each wildfire factor and the remaining wildfire factors, take the wildfire factor with the smallest sum as the object to be excluded, and use the wildfire sample data set of all the remaining wildfire factors as the input of the Naive Bayes network model. Delete the wildfire factor with the smallest sum of correlation confidence levels one by one until the model accuracy no longer improves.
[0025] Among them, the model accuracy is determined by the confusion matrix.
[0026] This application has at least the following beneficial effects:
[0027] This application first assigns different category weights to multiple wildfire factors to reduce the impact of data imbalance on analyzing the wildfire risks caused by different factors and improve the accuracy of using wildfire factors during the training of the wildfire risk assessment model. At the same time, use the out-of-bag data set to test the decision trees established for each resampled data set to obtain the factor importance, enhance the optimization effect of different wildfire factors, and reduce the noise introduced by redundant wildfire factors during the subsequent training process of the wildfire risk assessment model. Further, by analyzing the correlation characteristics between different wildfire factors, obtain the correlation confidence level, enhance the assessment accuracy of the model for the wildfire risks caused by wildfire factors, and avoid mis-excluding wildfire factors. Finally, evaluate the training results of the wildfire risk assessment model and optimize the factor selection to improve the performance of the trained wildfire risk assessment model. Description of the Drawings
[0028] Figure 1 It is a flowchart of the wildfire risk distribution spatio-temporal assessment method based on the full-process data processing provided by this application;
[0029] Figure 2 It is the acquisition process of the wildfire risk probability provided by this application. Detailed Embodiments
[0030] In the description of the embodiments of this application, words such as "exemplary", "or", "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary", "or", "for example" aims to present relevant concepts in a specific way.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0032] In addition, it should be noted that the terms "first" and "second" in this application and its accompanying drawings are used to distinguish similar objects, rather than to describe a specific order or sequence. For the methods disclosed in the embodiments of this application or shown in the flowcharts, which include one or more steps for implementing the methods, without departing from the scope of protection of this application, the execution order of multiple steps can be interchanged with each other, and some steps can also be deleted.
[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.
[0034] This application proposes a spatio-temporal assessment method for wildfire risk distribution based on the whole-process data processing, which is applied to the technical field of forest fire prevention and control. Referring to the attached Figure 1 , the method includes the following steps:
[0035] S1: Obtain all wildfire factors and fire point sequences of each sample point to form a wildfire sample data set.
[0036] Considering the complex influence of the environment around the transmission line on the wildfire risk level, this application enhances the reliability of training the wildfire risk assessment model by obtaining multi-source wildfire risk data and establishing a wildfire sample data set including three dimensions of geographical feature data, vegetation type factors, and meteorological condition factors.
[0037] First, collect the wildfire data of the area where the transmission line is located for one year from the meteorological department and the power department. The wildfire data includes: geographical features, vegetation types, and meteorological condition data. In this embodiment, box plots are used to detect and remove the outliers in the data set, and at the same time, the missing values are filled. It should be noted that the geographical feature data in this application includes three wildfire factors: altitude, slope, and aspect; the vegetation type data includes two wildfire factors: combustible risk level and vegetation index; the meteorological condition data includes two wildfire factors: annual precipitation and average annual temperature. For the convenience of subsequent analysis, all the values of each factor are normalized. In this embodiment, the maximum-minimum normalization method is used.
[0038] Secondly, divide the area where the transmission line is located into grids according to a preset size, and each grid is recorded as a sample point; in this embodiment, the preset size is specifically 1 km × 1 km, and the implementer can adjust it according to the actual situation. Considering that the environmental factors between regions change little in a short period of time, for the spatio-temporal assessment of wildfire risk, arrange the occurrence times of the fire points in each grid in chronological order throughout the year to obtain the fire point sequence of each grid, which reflects the fire point distribution at different times. The lengths of the fire point sequences of different sample points are different. The more historical fire points, the longer the length, and the length of the fire point sequence of non-fire points is 0.
[0039] Finally, the grid where the fire point is located is used as the fire point sample point, and other grids are used as non-fire point sample points to obtain the fire point distribution in different spaces. The wildfire factors are extracted to the corresponding sample points using geographic information software, and the geographic feature data, vegetation type data, meteorological condition data of the sample points and the fire point sequence of the sample points are combined in a fixed order to form the wildfire sample data of the sample points. Further, the sample data of all sample points in the area where the transmission line is located constitute the original wildfire sample data set. It should be noted that in this embodiment, the arrangement of the wildfire factors in a fixed order is specifically altitude, slope, aspect, combustible risk level, vegetation index, annual precipitation, and annual average temperature, and the implementer can set the arrangement order of the wildfire factors according to the actual situation.
[0040] S2: Divide the wildfire sample data set into a resampled data set and an out-of-bag data set through resampling technology; for a resampled data set, determine all subsets of each wildfire factor; analyze the length characteristics of the fire point sequence of each wildfire factor, and combine the proportion of the number of values of the corresponding subset in the resampled data set to obtain the class weight of each wildfire factor in each resampled data set.
[0041] There are many wildfire factors in the original wildfire sample data set, and redundant wildfire factors will introduce noise in the training process of the wildfire risk assessment model, thereby affecting the performance of the model.
[0042] This application uses the random forest algorithm for wildfire factor optimization. The wildfire sample data set is divided into a resampled data set and an out-of-bag data set through the Bootstrap resampling technology. In this embodiment, the number of resampling times is set to 100 times, and the implementer can adjust it by himself. So far, 100 resampled data sets and their corresponding out-of-bag data sets can be obtained.
[0043] Considering that the risk of wildfires caused by extreme situations of some environmental factors is relatively high, but the occurrence frequency is relatively low, which makes the data set have strong imbalance, thereby affecting the reliability of wildfire factor optimization.
[0044] This application assigns class weights to different wildfire factors to reduce the impact of data imbalance on wildfire factor optimization. The specific process is as follows:
[0045] First, since the values of different wildfire factors are different, it is difficult to distinguish the imbalance between the values of different wildfire factors among the sample data only from the numerical size characteristics of the wildfire factors. Therefore, the values of the wildfire factors in a resampled data set are divided into 4 subsets.
[0046] For the kth wildfire factor, calculate the absolute value of the difference between each value of the wildfire factor and the mean value of all values of the wildfire factor in each resampled data set; the absolute value of the difference is in the interval [0,σ k)All the values of the k-th wildfire factor corresponding to it are divided into the first subset of the k-th wildfire factor in each resampled dataset; the absolute value of the difference is in the interval [σ k , 2σ k )All the corresponding values are divided into the second subset of the k-th wildfire factor in each resampled dataset; the absolute value of the difference is in the interval [2σ k , 3σ k )All the corresponding values are divided into the third subset of the k-th wildfire factor in each resampled dataset; the absolute value of the difference is in the interval [3σ k , +∞)All the corresponding values are divided into the fourth subset of the k-th wildfire factor in each resampled dataset; where σ k represents the standard deviation of all the values of the k-th factor in each resampled dataset.
[0047] Next, analyze the distribution of the values of each wildfire factor among different subsets in each resampled dataset, and calculate the class weight w k,n of the k-th wildfire factor in the n-th resampled dataset, and its formula form is: where T represents the number of subsets of the wildfire factor in the resampled dataset, and the value in this embodiment is 4; t represents the subset serial number, and the larger the subset serial number, the greater the difference between the factor value in the subset and the mean value of the factor in the resampled dataset, reflecting the greater the degree of dispersion of the value of the k-th wildfire factor in the n-th resampled dataset; represents the mean value of the lengths of the fire point sequences of all sample points in the t-th subset of the k-th wildfire factor in the n-th resampled dataset; r t represents the proportion of the number of values of the k-th wildfire factor in the t-th subset in the n-th resampled dataset, specifically the ratio of the number of values in the t-th subset to the number of values in the n-th resampled dataset; ε is a preset value, and the value is 0.1, and its purpose is to avoid the denominator being zero.
[0048] It should be understood that in the subset corresponding to the wildfire factor, the greater the difference between the value of the wildfire factor and the mean value of the wildfire factor in the resampled dataset, and the smaller the proportion of the number of wildfire factor values in the subset in the resampled dataset, the greater the degree of dispersion of the distribution of the wildfire factor values, indicating the stronger the imbalance of the wildfire factor in the sample point data, and a larger class weight should be set.
[0049] Secondly, considering the distribution of fire points in time, the greater the mean value of the lengths of the fire point sequences of the corresponding sample data in the subset of the wildfire factor, the greater its role in judging the wildfire risk, so the obtained class weight is greater to improve the evaluation performance of the wildfire risk caused by extreme environmental factors.
[0050] S3: Based on the set values of each wildfire factor in each resampled dataset, establish each decision tree, and test the decision tree based on the out-of-bag dataset corresponding to the resampled dataset to obtain the out-of-bag error rate of each decision tree; analyze the distribution difference characteristics of the out-of-bag error rates of each wildfire factor and the remaining wildfire factors on all decision trees, and combine the corresponding class weights to obtain the factor importance of each wildfire factor.
[0051] Take the wildfire sample data of all sample points in the resampled dataset and the class weights of different wildfire factors as inputs, use the random forest algorithm to synchronously establish decision trees on a preset number of resampled datasets, randomly select a set number of wildfire factors as decision tree features, and predict the sample point data. In this embodiment, the value of the preset number is 100, and the value of the set number is 3. The implementer can adjust it according to the actual situation. In this application, the prediction results of the decision tree are fire point and non-fire point sample points. Based on this, one resampled dataset corresponds to one decision tree, and each decision tree corresponds to 3 decision tree features. It should be noted that in this embodiment, it is necessary to ensure that all wildfire factors are used as decision tree features at least once.
[0052] Use the out-of-bag dataset corresponding to each resampled dataset as the test set to test the established decision tree and obtain the out-of-bag error rate of each decision tree.
[0053] Calculate the factor importance according to the distribution of the out-of-bag error rate: where, w k,n represents the class weight of the kth wildfire factor in the nth resampled dataset; N k represents the number of resampled datasets with the kth wildfire factor as the decision tree feature; exp() represents the exponential function with the natural constant as the base; μ k,n represents the mean value of the average out-of-bag error rates of the remaining wildfire factors except the kth wildfire factor in the nth decision tree; er k,n represents the average value of the out-of-bag error rates of the kth wildfire factor in the nth decision tree among all decision trees, and it is denoted as the average out-of-bag error rate of the kth wildfire factor; s k represents the factor importance of the kth wildfire factor, reflecting its role in judging wildfire risks.
[0054] It should be understood that in the resampled dataset with the k-th wildfire factor as the decision tree feature, the smaller the average out-of-bag error rate of the k-th wildfire factor is relative to the mean of the average out-of-bag error rates of all other wildfire factors, the greater its role in the subsequent training of the wildfire risk assessment model, and the greater the calculated factor importance. On the other hand, to prevent mis-elimination and further improve the role of wildfire factors with extreme factor values causing high wildfire risks in subsequent model training, for wildfire factors with greater class weights, the calculated factor importance is greater.
[0055] S4: Preset the training set; according to the numerical difference characteristics and distance characteristics between pairwise combined sample points of any two wildfire factors in the training set, combined with the factor importance of each wildfire factor and the conditional probability obtained from the Bayesian network model, obtain the association confidence of the any two wildfire factors.
[0056] To evaluate the influence of different factors on the training of the wildfire risk assessment model, 70% of the sample point data is taken as the training set, and the remaining 30% of the sample point data is used as the test set to evaluate the model training effect.
[0057] First, take the wildfire sample dataset containing all wildfire factors as the input of the naive Bayesian network model, use the maximum likelihood estimation method to learn the initial network and the transition network, and obtain the conditional probabilities of different wildfire factors, which characterize the risk of a single wildfire factor causing a wildfire.
[0058] The closer the values of two sample points are for different wildfire factors, the greater the probability of the wildfire risk caused by the factor combination, and the greater the calculated association feature value. In this application, the values of any two wildfire factors in a sample point are combined into a factor vector, and the Euclidean distance between factor vectors of different sample points is calculated, which is used as a reflection of the numerical difference between two wildfire factors.
[0059] According to the value distribution of wildfire factors in the training set, calculate the association feature value between any two factors at sample points: λ ij =ln(δ + D)×E ij ; where E ij represents the Euclidean distance between the factor vectors formed by the i-th and j-th factors in two sample points; D represents the actual distance between two sample points; ln() represents the logarithmic function with the natural constant as the base, and its purpose is to control the influence degree of the actual distance between sample points on the weight of calculating the association feature value; δ is a preset parameter with a value of 1, and its purpose is to avoid the calculation result of the association feature value being negative; λ ij represents the association feature value between the i-th and j-th factors at two sample points, reflecting the association of the two factors causing wildfire risk at sample points.
[0060] It should be understood that during the process of the development, spread, and extinction of wildfires, there is a strong geographical continuity. Among the sample points of fire spots that are relatively close to each other, the probability of wildfires occurring at some sample points due to the spread of wildfires is relatively high, rather than being caused by the wildfire factors of the sample points themselves. Therefore, the closer the distance between two sample points, the smaller the correlation of the risk of wildfires jointly triggered by different factors.
[0061] Furthermore, according to the factor importance and the conditional probability corresponding to each wildfire factor, calculate the correlation confidence of two wildfire factors: Among them, represents the average value of the correlation eigenvalues between the i-th and j-th wildfire factors among all sample points in the training set, reflecting the correlation of the risk of wildfires triggered by the two factors within the sample points of the training set; s i and s j respectively represent the factor importance of the i-th and j-th wildfire factors; P i and P j respectively represent the conditional probabilities of the i-th and j-th wildfire factors; norm[] represents the normalization function, and in this embodiment, the arctangent normalization function is used; CI ij represents the correlation confidence of the i-th and j-th wildfire factors, reflecting the probability of the risk of wildfires jointly triggered by the two wildfire factors.
[0062] S5: According to the numerical characteristics of the correlation confidence between each wildfire factor and the remaining wildfire factors, train the wildfire risk assessment model to obtain the wildfire risk probability of the sample point data.
[0063] Taking the test set as the input and using the confusion matrix to calculate the model accuracy rate, evaluate the training result of the wildfire risk assessment model. Among them, the confusion matrix is a well-known existing technology, and this application will not elaborate on it. This application uses the naive Bayesian network model as the wildfire risk assessment model.
[0064] The higher the risk of the two wildfire factors jointly triggering the wildfire risk, the greater the correlation confidence, and the greater the role of the wildfire factor in assessing the wildfire risk. Calculate the sum of the correlation confidences between each wildfire factor and the remaining wildfire factors, and take the wildfire factor with the smallest sum of the correlation confidences as the object to be excluded. And use the wildfire sample data set containing all the remaining wildfire factors as the input of the naive Bayesian network model. At the same time, use the above method to train the model and delete the factor with the smallest sum of the correlation confidences one by one until the model accuracy rate no longer improves. Thus, the training of the wildfire risk assessment model is completed, and according to the wildfire risk probability of the sample points output by the model, guide the operation and maintenance personnel to carry out differentiated wildfire prevention work.
[0065] Among them, the flow chart for obtaining the wildfire risk probability is as Figure 2 shown.
[0066] The present application provides a spatio-temporal assessment method for wildfire risk distribution based on full-process data processing. The method includes: First, by assigning different category weights to multiple wildfire factors, the impact of data imbalance on analyzing the wildfire risk caused by different factors is reduced, and the accuracy of using wildfire factors during the training of the wildfire risk assessment model is improved; at the same time, the out-of-bag dataset is used to test the decision trees established for each resampled dataset to obtain the factor importance, enhancing the optimization effect on different wildfire factors and reducing the noise introduced by redundant wildfire factors during the subsequent training process of the wildfire risk assessment model; further, by analyzing the correlation characteristics between different wildfire factors, the correlation confidence is obtained, enhancing the assessment accuracy of the model for the wildfire risk caused by wildfire factors and avoiding the mis-elimination of wildfire factors. Finally, the training results of the wildfire risk assessment model are evaluated, and the factor selection is optimized to improve the performance of the trained wildfire risk assessment model.
[0067] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the block may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description. Sometimes, there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0068] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A spatio-temporal assessment method for wildfire risk distribution based on the whole-process data processing, characterized in that, The method includes the following steps: Obtain all wildfire factors and fire point sequences of each sample point to form a wildfire sample data set; Divide the wildfire sample data set into a resampled data set and an out-of-bag data set through resampling technology; for a resampled data set, determine all subsets of each wildfire factor; analyze the length characteristics of the fire point sequence of each wildfire factor, and combine the proportion of the number of values of the corresponding subset in the resampled data set to obtain the class weight of each wildfire factor in each resampled data set; Establish decision trees based on a set number of wildfire factors in each resampled data set, test the decision trees based on the out-of-bag data set corresponding to the resampled data set, and obtain the out-of-bag error rate of each decision tree; analyze the distribution difference characteristics of the out-of-bag error rates of each wildfire factor and the other wildfire factors on all decision trees, and combine the corresponding class weights to obtain the factor importance of each wildfire factor; Preset a training set; according to the numerical difference characteristics and distance characteristics between pairwise combined sample points of any two wildfire factors in the training set, combine the factor importance of each wildfire factor and the conditional probability obtained from the Bayesian network model to obtain the association confidence of the any two wildfire factors; Train a wildfire risk assessment model according to the numerical characteristics of the association confidence between each wildfire factor and the other wildfire factors to obtain the wildfire risk probability of the sample point.
2. The spatio-temporal assessment method for wildfire risk distribution based on the whole process of data processing according to claim 1, characterized in that All subsets of each wildfire factor are determined based on the 3σ principle.
3. The spatio-temporal assessment method for wildfire risk distribution based on the full-process data processing according to claim 1, wherein, The formula for the class weight of each wildfire factor in each resampled dataset is specifically as follows: Denote the class weight of the k-th wildfire factor in the n-th resampled dataset as w k,n , and its formula form is: where T represents the number of subsets of the wildfire factor in the resampled dataset; t represents the subset serial number; represents the mean length of the fire point sequences of all sample points in the t-th subset of the k-th wildfire factor in the n-th resampled dataset; r t represents the proportion of the number of values in the t-th subset of the k-th wildfire factor in the n-th resampled dataset; ε is a preset value.
4. The method for spatio-temporal assessment of wildfire risk distribution based on full-process data processing according to claim 1, wherein, When establishing a decision tree, a set number of wildfire factors are used as decision tree features.
5. The spatio-temporal assessment method for wildfire risk distribution based on the full-process data processing according to claim 4, characterized in that The factor importance of each wildfire factor is obtained, and the specific formula is as follows: where s k represents the factor importance of the k-th wildfire factor; w k,n represents the class weight of the k-th wildfire factor in the n-th resampled dataset; N k represents the number of resampled datasets with the k-th wildfire factor as the decision tree feature; exp() represents the exponential function with the natural constant as the base; μ k,n represents the mean of the out-of-bag error rates of the remaining wildfire factors except the k-th wildfire factor in the n-th decision tree; er k,n represents the out-of-bag error rate of the k-th wildfire factor in the n-th decision tree.
6. The spatio-temporal assessment method for wildfire risk distribution based on full-process data processing according to claim 5, characterized in that, The specific average out-of-bag error rate is the mean of the out-of-bag error rates of all decision trees with each wildfire factor as a decision tree feature.
7. The spatio-temporal assessment method for wildfire risk distribution based on the full-process data processing according to claim 1, wherein The process of the association confidence of any two wildfire factors is specifically as follows: Compare the difference characteristics between vectors composed of different sample points of any two wildfire factors, and combine the distance characteristics between sample points to obtain the association feature value of the any two wildfire factors between sample points; Denote the correlation confidence of the \(i\)-th and \(j\)-th wildfire factors as \(CI\). ij , and its formula form is as follows: Among them, represents the average value of the correlation eigenvalues between all sample points of the \(i\)-th and \(j\)-th wildfire factors in the training set; \(s\) i and \(s\) j respectively represent the factor importance of the \(i\)-th and \(j\)-th wildfire factors; \(P\) i and \(P\) j respectively represent the conditional probabilities of the \(i\)-th and \(j\)-th wildfire factors; \(norm[]\) represents the normalization function.
8. The spatio-temporal assessment method for wildfire risk distribution based on full-process data processing according to claim 7, characterized in that The process of obtaining the association feature value of any two wildfire factors between sample points is specifically as follows: Denote the vector composed of any two wildfire factors at a sample point as a factor vector; Obtain the positive correlation mapping result of the distance metric between two sample points, and combine it with the distance metric between the factor vectors of the any two wildfire factors in the two sample points to obtain the association feature value of the any two wildfire factors between sample points.
9. The spatio-temporal assessment method for wildfire risk distribution based on the full-process data processing according to claim 1, wherein, The training of the wildfire risk assessment model includes: Calculate the cumulative sum of the association confidence between each wildfire factor and the other wildfire factors, use the wildfire factor with the smallest cumulative sum as the object to be excluded, and use the wildfire sample data set of all the remaining wildfire factors as the input of the Naive Bayesian network model. Delete the wildfire factor with the smallest sum of association confidence one by one until the model accuracy no longer improves.
10. The method for spatio-temporal assessment of wildfire risk distribution based on the whole process of data processing as claimed in claim 9, wherein The model accuracy is determined through a confusion matrix.
Citation Information
Patent Citations
Hill fire risk prediction method based on stacking algorithm
CN109447331A
Forest fire risk prediction method based on a probability programming technology
CN109829583A
Wireless sensor network abnormal event detecting method based on multi-attribute correlation
CN105764162A
Weight clustering and under-sampling-based unbalanced data classification method
CN106778853A
Target node key information filling method and system based on association network
CN110706095A