Conflict Event Data Processing Method Based on Improved Bayesian Additive Regression Trees

By improving the Bayesian additive regression tree model and integrating multi-criteria decision model, combining susceptibility, exposure and vulnerability indicators, the conflict risk distribution map is generated, and the problem of difficulty in comprehensive conflict risk prediction in the existing technology is solved, and the accuracy and reliability of the prediction are improved.

CN117743983BActive Publication Date: 2025-07-01INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311807587.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-07-01
Estimated Expiration
2043-12-26

AI Technical Summary

Technical Problem

It is difficult for the prior art to conduct comprehensive conflict risk predictions in a large area, and the resulting model cannot provide a reasonable regional risk distribution map.

Method used

Using a method based on the improved Bayesian additive regression tree model and an integrated multi-criteria decision model, a conflict risk distribution map is constructed by obtaining multiple indicators that affect conflict risk, and a conflict risk distribution map is generated.

Benefits of technology

It improves the accuracy and reliability of conflict risk prediction, can capture potential conflict risks more comprehensively, provide more accurate and comprehensive risk prediction data, and provide data support for risk response decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117743983B_ABST
    Figure CN117743983B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method for processing conflict event data based on an improved Bayesian additive regression tree, including: obtaining multiple indicators affecting conflict risk in a target area, and dividing the multiple indicators into vulnerability indicators, exposure indicators, and fragility indicators of conflict events; extracting data of each indicator in units of geographical grids in the target area, and constructing vulnerability samples by using the data of each vulnerability indicator; training an improved Bayesian additive regression tree model with the vulnerability sample set, and predicting the future vulnerability distribution of conflict events in the target area through the trained model; integrating multiple multi-criteria decision-making models by using the data of each exposure and fragility indicator to evaluate the exposure and fragility distributions of the target area; generating a future conflict risk distribution map of the target area according to the vulnerability distribution, exposure distribution, and fragility distribution. This embodiment can give a reasonable regional risk distribution map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of data processing, and in particular, to a method for processing conflict event data based on an improved Bayesian additive regression tree. Background Art

[0002] The adverse impacts of conflict events on society, the environment, and the economy are increasing, such as armed conflict events, violent conflict events, etc. How to conduct comprehensive conflict risk prediction in a large-scale area and provide data support for risk response decision-making is an urgent problem to be solved.

[0003] In the prior art, machine learning algorithms are usually used to process historical conflict event data and construct a prediction model for future conflict risks. For example, the methods described in the following two references can be used for data processing and model construction: Feng Changqiang, et al., Research on the Prediction of Internal Armed Conflict Risks in Countries Based on the RF Algorithm. Journal of Information Engineering University, 2022, 23(03): p. 373-378; Halkia, M., et al., The Global Conflict Risk Index: A Quantitative Tool for Policy Support on Conflict Prevention. Progress in Disaster Science, 2020, 6: p. 100069. However, when using the above methods to process conflict event data in multiple regions, the obtained model cannot give a relatively reasonable regional risk distribution map.

[0004] In view of this, the present invention is proposed. Summary of the Invention

[0005] The embodiments of the present invention provide a method for processing conflict event data based on an improved Bayesian additive regression tree to solve the above technical problems.

[0006] In a first aspect, the embodiments of the present invention provide a method for processing conflict event data based on an improved Bayesian additive regression tree, including:

[0007] Obtain multiple indicators affecting conflict risks in a target area, and divide the multiple indicators into vulnerability indicators, exposure indicators, and fragility indicators of conflict events; wherein, the multiple indicators include several categories such as conflict event history, geographical features, geopolitical environment, natural disasters, social economy, and climate change;

[0008] Extract the data of each indicator in units of geographical grids in the target area, and use the data of each vulnerability indicator to construct a vulnerability sample;

[0009] Train an improved Bayesian additive regression tree model using a set of susceptibility samples, and predict the future conflict event susceptibility distribution in the target area through the trained model; wherein, the improved Bayesian additive regression tree model takes the multiple indicators as independent variables, combines the mixed effects of region types to emphasize the impact differences of region types on conflict events, and combines the Gaussian process of the geographical grid to ensure the spatial correlation of conflict event susceptibility;

[0010] Integrate multiple multi-criteria decision-making models using the data of each exposure index to evaluate the exposure distribution of the target area; integrate multiple multi-criteria decision-making models using the data of each vulnerability index to evaluate the vulnerability distribution of the target area;

[0011] Generate a future conflict risk distribution map of the target area according to the susceptibility distribution, exposure distribution and vulnerability distribution.

[0012] In a second aspect, an embodiment of the present invention provides an electronic device, which includes:

[0013] One or more processors;

[0014] A memory for storing one or more programs,

[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the conflict event data processing method based on the improved Bayesian additive regression tree according to any embodiment.

[0016] The method provided by the embodiment of the present invention uses an improved Bayesian additive regression tree model and an integrated multi-criteria decision-making model for integrated modeling, and finally draws a conflict risk distribution map in a large-scale area. The method can achieve the following beneficial effects:

[0017] 1. Improve the accuracy and reliability of risk prediction: This method uses an improved Bayesian additive regression tree model that combines Gaussian processes and mixed effects for susceptibility assessment. This model can complete automatic modeling of non-linear, continuous, and complex high-order interactions, better handle the heterogeneity and correlation in the data, and produce spatially continuous results. In addition, the integrated multi-criteria decision-making model can help decision-makers compare the advantages and disadvantages of different solutions and select the optimal solution; during the process, it can not only reduce the bias and variance of a single model, improve the overall prediction accuracy, but also reduce the risk of overfitting, and improve the robustness and generalization ability of the model.

[0018] 2. Introduction of the comprehensive assessment method: In this embodiment, the IPCC framework and the comprehensive assessment method in climate change are extended to the processing of conflict event data, enabling conflict risk prediction to not only focus on the probability of conflict occurrence but also comprehensively consider other important factors such as vulnerability and exposure, thereby better taking into account the impact of climate change and environmental factors on conflict risk. This comprehensive assessment method can capture potential conflict risks more comprehensively, provide more accurate and comprehensive risk prediction data, and provide data support for risk response decisions.

[0019] 3. Enhancement of global perspective and comprehensive analysis: This embodiment can expand the prediction area to a vast scope including multiple regional types, even a global scale, with the ability of global perspective and comprehensive analysis. Through the analysis of conflict time risks in a large-scale area, it can better display the influencing factors, grading, and change trends in a large-scale area, helping decision-makers better understand the dynamic changes of risks and response strategies, and reducing the environmental and social impacts brought by conflicts. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 It is a flowchart of a conflict event data processing method based on an improved Bayesian additive regression tree provided by an embodiment of the present invention;

[0022] Figure 2 It is a flowchart of another conflict event data processing method based on an improved Bayesian additive regression tree provided by an embodiment of the present invention;

[0023] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present invention.

[0025] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific

[0026] orientation, and thus should not be construed as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0027] In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0028] As described in the background art, when processing the conflict event data involving multiple regions using the methods in the prior art, the obtained model cannot give a relatively reasonable regional risk distribution map and cannot provide strong data support for risk response decisions. In this embodiment, the unreasonableness of its prediction results is summarized into several aspects:

[0029] 1. Ignoring spatial continuity and regional differences: The machine learning methods in the prior art usually only focus on the correlation between independent variables, lacking the consideration of the spatial correlation of dependent variables, thus resulting in discontinuous spatial predictions. At the same time, it cannot capture the differences between different regions and the correlation within groups.

[0030] 2. Lack of comprehensive evaluation methods: The machine learning methods in the prior art usually only focus on the occurrence probability of conflicts, while ignoring other important factors such as the vulnerability and exposure of conflict events. The prediction method from a single angle cannot capture potential conflict risks, leading to the neglect of some important factors.

[0031] 3. Lack of global perspective and comprehensive analysis: When conducting conflict risk analysis in multiple regions or even globally, the prior art lacks a global perspective and comprehensive analysis capabilities, and cannot fully consider the influencing factors, risk grading, and trends within a large regional scope, making it difficult to present the overall picture of risk data in a large region.

[0032] An embodiment of the present invention provides a method for processing conflict event data based on an improved Bayesian additive regression tree to overcome the above-mentioned defects. Generally speaking, risk refers to the possibility that people or valuable things are in danger and the result is uncertain, usually expressed as the probability of a harmful event or trend occurring multiplied by the consequences caused by the occurrence of these events or trends. In 2014, the Working Group II contribution to the Fifth Assessment Report of the Intergovernmental Panel on Climate Change (IPCC), "Climate Change 2014: Impacts, Adaptation, and Vulnerability", drew on previous research results and used the interaction of three variables, namely vulnerability, exposure, and susceptibility, as the three elements for risk prediction. In order to better predict conflict risks, this embodiment applies the above-mentioned assessment idea of climate change to conflict risk prediction, comprehensively considers the susceptibility, exposure, and vulnerability of conflict events, and proposes a method for processing conflict event data based on an improved Bayesian additive regression tree model and an integrated multi-criteria decision-making model, thereby constructing a comprehensive conflict risk prediction model. Among them, susceptibility describes the probability of a conflict event occurring, exposure describes the degree to which a region or system is exposed to these hazards, and vulnerability describes the degree to which a region or system is vulnerable to these hazards when facing these hazards.

[0033] More specifically, Figure 1 is a flowchart of a method for processing conflict event data based on an improved Bayesian additive regression tree provided by an embodiment of the present invention. This method is applicable to the situation of risk prediction for a large-scale area covering multiple types of regions and is executed by an electronic device. Figure 2 is a flowchart of a specific implementation manner of this method. Combining Figure 1 and Figure 2 , this method specifically includes the following steps:

[0034] S110. Obtain data of multiple indicators affecting conflict risks in the target area, and divide the multiple indicators into susceptibility indicators, exposure indicators, and vulnerability indicators of conflict events.

[0035] Here, the type of region refers to a specific area on the earth's surface, usually divided and defined according to factors such as physical geography, human geography, or administrative divisions. For example, the Middle East region, Southeast Asia, and South America, etc. Here, the target area refers to the large-scale area to be predicted, including multiple types of regions. For example, the target area is a country, a geopolitical block, or even the world.

[0036] The multiple indicators include categories such as conflict event history, geographical features, geopolitical environment, social economy, and climate change. As described above, in this embodiment, the risk of conflict events is evaluated from three perspectives: susceptibility, exposure, and vulnerability. Therefore, the indicators that affect the susceptibility, exposure, and vulnerability of conflict events are first divided from these indicators. Among them, the relevant influencing indicators of susceptibility include the disaster-causing factors and formation environment of conflict events, which are involved in several categories such as conflict event history, geographical features, geopolitical environment, social economy, and natural climate. Exemplarily, the relevant indicators of susceptibility are shown in Table 1.

[0037] Table 1 Relevant indicators of the susceptibility of conflict events

[0038]

[0039]

[0040] Conflict events may cause casualties, infrastructure losses, and environmental damage. Therefore, the exposure indicator reflects the characteristics of the social and natural environment affected by conflict events. The vulnerability indicator reflects the potential loss degree or risk resistance ability of the disaster-stricken area in the face of conflict events. Exemplarily, the specific indicators of exposure and vulnerability are shown in Table 2.

[0041] Table 2

[0042]

[0043] S120. Extract the data of each indicator in units of geographical grids in the target area, and construct a susceptibility sample using the data of each susceptibility indicator.

[0044] In this embodiment, the target area is divided into unified geographical grids, the original data of each indicator is extracted based on each geographical grid, and a susceptibility sample is constructed. Optionally, the resolution of the geographical grid can be 0.5°.

[0045] In a specific embodiment, after extracting the original data in units of geographical grids, the original data can be normalized. That is, considering the maximum and minimum values of each indicator in all grids, all the original data is standardized to the same range of 0-1. The formula for normalization calculation is:

[0046]

[0047] Among them, and X i respectively represent the normalized indicator value and the original indicator value of grid i, X max and X minThey represent the maximum and minimum values of all the original grid metrics respectively. All operations in the subsequent steps use the normalized data as the processing object.

[0048] Vulnerability can be quantified separately in the following way: Vulnerability represents whether a conflict event occurred in a certain year. If a conflict event occurred in a geographical grid in a certain year, the vulnerability of the conflict event in the geographical grid is quantified as 1; if no conflict event occurred in the geographical grid in a certain year, the vulnerability of the conflict event in the geographical grid is quantified as 0. Optionally, when the prediction unit is a year, vulnerability represents whether a conflict event occurred within a year. When an armed conflict occurred, it is marked as 1 (high vulnerability), otherwise 0 (low vulnerability). The formula is as follows:

[0049]

[0050] After quantification, the vulnerability values of the same geographical grid and the values of each vulnerability index shown in Table 1 together constitute a vulnerability sample, and multiple vulnerability samples constitute a vulnerability sample set. Among them, the sample with vulnerability = 1 can be called a positive vulnerability sample, and the sample with vulnerability = 0 can be called a negative vulnerability sample. Further, an equal number of positive vulnerability samples and negative vulnerability samples can be randomly selected to construct a one-year sample and a final multi-year sample for the final model training.

[0051] S130. Use the vulnerability sample set to train the improved Bayesian additive regression tree model, and predict the future conflict event vulnerability distribution in the target area through the trained model; wherein, the improved Bayesian additive regression tree model uses some or all of the multiple metrics as independent variables, combines the mixed effect of the region type to emphasize the impact difference of the region type on conflict events, and combines the Gaussian process of the geographical grid to ensure the spatial correlation of conflict event vulnerability.

[0052] To illustrate the vulnerability prediction model of this embodiment, first introduce the conventional Bayesian additive regression tree model. The conventional Bayesian additive regression tree model can be expressed as:

[0053]

[0054] where, y i is the dependent variable, P represents the number of tree models in the Bayesian additive regression tree model, X i is the input independent variable, represents the structure of the p-th tree model, Θ p represents the node parameter set of the p-th tree model, the G function returns the prediction result of each tree model, ∈ i is the error term.

[0055] In this embodiment, the above basic model is applied to vulnerability prediction, and the model is improved according to the characteristics of the entire target area in terms of area type and spatial continuity, resulting in the model shown in formula (2):

[0056]

[0057] Corresponding to the scenario of vulnerability prediction, y in the above formula i represents the vulnerability value of conflict events within the geographical grid s i , P represents the number of tree models in the Bayesian additive regression tree model, p represents the index of the tree model, and X i represents the values of each vulnerability index within the geographical grid s i ; represents the structure of the p-th tree model, Θ p represents the set of node parameter values of the p-th tree model, the G function returns the prediction result of each tree model, and ∈ i is the error term; z i represents the area type where the geographical grid s i is located. Different z i correspond to different sets of tree model node parameter values, but the values of the same node parameter for different z i satisfy the same prior distribution to achieve the mixed effect of z i in the model; f(s i ) represents a Gaussian process with respect to the geographical grid s i . Assume that the prior of the Gaussian process is:

[0058] f~GP(μ(s),k(s,s'))

[0059] where GP represents the Gaussian process, μ(s) is the mean function, k(s,s') is the covariance function, and s and s' represent two different geographical grids respectively. That is:

[0060]

[0061] where represents the expectation.

[0062] Compared with formula (1), the above model uses the categorical variable z iThe mixed effect of region types is added to emphasize the differential impact of region types on conflict events. Since the outbreak patterns of conflict events vary among different region types, the improved model better conforms to the actual data patterns, facilitating the identification of potential patterns and trends in local areas. However, regardless of the number of region types involved, the overall spatial distribution of the region should maintain a certain degree of continuity. Excessive abrupt changes do not conform to the actual patterns. Therefore, based on the random effect, a Gaussian distribution of geographical grids is added to the above model to ensure the spatial continuity of the prediction results. Through the combined action of the mixed effect and the Gaussian process, an overall prediction model can be constructed for large-scale regions with multiple regions, taking into account regional differences and overall smoothness, and obtaining a more reasonable vulnerability distribution map of conflict events.

[0063] The implementation process of the above random effect will be described below. For ease of understanding and description, the node parameters in this embodiment specifically refer to the parameters in the Bayesian additive regression tree model, that is, the model parameters in (∈ i is equivalent to the intercept of G and can also be regarded as a node parameter)), excluding the parameters in the Gaussian process f(s i ). Based on the above definitions, the random effect in formula (2) can be understood as: constructing a Bayesian additive regression tree model for different region types respectively. Each region type corresponds to a set of node parameters, but the same node parameters under different region types satisfy a unified prior distribution.

[0064] In a specific implementation, the training process of model (2) may include the following steps:

[0065] Step 1. Model setting: Determine the parameter settings and hyperparameters of the Bayesian additive regression tree model, such as the number of trees, the depth of the trees, the distribution of the base model, etc.; determine the form of the mixed effect model, and select the forms of the fixed effect and the random effect: the fixed effect is each vulnerability index, and the random effect is the region type; determine the form of the Gaussian process model, and select an appropriate kernel function or covariance function, such as a linear kernel, a polynomial kernel, a Gaussian kernel, etc., to describe the spatial correlation of the data.

[0066] Step 2. Parameter setting: Set the prior distribution and hyperparameters for the parameters in the model. According to domain knowledge or experience, select an appropriate prior distribution. Generally, it can be assumed to follow a normal distribution or a uniform distribution.

[0067] Step 3. Initialize the parameters: Set the initial values for the model parameters, and sample according to the region type respectively. Generally, use the mean or random value of the prior distribution as the initial value.

[0068] Step 4. Iteratively optimize the parameters using the samples: In each sample iteration, generate candidate parameter values according to the current parameter values, and calculate the acceptance probability; if the acceptance probability meets certain conditions, accept the candidate parameter values as the parameter values for the next iteration; otherwise, retain the current parameter values.

[0069] Step 5. Collect parameter samples: Since the prediction errors in the previous several iterations are relatively large, discard the parameter samples from the previous several iterations, and then start collecting parameter samples. Collect a sufficient number of parameter samples to obtain a reliable estimate and perform convergence diagnosis.

[0070] In another specific embodiment, the training process of model (2) may also include the following steps:

[0071] Step 1. Preset the prior distributions satisfied by the node parameters in each tree model. Optionally, for each node parameter, different z i correspondingly, this parameter satisfies a Gaussian distribution, and the mean and standard deviation of the given Gaussian distribution are provided.

[0072] Step 2. Extract a sample from the susceptibility sample set, and determine the region type where the sample is located; construct the model shown in formula (2) corresponding to the region type. At this time, all parameters in this model (including the node parameters of G and the parameters of f) are not assigned values. Specifically, for a specific region type, the model of formula (2) can be expressed in the following form:

[0073]

[0074] That is, for the current region type, a model as shown in formula (3) can be constructed. After construction, for any node parameter of G, randomly extract a value from the prior distribution satisfied by this parameter as the initial value of this node parameter; after performing the above operations on all node parameters of G, the initial values of all node parameters can be obtained. The initial values of the parameters of f can be set as needed, and no specific limitations are imposed in this embodiment. Substitute the initial values of the parameters of G and f into the current model, and then use the current sample to train this model to optimize the values of all parameters in the model.

[0075] Step 3. Extract a new sample from the susceptibility sample set. If the new sample has the same region type as the previously extracted sample, automatically match the model corresponding to the region type, and based on the latest parameter values (including all parameters) of this model, substitute the new sample into this model for training to further optimize the values of all parameters in the model. If the region type of the new sample is different from that of all previously extracted samples, construct a new model shown in formula (3) corresponding to the current region type. The f(s i) Still use the parameter values determined in the previous loop. For each node parameter, draw a value from the corresponding prior distribution as the initial value of each node parameter (the specific process is the same as in Step 2). Substitute each initial value into the new model, and then use the new sample to train the model to optimize the values of all parameters in the model.

[0076] Step 4: Return to the operation of drawing a new sample (i.e., Step 3), enter the next loop until the set termination condition is reached. Optionally, the termination condition can be that the models of each region type have reached the set number of training times, or have reached the set prediction accuracy, or all samples have been drawn, etc. This embodiment does not make specific limitations.

[0077] Through the above steps, the finally obtained model (2) can be understood as: for each region type, there is a corresponding set of parameters of G, and the initial parameter values satisfy the set distribution, so as to consider the grouped random effects of different regional conflict differences. At the same time, the parameter values will also be updated and adjusted according to the new data to match the data law; for f(s i ), all region types share a set of parameters of f, which will be continuously updated and adjusted according to each sample to constrain the overall spatial continuity.

[0078] After obtaining the final model through the above method, use the model to predict the future conflict risk probability at each location in the target area. Specifically, input the index data X i of each geographical grid s i into the model. The model automatically matches the corresponding i node parameter group according to the region type z i of s , substitute this group of parameters into formula (2) or (3), and calculate the susceptibility prediction value y i at s i . It should be noted that although the values of y i in the training samples are all 0 or 1, due to the constraint of spatial continuity, after training, the model predicts that y i can be any value in the interval [0, 1], representing the probability of a conflict event occurring in the future in grid s i . Arrange the predicted values y i of all grids into a two-dimensional image according to the geographical grid layout, which is the future conflict event susceptibility distribution map of the target area. This map maintains spatial continuity and can reflect the differences of different region types, and can present the hotspot effect of multi-hotspot distribution.

[0079] S140. Integrate multiple multi-criteria decision-making models using the data of each exposure index to evaluate the exposure distribution of the target area; integrate multiple multi-criteria decision-making models using the data of each vulnerability index to evaluate the vulnerability distribution of the target area.

[0080] In this embodiment, two sets of multi-criteria decision-making models are respectively set up to process the data of the exposure index and the data of the vulnerability index, and the exposure distribution and the vulnerability distribution of the target area are respectively evaluated. Among them, each set of multi-criteria decision-making models may include three models: CODAS (COmbinative Distance-based Assessment), EDAS (Evaluation based on Distance from Average Solution), and MOOSRA (Multi-Objective Optimization on the basis of Simple Ratio Analysis). By weighted averaging the evaluation results of the three multi-criteria decision-making models, the exposure distribution and the vulnerability distribution of the target area can be obtained respectively. Taking the evaluation of the exposure distribution as an example, in a specific embodiment, this process specifically includes the following steps:

[0081] Step 1. Use the CODAS multi-criteria decision-making model to calculate the values of each exposure index in each geographical grid to obtain the first value of the exposure of conflict events in each geographical grid. Optionally, first, for every two geographical grids, calculate the distance l1 between the current two geographical grids on each exposure evaluation index according to the values of each exposure evaluation index, and then perform weighted summation on the distance l1 on each exposure evaluation index with different weights as the distance between the current two geographical grids. Then, for each geographical grid, perform weighted summation on the distances between the current geographical grid and other geographical grids to obtain the first value of the exposure of conflict events in the current geographical grid.

[0082] Step 2. Use the EDAS multi-criteria decision-making model to calculate the values of each exposure index in the current geographical grid to obtain the second value of the exposure of conflict events in the current geographical grid. Optionally, first, calculate the average value of all geographical grids on each exposure index, and form an average solution from the average values; then, for each geographical grid, calculate the distance l2 between the current geographical grid and the average solution on each exposure evaluation index according to the values of each exposure index in the current geographical grid, and then perform weighted summation on the distance l2 on each exposure evaluation index with different weights as the second value of the exposure of conflict events in the current geographical grid.

[0083] Step 3: Use the MOOSRA multi-criteria decision-making model to calculate the values of each exposure index within the current geographical grid, and obtain the third value of the conflict event exposure within the current geographical grid. Optionally, first, calculate the maximum value, minimum value, and average value of all geographical grids on each exposure index, and take a positive ideal value between the maximum value and the average value of each exposure index. The positive ideal solutions are composed of the positive ideal values of each exposure index; take a negative ideal value between the average value and the minimum value of each exposure index, and the negative ideal solutions are composed of the negative ideal values of each exposure index. Then, for each geographical grid, according to the values of each exposure index within the current geographical grid, calculate the distance l3 between the current geographical grid and the positive ideal solution on each exposure evaluation index respectively, and then perform weighted summation on the distances l3 on each exposure evaluation index with different weights to obtain the positive distance; calculate the distance l4 between the current geographical grid and the negative ideal solution on each exposure evaluation index respectively, and then perform weighted summation on the distances l4 on each exposure evaluation index with different weights to obtain the negative distance; average the positive distance and the negative distance to obtain the third value of the conflict event exposure within the current geographical grid. Exemplarily, the ratio of the distance between the positive ideal value and the average value to the distance between the positive ideal value and the maximum value can be taken as 0.8; similarly, the ratio of the distance between the average value and the negative ideal value to the distance between the negative ideal value and the minimum value can also be taken as 0.8; the two ratios can also be flexibly adjusted according to actual needs, and no specific limitation is made in this embodiment.

[0084] Step 4: Perform weighted averaging on the three values of each geographical grid to obtain the final exposure value of each geographical grid; the final exposure values of each geographical grid together constitute the exposure distribution of the target area. Optionally, assign corresponding weights to each multi-criteria decision-making model; combine or adjust the weights of different multi-criteria decision-making models to obtain the final integrated weight; according to the integrated weight, perform weighted averaging on the evaluation results of different multi-criteria decision-making models to obtain the final evaluation result. Among them, the weights involved in the calculation steps of the multi-criteria decision-making model can be determined by means such as expert judgment, questionnaire survey, and analytic hierarchy process.

[0085] The evaluation of the vulnerability distribution is similar, and the expression about "exposure" in the above specific implementation manner can be replaced with "vulnerability".

[0086] In addition, it should be noted that the exposure values and vulnerability values of the above geographical grids can also be directly determined by means of expert judgment, questionnaire survey, analytic hierarchy process, etc. Exemplarily, taking the area composed of multiple geographical grids as a unit, through expert judgment and questionnaire survey, the exposure value of the area is evaluated according to the data of each exposure index, which is used as the exposure value of all grids in the area, and the vulnerability value of the area is evaluated according to the data of each vulnerability index, which is used as the vulnerability value of all grids in the area. The exposure distribution and vulnerability distribution determined by this method also belong to the protection scope of this application.

[0087] S150. Generate a future conflict risk distribution map of the target area according to the susceptibility distribution, exposure distribution and vulnerability distribution.

[0088] After obtaining the above-mentioned susceptibility distribution map, exposure distribution map and vulnerability distribution map, the future conflict risk values of each geographical grid can be calculated according to the following formula:

[0089] Conflict risk spatial value = Susceptibility w1 × Exposure w2 × Vulnerability w3

[0090] Among them, w1, w2 and w3 respectively represent the weights of susceptibility, exposure and vulnerability, and w1 > w2, w1 > w3. Preferably, w1 = 0.5, w2 = w3 = 0.25 can be taken. Arranging each conflict risk value in a two-dimensional image according to the geographical grid, the future conflict risk distribution map of the target area is obtained. To sum up, this embodiment provides a new integrated modeling method based on the improved Bayesian additive regression tree model and the integrated multi-criteria decision-making model to draw the conflict risk distribution map in a large-scale area. This method can achieve the following beneficial effects:

[0091] 1. Improve the accuracy and reliability of risk prediction: This uses an improved Bayesian additive regression tree model combining Gaussian process and mixed effects for susceptibility assessment. This model can complete the automatic modeling of non-linear, continuous and complex high-order interactions, better cope with the heterogeneity and correlation in the data, and produce spatially continuous results. In addition, the integrated multi-criteria

[0092] decision-making model can help decision-makers compare the advantages and disadvantages of different solutions and select the optimal solution; during the process, it can not only reduce the bias and variance of a single model, improve the overall prediction accuracy, but also reduce the risk of overfitting, and improve the robustness and generalization ability of the model.

[0093] 2. Introduction of the comprehensive evaluation method: In this embodiment, the IPCC framework and the comprehensive evaluation method in climate change are extended to the processing of conflict event data, enabling the conflict risk prediction to not only focus on the probability of conflict occurrence but also comprehensively consider other important factors such as vulnerability and exposure, thereby better considering the impact of climate change and environmental factors on conflict risks. This comprehensive evaluation method can more comprehensively capture potential conflict risks, provide more accurate and comprehensive risk prediction data, and provide data support for risk response decisions.

[0094] 3. Enhancement of the global perspective and comprehensive analysis: This embodiment can expand the prediction area to a vast range including multiple regional types, even the global scale, with the ability of global perspective and comprehensive analysis. Through the analysis of conflict time risks in a large-scale area, it can better show the influencing factors, grading, and changing trends in the large-scale area, helping decision-makers better understand the dynamic changes of risks and response strategies, and reducing the environmental and social impacts brought by conflicts.

[0095] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 3 shown, the device includes a processor 60, a memory 61, an input device 62, and an output device 63; the number of processors 60 in the device can be one or more, Figure 3 taking one processor 60 as an example; the processor 60, memory 61, input device 62, and output device 63 in the device can be connected through a bus or other means, Figure 3 taking connection through a bus as an example.

[0096] The memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the conflict event data processing method based on the improved Bayesian additive regression tree in the embodiment of the present invention. The processor 60 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 61, that is, implements the above-mentioned conflict event data processing method based on the improved Bayesian additive regression tree.

[0097] The memory 61 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 61 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 61 may further include a memory remotely disposed relative to the processor 60, and these remote memories may be connected to the device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0098] The input device 62 may be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the device. The output device 63 may include display devices such as a display screen.

[0099] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the conflict event data processing method based on the improved Bayesian additive regression tree in any embodiment.

[0100] The computer storage medium of the embodiment of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or

[0101] a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.

[0102] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0103] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0104] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the C language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A method for processing conflict event data based on an improved Bayesian additive regression tree, characterized in that, include: Obtain multiple indicators that affect conflict risk in the target area, and divide the multiple indicators into conflict event susceptibility indicators, exposure indicators, and vulnerability indicators; wherein the multiple indicators include conflict event history, geography, geo-environment, natural disasters, socio-economics, and climate change; Extracting data of each indicator based on the geographic grid in the target area, and constructing a susceptibility sample using the data of each susceptibility indicator; The improved Bayesian additive regression tree model is trained using the susceptibility sample set, and the future conflict event susceptibility distribution of the target area is predicted by the trained model; wherein the improved Bayesian additive regression tree model uses the multiple indicators as independent variables and combines the mixed effects of regional types to emphasize the differences in the impact of regional types on conflict events, and combines the Gaussian process of the geographic grid to ensure the spatial correlation of the susceptibility of conflict events; The data of each exposure index is integrated into a plurality of multi-criteria decision-making models to evaluate the exposure distribution of the target area; the data of each vulnerability index is integrated into a plurality of multi-criteria decision-making models to evaluate the vulnerability distribution of the target area; A future conflict risk distribution map of the target area is generated based on the susceptibility distribution, exposure distribution and vulnerability distribution.

2. The method according to claim 1, characterized in that, The data for any indicator include the value of the indicator in each geographic grid of the target area; The method of constructing a susceptibility sample set using the data of each susceptibility index includes: if a conflict event has occurred in a geographic grid within a year, the susceptibility of the conflict event in the geographic grid is quantified as 1; if a conflict event has not occurred in a geographic grid within a year, the susceptibility of the conflict event in the geographic grid is quantified as 0; the susceptibility value of the same geographic grid and the value of each susceptibility index together constitute a susceptibility sample; each year, an equal number of susceptibility positive samples and susceptibility negative samples are randomly selected to construct a one-year sample, and all one-year samples in the research period are aggregated into multi-year samples; The method of using the susceptibility sample set to train the improved Bayesian additive regression tree model and predicting the future conflict event susceptibility distribution of the target area through the trained model includes: using multi-year samples to train the improved Bayesian additive regression tree model and predicting the conflict event susceptibility distribution of the target area in the next one or more years through the trained model.

3. The method according to claim 1, wherein The improved Bayesian additive regression tree model is: Among them, y i represents the vulnerability value of conflict events within the geographical grid s i , P represents the number of tree models in the Bayesian additive regression tree model, p represents the index of the tree model, and X i represents the values of each vulnerability index within the geographical grid s i . represents the structure of the p-th tree model, and Θ p represents the set of node parameter values of the p-th tree model. The G function returns the prediction result of each tree model, and ∈ i is the error term; f(s i ) represents the Gaussian process with respect to the geographical grid s i . z i represents the region type where the geographical grid s i is located. Different z i correspond to different sets of node parameter values of the tree model, but the values of the same node parameter for different z i satisfy the same prior distribution to achieve the mixed effect of z i in the model.

4. The method according to claim 3, wherein The method of training the improved Bayesian additive regression tree model using the susceptibility sample set comprises: Preset the prior distribution that the values ​​of each node parameter satisfy; A sample is extracted from the susceptibility sample set to determine the region type where the sample is located; a model shown in formula (2) corresponding to the region type is constructed, and a value is extracted from each prior distribution as the initial value of each node parameter in the model, and the sample is used to optimize the parameter values ​​of the model based on each initial value; Extract a new sample from the set of samples with high susceptibility. If the new sample has the same regional type as the samples that have been extracted, substitute the new sample into the model corresponding to the regional type to continue optimizing the parameter values of the model. If the regional type of the new sample is different from that of all the samples that have been extracted, construct a new model as shown in formula (2) corresponding to the current regional type. The parameters of f(s i ) in the new model still adopt the values in the previous cycle. Extract a value from each prior distribution as the initial value of each node parameter in the new model, and use the new sample to optimize the parameter values of the new model based on each initial value; Return to the operation of extracting new samples and enter the next cycle until the set termination condition is reached.

5. The method according to claim 1, wherein The data for any indicator include the value of the indicator in each geographic grid of the target area; Integrating multiple multi-criteria decision-making models using the data of each exposure index to evaluate the exposure distribution of the target area, including: Using the CODAS multi-criteria decision-making model to calculate the values of each exposure index within each geographical grid, obtaining the first value of the exposure of conflict events within each geographical grid; Using the EDAS multi-criteria decision-making model to calculate the values of each exposure index within each geographical grid, obtaining the second value of the exposure of conflict events within each geographical grid; Using the MOOSRA multi-criteria decision-making model to calculate the values of each exposure index within each geographical grid, obtaining the third value of the exposure of conflict events within each geographical grid; Performing weighted averaging on the above three values of each geographical grid to obtain the final exposure value of each geographical grid; the final exposure values of each geographical grid together constitute the exposure distribution of the target area.

6. The method according to claim 5, wherein The using of the CODAS multi-criteria decision-making model to calculate the values of each exposure index within each geographical grid, obtaining the first value of the exposure of conflict events within each geographical grid, includes: For every two geographical grids, calculating the distances between the current two geographical grids on each exposure evaluation index according to the values of each exposure evaluation index, and then performing weighted summation by assigning different weights to the distances on each exposure evaluation index as the distance between the current two geographical grids; For each geographical grid, averaging the distances between the current geographical grid and other geographical grids to obtain the first value of the exposure of conflict events within the current geographical grid.

7. The method according to claim 5, wherein The using of the EDAS multi-criteria decision-making model to calculate the values of each exposure index within each geographical grid, obtaining the second value of the exposure of conflict events within each geographical grid, includes: Calculating the average value of all geographical grids on each exposure index, and forming an average solution from the average values; For each geographical grid, calculating the distances between the current geographical grid and the average solution on each exposure evaluation index according to the values of each exposure index within the current geographical grid, and then performing weighted summation by assigning different weights to the distances on each exposure evaluation index as the second value of the exposure of conflict events within the current geographical grid.

8. The method according to claim 5, characterized in that The using of the MOOSRA multi-criteria decision-making model to calculate the values of each exposure index within each geographical grid, obtaining the third value of the exposure of conflict events within each geographical grid, includes: Calculating the maximum value, minimum value and average value of all geographical grids on each exposure index; Taking a positive ideal value between the maximum value and the average value of each exposure index, and forming a positive ideal solution from the positive ideal values of each exposure index; taking a negative ideal value between the average value and the minimum value of each exposure index, and forming a negative ideal solution from the negative ideal values of each exposure index; For each geographical grid, according to the values of each exposure index within the current geographical grid, calculate the distances between the current geographical grid and the positive ideal solution on each exposure evaluation index respectively, and then assign different weights to the distances on each exposure evaluation index and perform weighted summation to obtain the positive distance; calculate the distances between the current geographical grid and the negative ideal solution on each exposure evaluation index respectively, and then assign different weights to the distances on each exposure evaluation index and perform weighted summation to obtain the negative distance; average the positive distance and the negative distance to obtain the third value of the exposure of conflict events within the current geographical grid.

9. The method according to claim 1, wherein Generate the future conflict risk distribution map of the target area according to the vulnerability distribution, exposure distribution and fragility distribution, including: Calculate the conflict risk value of each geographical grid in a future year according to the following formula: Conflict risk space value = Vulnerability w1 × Exposure w2 × Fragility w3 Wherein, w1, w2 and w3 respectively represent the weights of vulnerability, exposure and fragility, and w1>w2, w1>w3; Generate the future conflict risk distribution map of the target area according to each conflict risk value.

10. An electronic device, characterized in that, Including: One or more processors; A memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the conflict event data processing method based on the improved Bayesian additive regression tree according to any one of claims 1-8.

Citation Information

Patent Citations

  • Crop heavy metal enrichment risk prediction method and system based on Bayesian theory

    CN115775042A

  • General form of the tree alternating optimization (TAO) for learning decision trees

    WO2020247949A1