Bias supervision method for risk assessment algorithm based on geological safety big data

By constructing the conditional risk effect rank matrix and rank augmentation matrix, the bias of the geological safety risk assessment algorithm is identified and corrected, which solves the problem of algorithm bias in the existing technology and improves the accuracy and reliability of risk assessment.

CN120410230BActive Publication Date: 2025-09-09江苏省地质局第一地质大队
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510905771.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-09-09
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

Existing geological safety risk assessment algorithms are biased, resulting in outputs that are inconsistent with reality, making them difficult to effectively monitor and correct, and affecting the accuracy and reliability of geological safety risk assessments.

Method used

By constructing the conditional risk effect rank matrix and rank augmentation matrix, the bias risk of the geological safety risk assessment algorithm is calculated, and hierarchical grouping and non-parametric testing are performed using multi-source heterogeneous data to identify and correct algorithm bias.

Benefits of technology

It provides a fact-based correction benchmark, improves the generalization ability and interpretability of geological safety risk assessment algorithms, reduces the risk of misjudgment, and helps competent authorities effectively implement smart management and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120410230B_ABST
    Figure CN120410230B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for monitoring bias in risk assessment algorithms based on geological safety big data, relating to the field of geological disaster monitoring technology. The method comprises obtaining a single disaster event case sample from a geological safety big data sample library for a target area, calculating the single-case geological safety risk value according to a geological safety risk assessment algorithm, constructing a conditional risk effect rank matrix using the disaster event case sample, and calculating the risk of bias in the geological safety risk assessment algorithm for the conditional risk effect. The method is applicable to different types of risk estimation conditions and can, based on geological safety big data, identify errors in the assessment of actual risk by the calculated risk derived from the geological safety risk assessment algorithm. The method monitors and quantitatively evaluates the risk of bias in the geological safety risk assessment algorithm, reduces the risk of misjudgment by the geological safety risk assessment algorithm, and improves the interpretability of the generalization capability of the geological safety risk assessment algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of geological disaster monitoring, and in particular to a bias supervision method for a risk assessment algorithm based on geological safety big data. Background Art

[0002] With the increasing concentration of population and industry, the construction and operation of infrastructure, especially urban lifeline projects, are becoming increasingly difficult. Geological safety is a prerequisite for the safety, economy, and sustainability of infrastructure construction and operation.

[0003] The prerequisite for geological safety risk management and control is the accurate quantitative assessment of geological safety risks. A widely used geological safety risk assessment method uses data on disaster risk factors and basic data on the target area as input, and then applies a geological safety risk assessment algorithm to obtain a geological safety risk assessment value for the target area. There are two typical types of geological safety risk assessment algorithms: one that assigns scores to risk factors and then explicitly weights them; the other that relies on implicit statistical learning methods. When there are errors in the assignment of scores or weights, the output of the scoring and weighting system algorithm will systematically exhibit an "output-to-actual inversion" problem, with large actual losses but small assessed risk values. This same "output-to-actual inversion" problem can also occur when the statistical learning model's factor variable characteristics or hyperparameters are improperly set. When this "output-to-actual inversion" occurs, the calculated risk derived from the algorithm incorrectly predicts actual losses, indicating bias in the assessment algorithm. Therefore, there is still a lack of data processing methods to effectively supervise the bias of geological safety risk assessment algorithms. This limitation makes it difficult for the competent authorities to supervise the bias in the output results of various geological safety risk assessment algorithms, and thus difficult to select low-bias geological safety risk assessment algorithms as the basis for geological safety risk evaluation. This leads to bias in geological safety risk assessment results and difficulties in selecting assessment method standards, which restricts the competent authorities from effectively implementing intelligent management and control of geological safety risks. Summary of the Invention

[0004] The present invention is proposed in view of the problems existing in existing risk assessment algorithm bias supervision methods based on geological safety big data. Therefore, the problem to be solved by the present invention is how to provide a risk assessment algorithm bias supervision method based on geological safety big data.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0006] In a first aspect, the present invention provides a bias supervision method for a risk assessment algorithm based on geological safety big data, which includes obtaining a single disaster event case sample from a geological safety big data sample library of a target area, and calculating a single case geological safety risk value according to a geological safety risk assessment algorithm;

[0007] The conditional risk effect rank matrix is ​​constructed through disaster event case samples;

[0008] The risk of bias of geological safety risk assessment algorithm to conditional risk effect is calculated based on the conditional risk effect rank augmentation matrix;

[0009] Rank the risks of conditional risk effect bias and complete bias supervision of risk assessment algorithms based on geological safety big data.

[0010] As a preferred solution of the bias supervision method of the risk assessment algorithm based on geological safety big data described in the present invention, the single disaster event case sample includes the disaster name, data on factors causing the disaster, basic data on the target area, data on casualties caused by the disaster, direct economic loss data and data on the impact on infrastructure operation.

[0011] As a preferred solution of the bias supervision method of the risk assessment algorithm based on geological safety big data of the present invention, the method of constructing the conditional risk effect rank matrix through disaster event case samples includes:

[0012] Obtain data on casualties, direct economic losses, and impacts on infrastructure operations caused by disasters in a sample of individual disaster events;

[0013] The total number of deaths caused by disasters is used as the representative value of personnel losses in a single case, and the number of cases in the geological safety big data sample database is calculated. The rank of the representative value of the single case loss of personnel , will rank The same cases are grouped into the same loss subgroup, which is numbered as , , is the total number of personnel loss subgroups;

[0014] The direct economic losses caused by the disaster are taken as the representative value of direct economic losses of a single case, and the cases are calculated within the scope of the single personnel loss subgroup. The rank of the representative value of direct economic loss of a single case ;

[0015] Within a single personnel loss subgroup, cases were clustered into direct economic loss subsets based on direct economic loss. Under the condition of single personnel loss subgroup is divided into The direct economic loss subsets are expressed as follows:

[0016] ;

[0017] in, for The median distance between adjacent direct economic loss subsets under the conditions is: for The median of the elements in the kth direct economic loss subset under the condition, Code the subset of direct economic losses within the individual human loss subgroups, is the target personnel loss subgroup number, 、 are all positive integers, , ;

[0018] Within a single subset of direct economic losses, the conditional risk effect rank matrix is ​​obtained.

[0019] As a preferred solution of the risk assessment algorithm bias supervision method based on geological safety big data of the present invention, wherein: the risk of the geological safety risk assessment algorithm bias for the conditional risk effect calculated according to the conditional risk effect rank augmentation matrix includes:

[0020] Input the data of the disaster-causing factors in the case and the basic data of the target area into the geological safety risk assessment algorithm, and output the geological safety risk value of the case;

[0021] For the geological safety risk assessment algorithm, the rank of the geological safety risk value of the cases in the direct economic loss subset is calculated to generate the conditional risk effect rank augmented matrix;

[0022] The elements in the first column of the conditional risk effect rank augmentation matrix are the ranks of the geological safety risk values;

[0023] When there are missing values ​​in the conditional risk effect rank matrix and the missing values ​​cannot be estimated, the columns from the second column to the last column in the conditional risk effect rank augmented matrix are the columns from the first column to the last column in the conditional risk effect rank augmented matrix that do not contain missing values;

[0024] When there are missing values ​​in the conditional risk effect rank matrix but the missing values ​​can be estimated, the corresponding missing values ​​are filled with the estimated missing values. The columns from the second column to the last column in the conditional risk effect rank augmented matrix are the columns from the first column to the last column in the conditional risk effect rank augmented matrix that do not contain missing values ​​after the corresponding missing values ​​are filled with the estimated missing values.

[0025] When there are no missing values ​​in the conditional risk effect rank matrix, the columns from the second column to the last column in the conditional risk effect rank augmented matrix are the columns from the first column to the last column in the conditional risk effect rank matrix corresponding to the conditional risk effect rank augmented matrix;

[0026] The risk of conditional risk effect bias in geological safety risk assessment algorithms is calculated based on the conditional risk effect rank augmentation matrix.

[0027] As a preferred embodiment of the method for bias supervision of risk assessment algorithms based on geological safety big data of the present invention, the method of calculating the risk of bias of the geological safety risk assessment algorithm for conditional risk effects according to the conditional risk effect rank augmentation matrix includes:

[0028] When the number of elements in the direct economic loss subset is less than 100, the elementary row transformation method is used to solve the rank of the conditional risk effect rank augmentation matrix. The calculation formula for the risk of conditional risk effect bias is:

[0029] ;

[0030] Where, Geological safety risk assessment algorithm For The risk of bias in the conditional risk effect for the kth subset of direct economic losses under the condition; is the rank of the conditional risk effect rank augmented matrix; for and The minimum value of the number of rows and columns of the conditional risk effect rank augmented matrix when both are fixed.

[0031] As a preferred embodiment of the method for bias supervision of risk assessment algorithms based on geological safety big data of the present invention, the method of calculating the risk of bias of the geological safety risk assessment algorithm for conditional risk effects according to the conditional risk effect rank augmentation matrix further comprises:

[0032] When the number of elements in the direct economic loss subset is not less than 100, the column vector consisting of the first column of the conditional risk effect rank augmented matrix is ​​subtracted from itself to obtain the basis value column vector;

[0033] The column vectors formed by the columns from the second to the last column in the conditional risk effect rank augmented matrix are respectively subtracted from the column vectors formed by the first column in the conditional risk effect rank augmented matrix to obtain the deviation column vectors corresponding to the columns from the second to the last column in the conditional risk effect rank augmented matrix;

[0034] The two-sample Kolmogorov-Smirnov hypothesis test is performed on the two samples consisting of the base value column vector element set and the single deviation column vector element set corresponding to the second to last columns in the conditional risk effect rank augmented matrix. The risk calculation formula for the conditional risk effect bias is:

[0035] ;

[0036] Where, Geological safety risk assessment algorithm For The risk of bias in the conditional risk effect for the kth subset of direct economic losses under the condition; for The number of times the two samples in the two-sample Kolmogorov-Smirnov hypothesis test come from the same population, is the number of columns in the conditional risk effect rank augmentation matrix.

[0037] As a preferred solution of the bias supervision method of the risk assessment algorithm based on geological safety big data of the present invention, the risk ranking of the conditional risk effect bias includes:

[0038] Calculate the risk of bias of the conditional risk effect of the geological safety risk assessment algorithm for different direct economic loss subsets in the same human loss subgroup, rank the risk of bias of the conditional risk effect, and determine the risk of bias of the geological safety risk assessment algorithm for different direct economic loss subsets;

[0039] Calculate the risk of conditional risk effect bias of different geological safety risk assessment algorithms for the same direct economic loss subset in the same personnel loss subgroup, rank the risk of conditional risk effect bias, and determine the risk of bias of different geological safety risk assessment algorithms for the same direct economic loss subset.

[0040] The beneficial effects of the present invention are as follows: the theoretical basis of the method is reliable and the derivation is clear, multi-source heterogeneous data achieves measurement unification through rank operations under hierarchical grouping conditions, matrix analysis and non-parametric tests are respectively applicable to different types of risk estimation conditions, and can identify the assessment errors of the calculated risk obtained by the geological safety risk assessment algorithm against the actual risk based on geological safety big data, providing a fact-based correction benchmark for supervising the output bias of the geological safety risk assessment algorithm, thereby supervising and quantitatively evaluating the risk of bias in the geological safety risk assessment algorithm, improving the interpretability of the generalization ability of the geological safety risk assessment algorithm, and helping to reduce the risk of misjudgment of the geological safety risk assessment algorithm, providing a reliable data processing method for the competent authorities to supervise the algorithm bias of various geological safety risk assessment algorithms, and providing an effective algorithm bias supervision means for the competent authorities to effectively implement intelligent management and control of geological safety risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0042] Figure 1 This is an example diagram of the data main structure of the geological safety big data sample library. DETAILED DESCRIPTION

[0043] To make the above-mentioned objects, features, and advantages of the present invention more easily understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.

[0044] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0045] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it constitute a separate or selective embodiment that is mutually exclusive with other embodiments.

[0046] Example 1

[0047] Reference Figure 1 Table 1 is the first embodiment of the present invention, which provides a bias supervision method for a risk assessment algorithm based on geological safety big data, including step 1, obtaining a single disaster event case sample in the geological safety big data sample library of the target area, and calculating the single case geological safety risk value according to the geological safety risk assessment algorithm.

[0048] Specifically, within the target area, case data related to geological safety are extracted and integrated from the databases of departments responsible for natural resources management, emergency management, municipal pipelines, road traffic, rail transit, urban and rural and housing construction, meteorology, water conservancy, finance, and national defense to obtain a geological safety big data sample library for the target area. The individual disaster event case samples in the geological safety big data sample library include the name of the disaster, data on factors leading to the occurrence of the disaster, basic data of the target area, data on casualties caused by the disaster, data on direct economic losses caused by the disaster, and data on the impact of the disaster on infrastructure operations. The basic data of the target area include population density, proportion of housing area, infrastructure density, and economic density.

[0049] The case The data of disaster-causing factors and basic data of the target area are used as input, and the geological safety risk assessment algorithm is used to , output case Single case geological safety risk value ;

[0050] Case Code , the number of cases in the geological safety big data sample library is , the case code of a single case is unique.

[0051] Table 1 shows the data structure of direct economic loss subset 1 in personnel loss subgroup 1.

[0052]

[0053] The labels in Table 1 mean: 1) direct economic loss subset code in a single loss of life subgroup; 2) case code; 3) disaster name; 4) data on factors causing the disaster; 5) factor 1 causing the disaster; 6) factor 2 causing the disaster; 7) basic data of the target area; 8) population density; 9) proportion of housing area; 10) infrastructure density; 11) economic density; 12) rank of geological safety risk value of cases in the direct economic loss subset; 13) total number of deaths caused by the disaster; 14) direct economic losses caused by the disaster; 15) data on the impact of the disaster on infrastructure operation; 16) infrastructure type 1 (for example, the types of transport media include personnel, vehicles, gas, and digital information); 17) rank of the affected characteristics of infrastructure function indicators 1; 18) infrastructure 1) Rank of affected factors of infrastructure function indicator 2; 2) Rank of affected factors of infrastructure function indicator 3; 3) Rank of affected factors of infrastructure function indicator 4; 4) Rank of affected factors of infrastructure function indicator 5; 5) Rank of affected factors of infrastructure function indicator 6; 6) Rank of affected factors of infrastructure function indicator 1; 7) Rank of affected factors of infrastructure function indicator 2; 8) Rank of affected factors of infrastructure function indicator 3; 9) Rank of affected factors of infrastructure function indicator 6; ** indicates non-missing values ​​filled in according to actual situation.

[0054] Step 2: Construct the conditional risk effect rank matrix through disaster event case samples.

[0055] Specifically, the loss composition in the case is diverse, and it is necessary to perform hierarchical clustering on two of the three actual loss levels: casualties caused by disasters, direct economic losses caused by disasters, and impacts of disasters on infrastructure operations, so as to ensure that the final case set division realizes control variables for two of the three hierarchical variables: casualties caused by disasters, direct economic losses caused by disasters, and impacts of disasters on infrastructure operations, so as to ensure that the actual loss level variable used for comparison with the calculated risk variable obtained according to the geological safety risk assessment algorithm is screened under the condition that the other two actual loss level variables are fixed, thereby eliminating the influence of confounding factors (i.e., the two level variables used for hierarchical clustering) on ​​the results in subsequent non-parametric statistics and matrix analysis;

[0056] For a single case, the total number of deaths caused by the disaster is calculated as the representative value of the loss of personnel in the single case, and the number of cases in the geological safety big data sample library is calculated. The rank of the representative value of the single case loss of personnel ,Will The same cases are grouped into the same loss subgroup, which is numbered as , , is the total number of personnel loss subgroups, , The larger the Corresponding The bigger;

[0057] The direct economic loss caused by the disaster in a single case is taken as the representative value of the direct economic loss of a single case, and the case is calculated within the scope of the single personnel loss subgroup. The rank of the representative value of direct economic loss of a single case ;

[0058] Within a single personnel loss subgroup, cases with similar direct economic losses are clustered into the same direct economic loss subset, thus Under the condition of single personnel loss subgroup is divided into A subset of direct economic losses, and is a positive integer, and the clustering method is shown in the optimization problem in formula (1):

[0059] (1)

[0060] in, Indicates All conditions The largest positive integer in , for The median distance between adjacent direct economic loss subsets under the conditions is: for The median of the elements in the kth direct economic loss subset under the condition, for All within the scope The algebraic sum of 、 All are positive integers, the target personnel loss subgroup number , direct economic loss subset coding in a single personnel loss subgroup ;

[0061] Within a single subset of direct economic losses, the conditional risk effect rank matrix is ​​obtained;

[0062] The data on the impact of disasters on infrastructure operations include the ranks of the affected characteristics of the infrastructure function indicators and all data used to calculate the ranks of the affected characteristics of the infrastructure function indicators;

[0063] When the number of infrastructure in a single type of infrastructure is 1, the single indicator impact ratio is the absolute value of the difference between the representative value of a single functional indicator of a single infrastructure after the disaster and the representative value of the indicator before the disaster and 1;

[0064] The representative value of a single functional indicator before a disaster is the sample mean of the indicator on the time scale from the official operation of the infrastructure to the last time the representative value was recorded before the disaster occurred. The representative value of a single functional indicator after a disaster is the sample mean of the indicator on the time scale from the disaster occurrence to the last time the representative value was recorded before emergency repair measures were taken on the damaged parts of the infrastructure. The infrastructure functional indicator is the flow rate of various types of transport media carried by the infrastructure, which include personnel, vehicles, liquids, electricity, gas, and digital information.

[0065] For Cases in the kth direct economic loss subset under the condition , P is the total number of infrastructure types, p is the infrastructure type code, Q is the total number of types of infrastructure function indicators of a single type of infrastructure, when the number of infrastructure in a single type of infrastructure exceeds 1, q is the code of the infrastructure function indicator of a single type of infrastructure, when the number of infrastructure in a single type of infrastructure is 1, q is the code of the infrastructure function indicator of a single infrastructure;

[0066] In the target area Cases in the kth direct economic loss subset under the condition The assignment rule of p is: the types of transport media of a single type of infrastructure include one, two, three, four, five or six of the following: personnel, transportation, liquid, electricity, gas, and digital information. The types and quantities of transport media of the same type of infrastructure are the same, and the types of transport media of different types of infrastructure are different. Each positive integer in [1, P] is mapped one by one to each infrastructure type code p. The assignment rule of q corresponding to a single p is: the codes of the functional indicators of personnel, transportation, liquid, electricity, gas, and digital information infrastructure are assigned to 1, 2, 3, 4, 5, and 6 in sequence. When p and q are fixed, within the target area Cases in the kth direct economic loss subset under the condition The rank of the affected characteristics of the qth infrastructure function indicator of the pth type of infrastructure , p∈[1,P], q∈[1,Q];

[0067] When the number of infrastructure in a single type of infrastructure is 1, for the target area Cases in the kth direct economic loss subset under the condition ,The affected characteristics of the qth infrastructure function indicator of the pth infrastructure are as follows: the single indicator impact ratio of the qth infrastructure function indicator of the infrastructure in the pth infrastructure;

[0068] When the number of infrastructures in a single type of infrastructure exceeds 1, Cases in the kth direct economic loss subset under the condition , the affected characteristics of the qth infrastructure function indicator of the pth type of infrastructure are: the single indicator comprehensive impact ratio DZBZHB of the qth infrastructure function indicators of all infrastructures in the pth type of infrastructure, DZBZHB is the absolute value of the difference between ZaH divided by ZbH and 1; ZaH is the algebraic sum of the representative index values ​​of the qth infrastructure function of all individual infrastructures in the pth type of infrastructure after the disaster, and ZbH is the algebraic sum of the representative index values ​​of the qth infrastructure function of all individual infrastructures in the pth type of infrastructure before the disaster.

[0069] Conditional risk effect rank matrix In Each case in the kth direct economic loss subset under the condition The rows of the conditional risk effect rank matrix are elements of Arrange in ascending order, the columns of the conditional risk effect rank matrix are arranged in ascending order of q under the premise that the elements are arranged in ascending order of p, and the number of columns of a single conditional risk effect rank matrix is ​​equal to P multiplied by Q.

[0070] Step 3: Calculate the risk of bias of the geological safety risk assessment algorithm against the conditional risk effect based on the conditional risk effect rank augmentation matrix.

[0071] Specifically, in In the kth direct economic loss subset under the condition, perform this step operation;

[0072] For the geological safety risk assessment algorithm h, the calculation Cases in the kth direct economic loss subset under the condition Geological safety risk value Rank ;

[0073] For the geological safety risk assessment algorithm h, the conditional risk effect rank augmented matrix is ​​generated , the elements in the first column of the conditional risk effect rank augmentation matrix are , the rows of the conditional risk effect rank augmentation matrix are based on the elements Arrange in ascending order, the elements in the same row of the conditional risk effect rank augmentation matrix same;

[0074] When there are missing values ​​in the conditional risk effect rank matrix and the missing values ​​cannot be estimated, the columns from the second column to the last column of the conditional risk effect rank augmented matrix are the columns from the first column to the last column of the conditional risk effect rank matrix corresponding to the conditional risk effect rank augmented matrix that do not contain missing values;

[0075] When there are missing values ​​in the conditional risk effect rank matrix but the missing values ​​can be estimated, the corresponding missing values ​​should be filled with the estimated missing values. The columns from the second column to the last column of the conditional risk effect rank augmented matrix are the columns from the first column to the last column of the conditional risk effect rank augmented matrix that do not contain missing values ​​after the missing values ​​are filled with the estimated missing values.

[0076] When there are no missing values ​​in the conditional risk effect rank matrix, the columns ranging from the second column to the last column of the conditional risk effect rank augmented matrix are respectively the columns ranging from the first column to the last column of the conditional risk effect rank matrix corresponding to the conditional risk effect rank augmented matrix.

[0077] (1) When When the number of elements in the kth direct economic loss subset is less than 100 under the following conditions:

[0078] When the linear correlation degree of each row in the conditional risk effect rank augmented matrix is ​​higher, the linear correlation degree of the first column of the conditional risk effect rank augmented matrix and the columns from the second column to the last column of the conditional risk effect rank augmented matrix is ​​higher, then the geological safety risk assessment algorithm h used as the calculation basis for the first column of the conditional risk effect rank augmented matrix is ​​more accurate for the situation in The lower the risk of bias in the conditional risk effect of the kth subset of direct economic losses under the condition, the lower the degree to which the calculated risk incorrectly reflects the actual losses.

[0079] Using elementary row transformation method to solve the rank of conditional risk effect rank augmented matrix , then the geological safety risk assessment algorithm h is The risk of bias in the conditional risk effect for the kth subset of direct economic losses conditional on As shown in formula (2):

[0080] (2)

[0081] Where, for The minimum value of the number of rows and columns of the conditional risk effect rank augmented matrix when and are both fixed;

[0082] (2) When When the number of elements in the kth direct economic loss subset is not less than 100:

[0083] Subtract the column vector consisting of the first column of the conditional risk effect rank augmented matrix from itself to obtain the basis value column vector;

[0084] Subtract the column vectors formed by the first column of the conditional risk effect rank augmented matrix from the column vectors formed by the second to the last columns of the conditional risk effect rank augmented matrix to obtain the deviation column vectors corresponding to the columns from the second to the last columns of the conditional risk effect rank augmented matrix;

[0085] The double-sample Kolmogorov-Smirnov hypothesis test is performed on the double-sample consisting of the base value column vector element set and the single deviation column vector element set corresponding to the second to the last column of the conditional risk effect rank augmented matrix. The geological safety risk assessment algorithm h is The risk of bias in the conditional risk effect for the kth subset of direct economic losses conditional on As shown in formula (3):

[0086] (3)

[0087] Where, for The number of times the two samples in the two-sample Kolmogorov-Smirnov hypothesis test come from the same population, is the number of columns in the conditional risk effect rank augmentation matrix.

[0088] Step 4: Rank the risks of conditional risk effect bias and complete the bias supervision of the risk assessment algorithm based on geological safety big data.

[0089] Specifically, the risk of bias of the conditional risk effect of the geological safety risk assessment algorithm for different direct economic loss subsets in the same personnel loss subgroup is calculated, and the risk of bias of the conditional risk effect is ranked, so as to determine the risk of bias of the geological safety risk assessment algorithm for different direct economic loss subsets;

[0090] Calculated The larger the value is, the more effective the geological safety risk assessment algorithm h is. The higher the risk of conditional risk effect bias for the kth subset of direct economic losses under the condition;

[0091] Calculate the risk of conditional risk effect bias of different geological safety risk assessment algorithms for the same direct economic loss subset in the same personnel loss subgroup, rank the risk of conditional risk effect bias, and thus determine the risk of bias of different geological safety risk assessment algorithms for the same direct economic loss subset;

[0092] When each algorithm in the geological safety risk assessment algorithm set H is The kth direct economic loss subset is conducted under the condition When solving, The larger the value of the element, the more accurate the geological safety risk assessment algorithm corresponding to the element. The higher the risk of bias in the kth direct economic loss subset under the condition.

[0093] Example 2

[0094] As shown in Table 1, when ζ=1 and i=1, and C1=2, the first direct economic loss subset includes Case 68, Case 23, and Case 45. When p=1, since the transmission medium only includes personnel, vehicles, gas, and digital information but not liquid and electricity, the columns corresponding to liquid and electricity, that is, the columns corresponding to q=3 and q=4, are both missing values. Since there is no first-class infrastructure in Case 45, the cells in the row of Case 45 when p=1 are all missing values.

[0095] Example 3

[0096] For the geological safety risk assessment algorithm h, if The three cases in the kth direct economic loss subset under the condition are case 403, case 19, and case 26, and the missing values ​​cannot be estimated, and 、 、 , then in Under the condition that the number of elements in the kth direct economic loss subset is 3, 3 < 100, if we have formula (4), then we have formula (5),

[0097] (4)

[0098] (5)

[0099] Using elementary row transformation method to solve the rank of conditional risk effect rank augmented matrix , , according to formula (2), we can get ;

[0100] In this subset, the calculated risk variables are ranked according to the geological safety risk assessment algorithm h, i.e., the column vector [1.5 3 1.5] T , which is completely inconsistent with the actual impact ranking of infrastructure function indicators under the conditions of the same personnel losses and similar direct economic losses, that is, the column vectors in the conditional risk effect rank augmentation matrix except the first column. That is, the calculated risk obtained by algorithm h incorrectly predicts the actual loss, and algorithm h has a bias risk for the cases in this subset.

[0101] Example 4

[0102] For the geological safety risk assessment algorithm h, if The three cases in the kth direct economic loss subset under the condition are Case 2, Case 61 and Case 88, and 、 、 , then in Under the condition that the number of elements in the kth direct economic loss subset is 3, 3 < 100, if we have formula (6), we have formula (7):

[0103] (6)

[0104] (7)

[0105] Using elementary row transformation method to solve the rank of conditional risk effect rank augmented matrix , , according to formula (2), we can get ;

[0106] In this subset, the calculated risk variables are ranked according to the geological safety risk assessment algorithm, i.e., the column vector [1 2 3] T, which is completely consistent with the actual impact ranking of infrastructure function indicators under the conditions of the same personnel loss and similar direct economic losses, that is, the column vectors in the conditional risk effect rank augmentation matrix except the first column, that is, the calculated risk obtained by algorithm h correctly predicts the actual loss, and algorithm h has no biased risk for the cases in this subset.

[0107] Example 5

[0108] When Under the condition that the number of elements in the kth direct economic loss subset is 109, 109>100, if is 4, and the results of the two-sample Kolmogorov-Smirnov hypothesis test on the basis value column vector element set and the deviation column vector element set corresponding to the 2nd, 3rd, and 4th columns of the conditional risk effect rank augmented matrix are respectively:

[0109] Accept the null hypothesis that the base value column vector element set and the deviation column vector element set corresponding to the second column of the conditional risk effect rank augmented matrix come from the same population, and the alternative hypothesis that the base value column vector element set and the deviation column vector element set corresponding to the second column of the conditional risk effect rank augmented matrix do not come from the same population;

[0110] Accept the null hypothesis that the base value column vector element set and the deviation column vector element set corresponding to the third column of the conditional risk effect rank augmented matrix come from the same population, and the alternative hypothesis that the base value column vector element set and the deviation column vector element set corresponding to the third column of the conditional risk effect rank augmented matrix do not come from the same population;

[0111] Under the null hypothesis that the base value column vector element set and the deviation column vector element set corresponding to the 4th column of the conditional risk effect rank augmented matrix are from the same population, and the alternative hypothesis that the base value column vector element set and the deviation column vector element set corresponding to the 4th column of the conditional risk effect rank augmented matrix are not from the same population, the null hypothesis is rejected;

[0112] but is 2, so .

[0113] Example 6

[0114] If each algorithm in the geological safety risk assessment algorithm set H is The kth direct economic loss subset is conducted under the condition When solving, we get for ,but:

[0115] The corresponding geological safety risk assessment algorithm The risk of bias in the kth direct economic loss subset is the highest under the condition

[0116] The corresponding geological safety risk assessment algorithm Under the condition, the risk of bias in the kth direct economic loss subset is second,

[0117] The corresponding geological safety risk assessment algorithm The risk of bias in the kth subset of direct economic losses under the condition is the lowest.

[0118] In summary, the theoretical basis of this method is reliable and the derivation is clear. Multi-source heterogeneous data achieves measurement unification through rank operations under hierarchical grouping conditions. Matrix analysis and non-parametric tests are applicable to different types of risk estimation conditions respectively. It can identify the assessment errors of the calculated risk obtained by the geological safety risk assessment algorithm against the actual risk based on geological safety big data, and provide a fact-based correction benchmark for supervising the output bias of the geological safety risk assessment algorithm, thereby supervising and quantitatively evaluating the risk of bias in the geological safety risk assessment algorithm, improving the interpretability of the generalization ability of the geological safety risk assessment algorithm, and helping to reduce the risk of misjudgment of the geological safety risk assessment algorithm. It provides a reliable data processing method for the competent authorities to supervise the algorithmic bias of various geological safety risk assessment algorithms, and provides an effective algorithmic bias supervision means for the competent authorities to effectively implement intelligent management and control of geological safety risks.

[0119] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A bias supervision method for risk assessment algorithms based on geological safety big data, characterized by: include, Obtain a single disaster event case sample from the geological safety big data sample library of the target area, and calculate the geological safety risk value of the single case according to the geological safety risk assessment algorithm; The conditional risk effect rank matrix is ​​constructed through disaster event case samples; The construction of the conditional risk effect rank matrix through disaster event case samples includes: Obtain data on casualties, direct economic losses, and impacts on infrastructure operations caused by disasters in a sample of individual disaster events; The total number of deaths caused by disasters is used as the representative value of personnel losses in a single case, and the number of cases in the geological safety big data sample database is calculated. The rank of the representative value of the single case loss of personnel , will rank The same cases are grouped into the same loss subgroup, which is numbered as , , is the total number of personnel loss subgroups; The direct economic losses caused by the disaster are taken as the representative value of direct economic losses of a single case, and the cases are calculated within the scope of the single personnel loss subgroup. The rank of the representative value of direct economic loss of a single case ; Within a single personnel loss subgroup, cases were clustered into direct economic loss subsets based on direct economic loss. Under the condition of single personnel loss subgroup is divided into The optimization problem in the clustering method is expressed as follows: in, for The median distance between adjacent direct economic loss subsets under the conditions is: for The median of the elements in the kth direct economic loss subset under the condition, Code the subset of direct economic losses within the individual human loss subgroups, is the target personnel loss subgroup number, 、 are all positive integers, , ; Within a single subset of direct economic losses, the conditional risk effect rank matrix is ​​obtained; The risk of bias of geological safety risk assessment algorithm to conditional risk effect is calculated based on the conditional risk effect rank augmentation matrix; Ranking the risks of conditional risk effect bias and completing bias supervision of risk assessment algorithms based on geological safety big data; The ranking of risk of conditional risk effect bias includes: Calculate the risk of bias of the conditional risk effect of the geological safety risk assessment algorithm for different direct economic loss subsets in the same human loss subgroup, rank the risk of bias of the conditional risk effect, and determine the risk of bias of the geological safety risk assessment algorithm for different direct economic loss subsets; Calculate the risk of conditional risk effect bias of different geological safety risk assessment algorithms for the same direct economic loss subset in the same personnel loss subgroup, rank the risk of conditional risk effect bias, and determine the risk of bias of different geological safety risk assessment algorithms for the same direct economic loss subset.

2. The method for bias supervision of risk assessment algorithms based on geological safety big data according to claim 1, characterized in that: The sample of single disaster event cases includes the name of the disaster, data on factors causing the disaster, basic data on the target area, as well as data on casualties caused by the disaster, direct economic losses, and impact on infrastructure operations.

3. The method for bias supervision of risk assessment algorithms based on geological safety big data according to claim 2, characterized in that: The risk of bias in the conditional risk effect of the geological safety risk assessment algorithm calculated based on the conditional risk effect rank augmentation matrix includes: Input the data of the disaster-causing factors in the case and the basic data of the target area into the geological safety risk assessment algorithm, and output the geological safety risk value of the case; For the geological safety risk assessment algorithm, the rank of the geological safety risk value of the cases in the direct economic loss subset is calculated to generate the conditional risk effect rank augmented matrix; The elements in the first column of the conditional risk effect rank augmentation matrix are the ranks of the geological safety risk values; When there are missing values ​​in the conditional risk effect rank matrix and the missing values ​​cannot be estimated, the columns from the second column to the last column in the conditional risk effect rank augmented matrix are the columns from the first column to the last column in the conditional risk effect rank augmented matrix that do not contain missing values; When there are missing values ​​in the conditional risk effect rank matrix but the missing values ​​can be estimated, the corresponding missing values ​​are filled with the estimated missing values. The columns from the second column to the last column in the conditional risk effect rank augmented matrix are the columns from the first column to the last column in the conditional risk effect rank augmented matrix that do not contain missing values ​​after the corresponding missing values ​​are filled with the estimated missing values. When there are no missing values ​​in the conditional risk effect rank matrix, the columns from the second column to the last column in the conditional risk effect rank augmented matrix are the columns from the first column to the last column in the conditional risk effect rank matrix corresponding to the conditional risk effect rank augmented matrix; The risk of conditional risk effect bias in geological safety risk assessment algorithms is calculated based on the conditional risk effect rank augmentation matrix.

4. The method for bias supervision of risk assessment algorithms based on geological safety big data according to claim 3, characterized in that: The risk of bias of the geological safety risk assessment algorithm for conditional risk effects calculated according to the conditional risk effect rank augmentation matrix includes: When the number of elements in the direct economic loss subset is less than 100, the elementary row transformation method is used to solve the rank of the conditional risk effect rank augmentation matrix. The calculation formula for the risk of conditional risk effect bias is: Where, Geological safety risk assessment algorithm For The risk of bias in the conditional risk effect for the kth subset of direct economic losses under the condition; is the rank of the conditional risk effect rank augmented matrix; for and The minimum value of the number of rows and columns of the conditional risk effect rank augmented matrix when both are fixed.

5. The method for bias supervision of risk assessment algorithms based on geological safety big data according to claim 4, characterized in that: The risk of bias of the geological safety risk assessment algorithm for conditional risk effects calculated according to the conditional risk effect rank augmentation matrix also includes: When the number of elements in the direct economic loss subset is not less than 100, the column vector consisting of the first column of the conditional risk effect rank augmented matrix is ​​subtracted from itself to obtain the basis value column vector; The column vectors formed by the columns from the second to the last column in the conditional risk effect rank augmented matrix are respectively subtracted from the column vectors formed by the first column in the conditional risk effect rank augmented matrix to obtain the deviation column vectors corresponding to the columns from the second to the last column in the conditional risk effect rank augmented matrix; The two-sample Kolmogorov-Smirnov hypothesis test is performed on the two samples consisting of the base value column vector element set and the single deviation column vector element set corresponding to the second to last columns in the conditional risk effect rank augmented matrix. The risk calculation formula for the conditional risk effect bias is: Where, Geological safety risk assessment algorithm For The risk of bias in the conditional risk effect for the kth subset of direct economic losses under the condition; for The number of times the two samples in the two-sample Kolmogorov-Smirnov hypothesis test come from the same population, is the number of columns in the conditional risk effect rank augmentation matrix.

Citation Information

Patent Citations

  • Geological disaster risk assessment method and system based on fuzzy comprehensive evaluation method

    CN117788245A