A method and system for monitoring comprehensive utilization of hazardous waste
By modifying the Shapley value method and combining the Graph Laplace matrix and the phase synchronization index, the problem of neglecting indicator synergy and temporal synchronization in hazardous waste monitoring by the traditional Shapley value is solved, and more accurate contribution assessment and risk assessment are achieved.
Patent Information
- Application Number
- CN202511431582.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-10-09
AI Technical Summary
The traditional Shapley value method fails to effectively consider the synergy and temporal synchronization of monitoring indicators in the comprehensive utilization monitoring of hazardous waste, resulting in inaccurate contribution assessment.
By calculating the mutual information and phase synchronization index among monitoring indicators, the Shapley value method is modified to obtain improved Shapley values for monitoring indicators. Combined with the graphical Laplacian matrix and the phase synchronization index, the correlation strength and time series change trend among indicators are reflected.
This improves the accuracy and interpretability of the contribution assessment of monitoring indicators, providing a reliable basis for risk tracing and control decisions in the comprehensive utilization of hazardous waste.
Smart Images

Figure CN120931162B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric data processing, and in particular to a hazardous waste comprehensive utilization monitoring method and system. BACKGROUND
[0002] In order to ensure production safety, prevent environmental pollution and improve resource efficiency, it is necessary to monitor and assess the risk of hazardous waste comprehensive utilization process in real time. The existing monitoring methods mostly rely on setting threshold alarms for a single or a few key indicators. When the system is in a sub-health or risk budding state, the single indicator may still fluctuate within the normal range, resulting in missed reports.
[0003] In recent years, although multi-index fusion risk assessment models (such as machine learning methods based on neural networks or support vector machines) have been introduced, they have improved the accuracy of risk identification to some extent. However, the "black box" nature of the model makes it less interpretable, which is not conducive to targeted control by operators.
[0004] The Shapley value method derived from cooperative game theory has been gradually applied to the contribution analysis of multi-index systems due to its unique, fair and marginal contribution summation properties. The core idea is to regard each monitoring indicator as a participant and the monitoring system composed of all monitoring indicators as a cooperative alliance. By calculating the expected value of the marginal contribution brought by each indicator to different sub-alliances, the total risk of the system is allocated fairly.
[0005] However, the traditional Shapley value has some limitations in calculating marginal contribution: the Shapley value only calculates the independent marginal contribution of the indicator, without considering the synergistic or antagonistic effects that may occur after the combination of multiple indicators, resulting in an evaluation result deviating from the actual system behavior; monitoring data exists in the form of time series, and there are dynamic correlations such as trend synchronization, phase lag or cycle coupling between indicators. The traditional Shapley value model lacks sensitivity to such time series correlations, thus affecting the accuracy of contribution assessment. SUMMARY
[0006] The present application provides a hazardous waste comprehensive utilization monitoring method and system to solve the technical problem of inaccurate contribution assessment caused by the traditional Shapley value ignoring indicator synergy and time series synchronization.
[0007] In a first aspect, the present application provides a hazardous waste comprehensive utilization monitoring method, comprising the following steps:
[0008] S1, obtain time series data of N monitoring indicators, input time series data of any subset S of the N monitoring indicators into n basic evaluation models, obtain n risk state probability distributions, and calculate respective Shannon entropies to obtain an uncertainty vector, determine an n-dimensional weight vector according to a system risk preference, and obtain a characteristic function value of the subset S by calculating an inner product of the uncertainty vector and the weight vector;
[0009] S2, calculate mutual information between monitoring indicators according to time series data of the N monitoring indicators to obtain a graph Laplacian matrix; extract a subgraph Laplacian matrix corresponding to the subset S, and take a second smallest eigenvalue of the subgraph Laplacian matrix as a correlation coefficient of the subset S;
[0010] S3, for any monitoring indicator i and any subset S not containing the monitoring indicator i, calculate a marginal contribution, fuse time series of monitoring indicators in the subset S into a comprehensive time series of the subset S, and calculate a phase synchronization index between a time series of the monitoring indicator i and the comprehensive time series; multiply the marginal contribution, the correlation coefficient of the subset S, and the phase synchronization index to obtain a modified marginal contribution;
[0011] S4, according to a Shapley value formula, weighted sum the modified marginal contributions of all subsets S not containing the monitoring indicator i to obtain an improved Shapley value of the monitoring indicator i.
[0012] Further, in S1, the process of obtaining the characteristic function value of the subset S includes the following steps:
[0013] S11, input time series data of the subset S into three basic evaluation models of a long short-term memory network, a gated recurrent unit and a support vector machine, and each model outputs a probability distribution about three states of {low risk, medium risk, high risk};
[0014] S12, calculate Shannon entropy of each probability distribution to obtain three entropy values , which constitute a three-dimensional uncertainty vector ;
[0015] S13, according to the system risk preference, set a weight vector W for the three basic evaluation models;
[0016] S14, obtain the characteristic function value of the subset S by calculating an inner product of the uncertainty vector U and the weight vector W .
[0017] Further, in S2, the process of taking the second smallest eigenvalue of the subgraph Laplacian matrix as the correlation coefficient of the subset S includes the following steps:
[0018] S21. For any two monitoring indicators i and j among the N monitoring indicators, calculate the mutual information value MI(i,j) between their time series, and use the mutual information value MI(i,j) as an element of the adjacency matrix A. Construct an N×N adjacency matrix A;
[0019] S22, calculate the degree matrix D, where the degree matrix D is a diagonal matrix, and its diagonal elements are... Let A be the sum of all elements in the i-th row of the adjacency matrix A;
[0020] S23, construct the graph Laplace matrix L according to the formula L=DA;
[0021] S24, extract the rows and columns corresponding to each monitoring indicator in subset S from the graph Laplacian matrix L to form a subgraph Laplacian matrix. ;
[0022] S25, when subset S contains at least two monitoring indicators, calculate We collect all the feature values and sort them, then take the second smallest feature value as the correlation coefficient of subset S. When subset S contains only one monitoring indicator, the correlation coefficient is... The value is 1.
[0023] Furthermore, in S3, the process of calculating the phase synchronization index includes the following steps:
[0024] When subset S is an empty set, the phase synchronization index C(i, S) is 1;
[0025] When the subset S is a non-empty set, the phase synchronization index C(i, S) is calculated through the following steps:
[0026] S31, using principal component analysis, integrates the time series data of each monitoring indicator in subset S into a comprehensive time series, and selects the first principal component as the comprehensive time series;
[0027] S32, apply Hilbert transform to the time series of monitoring indicator i and the comprehensive time series of subset S respectively, and extract the instantaneous phase at their respective time points t. and ;
[0028] S33, Calculate the instantaneous phase difference ;
[0029] S34, the phase synchronization index is obtained by calculating the modulus of the average value of the complex exponent over time. : Where j represents the imaginary unit, satisfying , This represents the average value over time t.
[0030] Furthermore, in S4, the process of obtaining the improved Shapley value of monitoring indicator i includes the following steps:
[0031] For any subset S that does not contain monitoring indicator i, the corrected marginal contribution is calculated as follows. : ;
[0032] in, and Let S be the characteristic functions of subsets S that contain and exclude monitoring index i, respectively. A subset containing monitoring indicator i The relevant numerical values, The phase synchronization index between the monitoring index i and the subset S;
[0033] The improved Shapley value of monitoring indicator i is obtained by weighted summing of the modified marginal contributions of all subsets S that do not contain monitoring indicator i according to the following formula. : ;
[0034] Where |S| is the number of monitoring indicators in subset S, and N is the total number of monitoring indicators. It is any subset of the set of all monitoring indicators except for monitoring indicator i.
[0035] Secondly, the present invention provides a monitoring system for the comprehensive utilization of hazardous waste, comprising the following modules:
[0036] The feature function value calculation module is used to acquire time series data of N monitoring indicators, input the time series data of any subset S of the N monitoring indicators into n basic evaluation models, obtain n risk state probability distributions, calculate the Shannon entropy of each to obtain the uncertainty vector, determine the n-dimensional weight vector according to the system risk preference, and obtain the feature function value of subset S by calculating the inner product of the uncertainty vector and the weight vector.
[0037] The correlation coefficient calculation module is used to calculate the mutual information between monitoring indicators based on the time series data of N monitoring indicators to obtain the graph Laplacian matrix; extract the subgraph Laplacian matrix corresponding to subset S, and calculate the second smallest eigenvalue of the subgraph Laplacian matrix as the correlation coefficient of subset S;
[0038] a marginal contribution calculation module, configured to calculate a marginal contribution for any monitoring index i and any subset S not containing the monitoring index i, fuse time series of each monitoring index in the subset S into a comprehensive time series of the subset S, and calculate a phase synchronization index between a time series of the monitoring index i and the comprehensive time series; and multiply the marginal contribution, a correlation coefficient of the subset S, and the phase synchronization index to obtain a revised marginal contribution;
[0039] an optimization module, configured to perform weighted summation on the revised marginal contributions according to a Shapley value formula for all subsets S not containing the monitoring index i, to obtain an improved Shapley value of the monitoring index i.
[0040] Further, in the characteristic function value calculation module, when obtaining the characteristic function value of the subset S, time series data of the subset S are input into three basic evaluation models of a long short-term memory network, a gated recurrent unit, and a support vector machine, each model outputs a probability distribution about three states of {low risk, medium risk, high risk};
[0041] Shannon entropy of each probability distribution is calculated to obtain three entropy values , which constitute a three-dimensional uncertainty vector ;
[0042] According to a system risk preference, a weight vector W is set for the three basic evaluation models;
[0043] The characteristic function value of the subset S is obtained by calculating an inner product of the uncertainty vector U and the weight vector W .
[0044] Further, in the correlation coefficient calculation module, when the second smallest eigenvalue of the subgraph Laplacian matrix is taken as the correlation coefficient of the subset S, mutual information values MI(i, j) between time series of any two monitoring indices i and j in the N monitoring indices are calculated, and the mutual information values MI(i, j) are taken as elements of an adjacency matrix A to construct an N×N adjacency matrix A;
[0045] A degree matrix D is calculated, wherein the degree matrix D is a diagonal matrix, diagonal elements of the degree matrix D are sums of all elements in the i-th row of the adjacency matrix A;
[0046] A graph Laplacian matrix L is constructed according to a formula L=D-A;
[0047] Rows and columns corresponding to monitoring indices in the subset S are extracted from the graph Laplacian matrix L to construct a subgraph Laplacian matrix ;
[0048] When the subset S contains at least two monitoring indices, a value of all eigenvalues of the characteristic function of the subset S and rank them, and take the second smallest eigenvalue as the correlation index of the subset S ; when the subset S contains only one monitoring index, the correlation index is 1.
[0049] Further, in the marginal contribution calculation module, when calculating the phase synchronization index C(i, S), if the subset S is an empty set, the phase synchronization index C(i, S) is 1;
[0050] When the subset S is a non-empty set, the phase synchronization index C(i, S) is calculated by the following steps:
[0051] S31, using the principal component analysis method, the time series data of each monitoring index in the subset S is fused into a comprehensive time series, and the first principal component is selected as the comprehensive time series;
[0052] S32, the Hilbert transform is applied to the time series of the monitoring index i and the comprehensive time series of the subset S respectively, and the instantaneous phases and at the time point t are extracted;
[0053] S33, the instantaneous phase difference is calculated;
[0054] S34, the phase synchronization index C(i, S) is obtained by calculating the modulus of the average value of the complex number in time: where j represents the imaginary unit, and satisfies , represents the average value at time t.
[0055] Further, in the optimization module, when obtaining the improved Shapley value of the monitoring index i, for any subset S that does not contain the monitoring index i, the following calculation is used to correct the marginal contribution : ;
[0056] where and are the characteristic functions of the subset S containing and not containing the monitoring index i, respectively, is the correlation index value of the subset containing the monitoring index i, is the phase synchronization index of the monitoring index i and the subset S;
[0057] The improved Shapley value of the monitoring index i is obtained by weighted summing the corrected marginal contributions of all subsets S not containing the monitoring index i according to the following formula: : ;
[0058] Wherein, |S| is the number of monitoring indicators in the subset S, N is the total number of monitoring indicators, is any one subset of the set composed of all other monitoring indicators except monitoring indicator i.
[0059] Beneficial effects are: the present application integrates the intrinsic structure information and time information of monitoring indicators when calculating the marginal contribution of each monitoring indicator. On the one hand, by constructing the graph Laplacian matrix and calculating the eigenfunction value to obtain the correlation strength of the indicator subset, the contribution degree evaluation can reflect the process correlation characteristics. On the other hand, by using the phase synchronization index to obtain the change trend of a single monitoring indicator and the indicator subset in the time series, by combining the marginal contribution, correlation coefficient and phase synchronization index, the improved Shapley value can better represent the influence of each monitoring indicator on the overall risk, improve the accuracy and interpretability of the contribution degree result, and provide a more reliable decision basis for risk tracing and regulation of the comprehensive utilization process of hazardous waste. BRIEF DESCRIPTION OF DRAWINGS
[0060] Figure 1 is a flow chart for a hazardous waste comprehensive utilization monitoring method;
[0061] Figure 2 is a schematic diagram of a basic evaluation model;
[0062] Figure 3 is a schematic diagram of a hazardous waste comprehensive utilization monitoring system. DETAILED DESCRIPTION
[0063] The embodiments of the hazardous waste comprehensive utilization monitoring method provided by the present application include the following steps:
[0064] As shown in the figure, a hazardous waste comprehensive utilization monitoring method includes the following steps: Figure 1
[0065] S1, obtain the time series data of N monitoring indicators, input the time series data of any subset S of N monitoring indicators into n basic evaluation models, obtain n risk state probability distributions, and calculate the respective Shannon entropy to obtain an uncertainty vector, determine an n-dimensional weight vector according to the system risk preference, and obtain the eigenfunction value of the subset S by calculating the inner product of the uncertainty vector and the weight vector.
[0066] By a distributed control system (DCS) or a supervisory control and data acquisition system (SCADA), N key process parameters (such as reaction kettle temperature, feed flow rate, tail gas component concentration, pH value, etc.) in the comprehensive utilization process of hazardous waste are collected, N time series with a length of T are constructed, and a monitoring data matrix is formed. n heterogeneous machine learning models are selected as the basic evaluation models, which can be a logistic regression model, a support vector machine model, and a long short-term memory network model. For any subset S, the time series data of the corresponding key process parameters contained in the subset S are input into each model as input. Each model evaluates the data at each time point and outputs the probability of the system being in different risk states, thereby obtaining n risk state probability distributions, as shown in Figure 2 , wherein the different risk states include normal, pre-warning, alarm, etc. The information entropy of each model output risk state probability distribution is calculated according to the Shannon entropy formula. The n information entropy values calculated by the n models are combined into an n-dimensional uncertainty vector. According to the performance of the n basic evaluation models evaluated by experts or historical data, an n-dimensional weight vector is set, and the weight vector represents the preference for different types of evaluation results. The inner product of the uncertainty vector and the weight vector is calculated to obtain the characteristic function value of the subset S , and the characteristic function value integrates the evaluation results of multiple models.
[0067] More specifically, in S1, the process of obtaining the characteristic function value of the subset S includes the following steps:
[0068] S11, input the time series data of the subset S into three basic evaluation models of long short-term memory network (LSTM), gated recurrent unit (GRU) and support vector machine (SVM), and each model outputs the probability distribution of {low risk, medium risk, high risk} three states;
[0069] S12, calculate the Shannon entropy of each probability distribution to obtain three entropy values , which constitute a three-dimensional uncertainty vector ;
[0070] S13, according to the system risk preference, set the weight vector for the three basic evaluation models;
[0071] S14, the inner product of the uncertainty vector U and the weight vector W is calculated to obtain the characteristic function value of the subset S .
[0072] Suppose a monitoring system contains multiple process parameter indicators, two monitoring indicators, temperature and pressure, are selected to form a subset S. The monitoring data of the two monitoring indicators at the past 100 time points are input into the pre-trained long short-term memory network, gated recurrent unit and support vector machine as a whole. After the long short-term memory network model analyzes the data, the output risk state probability distribution is low risk 0.1, medium risk 0.2 and high risk 0.7; the probability distribution output by the gated recurrent unit model is low risk 0.2, medium risk 0.6 and high risk 0.2, and the probability output by the support vector machine model is low risk 0.15, medium risk 0.35 and high risk 0.5.
[0073] The Shannon entropy of each of the above three probability distributions is calculated to measure the uncertainty of the output results of each model, for example, the entropy value of the long short-term memory network model is 0.88, the entropy value of the gated recurrent unit model is 1.03, and the entropy value of the support vector machine model is 1.06. The three entropy values form an uncertainty vector , that is , according to the evaluation of the reliability of different models and the preference of risk management, it is assumed that the weight vector = [0.4, 0.3, 0.3]. The inner product operation is performed on the uncertainty vector and the weight vector to obtain the characteristic function value 0.979.
[0074] S2, according to the time series data of N monitoring indicators, the mutual information between the monitoring indicators is calculated to obtain a graph Laplacian matrix; the subgraph Laplacian matrix corresponding to the subset S is extracted, and the second smallest eigenvalue of the subgraph Laplacian matrix is taken as the correlation coefficient of the subset S.
[0075] For any two monitoring indicators i and j in the N monitoring indicators, the mutual information value MI(i, j) between the time series data thereof is calculated. All the mutual information values form an N×N adjacency matrix A, wherein is equal to MI(i, j). The degree matrix D is a diagonal matrix, and the diagonal elements are the sum of all elements in the i-th row of the adjacency matrix A. The N×N graph Laplacian matrix L is calculated by the formula L=D-A.
[0076] For any subset S containing k monitoring indicators, the rows and columns corresponding to the k monitoring indicators are extracted from the complete graph Laplacian matrix L to form a k×k subgraph Laplacian matrix . Further, the second smallest eigenvalue of the subgraph Laplacian matrix Eigenvalue decomposition is performed, all eigenvalues are sorted from small to large, and the second eigenvalue is taken as the correlation number of the subset S .
[0077] More specifically, in S2, the process of taking the second smallest eigenvalue of the subgraph Laplacian matrix as the correlation number of the subset S includes the following steps:
[0078] S21, for any two monitoring indicators i and j in the N monitoring indicators, the mutual information value MI(i, j) between their time series is calculated, and the mutual information value MI(i, j) is taken as the element of the adjacency matrix A , an N×N adjacency matrix A is constructed;
[0079] S22, the degree matrix D is calculated, wherein the degree matrix D is a diagonal matrix, and the diagonal elements are the sum of all elements in the i-th row of the adjacency matrix A;
[0080] S23, the graph Laplacian matrix L is constructed according to the formula L=D-A;
[0081] S24, the rows and columns corresponding to the monitoring indicators in the subset S are extracted from the graph Laplacian matrix L to form a subgraph Laplacian matrix ;
[0082] S25, when the subset S contains at least two monitoring indicators, the eigenvalues of are calculated and sorted, and the second smallest eigenvalue is taken as the correlation number of the subset S ; when the subset S contains only one monitoring indicator, the correlation number is 1.
[0083] Suppose there are three monitoring indicators, monitoring indicator 1 (temperature), monitoring indicator 2 (pressure) and monitoring indicator 3 (vibration), and the target is to calculate the correlation number of the subset S composed of monitoring indicator 1 and monitoring indicator 2, the specific steps are as follows:
[0084] The time series data of the three monitoring indicators is analyzed, and the strength of information correlation between any two monitoring indicators, i.e. the mutual information value, is calculated. Suppose the mutual information value between monitoring indicator 1 and monitoring indicator 2 is 0.5, the mutual information value between monitoring indicator 1 and monitoring indicator 3 is 0.2, and the mutual information value between monitoring indicator 2 and monitoring indicator 3 is 0.6. A 3×3 adjacency matrix A is constructed, wherein the diagonal elements are zero.
[0085] The degree matrix D is a diagonal matrix. The first element on the diagonal is the sum of all elements in the first row of the adjacency matrix A, i.e., 0.5 + 0.2 = 0.7. The second element is 1.1 (0.5 + 0.6), and the third element is 0.8 (0.2 + 0.6). Subtracting the adjacency matrix A from the degree matrix D yields the graph Laplacian matrix L. Since the target subset S contains monitoring index 1 and monitoring index 2, the first and second rows, as well as the first and second columns, are extracted from matrix L to form a 2×2 subgraph Laplacian matrix L. S The second smallest eigenvalue of the graph Laplacian matrix is also called the algebraic connectivity. The Laplacian matrix L of a subgraph is calculated as follows: S The eigenvalues were obtained as two values, 0.44 and 1.36. The second smallest eigenvalue, 1.36, was selected as the correlation coefficient τ(S) of subset S, representing the degree of connection between monitoring index 1 and monitoring index 2 in the whole structure.
[0086] S3. For any monitoring indicator i and any subset S that does not contain the monitoring indicator i, calculate the marginal contribution, merge the time series of each monitoring indicator in subset S into a comprehensive time series of subset S, calculate the phase synchronization index between the time series of monitoring indicator i and the comprehensive time series, and multiply the marginal contribution, the correlation coefficient of subset S and the phase synchronization index to obtain the corrected marginal contribution.
[0087] Based on the correlation coefficient determined in S2 Calculate the subsets containing monitoring indicator i respectively. Relationship values The correlation values between the subset S and the subset S that does not contain monitoring indicator i Subtracting the two yields the traditional marginal contribution of monitoring indicator i when it is added to subset S. The time series of all monitoring indicators contained in subset S are merged into a comprehensive time series through methods such as weighted averaging. The time series of monitoring indicator i were analyzed respectively. and composite time series Perform Hilbert transform to obtain their respective analytic signals, and extract the instantaneous phase from them. and Calculate the phase difference between the two at all time points t. The phase synchronization index C(i, S) is calculated based on the distribution of the phase difference, with a value between 0 and 1, where 1 represents complete synchronization. A subset is then obtained from the calculation results. correlation coefficient Multiplying the three together, we get... C(i, S) and Multiplying the results yields the corrected marginal contribution.
[0088] More specifically, in S3, the process of calculating the phase synchronization index C(i, S) includes the following steps:
[0089] When the subset S is an empty set, the phase synchronization index C(i, S) is 1;
[0090] When the subset S is a non-empty set, the phase synchronization index C(i, S) is calculated by the following steps:
[0091] S31, using principal component analysis method, the time series data of each monitoring index in the subset S is fused into a comprehensive time series, and the first principal component is selected as the comprehensive time series;
[0092] S32, applying Hilbert transform to the time series of monitoring index i and the comprehensive time series of subset S respectively, extracting the instantaneous phase at time point t and ;
[0093] S33, calculating the instantaneous phase difference ;
[0094] S34, by calculating the modulus of the average value of complex number in time, the phase synchronization index is obtained: where j represents the imaginary unit, satisfying , represents the average value at time t.
[0095] For example, assuming that each monitoring index has a time series containing 1000 data points. When the subset S is not empty, principal component analysis is performed on the two time series of temperature and pressure in the subset S, and the component that best represents their common trend of change, i.e. the first principal component, is extracted. The first principal component constitutes a new comprehensive time series of 1000 data points, representing the overall behavior of the subset S. Hilbert transform is applied to the original time series of vibration index and the comprehensive time series of subset S respectively, and the instantaneous phase varying with time is analyzed from the signal, obtaining two instantaneous phases, one corresponding to the vibration index and the other corresponding to the comprehensive dynamics of temperature and pressure.
[0096] At each time point, the difference between the two instantaneous phases is calculated to obtain an instantaneous phase difference. According to the formula, each instantaneous phase difference is taken as the index to calculate the complex value in Euler's formula, and then the average of the complex values of all 1000 data points is calculated. The average result is a complex number, and the modulus (i.e. absolute value) of it is taken to obtain the phase synchronization index. For example, if the calculation result is 0.85, it means that the change of the vibration index is synchronized with the change of the temperature and pressure subset.
[0097] S4. For all subsets S that do not contain monitoring index i, the modified marginal contributions are weighted and summed according to the Shapley value formula to obtain the improved Shapley value of monitoring index i.
[0098] For each monitoring indicator i, iterate through all subsets S that do not contain it. Calculate the result using the classic Shapley value weighted summation formula. This is the improved Shapley value of monitoring indicator i, representing the comprehensive contribution of monitoring indicator i to the overall risk of the system.
[0099] More specifically, in S4, the process of obtaining the improved Shapley value of monitoring indicator i includes the following steps:
[0100] For any subset S that does not contain monitoring indicator i, the corrected marginal contribution is calculated as follows. : ;
[0101] in, and Let S be the characteristic functions of subsets S that contain and exclude monitoring index i, respectively. A subset containing monitoring indicator i The relevant numerical values, The phase synchronization index between the monitoring index i and the subset S;
[0102] The improved Shapley value of monitoring indicator i is obtained by weighted summing of the modified marginal contributions of all subsets S that do not contain monitoring indicator i according to the following formula. : ;
[0103] Where |S| is the number of monitoring indicators in subset S, and N is the total number of monitoring indicators. It is any subset of the set of all monitoring indicators except for monitoring indicator i.
[0104] Continuing with the example of temperature, pressure, and vibration, to assess the importance of temperature—that is, to calculate its improved Shapley value—it is necessary to examine all possible combinations (subsets) that do not include temperature. These subsets include the empty set, subsets containing only pressure, subsets containing only vibration, and subsets containing both pressure and vibration. For each of these subsets, a corrected marginal contribution value needs to be calculated.
[0105] Taking a subset S containing only pressure as an example, we calculate the marginal contribution of adding temperature to this subset. Specifically, we calculate the characteristic function values of subset S itself and the subset S combined with temperature. For example, the characteristic function value of the pressure subset is 0.8, while the characteristic function value of the pressure and temperature combination is 1.1, and the difference of 0.3 is the basic marginal contribution. We further calculate the correlation coefficient of the pressure and temperature combination, assumed to be 1.3; and the phase synchronization index between the temperature index and the pressure subset, assumed to be 0.9. Multiplying these three terms together gives the corrected marginal contribution of 0.351.
[0106] The above calculation is repeated for all subsets that do not include temperature, and the resulting corrected marginal contribution values are weighted and summed. The weights are determined by the combination formula to ensure fairness. For example, a pressure subset containing one member has a weight of one-sixth in a system with three monitoring indicators. The corrected marginal contribution value of each subset is multiplied by its corresponding weight, and all results are summed to obtain the improved Shapley value for the temperature indicator.
[0107] Embodiments of the monitoring system for comprehensive utilization of hazardous waste provided by the present invention:
[0108] like Figure 3 As shown, the monitoring system for comprehensive utilization of hazardous waste includes a characteristic function value calculation module, a correlation coefficient calculation module, a marginal contribution calculation module, and an optimization module.
[0109] The feature function value calculation module is used to acquire time series data of N monitoring indicators. It inputs the time series data of any subset S of the N monitoring indicators into n basic evaluation models to obtain n risk state probability distributions and calculates the Shannon entropy of each to obtain the uncertainty vector. It determines the n-dimensional weight vector according to the system risk preference and obtains the feature function value of subset S by calculating the inner product of the uncertainty vector and the weight vector.
[0110] The correlation coefficient calculation module is used to calculate the mutual information between monitoring indicators based on the time series data of N monitoring indicators to obtain the graph Laplacian matrix; extract the subgraph Laplacian matrix corresponding to subset S, and calculate the second smallest eigenvalue of the subgraph Laplacian matrix as the correlation coefficient of subset S;
[0111] The marginal contribution calculation module is used to calculate the marginal contribution for any monitoring indicator i and any subset S that does not contain the monitoring indicator i, merge the time series of each monitoring indicator in subset S into a comprehensive time series of subset S, calculate the phase synchronization index between the time series of monitoring indicator i and the comprehensive time series, and multiply the marginal contribution, the correlation coefficient of subset S and the phase synchronization index to obtain the corrected marginal contribution.
[0112] The optimization module is configured to calculate the improved Shapley value of the monitoring indicator i by performing a weighted summation of the modified marginal contributions of all subsets S that do not contain the monitoring indicator i according to the Shapley value formula.
[0113] In an optional embodiment, in the feature function value calculation module, when the feature function value of the subset S is calculated, the time series data of the subset S is input into three basic evaluation models, i.e., a long short-term memory network, a gated recurrent unit and a support vector machine, and each model outputs a probability distribution about three states, i.e., low risk, medium risk and high risk.
[0114] The Shannon entropy of each probability distribution is calculated to obtain three entropy values , which constitute a three-dimensional uncertainty vector .
[0115] According to the risk preference of the system, a weight vector W is set for the three basic evaluation models.
[0116] The feature function value of the subset S is calculated by calculating the inner product of the uncertainty vector U and the weight vector W .
[0117] In an optional embodiment, in the correlation coefficient calculation module, when the second smallest eigenvalue of the subgraph Laplacian matrix is taken as the correlation coefficient of the subset S, for any two monitoring indicators i and j in the N monitoring indicators, the mutual information value MI(i, j) between the time series of the two monitoring indicators is calculated, and the mutual information value MI(i, j) is taken as an element of the adjacency matrix A to construct an N×N adjacency matrix A.
[0118] A degree matrix D is calculated, wherein the degree matrix D is a diagonal matrix, and diagonal elements of the degree matrix D are the sum of all elements in the i-th row of the adjacency matrix A.
[0119] A graph Laplacian matrix L is constructed according to the formula L=D-A.
[0120] The rows and columns corresponding to the monitoring indicators in the subset S are extracted from the graph Laplacian matrix L to construct a subgraph Laplacian matrix .
[0121] When the subset S contains at least two monitoring indicators, all eigenvalues of are calculated and sorted, and the second smallest eigenvalue is taken as the correlation coefficient of the subset S ; when the subset S contains only one monitoring indicator, the correlation coefficient is 1.
[0122] In an optional embodiment, in the marginal contribution calculation module, when the subset S is an empty set, the phase synchronization index C(i, S) is 1 when calculating the phase synchronization index;
[0123] When the subset S is a non-empty set, the phase synchronization index C(i, S) is calculated by the following steps:
[0124] S31, using the principal component analysis method, the time series data of each monitoring index in the subset S is fused into a comprehensive time series, and the first principal component is selected as the comprehensive time series;
[0125] S32, the Hilbert transform is applied to the time series of the monitoring index i and the comprehensive time series of the subset S respectively, and the instantaneous phase at the time point t is extracted and ;
[0126] S33, the instantaneous phase difference is calculated;
[0127] S34, the phase synchronization index is obtained by calculating the modulus of the average value of the complex number in time: , where j represents the imaginary unit, and satisfies , represents the average value at time t.
[0128] In an optional embodiment, in the optimization module, when obtaining the improved Shapley value of the monitoring index i, for any subset S not containing the monitoring index i, the following is used to calculate the modified marginal contribution : ;
[0129] , where and are the characteristic functions of the subset S containing and not containing the monitoring index i, is the correlation value of the subset containing the monitoring index i, is the phase synchronization index of the monitoring index i and the subset S;
[0130] The modified marginal contributions of all subsets S not containing the monitoring index i are weighted and summed according to the following formula to obtain the improved Shapley value of the monitoring index i : ;
[0131] , where |S| is the number of monitoring indices in the subset S, N is the total number of monitoring indices, is any subset of the set consisting of all other monitoring indices except the monitoring index i.
Claims
1. A method for monitoring the comprehensive utilization of hazardous waste, characterized in that, The method comprises the following steps: S1, obtaining time series data of N monitoring indexes, inputting time series data of any subset S of the N monitoring indexes into n basic evaluation models, obtaining n risk state probability distributions, calculating respective Shannon entropies to obtain an uncertainty vector, determining an n-dimensional weight vector according to system risk preference, and obtaining a characteristic function value of the subset S by calculating the inner product of the uncertainty vector and the weight vector, comprising: S11, inputting the time series data of the subset S into three basic evaluation models of long short-term memory network, gated recurrent unit and support vector machine, and each model outputs a probability distribution about three states of {low risk, medium risk, high risk}; S12, calculate the Shannon entropy of each probability distribution, to obtain three entropy values , to form a three-dimensional uncertainty vector ; S13, setting a weight vector W for the three basic evaluation models according to the system risk preference; S14, obtain the characteristic function value of the subset S by calculating the inner product of the uncertainty vector U and the weight vector W ; S2, according to the time series data of the N monitoring indexes, calculating mutual information between the monitoring indexes to obtain a graph Laplacian matrix; extracting a subgraph Laplacian matrix corresponding to the subset S, and taking the second smallest eigenvalue of the subgraph Laplacian matrix as the correlation coefficient of the subset S, comprising: S21, for any two monitoring indexes i and j in the N monitoring indexes, calculating the mutual information value MI(i, j) between their time series, taking the mutual information value MI(i, j) as an element of an adjacency matrix A , constructing an N×N adjacency matrix A; S22, computing a degree matrix D, where the degree matrix D is a diagonal matrix whose diagonal elements are is the sum of all elements of the ith row of the adjacency matrix A. S23, constructing a graph Laplacian matrix L according to the formula L=D-A; S24, extracting the rows and columns corresponding to the monitoring indexes in the subset S from the graph Laplacian matrix L to form a subgraph Laplacian matrix ; S25, when the subset S contains at least two monitoring indexes, calculate all eigenvalues of the matrix and rank them, take the second smallest eigenvalue as the correlation index of the subset S ; when the subset S contains only one monitoring index, the correlation index is 1; S3, for any monitoring index i and any subset S not containing the monitoring index i, calculating a marginal contribution, fusing time series of monitoring indexes in the subset S into a comprehensive time series of the subset S, calculating a phase synchronization index between the time series of the monitoring index i and the comprehensive time series, and multiplying the marginal contribution, the correlation coefficient of the subset S and the phase synchronization index to obtain a revised marginal contribution; S4, for all subsets S not containing the monitoring indicator i, a weighted sum of the modified marginal contributions according to the Shapley value formula is performed to obtain the improved Shapley value of the monitoring indicator i, including: for any one subset S not containing the monitoring indicator i, the modified marginal contribution is calculated by : ; wherein, and respectively are characteristic functions of the subset S of monitoring indices i, including and not including respectively, is a correlation value of the subset of monitoring indices i, is a phase synchronization index of the monitoring index i with the subset S. The improved Shapley value for monitoring indicator i is obtained by a weighted sum of the modified marginal contributions for all subsets S not containing monitoring indicator i according to the following formula : ; where |S| is the number of monitoring indicators in the subset S, and N is the total number of monitoring indicators, is any subset of the set of all other monitoring indicators except monitoring indicator i.
2. The method for monitoring comprehensive utilization of hazardous waste according to claim 1, characterized in that, In S3, the process of calculating the phase synchronization index comprises the following steps: When the subset S is an empty set, the phase synchronization index C(i, S) is 1; When the subset S is a non-empty set, the phase synchronization index C(i, S) is calculated by the following steps: S31, using principal component analysis method to fuse time series data of monitoring indexes in the subset S into a comprehensive time series, and selecting a first principal component as the comprehensive time series; S32, applying a Hilbert transform to the time series of the monitoring indicators i and to the integrated time series of the subset S, respectively, to extract the instantaneous phases at the respective time points t and ; S33, calculate the instantaneous phase difference ; S34, the phase synchronization index is obtained by calculating the modulus of the average of the complex exponentials over time : where j represents the imaginary unit, satisfying , denotes the average over time t.
3. A hazardous waste comprehensive utilization monitoring system for performing the method of claim 1, characterized by, The method comprises the following modules: A characteristic function value calculation module is configured to obtain time series data of N monitoring indexes, input time series data of any subset S of the N monitoring indexes into n basic evaluation models, obtain n risk state probability distributions, calculate respective Shannon entropies to obtain an uncertainty vector, determine an n-dimensional weight vector according to system risk preference, and obtain a characteristic function value of the subset S by calculating the inner product of the uncertainty vector and the weight vector; A correlation coefficient calculation module is configured to calculate mutual information between monitoring indexes to obtain a graph Laplacian matrix according to time series data of the N monitoring indexes, extract a subgraph Laplacian matrix corresponding to the subset S, and calculate a second smallest eigenvalue of the subgraph Laplacian matrix as a correlation coefficient of the subset S; A marginal contribution calculation module is configured to calculate a marginal contribution for any monitoring index i and any subset S not containing the monitoring index i, fuse time series of monitoring indexes in the subset S into a comprehensive time series of the subset S, calculate a phase synchronization index between the time series of the monitoring index i and the comprehensive time series, and multiply the marginal contribution, the correlation coefficient of the subset S and the phase synchronization index to obtain a revised marginal contribution; An optimization module is configured to perform weighted summation on the revised marginal contributions according to a Shapley value formula for all subsets S not containing the monitoring index i, to obtain an improved Shapley value of the monitoring index i.
4. The hazardous waste comprehensive utilization monitoring system according to claim 3, wherein In the characteristic function value calculation module, when the characteristic function value of the subset S is obtained, the time series data of the subset S is input into three basic evaluation models of long short-term memory network, gated recurrent unit and support vector machine, and each model outputs a probability distribution about three states of {low risk, medium risk, high risk}; Shannon entropy of each probability distribution is calculated to obtain three entropy values , which constitute a three-dimensional uncertainty vector ; According to the system risk preference, a weight vector W is set for the three basic evaluation models; The characteristic function value of the subset S is obtained by calculating the inner product of the uncertainty vector U and the weight vector W .
5. The hazardous waste comprehensive utilization monitoring system according to claim 3, wherein In the correlation coefficient calculation module, when taking the second smallest eigenvalue of the subgraph Laplace matrix as the correlation coefficient of the subset S, for any two monitoring indexes i and j in the N monitoring indexes, the mutual information value MI(i, j) between their time series is calculated, and the mutual information value MI(i, j) is taken as an element of the adjacency matrix A , and the N*N adjacency matrix A is constructed. computing a degree matrix D, wherein the degree matrix D is a diagonal matrix whose diagonal elements are the sum of all elements of the ith row of the adjacency matrix A; A graph Laplacian matrix L is constructed according to the formula L=D-A; Extracting the rows and columns corresponding to the monitoring indexes in the subset S from the Laplacian matrix L, a sub-Laplacian matrix Ls is constructed ; When the subset S contains at least two monitoring indicators, all eigenvalues of the matrix are calculated and sorted, and the second smallest eigenvalue is taken as the correlation index of the subset S ; when the subset S contains only one monitoring indicator, the correlation index is 1.
6. The hazardous waste comprehensive utilization monitoring system according to claim 3, wherein In the marginal contribution calculation module, when the phase synchronization index is calculated, the phase synchronization index C(i, S) is 1 when the subset S is an empty set; When the subset S is a non-empty set, the phase synchronization index C(i, S) is calculated by the following steps: S31, using the principal component analysis method, the time series data of each monitoring index in the subset S is fused into a comprehensive time series, and the first principal component is selected as the comprehensive time series; S32, applying a Hilbert transform to the time series of the monitoring indicators i and to the integrated time series of the subset S, respectively, to extract the instantaneous phase at the respective time points t and ; S33, calculate the instantaneous phase difference ; S34, the phase synchronization index is obtained by calculating the modulus of the average of the complex exponentials over time : where j represents the imaginary unit and satisfies , denotes the average over time t.
7. The hazardous waste comprehensive utilization monitoring system according to any one of claims 3-6, characterized in that, In the optimization module, when obtaining the improved Shapley value of the monitoring index i, for any subset S not containing the monitoring index i, the following calculation is used to correct the marginal contribution : ; wherein, and respectively are the characteristic function of the subset S containing and not containing the monitoring indicator i, is the correlation value of the subset containing the monitoring indicator i, is the phase synchronization index of the monitoring indicator i with the subset S; The improved Shapley value for monitoring indicator i is obtained by a weighted sum of the modified marginal contributions for all subsets S not containing monitoring indicator i according to the following formula : ; where |S| is the number of monitoring indicators in the subset S, and N is the total number of monitoring indicators, is any subset of the set of all other monitoring indicators except monitoring indicator i.
Citation Information
Patent Citations
Tunnel monitoring system based on big data analysis
CN114066183A
Unmanned aerial vehicle prediction model health condition monitoring method based on boundary measurement
CN118605295A