Emission concentration anomaly backtracking diagnosis method and system
Patent Information
- Application Number
- CN202610907033.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-11
AI Technical Summary
这种缺陷使得诊断结果混杂大量虚假关联,无法有效定位真正的异常起源参数,降低了现场运维的决策效率
Smart Images

Figure CN122734599A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial emission monitoring technology, and in particular to a method and system for tracing and diagnosing abnormal emission concentrations. Background Technology
[0002] In the field of tracing and diagnosing anomalies in emission concentrations, existing technologies typically rely on statistical correlation analysis based on historical data or simple material balance models. The conventional approach involves first detecting anomalies in the concentration sequence of the target emission point, and then using indicators such as Pearson correlation coefficient or mutual information to correlate the abnormal changes in emission point concentrations with the operating parameters or emission data of upstream process nodes. Some schemes also employ principal component analysis or independent component analysis to perform blind source separation of mixed signals, attempting to infer the contributions of each potential emission source.
[0003] However, these conventional approaches have significant drawbacks. Firstly, they ignore the dynamic flow characteristics of the process network, such as the fluctuating flow distribution ratios across transmission paths and the non-uniform residence time distribution during material transport. This results in a complex time-varying coupling between the mixed concentration signal observed at the discharge outlet and the actual discharge time at each process node. Conventional static correlation or blind source separation methods cannot accurately eliminate this time lag and aliasing effect, leading to a phase misalignment in the source contribution allocation in the source tracing results and causing systematic bias in the determination of abnormal moments.
[0004] Secondly, conventional causal inference methods are susceptible to common cause confounding. In complex industrial processes, multiple process parameters may simultaneously become abnormal due to common external disturbances (such as fluctuations in raw material quality or changes in ambient temperature). In such cases, analyses based on Granger causality tests or simple conditional correlation coefficients often identify spurious correlations caused by common causes as direct causal relationships, thus incorrectly listing irrelevant parameters as root cause candidates. This deficiency results in diagnostic results mixed with a large number of spurious associations, making it impossible to effectively locate the true source parameters of the anomalies and reducing the efficiency of on-site operation and maintenance decision-making. Summary of the Invention
[0005] The present invention provides a method and system for tracing and diagnosing abnormal emission concentrations, which can solve the problems in the prior art.
[0006] A first aspect of the present invention provides a method for tracing and diagnosing abnormal emission concentrations, comprising:
[0007] Acquire emission concentration observation sequences and flow topology data of the process pipeline network at the target emission outlet;
[0008] Based on the flow topology data, a set of material transport paths from each process node to the target emission outlet is constructed. For abnormal moments in the emission concentration observation sequence, combined with the flow distribution ratio and residence time distribution of each transport path, the mixed concentration signal of the target emission outlet is subjected to inverse spatiotemporal deconvolution operation to decompose it into the independent emission contribution of each process node at historical moments, and the source intensity time series reconstruction result is obtained.
[0009] For the target process node with abnormal contribution in the source intensity time series reconstruction result, extract the time series data of the operating parameters of the node before the abnormal period, identify the direct causal dependency by constructing the conditional independence test matrix of its operating parameters and remove the pseudo-correlation edges of common cause confusion, construct the parameter coupling causal directed graph, and identify the set of shortest causal paths with abnormal emission as the endpoint as the root cause parameter candidate set.
[0010] Calculate the causal contribution strength between the state deviation of each parameter in the candidate set of root cause parameters and the abnormal amplitude of the source intensity, and sort them to generate the source tracing diagnosis results.
[0011] The steps to obtain the source intensity temporal reconstruction results include:
[0012] Based on the flow topology data, a set of material transport paths from each process node to the target discharge port is constructed;
[0013] For each transport path in the set of material transport paths, a transport kernel function is constructed based on the residence time distribution and flow distribution ratio of the transport path. The transport kernel function represents the mapping relationship from the unit emission of the starting process node of the transport path at a historical moment to the concentration contribution of the target emission outlet at the current moment. Its time dimension is determined by the residence time distribution, and its amplitude dimension is determined by the flow distribution ratio.
[0014] For abnormal moments in the emission concentration observation sequence, the mixed concentration signal of the target emission outlet at the abnormal moment is extracted as the observation value. Based on the transmission kernel function of all transmission paths in the material transmission path set, the independent emission contribution of each process node at the historical moment is convolved with the transmission kernel function of the corresponding transmission path. The convolution results of all transmission paths are accumulated to establish the multipath superposition equation of the mixed concentration signal of the target emission outlet at the abnormal moment.
[0015] The independent emission contribution of each process node at a historical moment in the multipath superposition equation is taken as the unknown to be solved. By performing inverse deconvolution operation on the multipath superposition equation, the independent emission contribution of each process node at a historical moment is obtained as the source intensity time series reconstruction result.
[0016] The steps for inverse deconvolution to obtain the temporal reconstruction results of source intensity include:
[0017] The mixed concentration signal of the target emission outlet at an abnormal moment is decomposed into a fast-changing anomaly component and a slow-changing baseline component. The fast-changing anomaly component represents the instantaneous concentration fluctuation of the sudden emission event, and the slow-changing baseline component represents the background concentration trend of the continuous emission process.
[0018] For the rapidly changing abnormal component, a subset of the transmission kernel functions in the material transmission path set whose residence time is shorter than a first preset threshold is selected for inverse deconvolution operation to solve the instantaneous emission contribution of each process node; for the slowly changing baseline component, a subset of the transmission kernel functions whose residence time is longer than a second preset threshold is selected for inverse deconvolution operation to solve the continuous emission contribution of each process node.
[0019] The instantaneous emission contribution and the continuous emission contribution are aligned by time and then superimposed to obtain the independent emission contribution of each process node at a historical moment, which is used as the source intensity time series reconstruction result.
[0020] For the target process node with abnormal contribution in the source intensity time series reconstruction results, the following steps are taken: extracting the time series data of the operating parameters of the node before the abnormal period; identifying direct causal dependencies by constructing a conditional independence test matrix of its operating parameters and eliminating pseudo-correlation edges that cause common cause confusion; constructing a parameter-coupled causal directed graph; and identifying the set of shortest causal paths ending with abnormal emissions as the root cause parameter candidate set.
[0021] For the time series data of the operation parameters of the target process node before the abnormal period, multiple time windows of different lengths are established by backtracking forward from the abnormal period. In each window, the conditional independence test value between the operation parameters is calculated and a conditional independence test matrix is constructed. Parameter pairs with monotonically increasing test values in the matrix are marked as direct causal dependencies, and parameter pairs with test values that first increase and then decrease are marked as indirect causal relationships.
[0022] For parameter pairs with indirect causal relationships, identify intermediate transit nodes, remove the direct causal edges of the parameter pair, and retain the indirect causal edges that pass through the intermediate transit nodes.
[0023] Based on direct causal dependencies and indirect causal edges, and combined with the occurrence time of intermediate transmission nodes, the operation parameters are divided into causal levels according to the time transmission order, and a parameter-coupled causal directed graph is constructed. Starting from the endpoint of the anomaly emission, the graph is traversed in reverse order of decreasing level to extract the set of causal paths that cross the fewest levels as the root cause parameter candidate set.
[0024] The steps for identifying intermediate transit nodes include:
[0025] For parameter pairs with indirect causal relationships, the earliest and latest response times of the predecessor and successor parameters in the parameter pair are extracted, and the operational parameters whose response times are between the two are used as a set of candidate intermediate transit nodes.
[0026] The candidate intermediate transit nodes are successively substituted as condition variables into the conditional independence test of the parameter pair. The candidate node with the largest decrease in test value is identified as the first-level intermediate transit node. The residual conditional independence test value after introducing the node is calculated. When the residual test value exceeds the decision threshold, the remaining candidate nodes are substituted again into the residual correlation test to identify the second-level intermediate transit node. The process is iterated until the residual test value is lower than the decision threshold to obtain a multi-level intermediate transit node sequence.
[0027] The conditional independence test value between adjacent node pairs in the multi-level intermediate transit node sequence is calculated as the causal strength of the transit segment. The causal strength is then labeled to the corresponding transit segment in the parameter-coupled causal directed graph, which is used to select the transit path with higher causal strength in the subsequent shortest causal path identification.
[0028] The steps for calculating the causal contribution strength between the state deviation of each parameter in the candidate set of root cause parameters and the abnormal amplitude of the source intensity, and sorting them to generate source tracing diagnosis results, include:
[0029] For each parameter in the root cause parameter candidate set, the causal contribution intensity is calculated by comparing the state deviation of the parameter before the abnormal period with the abnormal amplitude of the corresponding target process node in the source intensity time series reconstruction result, and the direct causal contribution value is obtained. In the parameter coupled causal directed graph, all causal paths starting from the parameter and ending with abnormal emission are extracted. The state deviation of the intermediate node parameter is extracted along each path. After calculating the path transmission attenuation, the transmission contribution of all paths is accumulated to obtain the indirect causal contribution value.
[0030] For the root cause parameter candidate set, multi-parameter combinations are screened based on the parameter coupling causal directed graph and process topology. The source intensity anomaly amplitude under the condition that each parameter in the parameter combination deviates simultaneously is calculated and compared with the sum of the individual deviation anomaly amplitudes of each parameter. The excess part is taken as the co-causal contribution value of the parameter combination.
[0031] The direct and indirect causal contribution values of each parameter are summed to obtain the total causal contribution value of a single parameter, and then sorted to generate a single parameter source tracing ranking result. The parameter combinations are sorted according to their synergistic causal contribution values to generate a synergistic parameter source tracing ranking result. The single parameter source tracing ranking result and the synergistic parameter source tracing ranking result are merged to generate a source tracing diagnosis result.
[0032] The steps for identifying co-causal contributions include:
[0033] In the parameter-coupled causal directed graph, parameter pairs with a common downstream node are identified, and parameter pairs whose state deviation of the common downstream node exceeds a preset deviation threshold are regarded as common convergent collaborative candidate combinations.
[0034] In the parameter-coupled causal directed graph, multiple non-intersecting causal paths from different parameters to the abnormal emission target node are identified. The starting parameters of each path are extracted to form a parameter combination. The sum of the causal transmission strength of each path is calculated. The parameter combination whose sum of strength exceeds the third preset threshold is taken as a multi-path superposition type collaborative candidate combination.
[0035] Based on the process topology, identify process nodes that are physically adjacent or have material flow exchange relationships, extract the corresponding operation parameters of the process nodes to form parameter combinations, and use the parameter combinations in which the time interval between the state deviation of each parameter in the combination is less than a preset time window as physical coupling type collaborative candidate combinations.
[0036] For the three types of collaborative candidate combinations, the difference between the source intensity anomaly amplitude under the condition that the parameters deviate simultaneously and the sum of the individual deviation anomaly amplitudes of each parameter is calculated. Combinations with a positive difference that exceeds the preset collaborative threshold are marked as valid collaborative combinations and the difference is recorded as the collaborative causal contribution value.
[0037] A second aspect of the present invention provides a source tracing and diagnostic system for abnormal emission concentrations, comprising:
[0038] The data acquisition unit is used to acquire the emission concentration observation sequence of the target emission outlet and the flow topology data of the process pipeline network;
[0039] The deconvolution unit is used to construct a set of material transport paths from each process node to the target emission outlet based on the flow topology data. For abnormal moments in the emission concentration observation sequence, combined with the flow distribution ratio and residence time distribution of each transport path, the mixed concentration signal of the target emission outlet is subjected to inverse spatiotemporal deconvolution operation to decompose it into the independent emission contribution of each process node at historical moments, and the source intensity time series reconstruction result is obtained.
[0040] The causal identification unit is used to extract the time series data of the operating parameters of the target process node that makes an abnormal contribution in the source intensity time series reconstruction result before the abnormal period, identify the direct causal dependency by constructing the conditional independence test matrix of its operating parameters and remove the pseudo-correlation edges that are confused by common causes, construct the parameter coupling causal directed graph, and identify the set of shortest causal paths ending with abnormal emissions as the root cause parameter candidate set.
[0041] The contribution ranking unit is used to calculate the causal contribution strength between the state deviation of each parameter in the root cause parameter candidate set and the source intensity anomaly amplitude, and to rank them to generate source tracing diagnosis results.
[0042] A third aspect of the present invention provides an electronic device, comprising:
[0043] processor;
[0044] Memory used to store processor-executable instructions;
[0045] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0046] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0047] This method can accurately separate the independent emission contributions of each node from complex process pipeline networks. By using inverse spatiotemporal deconvolution operations, it effectively eliminates signal overlap and delay interference caused by differences in transmission paths and residence time dispersion in mixed signals, significantly improving the spatial positioning accuracy of abnormal emission sources. Even in industrial scenarios with multi-branch pipeline networks and multi-node coupling, it can decompose the actual instantaneous emission of each process node based on the observation data of a single emission outlet, avoiding misjudgment or omission caused by signal aliasing.
[0048] When constructing the parameter-coupled causal directed graph, a conditional independence test matrix is used to automatically identify direct causal dependencies between operational parameters and eliminate spurious correlation edges introduced by codependent variables, significantly reducing the interference of false associations on root cause localization. This method only extracts the time-series data of operational parameters of the anomaly source before the anomaly period, and can accurately construct the optimal causal network with a limited amount of data, avoiding redundant calculations in full-time, full-parameter analysis, while ensuring the statistical reliability of causal relationships.
[0049] By calculating the causal contribution strength between the state deviation of each parameter in the candidate root cause parameter set and the abnormal amplitude of the source intensity, the root cause parameters of abnormal emissions are quantitatively ranked, and the most important control target is directly output. This ranking mechanism simplifies complex causal paths into intuitive contribution indicators, making it easier for operation and maintenance personnel to quickly identify abnormal parameters of equipment or process nodes, and significantly improving the decision-making efficiency and repair accuracy of anomaly diagnosis. Attached Figure Description
[0050] Figure 1 A flowchart illustrating the diagnostic method for tracing the source of abnormal emission concentrations;
[0051] Figure 2 Flowchart for constructing a parameter-coupled causal directed graph and identifying root cause candidate sets. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0053] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes will not be repeated in some embodiments.
[0054] Figure 1 This is a flowchart illustrating the method for tracing and diagnosing abnormal emission concentrations according to an embodiment of the present invention.
[0055] Methods for tracing and diagnosing abnormal emission concentrations include:
[0056] Acquire emission concentration observation sequences and flow topology data of the process pipeline network at the target emission outlet;
[0057] Based on the flow topology data, a set of material transport paths from each process node to the target emission outlet is constructed. For abnormal moments in the emission concentration observation sequence, combined with the flow distribution ratio and residence time distribution of each transport path, the mixed concentration signal of the target emission outlet is subjected to inverse spatiotemporal deconvolution operation to decompose it into the independent emission contribution of each process node at historical moments, and the source intensity time series reconstruction result is obtained.
[0058] For the target process node with abnormal contribution in the source intensity time series reconstruction result, extract the time series data of the operating parameters of the node before the abnormal period, identify the direct causal dependency by constructing the conditional independence test matrix of its operating parameters and remove the pseudo-correlation edges of common cause confusion, construct the parameter coupling causal directed graph, and identify the set of shortest causal paths with abnormal emission as the endpoint as the root cause parameter candidate set.
[0059] Calculate the causal contribution strength between the state deviation of each parameter in the candidate set of root cause parameters and the abnormal amplitude of the source intensity, and sort them to generate the source tracing diagnosis results.
[0060] In one optional implementation, the step of obtaining the source intensity temporal reconstruction result includes:
[0061] Based on the flow topology data, a set of material transport paths from each process node to the target discharge port is constructed;
[0062] For each transport path in the set of material transport paths, a transport kernel function is constructed based on the residence time distribution and flow distribution ratio of the transport path. The transport kernel function represents the mapping relationship from the unit emission of the starting process node of the transport path at a historical moment to the concentration contribution of the target emission outlet at the current moment. Its time dimension is determined by the residence time distribution, and its amplitude dimension is determined by the flow distribution ratio.
[0063] For abnormal moments in the emission concentration observation sequence, the mixed concentration signal of the target emission outlet at the abnormal moment is extracted as the observation value. Based on the transmission kernel function of all transmission paths in the material transmission path set, the independent emission contribution of each process node at the historical moment is convolved with the transmission kernel function of the corresponding transmission path. The convolution results of all transmission paths are accumulated to establish the multipath superposition equation of the mixed concentration signal of the target emission outlet at the abnormal moment.
[0064] The independent emission contribution of each process node at a historical moment in the multipath superposition equation is taken as the unknown to be solved. By performing inverse deconvolution operation on the multipath superposition equation, the independent emission contribution of each process node at a historical moment is obtained as the source intensity time series reconstruction result.
[0065] For example, after obtaining the emission concentration observation sequence of the target emission outlet and the flow topology data of the process pipeline network, all material transport paths from each process node in the process pipeline network to the target emission outlet are enumerated based on the flow topology data to construct a set of material transport paths. Each transport path in this set corresponds to a complete flow channel starting from a certain process node and connected to the target emission outlet via the pipeline topology. For complex pipeline networks with branching and merging structures, multiple parallel transport paths exist for the same process node, all of which must be included in the path set to ensure the integrity of subsequent concentration signal decomposition.
[0066] For each transport path in the set of material transport paths, a corresponding transport kernel function is constructed. The physical meaning of the transport kernel function is: the mapping relationship of the concentration contribution formed at the target discharge port when a unit emission is released at the starting process node of the transport path at a certain historical moment, flows through the transport path to the target discharge port. The construction of the transport kernel function requires two types of parameters: one is the residence time distribution of the transport path, and the other is the flow distribution ratio of the transport path. The residence time distribution describes the time delay and statistical dispersion characteristics experienced by the material from the starting node to the target discharge port in the transport path. It is usually derived from physical parameters such as pipe length, cross-sectional area, and flow velocity. It can be parameterized using an axial dispersion model or a channel string model to reflect the time broadening effect of the concentration signal caused by factors such as uneven flow velocity distribution and pipe bends in the actual pipeline network. The flow distribution ratio describes the proportion of the material flow from the transport path to the total flow at the target discharge port at the pipe network confluence node. It is calculated from the flow balance relationship of each pipe segment in the pipeline network. The transmission kernel function is determined by the dwell time distribution in the time dimension and scaled by the flow distribution ratio in the amplitude dimension, thus simultaneously reflecting the signal's time delay characteristics and intensity attenuation characteristics.
[0067] Let the first The transmission kernel function for each transmission path is ,in Indicates a time delay, then satisfy ,in For the first The traffic allocation ratio of each transmission path, Let the normalized dwell time distribution density function of this transmission path satisfy the following condition: Through the above parameterization method, the transport kernel function of each transport path can be calculated independently based on the physical parameters in the flow topology data, without relying on historical emission data for fitting, thus ensuring the physical interpretability of subsequent inversion calculations.
[0068] After determining the transport kernel functions for all transport paths, for anomaly moments identified in the emission concentration observation sequence, the mixed concentration signal at the target emission outlet at that anomaly moment is extracted as the observation value. The mixed concentration signal observed at the target emission outlet at the anomaly moment is the superposition result of the concentration contributions from all transport paths at that moment. For the ... The sequence of independent emission contributions of the starting process node at each historical moment along a transmission path is denoted as follows: ,in This represents the time coordinates of a historical moment. The transmission path is located at the target emission point at the current time. The resulting concentration contribution is equal to With transfer kernel function Convolution in The value at time, i.e. The convolution results of all transmission paths are accumulated at the target emission outlet to establish a multipath superposition equation for the mixed concentration signal at the target emission outlet at anomaly times. Its form is: ,in For the target emission outlet Observed mixed concentration at time, This represents the observed noise term. The multipath superposition equation explicitly correlates the observable mixed concentration signal at the target emission outlet with the independent emission contributions of each process node at different historical moments through a physical transmission model, providing a mathematical basis for subsequent inversion solutions.
[0069] During discretization, the time axis is set at sampling intervals. Discretize the data, and denote the historical time index as follows: The current index is denoted as Then the multi-path superposition equation is transformed into a discrete convolution summation form: ,in The dispersion number corresponding to the effective support length of the residence time distribution. For the first The transmission path transmits the kernel function in the first step. Discrete sampled values at each time step. The independent emission contributions of all process nodes at each historical moment. The expanded arrangement forms a vector of unknowns to be solved. The multi-path superposition equation can be organized into a system of linear equations. Its coefficient matrix is composed of discrete sampled values of the transmission kernel function of each transmission path arranged in a convolutional structure, and is called the transmission matrix.
[0070] Inverse deconvolution operations are performed on the multipath superposition equations, i.e., solving the above linear equation system, to inversely obtain the independent emission contributions of each process node at each historical time point. Since the product of the number of transmission paths and the number of historical time points is usually much larger than the number of observation equations, this problem is an underdetermined inverse problem, and direct solutions suffer from non-uniqueness. Therefore, regularization constraints are introduced during the solution process. Commonly used regularization methods include Tikhonov regularization and sparsity-promoted regularization. Tikhonov regularization, through the analysis of the unknown vector... Norm-based penalties suppress oscillations in solutions, suitable for scenarios where emission contributions are relatively smooth over time; sparsity-enhanced regularization works by applying a vector of unknowns... The norm imposes a penalty, encouraging sparsity of solutions, and is suitable for scenarios where only a few process nodes exhibit anomalous emissions in localized time periods. The regularization parameter can be selected using generalized cross-validation or the L-curve criterion, striking a balance between the fitting residuals of the solution and the regularization penalty term. The independent emission contributions of each process node at each historical time point, obtained from the solution, constitute the source intensity time-series reconstruction results, providing a quantitative basis for subsequent identification of target process nodes with anomalous contributions and their root cause parameter analysis.
[0071] In practical engineering applications, the flow rate in process piping networks fluctuates with changes in operating conditions. In such cases, the flow distribution ratio... It is not a constant, but a time-varying parameter that changes with time. For such cases, the abnormal period can be divided into several short time windows. Within each short time window, the traffic distribution ratio is approximately assumed to remain stable. The transmission kernel function is constructed segment by segment, and the establishment and inverse deconvolution operation of the multipath superposition equation are performed segment by segment. The inversion results of each short time window are spliced together to form a complete source intensity time series reconstruction result, thereby adapting to the source tracing and diagnosis needs under time-varying traffic conditions.
[0072] In one optional implementation, the step of inverse deconvolution to solve for the temporal reconstruction result of the source intensity includes:
[0073] The mixed concentration signal of the target emission outlet at an abnormal moment is decomposed into a fast-changing anomaly component and a slow-changing baseline component. The fast-changing anomaly component represents the instantaneous concentration fluctuation of the sudden emission event, and the slow-changing baseline component represents the background concentration trend of the continuous emission process.
[0074] For the rapidly changing abnormal component, a subset of transmission kernel functions in the material transmission path set whose dwell time is shorter than a first preset threshold is selected for inverse deconvolution operation to solve the instantaneous emission contribution of each process node;
[0075] For the slowly varying baseline component, a subset of transmission kernel functions with a dwell time longer than a second preset threshold is selected for inverse deconvolution operation to solve for the continuous emission contribution of each process node;
[0076] The instantaneous emission contribution and the continuous emission contribution are aligned by time and then superimposed to obtain the independent emission contribution of each process node at a historical moment, which is used as the source intensity time series reconstruction result.
[0077] For example, before performing inverse deconvolution on the mixed concentration signal of the target emission outlet, it is necessary to first separate the components of the original observation sequence. Emission anomalies in the process pipeline network typically include two types of concentration disturbances with different characteristics: one is short-term, high-amplitude concentration fluctuations caused by sudden events such as valve opening, instantaneous equipment leakage, or batch feeding, with a timescale often ranging from minutes to ten minutes; the other is background concentration trend changes caused by persistent factors such as slow process load drift, catalyst deactivation, or pipeline corrosion, with a timescale reaching several hours or even days. Directly deconvolving these two types of components together will cause interference between the transfer kernel functions at different timescales, resulting in a decrease in the accuracy of source intensity reconstruction. Therefore, before the inverse deconvolution operation, the mixed concentration observation sequence of the target emission outlet is first decomposed into a time-frequency component, separating it into a rapidly varying anomaly component and a slowly varying baseline component.
[0078] Component separation can be achieved using empirical mode decomposition or multi-scale filtering methods. Taking multi-scale filtering as an example, low-pass filters with different cutoff frequencies are applied to the mixed concentration observation sequence of the target emission outlet. The low-frequency component is the slowly varying baseline component, denoted as . The background concentration trend of the continuous emission process is characterized by the base concentration; the residual obtained by subtracting the slowly varying baseline component from the original mixed concentration signal is the rapidly varying anomaly component, denoted as . This characterizes the instantaneous concentration fluctuations of sudden emission events. The two components satisfy an additive relationship, meaning the mixed concentration signal equals the sum of the fast-changing anomaly component and the slow-changing baseline component, ensuring the integrity of the information in the decomposition process. In practical engineering scenarios, the cutoff frequency of the low-pass filter can be set according to the typical response time of the process flow. It is generally recommended that the timescale corresponding to the slow-changing baseline component be no less than three times the cycle time of a single batch operation to ensure sufficient separation between the two types of components in the frequency domain.
[0079] For rapidly changing anomalous components The inverse deconvolution process filters materials from the set of material transport paths whose dwell time is shorter than a first preset threshold. The residence time is a subset of the transmission paths. The quantification of residence time is based on a comprehensive calculation of the pipe segment length, pipe diameter, and flow velocity parameters for each transmission path, typically using the mean residence time of the transmission kernel function. As a basis for judgment, among them For the first The first moment of the normalized dwell time distribution density function for each transmission path. Satisfies... The transmission paths are classified into the fast-variant subset, and the corresponding transmission kernel function subset is denoted as . . use right Perform inverse deconvolution operations to solve for the instantaneous emission contribution of each process node. ,in Indicates the first The instantaneous emission contribution of the starting process node of the fast-changing path at a historical moment. The inverse deconvolution solution adopts the regularized least squares method. By introducing the Tikhonov regularization constraint to suppress the oscillation of the solution, the regularization parameter can be adaptively selected according to the generalized cross-validation criterion to ensure that a physically reasonable non-negative source intensity estimate can still be obtained under noise interference.
[0080] For slowly varying baseline components The inverse deconvolution process filters materials from the set of material transport paths whose dwell time exceeds a second preset threshold. A subset of transmission paths. Satisfying The transmission paths are classified into the slow-varying subset, and the corresponding subset of transmission kernel functions is denoted as . . use right Perform inverse deconvolution operations to solve for the continuous emission contribution of each process node. ,in Indicates the first The continuous emission contribution of the starting process node of the slowly varying path over a historical period. Because the timescale of the slowly varying baseline component is longer, its corresponding transport kernel function exhibits wider broadening characteristics in the time domain, resulting in a relatively lower ill-conditioning degree in the inverse deconvolution problem and better solution stability than the case of the rapidly varying component. Based on this, a smoothness constraint can be further imposed on the continuous emission contribution, requiring that the change in contribution between adjacent time steps does not exceed a set proportion, in order to conform to the prior physical knowledge of the slow evolution of the continuous emission process.
[0081] It is worth noting that the first preset threshold With the second preset threshold The settings need to be determined in conjunction with the specific topology and flow velocity characteristics of the process piping network. When At this time, there exist transmission paths with dwell times between the two. These paths are not included in either the fast-changing subset or the slow-changing subset. They can be assigned to the two subsets respectively, or treated as intermediate-scale paths, depending on actual needs. When some transmission paths simultaneously meet two screening criteria, the transmission kernel function of that path needs to participate in two inverse deconvolution operations, and the corresponding emission contribution will not be counted repeatedly in subsequent superposition. In engineering practice, it is recommended to... Set to 2 to 3 times the dwell time of the shortest transmission path in the process pipeline network, The dwell time of the longest transmission path in the process pipeline is set to 50% to 70% to ensure that the coverage of the material transmission network by the fast-changing subset and the slow-changing subset is reasonably complementary.
[0082] After performing independent inverse deconvolution operations on the two types of components, the contribution to instantaneous emissions is calculated. Contribution to continuous emissions Time alignment is performed. The core of time alignment is to ensure that the two types of contributions have a consistent reference benchmark on the same discrete time axis, avoiding systematic offset errors introduced due to different time index definitions in the two deconvolution operations. After time alignment, the instantaneous emission contribution and continuous emission contribution of the same process node at the same historical moment are directly superimposed to obtain the independent emission contribution of that node at that historical moment. ,in satisfy In other words, the total independent emission contribution of each process node equals the sum of its instantaneous emission contribution and its continuous emission contribution. The source intensity time-series reconstruction result is obtained by summing the independent emission contributions of all process nodes over a complete historical period. This result includes both the timing and intensity information of sudden emission events and the evolution trend information of continuous emission processes, providing a complete and accurate input basis for subsequent causal root cause analysis of anomalous contribution nodes.
[0083] In an optional implementation, for the target process node with abnormal contribution in the source intensity time-series reconstruction result, the following steps are taken: extracting the time-series data of the node's operating parameters before the abnormal period; identifying direct causal dependencies and eliminating spurious correlation edges that cause common cause confusion by constructing a conditional independence test matrix for its operating parameters; constructing a parameter-coupled causal directed graph; and identifying the set of shortest causal paths ending with abnormal emissions as the root cause parameter candidate set:
[0084] For the time series data of the operation parameters of the target process node before the abnormal period, multiple time windows of different lengths are established by backtracking forward from the abnormal period. In each window, the conditional independence test value between the operation parameters is calculated and a conditional independence test matrix is constructed. Parameter pairs with monotonically increasing test values in the matrix are marked as direct causal dependencies, and parameter pairs with test values that first increase and then decrease are marked as indirect causal relationships.
[0085] For parameter pairs with indirect causal relationships, identify intermediate transit nodes, remove the direct causal edges of the parameter pair, and retain the indirect causal edges that pass through the intermediate transit nodes.
[0086] Based on direct causal dependencies and indirect causal edges, and combined with the occurrence time of intermediate transmission nodes, the operation parameters are divided into causal levels according to the time transmission order, and a parameter-coupled causal directed graph is constructed. Starting from the endpoint of the anomaly emission, the graph is traversed in reverse order of decreasing level to extract the set of causal paths that cross the fewest levels as the root cause parameter candidate set.
[0087] Combination Figure 2The flowchart for constructing a parameter-coupled causal directed graph and identifying root cause candidate sets is explained. After completing the temporal reconstruction of the source intensity and locating the target process node contributing to the anomaly, it is necessary to delve deeper into the node and trace the root cause of the abnormal emissions from the perspective of its operating parameters. The time-series data of the operating parameters of the target process node before the anomaly period are extracted. These operating parameters typically include multi-dimensional process variables such as reaction temperature, pressure, feed flow rate, catalyst activity indicators, and valve opening. The time-series data reflects the evolution trajectory of the node's operating state in the period leading up to the anomaly.
[0088] Starting with the abnormal period as the endpoint, multiple time windows of varying lengths are established by backtracking. The different window lengths are designed to capture the causal coupling relationships between operational parameters at different time scales. Short windows focus on revealing rapidly transmitted direct causal effects, while long windows help identify lagged causal chains transmitted through multiple intermediate stages. Within each time window, for all operational parameters of the target process node, the conditional independence test value between any two parameters is calculated under a given set of third-party parameters. The conditional independence test value measures the residual statistical dependence between two parameters after controlling for the influence of other parameters; a larger test value indicates a stronger direct statistical association between the two. The conditional independence test values obtained for each parameter pair under different time windows are arranged according to window length to form a conditional independence test matrix. The rows of the matrix correspond to the parameter pair indices, and the columns correspond to the time window length levels.
[0089] Monotonicity analysis was performed on the test value sequence of each row in the conditional independence test matrix. If the test value of a parameter pair shows a monotonically increasing trend with the increase of the time window length, it indicates that the statistical dependence between the two parameters continues to deepen as the observation time range expands. This pattern is consistent with the theoretical expectation of direct causality—direct causality can accumulate more covariant evidence over a longer time window. Therefore, such parameter pairs are marked as direct causal dependencies. If the test value of a parameter pair shows an inverted U-shaped trend of first increasing and then decreasing with the increase of the window length, it indicates that the two parameters show a strong statistical correlation within a medium-length window. However, as the time range further expands, the correlation weakens. This characteristic is consistent with the theoretical expectation of indirect causality—two parameters are correlated through a common intermediate transit node. When the window is long enough, the independent variability of the intermediate node gradually dilutes the apparent correlation between the two parameters. Therefore, such parameter pairs are marked as indirect causal relationships.
[0090] For parameter pairs marked as having indirect causal relationships, intermediate transit nodes are further identified. The identification method is as follows: among all operational parameters other than the parameter pair, one candidate parameter is included in the condition set for conditional independence testing. If, after controlling for a candidate parameter, the test value of the original parameter pair significantly decreases and approaches the statistical independence threshold, then the candidate parameter is determined to be an intermediate transit node. After identifying the intermediate transit nodes, direct causal edges between the original parameter pairs are removed, retaining indirect causal edges from the first parameter to the intermediate transit node, and then from the intermediate transit node to the second parameter. This removal operation aims to eliminate spurious correlation edges introduced by common cause confusion, avoiding misjudging false direct associations between two parameters caused by common upstream driving factors as true causal dependencies, thereby ensuring that the subsequently constructed causal directed graph has physical interpretability at the topological level.
[0091] After confirming direct causal edges and eliminating indirect pseudo-correlated edges, the causal hierarchy of all operational parameters is determined by combining the occurrence time information of intermediate transit nodes and following the chronological order of causal transmission. Specifically, the parameter node that first deviates from its state in time and has no incoming edges is designated as the first causal level; the parameter node directly driven by the first-level parameters is designated as the second causal level, and so on, until the anomalous emission event itself is included as the endpoint node in the highest causal level. After the hierarchy is determined, all directed causal edges are connected in hierarchical order to construct a parameter-coupled causal directed graph. Each edge of this directed graph represents a direct causal dependency verified by conditional independence testing, and the overall topology of the graph reflects the complete causal transmission network from early operational parameter perturbations to the final anomalous emission event.
[0092] Starting with the endpoint of the abnormal emission, the algorithm traverses backward along the edges of the causal directed graph, from higher to lower levels, to extract all directed paths that can reach the endpoint of the abnormal emission from the first causal level via several intermediate nodes. Among all causal paths, the path length is measured by the number of causal levels traversed, and the set of paths with the fewest traversals is selected. The fewest traversals mean the shortest causal chain and the fewest intermediate links; the corresponding parameter perturbations can drive the abnormal emission in the most direct way, thus possessing the highest root cause confidence. The first-level parameter nodes involved in these shortest causal path sets are then aggregated to form a root cause parameter candidate set, which is used for subsequent causal contribution strength calculation and source tracing diagnosis result ranking.
[0093] In real-world engineering scenarios, time-series data of operational parameters often suffer from data quality issues such as inconsistent sampling frequencies, missing values, or sensor drift. Before constructing the conditional independence test matrix, the original time-series data needs to be preprocessed with alignment interpolation and outlier removal to ensure that all parameters are compared at the same time resolution. Furthermore, the statistical significance of conditional independence test values is affected by sample size; if the time window is too short, the test power is insufficient, leading to false positives being misclassified as statistically independent. Therefore, the minimum time window length must be reasonably configured in conjunction with the autocorrelation time scale of the operational parameters to ensure that the number of effective independent samples within each window meets the statistical power requirements of the test. For high-dimensional operational parameter scenarios, the number of combinations of condition sets increases exponentially with the parameter dimension. A fast conditional independence approximation test method based on partial correlation coefficients can be used to reduce computational complexity, ensuring the accuracy of causal identification while meeting the real-time requirements of industrial online diagnostics.
[0094] In one alternative implementation, the step of identifying intermediate transit nodes includes:
[0095] For parameter pairs with indirect causal relationships, the earliest and latest response times of the predecessor and successor parameters in the parameter pair are extracted, and the operational parameters whose response times are between the two are used as a set of candidate intermediate transit nodes.
[0096] The candidate intermediate transit nodes are successively substituted as condition variables into the conditional independence test of the parameter pair. The candidate node with the largest decrease in test value is identified as the first-level intermediate transit node. The residual conditional independence test value after introducing the node is calculated. When the residual test value exceeds the decision threshold, the remaining candidate nodes are substituted again into the residual correlation test to identify the second-level intermediate transit node. The process is iterated until the residual test value is lower than the decision threshold to obtain a multi-level intermediate transit node sequence.
[0097] The conditional independence test value between adjacent node pairs in the multi-level intermediate transit node sequence is calculated as the causal strength of the transit segment. The causal strength is then labeled to the corresponding transit segment in the parameter-coupled causal directed graph, which is used to select the transit path with higher causal strength in the subsequent shortest causal path identification.
[0098] For example, in the construction of a parameter-coupled causal directed graph, not all relationships between parameter pairs are direct causal relationships. For parameter pairs with indirect causal relationships identified through the conditional independence test matrix, it is necessary to further explore the intermediate transit nodes to make the implicit multi-level causal chain explicit, thus avoiding misjudging indirect associations as direct causal dependencies and introducing incorrect tracing paths.
[0099] For a pair of parameters with an indirect causal relationship, the earliest and latest response times of the precursor and successor parameters within the abnormal period are extracted. Since there is a temporal difference between the response time intervals of the precursor and successor parameters, other operational parameters whose response times fall after the earliest response time of the precursor parameter and before the latest response time of the successor parameter are included in the candidate intermediate transit node set. The physical meaning of this temporal constraint is that the true intermediate transit node must be later than the disturbance start time of the precursor parameter and earlier than the response completion time of the successor parameter, conforming to the temporal ordering principle of causal transit. If the response time of an operational parameter exceeds the above time window, regardless of its statistical correlation with the parameter pair, it does not meet the temporal condition of causal transit and should be excluded from the candidate set to reduce the computational burden of subsequent conditional independence tests and lower the probability of false positives.
[0100] After the set of candidate intermediate transit nodes is determined, each candidate node in the set is sequentially used as a condition variable and substituted into the conditional independence test between the predecessor and successor parameters. The conditional independence test value of the parameter pairs is calculated one by one when the candidate node is used as a condition. Let the predecessor parameter be... The successor parameter is The first in the set of candidate intermediate transit nodes The candidate nodes are The corresponding conditional independence test value is denoted as Without introducing any condition variables, and The original conditional independence test value between them is denoted as Introducing candidate nodes Afterwards, the decrease in the test value was Select The largest candidate node is designated as the first-level intermediate relay node, denoted as . Its physical meaning is: among all candidate nodes, with To provide the conditions that can be explained to the greatest extent possible and Statistical dependency between them, i.e. It carries the most important causal information between the parameter pairs.
[0101] Introducing the first-level intermediate transmission node Then, calculate the residual conditional independence test value. Compare it with the preset judgment threshold. Compare. If ,illustrate It has been fully explained and The indirect causal relationship between them terminates when intermediate transit nodes are identified, resulting in a single-level intermediate transit node sequence. If This indicates that based solely on The indirect dependency of this parameter pair cannot be fully explained yet; a deeper causal chain exists, requiring further identification of second-level intermediate nodes. At this point, [the following will be considered]. Given the known conditions, for each node in the remaining candidate node set... ( ) Calculate simultaneously and The residual conditional independence test value when the condition is . The node that causes the residual test value to decrease the most is selected as the second-level intermediate propagation node. The above iterative process continues, with each iteration adding a new node to the already identified set of intermediate transit nodes, until the current residual condition independence test value falls below the decision threshold. The iteration terminates, ultimately yielding an ordered sequence of multi-level intermediate transit nodes. ,in This represents the total number of intermediate transmission nodes identified.
[0102] Determination threshold The settings need to be determined comprehensively based on the sampling frequency, data length, and response characteristics of the operating parameters and time-series data. When sufficient data is available, statistical methods based on permutation tests or bootstrap resampling can be used to estimate the null hypothesis distribution of the conditional independence test values, using the critical value at a specific significance level as... In scenarios with limited data, empirical thresholds can be preset based on the statistical characteristics of historical process data, and calibration can be performed through cross-validation in practical applications.
[0103] After the multi-level intermediate transit node sequence is determined, the causal strength between adjacent node pairs in the sequence is quantified and labeled. Specifically, for the first node in the sequence... Level Node With the Level Node The transmission segment between nodes is calculated assuming that the remaining identified nodes in the sequence are taken into account. and The conditional independence test value between them is denoted as This serves as the causal strength of the transmission segment. Similarly, the precursor parameters... With the first-level intermediate transmission node The causal strength between them, and the final intermediate transit node. With successor parameters The causal strength between all segments is calculated in the same way. Causal strength of all transit segments. Label the corresponding directed edges in the parameter-coupled causal directed graph to complete a refined description of the indirect causal path.
[0104] In the subsequent shortest causal path identification stage, the causal strength of each transmission segment is used. The paths are weighted and evaluated. Higher causal strength indicates a stronger causal relationship within the transmission segment and a more significant contribution to the propagation of anomalous signals. During the path search process terminating in anomalous emissions, paths with higher causal strength in each transmission segment are prioritized. This allows for the selection of root cause propagation chains with clearer physical meaning and stronger causal relationships from multiple causal transmission paths, improving the accuracy and interpretability of the source tracing diagnosis results. For processes with multiple parallel intermediate transmission nodes, the multi-level intermediate transmission node sequence identification mechanism can distinguish between primary and secondary transmission paths, avoiding causal attribution bias caused by ignoring intermediate links, and ensuring that the final source tracing diagnosis results accurately reflect the propagation mechanism of abnormal process parameters.
[0105] In one optional implementation, the step of calculating the causal contribution strength between the state deviation of each parameter in the candidate set of root cause parameters and the abnormal amplitude of the source intensity, and sorting them to generate the source tracing diagnosis results includes:
[0106] For each parameter in the root cause parameter candidate set, the causal contribution intensity of the parameter is calculated by comparing the state deviation of the parameter before the abnormal period with the abnormal amplitude of the corresponding target process node in the source intensity time series reconstruction result, and the direct causal contribution value is obtained.
[0107] In the parameter-coupled causal directed graph, all causal paths starting from the parameter and ending at the abnormal emission are extracted. The state deviation of the intermediate node parameter is extracted along each path. After calculating the path propagation attenuation, the propagation contribution of all paths is accumulated to obtain the indirect causal contribution value.
[0108] For the root cause parameter candidate set, multi-parameter combinations are screened based on the parameter coupling causal directed graph and process topology. The source intensity anomaly amplitude under the condition that each parameter in the parameter combination deviates simultaneously is calculated and compared with the sum of the individual deviation anomaly amplitudes of each parameter. The excess part is taken as the co-causal contribution value of the parameter combination.
[0109] The direct and indirect causal contribution values of each parameter are summed to obtain the total causal contribution value of a single parameter, and then sorted to generate a single parameter source tracing ranking result. The parameter combinations are sorted according to their synergistic causal contribution values to generate a synergistic parameter source tracing ranking result. The single parameter source tracing ranking result and the synergistic parameter source tracing ranking result are merged to generate a source tracing diagnosis result.
[0110] For example, after completing the temporal reconstruction of source intensity and the construction of the parameter-coupled causal directed graph, the causal contribution intensity quantification calculation is performed on each parameter in the root cause parameter candidate set. For a certain parameter in the candidate set, the historical state sequence of the parameter in several sampling periods before the abnormal period is first extracted, and its deviation from the normal operation benchmark value is calculated, which is denoted as the state deviation of the parameter. ,in This is the index number of the parameter in the candidate set. State deviation. The standard mean square deviation can be used for calculation, which is to divide the difference between the mean of the parameter during the abnormal period and the benchmark mean by the benchmark standard deviation, thereby eliminating the numerical scale difference between operating parameters of different dimensions and making the deviation of each parameter comparable.
[0111] At the same time, extract the anomalous amplitude corresponding to the target process node to which the parameter belongs from the source intensity time series reconstruction results. This magnitude reflects the excess emission contribution of the target process node to the target emission outlet's mixed concentration during abnormal periods. The deviation from the target emission level is... Abnormal amplitude By performing a product operation and introducing a normalization coefficient, the direct causal contribution of this parameter to the abnormal emissions of the target process node is obtained. The physical meaning of the direct causal contribution value is: the strength of the ability to directly drive the target process node to produce abnormal emissions when the parameter deviates, without being transmitted through other intermediate parameters.
[0112] The calculation of indirect causal contribution values relies on a parameter-coupled causal directed graph. In this directed graph, parameters... Starting from the node of the abnormal emission event, enumerate all directed paths. For a given path, let the parameters of the intermediate nodes traversed by the path be as follows: ,in This represents the number of intermediate nodes in the path. The causal strength weights between adjacent nodes are extracted segment by segment along the path. The weights were obtained through conditional independence tests and partial correlation coefficient estimation when constructing the parameter-coupled causal directed graph. Path propagation attenuation is calculated by multiplying the causal strength weights of each segment to obtain the end-to-end propagation coefficient of the path. The path transfer coefficient is compared with the path origin parameter. State deviation Multiplying these yields the propagation contribution of the path. For parameters... The parameters are obtained by summing the propagation contributions of all paths from the origin. Indirect causal contribution value Indirect causal contribution values capture the parameters. The cumulative effect of influencing target emissions through multi-level transmission chains is often significant, especially in scenarios where process flows are tightly coupled and parameters have multi-level control relationships.
[0113] After quantifying the contribution of a single parameter, the calculation of the synergistic causal contribution of multiple parameters is further carried out. The existence of synergistic effects means that when multiple parameters deviate simultaneously, their combined impact on the source intensity anomaly exceeds the sum of the individual effects of each parameter, reflecting the nonlinear coupling between parameters or the convergence effect in the process topology. Based on the directed graph of parameter coupling causality and the process topology, parameter combinations that meet the following conditions are selected as synergistic analysis objects: the parameters in the combination have a common downstream convergence node in the directed graph, or jointly affect the material mixing process of the same process unit in the process topology. For the selected parameter combinations... (in (For combined indexes), calculate the source intensity anomaly amplitude of the target process node under the condition that all parameters in the combination are simultaneously in a deviated state. This value can be obtained by substituting the deviations of each parameter into the constructed parameter-emission response model and solving it together.
[0114] Combined abnormal amplitude The sum of the abnormal amplitudes corresponding to the individual deviations of each parameter in the combination. By comparing the two, the difference is the co-causal contribution value. The calculation relationship is as follows: .when When the value is positive, it indicates that the parameter combination has a positive synergistic amplification effect, meaning that the simultaneous deviation of multiple parameters will lead to an emission anomaly that significantly exceeds the expected linear superposition. When the value is negative, it indicates that there is an antagonistic effect, and the deviations of each parameter cancel each other out to some extent. Only parameter combinations with positive synergistic causal contribution values that exceed the preset synergistic judgment threshold are included in the synergistic parameter source ranking results to ensure the engineering practicality of the output results and avoid misjudging weak synergistic effects as significant root cause combinations.
[0115] For generating single-parameter causal ranking results, the direct causal contribution values of each parameter are used. and indirect causal contribution value Summing yields the total causal contribution value of a single parameter. .according to All parameters in the candidate set of root cause parameters are sorted from largest to smallest to generate a single-parameter source tracing ranking result. The parameters with higher rankings indicate that their deviations from the target emission outlet have the strongest driving force for anomalies in concentration, and they are the priority targets for verification and intervention.
[0116] The single-parameter source tracing ranking results are combined with the collaborative parameter source tracing ranking results to form the final source tracing diagnostic result. The merging strategy adopts a unified scoring framework, which combines the collaborative causal contribution values of the collaborative parameter combinations. The diagnostic results are weighted and fused with the sum of the total causal contribution values of each individual parameter within the combination, making the single-parameter diagnostic conclusions comparable and rankable on the same scale as the multi-parameter collaborative diagnostic conclusions. The source tracing diagnostic results are output in list format, including parameter names or parameter combination identifiers, corresponding total causal contribution values, contribution types (direct, indirect, or collaborative), and recommended process intervention priorities, providing operators with a clear basis for locating the root causes of emission anomalies and making appropriate decisions. In actual industrial scenarios, these diagnostic results can be further matched and verified against a historical anomaly case database. By comparing the historical root cause distribution under similar operating conditions, the confidence level of the current diagnostic conclusions is checked, thereby improving the engineering reliability of the source tracing diagnosis.
[0117] In one alternative implementation, the steps of co-causal contribution identification include:
[0118] In the parameter-coupled causal directed graph, parameter pairs with a common downstream node are identified, and parameter pairs whose state deviation of the common downstream node exceeds a preset deviation threshold are regarded as common convergent collaborative candidate combinations.
[0119] In the parameter-coupled causal directed graph, multiple non-intersecting causal paths from different parameters to the abnormal emission target node are identified. The starting parameters of each path are extracted to form a parameter combination. The sum of the causal transmission strength of each path is calculated. The parameter combination whose sum of strength exceeds the third preset threshold is taken as a multi-path superposition type collaborative candidate combination.
[0120] Based on the process topology, identify process nodes that are physically adjacent or have material flow exchange relationships, extract the corresponding operation parameters of the process nodes to form parameter combinations, and use the parameter combinations in which the time interval between the state deviation of each parameter in the combination is less than a preset time window as physical coupling type collaborative candidate combinations.
[0121] For the three types of collaborative candidate combinations, the difference between the source intensity anomaly amplitude under the condition that the parameters deviate simultaneously and the sum of the individual deviation anomaly amplitudes of each parameter is calculated. Combinations with a positive difference that exceeds the preset collaborative threshold are marked as valid collaborative combinations and the difference is recorded as the collaborative causal contribution value.
[0122] For example, after constructing the directed graph of parameter coupling causality and determining the candidate set of root cause parameters, it is necessary to further identify the synergistic causal relationships among multiple parameters. Single-parameter causal contribution analysis can only reveal the degree to which each parameter independently affects emission anomalies. However, in actual processes, multiple operating parameters often jointly drive abnormal increases in emission concentrations through structural coupling relationships. In this case, relying solely on single-parameter ranking will underestimate the actual impact of some parameter combinations. The synergistic causal contribution identification step addresses this problem by screening synergistic candidate combinations from three dimensions: graph structure features, path topology features, and physical coupling features, and then using a unified quantification criterion to determine effective synergies.
[0123] In parameter-coupled causal directed graphs, a typical structural feature exists: directed edges from two or more parameter nodes simultaneously point to the same downstream node, which acts as a "convergence point" in the causal graph. For this type of structure, all nodes in the directed graph are traversed, nodes with an in-degree greater than 1 are identified, and all their predecessor parameter pairs are extracted to form a set of common convergence candidate parameter pairs. For each pair of predecessor parameters, the state deviation of their common downstream node during the abnormal period is examined to see if it exceeds a preset deviation threshold. If the state deviation of the common downstream node exceeds this threshold, it indicates that the downstream node did indeed deviate significantly during the abnormal period, and the synergistic effect of its predecessor parameter pair has practical physical significance. This parameter pair is then included in the common convergence candidate combination. The preset deviation threshold can be set based on the statistical fluctuation range of the downstream node's state variables under historical normal operating conditions, determined by adding a certain number of standard deviations to the mean, thus ensuring that the selected candidate combinations are statistically significant.
[0124] In parameter-coupled causal directed graphs, another structural feature exists: starting from nodes with different initial parameters, each path can reach the anomalous emission target node along its own independent directed edge sequence, and these paths do not share any intermediate nodes, forming multiple disjoint causal paths. When identifying such paths, a depth-first search is performed in the directed graph, with the anomalous emission target node as the endpoint. All simple paths originating from each parameter in the root cause parameter candidate set and reaching the target node are enumerated. The enumerated path set is then filtered for disjointness, retaining path combinations where the intermediate node sets between paths do not overlap. For each disjoint path group, the initial parameters of each path are extracted to form a parameter combination, and the sum of the causal transmission strengths of each path in this combination is calculated. The calculation method for path causal transmission strength is consistent with the calculation method for the end-to-end transmission coefficient in single-parameter indirect causal contribution: the causal strength weights between adjacent nodes on the path are multiplied sequentially to obtain the end-to-end transmission coefficient of the path; then, the transmission coefficients of all paths within the combination are summed to obtain the sum of the path causal transmission strengths. Parameter combinations whose sum of intensities exceeds a third preset threshold are marked as multi-path superposition type collaborative candidate combinations. The third preset threshold can be determined based on the upper quantile of the distribution of path transmission intensity under historical normal operating conditions to distinguish between real collaborative effects and random fluctuations.
[0125] The identification of physically coupled collaborative candidate combinations relies on process topology information, rather than solely on the edge structure of the causal graph. Based on the flow topology data of the process piping network, it identifies physically adjacent process node pairs and process node pairs with direct material flow exchange relationships. Physical adjacency is determined by whether the spatial distance between nodes in the process layout diagram is below a preset distance threshold; the existence of a material flow exchange relationship is determined by whether there is a non-zero flow rate at the corresponding position in the piping network connection matrix. For process node pairs that satisfy any of the above conditions, the corresponding operating parameters of each node are extracted to form a physically coupled parameter combination. The timing of state deviations for each parameter in this combination is further calculated. If the maximum interval between the deviation times of each parameter is less than a preset time window, the abnormal deviations of each parameter are considered to be approximately synchronous in time, reflecting a physical coupling propagation effect, and this parameter combination is included in the physically coupled collaborative candidate combination. The length of the preset time window can be set with reference to the average transmission and residence time of materials between adjacent nodes to ensure the physical rationality of the time window.
[0126] After screening the three types of collaborative candidate combinations, a unified validity determination is performed on all candidate combinations. This is done for each candidate parameter combination. Calculate the combined source intensity anomaly amplitude of the target process node under the condition that all parameters in the combination deviate simultaneously. And the sum of the abnormal magnitudes when each parameter in the combination deviates individually. The combined source intensity anomaly magnitude is calculated by simultaneously substituting the deviations of all parameters within the combination into the constructed source intensity response model; the sum of the individual parameter deviation anomaly magnitudes is calculated by substituting the deviation of each parameter individually into the model and summing the results. The difference between the two is then calculated. If the difference is positive, it indicates that the parameter combination produced a superadditive synergistic enhancement effect when they deviated simultaneously, meaning that the emission anomaly caused by the simultaneous anomaly of multiple parameters exceeds the sum of the independent effects of each parameter. If the difference further exceeds a preset synergistic threshold, the combination is marked as an effective synergistic combination, and the difference is recorded as the synergistic causal contribution value. The setting of the preset synergy threshold should take into account the estimation error range of the source intensity response model to avoid misjudging the difference caused by the model estimation error as a real synergy effect.
[0127] By identifying and determining the effectiveness of the three types of synergistic candidate combinations, we can systematically capture the synergistic effects of parameters generated under different mechanisms during the process. Co-convergence type synergistic candidate combinations reflect the mechanism by which multiple upstream parameters jointly influence an intermediate process state variable, thereby jointly driving emission anomalies; multi-path superposition type synergistic candidate combinations reflect the superposition mechanism by which multiple parameters simultaneously exert influence on the target node through their respective independent propagation links; and physically coupled type synergistic candidate combinations reflect the spatiotemporal near-synchronous anomaly mechanism between adjacent process units caused by physical processes such as material flow or heat transfer. The synergistic causal contribution values of the three types of effective synergistic combinations are then analyzed. Compared with the total causal contribution value of a single parameter By incorporating these findings into a ranking system based on source tracing and diagnostic results, the true driving mechanisms of emission anomalies can be more comprehensively reflected, providing a more accurate basis for decision-making regarding process optimization and anomaly handling.
[0128] A second aspect of the present invention provides a source tracing and diagnostic system for abnormal emission concentrations, comprising:
[0129] The data acquisition unit is used to acquire the emission concentration observation sequence of the target emission outlet and the flow topology data of the process pipeline network;
[0130] The deconvolution unit is used to construct a set of material transport paths from each process node to the target emission outlet based on the flow topology data. For abnormal moments in the emission concentration observation sequence, combined with the flow distribution ratio and residence time distribution of each transport path, the mixed concentration signal of the target emission outlet is subjected to inverse spatiotemporal deconvolution operation to decompose it into the independent emission contribution of each process node at historical moments, and the source intensity time series reconstruction result is obtained.
[0131] The causal identification unit is used to extract the time series data of the operating parameters of the target process node that makes an abnormal contribution in the source intensity time series reconstruction result before the abnormal period, identify the direct causal dependency by constructing the conditional independence test matrix of its operating parameters and remove the pseudo-correlation edges that are confused by common causes, construct the parameter coupling causal directed graph, and identify the set of shortest causal paths ending with abnormal emissions as the root cause parameter candidate set.
[0132] The contribution ranking unit is used to calculate the causal contribution strength between the state deviation of each parameter in the root cause parameter candidate set and the source intensity anomaly amplitude, and to rank them to generate source tracing diagnosis results.
[0133] A third aspect of the present invention provides an electronic device, comprising:
[0134] processor;
[0135] Memory used to store processor-executable instructions;
[0136] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0137] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0138] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method of diagnosing the source of an anomaly in emissions concentration, characterized in that, include: Acquire emission concentration observation sequences and flow topology data of the process pipeline network at the target emission outlet; Based on the flow topology data, a set of material transport paths from each process node to the target emission outlet is constructed. For abnormal moments in the emission concentration observation sequence, combined with the flow distribution ratio and residence time distribution of each transport path, the mixed concentration signal of the target emission outlet is subjected to inverse spatiotemporal deconvolution operation to decompose it into the independent emission contribution of each process node at historical moments, and the source intensity time series reconstruction result is obtained. For the target process node with abnormal contribution in the source intensity time series reconstruction result, extract the time series data of the operating parameters of the node before the abnormal period, identify the direct causal dependency by constructing the conditional independence test matrix of its operating parameters and remove the pseudo-correlation edges of common cause confusion, construct the parameter coupling causal directed graph, and identify the set of shortest causal paths with abnormal emission as the endpoint as the root cause parameter candidate set. Calculate the causal contribution strength between the state deviation of each parameter in the candidate set of root cause parameters and the abnormal amplitude of the source intensity, and sort them to generate the source tracing diagnosis results.
2. The method of claim 1, wherein, The steps to obtain the source intensity temporal reconstruction results include: For each transport path in the set of material transport paths, a transport kernel function is constructed based on the residence time distribution and flow distribution ratio of the transport path. The transport kernel function represents the mapping relationship from the unit emission of the starting process node of the transport path at a historical moment to the concentration contribution of the target emission outlet at the current moment. Its time dimension is determined by the residence time distribution, and its amplitude dimension is determined by the flow distribution ratio. For abnormal moments in the emission concentration observation sequence, the mixed concentration signal of the target emission outlet at the abnormal moment is extracted as the observation value. Based on the transmission kernel function of all transmission paths in the material transmission path set, the independent emission contribution of each process node at the historical moment is convolved with the transmission kernel function of the corresponding transmission path. The convolution results of all transmission paths are accumulated to establish the multipath superposition equation of the mixed concentration signal of the target emission outlet at the abnormal moment. The independent emission contribution of each process node at a historical moment in the multipath superposition equation is taken as the unknown to be solved. By performing inverse deconvolution operation on the multipath superposition equation, the independent emission contribution of each process node at a historical moment is obtained as the source intensity time series reconstruction result.
3. The method of claim 2, wherein, The steps for inverse deconvolution to obtain the temporal reconstruction results of source intensity include: The mixed concentration signal of the target emission outlet at an abnormal moment is decomposed into a fast-changing anomaly component and a slow-changing baseline component. The fast-changing anomaly component represents the instantaneous concentration fluctuation of the sudden emission event, and the slow-changing baseline component represents the background concentration trend of the continuous emission process. For the rapidly changing abnormal component, a subset of the transmission kernel functions in the material transmission path set whose residence time is shorter than a first preset threshold is selected for inverse deconvolution operation to solve the instantaneous emission contribution of each process node; for the slowly changing baseline component, a subset of the transmission kernel functions whose residence time is longer than a second preset threshold is selected for inverse deconvolution operation to solve the continuous emission contribution of each process node. The instantaneous emission contribution and the continuous emission contribution are aligned by time and then superimposed to obtain the independent emission contribution of each process node at a historical moment, which is used as the source intensity time series reconstruction result.
4. The method of claim 1, wherein, The steps for constructing a parameter-coupled causal directed graph and identifying the set of shortest causal paths ending with anomalous emissions as candidate root cause parameters include: For the time series data of the operation parameters of the target process node before the abnormal period, multiple time windows of different lengths are established by backtracking forward from the abnormal period. In each window, the conditional independence test value between the operation parameters is calculated and a conditional independence test matrix is constructed. Parameter pairs with monotonically increasing test values in the matrix are marked as direct causal dependencies, and parameter pairs with test values that first increase and then decrease are marked as indirect causal relationships. For parameter pairs with indirect causal relationships, identify intermediate transit nodes, remove the direct causal edges of the parameter pair, and retain the indirect causal edges that pass through the intermediate transit nodes. Based on direct causal dependencies and indirect causal edges, and combined with the occurrence time of intermediate transmission nodes, the operation parameters are divided into causal levels according to the time transmission order, and a parameter-coupled causal directed graph is constructed. Starting from the endpoint of the anomaly emission, the graph is traversed in reverse order of decreasing level to extract the set of causal paths that cross the fewest levels as the root cause parameter candidate set.
5. The method of claim 4, wherein, The steps for identifying intermediate transit nodes include: For parameter pairs with indirect causal relationships, the earliest and latest response times of the predecessor and successor parameters in the parameter pair are extracted, and the operational parameters whose response times are between the two are used as a set of candidate intermediate transit nodes. The candidate intermediate transit nodes are successively substituted as condition variables into the conditional independence test of the parameter pair. The candidate node with the largest decrease in test value is identified as the first-level intermediate transit node. The residual conditional independence test value after introducing the node is calculated. When the residual test value exceeds the decision threshold, the remaining candidate nodes are substituted again into the residual correlation test to identify the second-level intermediate transit node. The process is iterated until the residual test value is lower than the decision threshold to obtain a multi-level intermediate transit node sequence. The conditional independence test value between adjacent node pairs in the multi-level intermediate transit node sequence is calculated as the causal strength of the transit segment, and the causal strength is labeled to the corresponding transit segment in the parameter-coupled causal directed graph.
6. The method of claim 1, wherein, The steps for calculating the causal contribution strength between the state deviation of each parameter in the candidate set of root cause parameters and the abnormal amplitude of the source intensity, and sorting them to generate source tracing diagnosis results, include: For each parameter in the root cause parameter candidate set, the causal contribution intensity is calculated by comparing the state deviation of the parameter before the abnormal period with the abnormal amplitude of the corresponding target process node in the source intensity time series reconstruction result, and the direct causal contribution value is obtained. In the parameter coupled causal directed graph, all causal paths starting from the parameter and ending with abnormal emission are extracted. The state deviation of the intermediate node parameter is extracted along each path. After calculating the path transmission attenuation, the transmission contribution of all paths is accumulated to obtain the indirect causal contribution value. For the root cause parameter candidate set, multi-parameter combinations are screened based on the parameter coupling causal directed graph and process topology. The source intensity anomaly amplitude under the condition that each parameter in the parameter combination deviates simultaneously is calculated and compared with the sum of the individual deviation anomaly amplitudes of each parameter. The excess part is taken as the co-causal contribution value of the parameter combination. The direct and indirect causal contribution values of each parameter are summed to obtain the total causal contribution value of a single parameter, and then sorted to generate a single parameter source tracing ranking result. The parameter combinations are sorted according to their synergistic causal contribution values to generate a synergistic parameter source tracing ranking result. The single parameter source tracing ranking result and the synergistic parameter source tracing ranking result are merged to generate a source tracing diagnosis result.
7. The method of claim 6, wherein, The steps for identifying co-causal contributions include: In the parameter-coupled causal directed graph, parameter pairs with a common downstream node are identified, and parameter pairs whose state deviation of the common downstream node exceeds a preset deviation threshold are regarded as common convergent collaborative candidate combinations. In the parameter-coupled causal directed graph, multiple non-intersecting causal paths from different parameters to the abnormal emission target node are identified. The starting parameters of each path are extracted to form a parameter combination. The sum of the causal transmission strength of each path is calculated. The parameter combination whose sum of strength exceeds the third preset threshold is taken as a multi-path superposition type collaborative candidate combination. Based on the process topology, identify process nodes that are physically adjacent or have material flow exchange relationships, extract the corresponding operation parameters of the process nodes to form parameter combinations, and use the parameter combinations in which the time interval between the state deviation of each parameter in the combination is less than a preset time window as physical coupling type collaborative candidate combinations. For the three types of collaborative candidate combinations, the difference between the source intensity anomaly amplitude under the condition that the parameters deviate simultaneously and the sum of the individual deviation anomaly amplitudes of each parameter is calculated. Combinations with a positive difference that exceeds the preset collaborative threshold are marked as valid collaborative combinations and the difference is recorded as the collaborative causal contribution value.
8. An emissions concentration anomaly back-tracing diagnostic system for implementing the method of any one of claims 1-7, characterized in that, include: The data acquisition unit is used to acquire the emission concentration observation sequence of the target emission outlet and the flow topology data of the process pipeline network; The deconvolution unit is used to construct a set of material transport paths from each process node to the target emission outlet based on the flow topology data. For abnormal moments in the emission concentration observation sequence, combined with the flow distribution ratio and residence time distribution of each transport path, the mixed concentration signal of the target emission outlet is subjected to inverse spatiotemporal deconvolution operation to decompose it into the independent emission contribution of each process node at historical moments, and the source intensity time series reconstruction result is obtained. The causal identification unit is used to extract the time series data of the operating parameters of the target process node that makes an abnormal contribution in the source intensity time series reconstruction result before the abnormal period, identify the direct causal dependency by constructing the conditional independence test matrix of its operating parameters and remove the pseudo-correlation edges that are confused by common causes, construct the parameter coupling causal directed graph, and identify the set of shortest causal paths ending with abnormal emissions as the root cause parameter candidate set. The contribution ranking unit is used to calculate the causal contribution strength between the state deviation of each parameter in the root cause parameter candidate set and the source intensity anomaly amplitude, and to rank them to generate source tracing diagnosis results.
9. An electronic device, comprising: include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having stored thereon computer program instructions, wherein, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.