A Manufacturing Quality Anomaly Handling Method Based on Temporal Causal Knowledge Graph
Patent Information
- Application Number
- CN202610883809.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-11
AI Technical Summary
[0004]针对现有技术所存在的处置动作的前置因果条件难以判定、设备安全与处置效率难以动态平衡、多异常并发时缺乏自动冲突协调及决策参数无法自适应进化等问题,本发明提出一种基于时序因果知识图谱的制造质量异常处置方法
[0050]The beneficial effects of this invention are as follows: By constructing a knowledge graph that integrates temporal causal subgraphs, the dynamic transmission relationship of operations, delays, and effects between devices is explicitly represented using causal delay time as an attribute. This expands the static association of traditional knowledge graphs into a causal network with a time dimension, providing a precise time benchmark for determining the causal readiness of handling actions. This ensures that handling actions are executed after the causal impact has actually been transmitted to the target device, fundamentally avoiding invalid operations during the delay period and improving the causal temporal security of anomaly handling. Regarding handling process guidance, a handling process subgraph is obtained through the handling scheme nodes associated with anomaly nodes. Utilizing the handling step nodes and constraint edges within the handling process subgraph, a structured process template is provided for sequential decision-making, ensuring that decisions are always made within the logical framework specified by the process, avoiding blind global searches. In terms of action space construction, a dual filtering approach is employed based on logical constraints and a temporal causal subgraph. First, logical constraints eliminate candidate actions that are technologically infeasible. Then, causal readiness is determined using causal delay time, eliminating candidate actions whose causal premises are not yet met. This composite constraint extends the traditional single logical constraint into a synergistic constraint of logic and causality, ensuring that subsequent sequential decisions are always optimized within a set of actions that meet the necessary conditions, thus eliminating the process risk of secondary anomalies caused by inappropriate timing of actions. Regarding decision optimization, the reward function is composed of a weighted average of equipment safety deviation penalty and action efficiency reward. By adjusting the weighting coefficients and penalty mechanism, the rigid constraint of prioritizing safety in the industrial setting is mathematically embedded into the decision objective. This guides the sequential decision model to dynamically balance yield and timeliness while ensuring equipment safety, avoiding a trade-off between safety and efficiency, and significantly enhancing the dynamic adaptability of the decision. In terms of knowledge self-evolution, based on the actual operating data returned by the manufacturing execution equipment, the decision parameters on the association edges between abnormal nodes and handling plan nodes are updated online. This upgrades the traditional static knowledge graph into a dynamic knowledge graph with self-evolution capabilities, overcoming the shortcomings of insufficient long-term effectiveness of fixed decision parameters. It can adapt to working condition drift and experience accumulation, ensuring continuous optimization and long-term reliability of manufacturing quality anomaly handling. In terms of executability, by converting the sequence of handling instructions obtained from optimization into executable instructions and issuing them to the manufacturing execution equipment, a direct connection from decision optimization to physical execution is achieved, forming a complete technical closed loop of perception, cognition, decision-making, and execution.
Smart Images

Figure CN122734698A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of manufacturing quality control technology, and in particular to a method for handling manufacturing quality anomalies based on a temporal causal knowledge graph. Background Technology
[0002] In discrete manufacturing and process industries, manufacturing processes are susceptible to quality anomalies due to multiple factors, and rapid and safe handling is crucial for production line yield and equipment safety. In recent years, knowledge graph technology has been used for correlation retrieval between anomalies and handling solutions; meanwhile, causal testing methods have been used for offline root cause diagnosis, but they have not yet been integrated with real-time handling decisions.
[0003] However, existing technical solutions still face the following limitations in practical industrial applications. First, existing technical solutions mainly focus on the static correlation between anomalies and solutions, making it difficult to accurately identify the dynamic timing conditions required for the action to take effect, which may lead to abnormal downstream equipment status or out-of-tolerance process parameters. Second, existing technical solutions rely heavily on fixed rules or human experience, making it difficult to achieve dynamic optimization between equipment safety constraints and disposal efficiency goals. Third, existing technical solutions lack an automatic coordination mechanism based on process coupling, heavily relying on human intervention, resulting in slow response and a high risk of errors. Fourth, existing technical solutions typically use statically set decision parameters, which cannot adaptively correct deviations based on execution feedback to cope with operating condition drift, nor can they provide reasonable initial decision-making basis for new or low-frequency anomalies, thus limiting long-term effectiveness. Summary of the Invention
[0004] To address the shortcomings of existing technologies, such as difficulty in determining the pre-existing causal conditions for handling actions, the challenge of dynamically balancing equipment safety and handling efficiency, the lack of automatic conflict coordination when multiple anomalies occur concurrently, and the inability of decision parameters to adaptively evolve, this invention proposes a manufacturing quality anomaly handling method based on a temporal causal knowledge graph. This method constructs a knowledge graph that integrates temporal causal subgraphs, utilizes causal delay time constraints to determine decision timing, dynamically optimizes handling strategies with a dual drive of safety and efficiency, and automatically resolves concurrent conflicts by incorporating the topological features of the knowledge graph. Decision parameters are updated online through execution feedback, forming a closed loop from causal discovery to self-evolving decision-making. This comprehensively improves the causal temporal safety, dynamic adaptability, and long-term effectiveness of manufacturing quality anomaly handling.
[0005] The present invention achieves the above objectives through the following technical solutions:
[0006] A method for handling manufacturing quality anomalies based on temporal causal knowledge graphs, comprising:
[0007] A knowledge graph for handling manufacturing quality anomalies is constructed. The knowledge graph includes anomaly nodes, handling plan nodes, equipment nodes, and associated edges and time-series causal subgraphs connecting the anomaly nodes and handling plan nodes. The associated edges carry decision parameters. The time-series causal subgraph is composed of time-series causal directed edges connecting equipment nodes, and the time-series causal directed edges have causal delay time as an attribute.
[0008] When a manufacturing quality anomaly is detected, the corresponding anomaly node in the knowledge graph is determined, the environmental parameters associated with the anomaly node are obtained, and the corresponding handling plan node is obtained through the associated edges connected to the anomaly node. The handling process subgraph is obtained from the handling plan node; the handling process subgraph includes handling step nodes and constraint edges connecting the handling step nodes.
[0009] Based on the logical constraints and temporal causal subgraph of the constraint edges, logical constraint elimination and causal readiness determination are performed to construct the available action space in the current state.
[0010] Based on the sequential decision model, the candidate disposal steps in the available action space are optimized and solved to generate a sequence of disposal instructions; the optimization objective of the sequential decision model is determined based on a reward function that includes a penalty term for equipment safety deviation and a reward term for disposal efficiency.
[0011] The sequence of processing instructions is converted into executable instructions and sent to the manufacturing execution equipment;
[0012] The decision parameters on the associated edges are updated based on the feedback returned by the manufacturing execution equipment.
[0013] As a preferred embodiment of the present invention, the method for constructing the temporal causal subgraph includes:
[0014] Based on the topological connection relationship between device nodes as a prior constraint, device node pairs with topological connections are selected, and for each device node pair, at least one causal direction to be tested is specified, with one device as the cause-end device and the other device as the result-end device.
[0015] Acquire historical operational data for device node pairs; the historical operational data includes continuous numerical data and discrete event-type operation records.
[0016] The size of the sliding window is determined, and multiple sliding windows are generated based on a preset sliding step size. When there is continuous numerical data in the device node pair, the spectrum analysis of the continuous numerical data of the result device is performed first. If there is no continuous numerical data in the result device, the spectrum analysis of the continuous numerical data of the cause device is performed. The size of the sliding window is determined according to the period of the dominant frequency component of the spectrum. And / or, the size of the sliding window is adjusted according to the business stage label.
[0017] Within each sliding window, the causal direction to be tested is subjected to Granger causality test and / or transitive entropy algorithm to perform causality test. The causal delay time between device node pairs within the sliding window is calculated and a significance test is performed. The causal delay time that passes the significance test is marked as a significant causal delay time value.
[0018] Determine whether the proportion of significant causal delay time values generated meets the preset conditions; if it does, calculate the median of all significant causal delay time values as the final causal delay time of the device node pair in the causal direction, and generate a time-series causal directed edge from the cause device to the result device with the final causal delay time as the attribute; if it does not meet the conditions, it is not a time-series causal directed edge for the device node pair in the causal direction.
[0019] Connect all generated temporal causal directed edges and their corresponding device nodes to form a temporal causal subgraph.
[0020] As a preferred embodiment of the present invention, the environmental parameters include at least one of equipment operating status parameters, process parameters, environmental condition parameters, and the trigger time of operation events of equipment nodes;
[0021] The logical constraints of the constraint edges include timing constraints, dependency constraints, and mutual exclusion constraints; the timing constraints indicate the execution order of two processing steps, the dependency constraints indicate that the execution of one processing step is contingent upon the completion of another processing step, and the mutual exclusion constraints indicate that two processing steps cannot be executed in parallel within the same processing cycle.
[0022] As a preferred embodiment of the present invention, constructing the available action space in the current state includes:
[0023] Traverse all candidate processing steps in the processing flow subgraph and eliminate candidate processing steps that violate logical constraints;
[0024] For the remaining candidate handling steps, query the temporal causal directed edges in the temporal causal subgraph that have the device node associated with the candidate handling step as the result end; if the current time is less than the sum of the causal delay time of the edge and the most recent operation trigger time of the corresponding cause end device node, then remove the candidate handling step from the available action space.
[0025] The remaining candidate processing steps after elimination constitute the available action space.
[0026] As a preferred embodiment of the present invention, the reward function is constructed as follows:
[0027] The current state is defined by the real-time operating status and process parameter set of the equipment nodes associated with the abnormal node, and the current action is defined by the candidate handling steps selected from the available action space.
[0028] The reward function is composed of a weighted average of a device safety deviation penalty term and a handling efficiency reward term, with the weight coefficient of the device safety deviation penalty term being greater than the weight coefficient of the handling efficiency reward term.
[0029] The equipment safety deviation penalty is determined based on the deviation between the measured value and the rated value of the equipment's safe operating parameters after the current action is performed. A segmented penalty mechanism is adopted: when the deviation does not exceed the safety allowable deviation threshold, the equipment safety deviation penalty is zero; when the deviation exceeds the safety allowable deviation threshold but does not exceed the critical fault deviation threshold, the equipment safety deviation penalty increases monotonically with the deviation; when the deviation reaches or exceeds the critical fault deviation threshold, the equipment safety deviation penalty takes a saturation value.
[0030] The efficiency reward is determined based on the yield statistics and execution time of the current action, and is composed of yield and timeliness indicators weighted by a quality and timeliness balance coefficient. The yield indicator is determined based on the pass rate within the sliding statistical window. The timeliness indicator is determined based on the ratio of actual execution time to the standard allowable processing time.
[0031] In a preferred embodiment of the present invention, the weight ratio of the equipment safety deviation penalty item to the handling efficiency reward item is dynamically adjusted according to the severity of the current abnormal node; the severity is determined by the ratio of the number of time-series causal directed edges with the equipment node associated with the current abnormal node as the cause to the number of time-series causal directed edges with the device node as the result; wherein, the greater the severity, the greater the weight ratio.
[0032] The quality and timeliness balance coefficient in the processing efficiency reward item is dynamically adjusted according to the total number of workpieces in the sliding statistical window; wherein, the smaller the total number of workpieces, the more the quality and timeliness balance coefficient tends to prioritize ensuring yield; the larger the total number of workpieces, the more the quality and timeliness balance coefficient tends to prioritize ensuring processing timeliness.
[0033] As a preferred embodiment of the present invention, when multiple quality anomalies are detected simultaneously and there are action conflicts in multiple sequences of handling instructions due to mutual exclusion constraints, conflict resolution is performed on the multiple sequences of handling instructions:
[0034] Using the safety threshold of the device node corresponding to each anomaly as a hard constraint, minimizing the sum of the device safety deviation penalty terms corresponding to each anomaly as the first objective, and maximizing the sum of the handling efficiency reward terms corresponding to each anomaly as the second objective, we solve for the Pareto optimal solution set.
[0035] Sort the abnormal node pairs involved in the conflict in the knowledge graph from smallest to largest topological distance, select the solution corresponding to the abnormal node pair with the smallest topological distance from the Pareto optimal solution set for joint optimization, and output the conflict-free action sequence as the disposal instruction sequence.
[0036] As a preferred embodiment of the present invention, the step of converting the sequence of processing instructions into executable instructions includes:
[0037] Based on the action type of the disposal step node in the disposal instruction sequence, the corresponding instruction template is matched from the predefined action template library; the action template library includes control instruction templates, work order generation templates and notification templates; using the attribute information of the device nodes associated with the disposal step nodes in the knowledge graph, the parameters in the instruction template are filled in to generate atomic instructions as executable instructions;
[0038] For processing step nodes with corresponding temporal causal directed edges, the causal delay time of the temporal causal directed edge is written as the execution pre-wait time into the corresponding atomic instruction.
[0039] As a preferred embodiment of the present invention, updating the decision parameters on the associated edge based on the feedback result returned by the manufacturing execution equipment includes:
[0040] Obtain feedback results returned by the manufacturing execution equipment; the feedback results include the measured values of the safe operating parameters of the equipment nodes after the execution action, the number of defective products and the total number of workpieces in the sliding statistics window, the actual execution time, and a flag indicating whether secondary anomalies have been triggered;
[0041] Based on the feedback results, the actual reward value is calculated using a reward function;
[0042] Based on the actual reward value, the Q value on the associated edge corresponding to the action is updated using the Bellman equation; the Q value is a specific form of the decision parameter, and the associated edges are classified according to the action type.
[0043] When the secondary anomaly flag is set, the causal delay time of the time-series causal directed edge between the device node associated with the secondary anomaly and the device node associated with the current anomaly node is corrected to a sliding window statistical value based on the difference between the occurrence time of the secondary anomaly and the trigger time of the cause.
[0044] Based on the corrected causal delay time, the reward function for subsequent decisions is recalculated.
[0045] As a preferred embodiment of the present invention, when the Q value on the associated edge is updated, the update result is transferred to the associated edges of other abnormal nodes with similar structures to the current abnormal node, including:
[0046] Determine the set of anomalous nodes in the knowledge graph that have structural similarity to the current anomalous node; structural similarity is measured by the Jaccard similarity coefficient between the current anomalous node and the set of first-order neighbor node types of any anomalous node in the set of anomalous nodes.
[0047] For any edge with action type k on any abnormal node in the abnormal node set, update the Q value using the following formula:
[0048]
[0049] in, Abnormal nodes before the update Upper Action-related edges of value; For abnormal nodes after the update The edge associated with the k-th type of action of value; The target abnormal node to be dealt with at present Upper Action-related edges The current value; For transfer learning rate, ; The target abnormal node to be dealt with at present with abnormal nodes Structural similarity between them ; Number the action type.
[0050] The beneficial effects of this invention are as follows: By constructing a knowledge graph that integrates temporal causal subgraphs, the dynamic transmission relationship of operations, delays, and effects between devices is explicitly represented using causal delay time as an attribute. This expands the static association of traditional knowledge graphs into a causal network with a time dimension, providing a precise time benchmark for determining the causal readiness of handling actions. This ensures that handling actions are executed after the causal impact has actually been transmitted to the target device, fundamentally avoiding invalid operations during the delay period and improving the causal temporal security of anomaly handling. Regarding handling process guidance, a handling process subgraph is obtained through the handling scheme nodes associated with anomaly nodes. Utilizing the handling step nodes and constraint edges within the handling process subgraph, a structured process template is provided for sequential decision-making, ensuring that decisions are always made within the logical framework specified by the process, avoiding blind global searches. In terms of action space construction, a dual filtering approach is employed based on logical constraints and a temporal causal subgraph. First, logical constraints eliminate candidate actions that are technologically infeasible. Then, causal readiness is determined using causal delay time, eliminating candidate actions whose causal premises are not yet met. This composite constraint extends the traditional single logical constraint into a synergistic constraint of logic and causality, ensuring that subsequent sequential decisions are always optimized within a set of actions that meet the necessary conditions, thus eliminating the process risk of secondary anomalies caused by inappropriate timing of actions. Regarding decision optimization, the reward function is composed of a weighted average of equipment safety deviation penalty and action efficiency reward. By adjusting the weighting coefficients and penalty mechanism, the rigid constraint of prioritizing safety in the industrial setting is mathematically embedded into the decision objective. This guides the sequential decision model to dynamically balance yield and timeliness while ensuring equipment safety, avoiding a trade-off between safety and efficiency, and significantly enhancing the dynamic adaptability of the decision. In terms of knowledge self-evolution, based on the actual operating data returned by the manufacturing execution equipment, the decision parameters on the association edges between abnormal nodes and handling plan nodes are updated online. This upgrades the traditional static knowledge graph into a dynamic knowledge graph with self-evolution capabilities, overcoming the shortcomings of insufficient long-term effectiveness of fixed decision parameters. It can adapt to working condition drift and experience accumulation, ensuring continuous optimization and long-term reliability of manufacturing quality anomaly handling. In terms of executability, by converting the sequence of handling instructions obtained from optimization into executable instructions and issuing them to the manufacturing execution equipment, a direct connection from decision optimization to physical execution is achieved, forming a complete technical closed loop of perception, cognition, decision-making, and execution. Attached Figure Description
[0051] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1This is a flowchart of a manufacturing quality anomaly handling method based on temporal causal knowledge graph proposed in this invention; Figure 2 This is a flowchart illustrating the construction of a temporal causal subgraph in an embodiment of the present invention. Figure 3 This is a flowchart illustrating the construction of the available action space in an embodiment of the present invention; Figure 4 This is a flowchart illustrating the instruction conversion process in an embodiment of the present invention; Figure 5 This is a flowchart of the decision parameter update process in an embodiment of the present invention; Figure 6 This is a flowchart illustrating the Q-value migration process in an embodiment of the present invention. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention are within the scope of protection of the present invention.
[0053] like Figure 1 As shown, this is an embodiment of the present invention, which provides a method for handling manufacturing quality anomalies based on a temporal causal knowledge graph, including:
[0054] S1, Construct a knowledge graph for handling manufacturing quality anomalies.
[0055] The knowledge graph includes abnormal nodes, disposal plan nodes, device nodes, and associated edges connecting abnormal nodes and disposal plan nodes, as well as a temporal causal subgraph. The associated edges carry decision parameters; the temporal causal subgraph consists of directed temporal causal edges connecting device nodes, with the causal delay time as an attribute.
[0056] Furthermore, this knowledge graph integrates a heterogeneous graph network of static topology and dynamic temporal causality in the manufacturing process. Each temporal causal directed edge represents the impact of changes in the cause-end device on the result-end device after a certain time delay.
[0057] The decision parameters on the associated edges are used to characterize the long-term expected benefit of selecting the corresponding handling plan or handling step under a specific abnormal state. The decision parameters can be Q values, and each Q value is jointly associated with the abnormal node, the handling plan node, the action type of the handling step, and the current state. Specifically, for abnormal-plan associated edges with historical handling data, the initial value of the decision parameter is the historical success rate of the corresponding plan for that associated edge; for newly added associated edges lacking historical data, the initial value of the decision parameter is set to a preset benchmark value, which is determined according to the safety level of the production line, with lower values for production lines with higher safety levels. For example, for a semiconductor production line with a safety level of 1, the preset benchmark value is set to 0.1; for a general processing production line with a safety level of 3, the preset benchmark value is 0.3.
[0058] Anomaly nodes correspond to quality anomaly events that occurred during the historical manufacturing process; each anomaly node is associated with at least one equipment node, and the association is determined based on the production equipment involved when the quality anomaly occurred; in addition, anomaly nodes are also associated with the quality anomaly type label and historical occurrence frequency corresponding to the anomaly node, which are used for subsequent anomaly severity assessment.
[0059] Each handling plan node corresponds to a pre-defined standard handling procedure for a specific anomaly node. Each handling plan node is connected to at least one anomaly node via an associated edge, which indicates that the handling plan is used to address that anomaly. The content of each handling plan node originates from the standard operating procedure and is structured and broken down into a set of ordered handling step nodes, forming a handling process subgraph. When an anomaly node is connected to multiple handling plan nodes via multiple associated edges, all handling plan nodes pointed to by these edges are retrieved, and the handling process subgraphs in each handling plan node are merged as candidate sources of available action space in subsequent steps. The decision parameters carried on each associated edge are used to distinguish the priority of different plans in subsequent sequential decision optimization.
[0060] Due to the physical coupling and process flow relationships between manufacturing equipment, the operation of a single piece of equipment often requires a certain delay before it triggers a change in the state of downstream equipment. To characterize this dynamic transmission relationship, a time-series causal subgraph was constructed.
[0061] like Figure 2 As shown, the methods for constructing time-series causal subgraphs include:
[0062] S11, based on the topological connection relationship between device nodes as a prior constraint, filter out device node pairs with topological connections, and for each device node pair, specify at least one causal direction to be tested, with one device as the cause end device and the other device as the result end device.
[0063] Furthermore, this step can filter out node pairs that are physically impossible to have a causal relationship, significantly reducing computational complexity.
[0064] S12, obtain historical operating data of device node pairs; historical operating data includes continuous numerical data and discrete event operation records.
[0065] In one specific embodiment, continuous numerical data includes sensor sample values such as temperature, pressure, and vibration frequency that change continuously over time; discrete event operation records include discrete events with timestamps, such as valve opening and closing, motor starting and stopping, and parameter setting changes.
[0066] S13, determine the size of the sliding window and generate multiple sliding windows based on the preset sliding step size; wherein, when there is continuous numerical data in the device node pair, the continuous numerical data of the result device is given priority for spectrum analysis; if there is no continuous numerical data in the result device, the continuous numerical data of the cause device is given spectrum analysis, and the size of the sliding window is determined according to the period of the dominant frequency component of the spectrum; and / or, the size of the sliding window is adjusted according to the business stage label.
[0067] Furthermore, since the manufacturing process is non-stationary, a fixed window cannot accurately capture dynamic causal relationships. Therefore, a multi-scale adaptive sliding window is adopted.
[0068] For continuous numerical data, spectral analysis such as Fast Fourier Transform is performed to extract the period of the dominant frequency component. The size of the sliding window is then set to an integer multiple of this period, such as 3 to 5 times, to ensure that the sliding window contains the complete dynamic fluctuation period. When the device node pair only contains discrete event-type operation records, different sliding window sizes are used for different business stages based solely on the business stage labels issued by the manufacturing execution system, ensuring that the data within the window is under the same steady-state condition.
[0069] The preset sliding step size is 1 / 4 to 1 / 2 of the sliding window size. For example, when the sliding window size is 200 sampling points, the preset sliding step size is set to 100 sampling points, with 50% data overlap between adjacent windows. This setting ensures that there are a sufficient number of significant windows that reflect the dynamic changes in the operating conditions in the subsequent median taking step, thus enhancing the statistical robustness of the final causal delay time.
[0070] S14. Within each sliding window, for the causal direction to be tested, the Granger causality test and / or the transitive entropy algorithm are used to perform causality testing. The causal delay time between the device node pairs within the sliding window is calculated and a significance test is performed. The causal delay time that passes the significance test is marked as a significant causal delay time value.
[0071] Furthermore, the purpose of causality testing is not only to determine whether a causal relationship exists, but also to measure the time required for a causal event to propagate to an outcome event, i.e., the causal delay time.
[0072] When historical operational data includes device node pairs containing both continuous numerical data and discrete event-type operation records, the discrete event-type operation records are encoded into binary pulse sequences. The event occurrence time is assigned a value of 1, and all other times are assigned a value of 0. The sampling frequency is aligned with that of the continuous numerical data. If only discrete event-type operation records exist, resampling is performed at a preset sampling frequency. After the encoded binary pulse sequences and continuous numerical data are aligned along a unified time axis within the same sliding window, they participate in Granger causality tests or transit entropy calculations, thus uniformly handling heterogeneous data in causality tests. The preset sampling frequency is set based on the typical time interval of discrete events. In one specific embodiment, the preset sampling frequency is set to 1Hz to ensure that the average event interval contains 5 to 20 sampling points.
[0073] Taking the Granger causality test as an example, for each causal direction to be tested, the causal delay time is calculated within each sliding window, as follows:
[0074] For the time series Y of the result terminal device, an autoregressive model using only its own historical data is constructed as the baseline prediction model, with the autoregressive order being... Determined according to the Akaike Information Criterion. Then, for each candidate time delay step... An extended prediction model is constructed, which, based on the baseline prediction model, further incorporates the effect of delay on the cause device X. , , …, Historical data of steps, i.e., the length of the added steps is The lag term window. The F-statistic is calculated by comparing the sum of squared predicted residuals of the baseline and extended prediction models. This F-statistic quantifies the impact of introducing causal devices. The degree of improvement in the predictive ability of the result-side equipment is calculated after a lag term window starting from the causal side. The autoregressive order is also used as the lag order of the causal side equipment; that is, the causal side lag order and the autoregressive order take the same value.
[0075] Traversal ∈[1,50], calculate each The F-statistic is used to predict the outcome device at the given delay starting point. When the F-statistic reaches its maximum value, it indicates that the cause-side device has the strongest predictive ability for the result-side device at that delay starting point, and the corresponding time delay step is... This refers to the causal delay time calculated within this window; the maximum value of the F-statistic is compared with the significance level threshold; if the maximum value of the F-statistic is greater than the significance level threshold, then a significant causal relationship from X to Y exists within this window, and this is recorded as... The significant causal delay time value for this window is marked. If the maximum value of the F-statistic is less than or equal to the significance level threshold, it is determined that there is no significant causal relationship between X and Y within this window, and no significant causal delay time value is generated in this window. The significance level threshold is determined by referring to the F-distribution table based on the numerator and denominator degrees of freedom. The numerator degrees of freedom are the lag order of the causal device, and the denominator degrees of freedom are the number of samples within the sliding window minus the autoregressive order, the causal lag order, and the constant term.
[0076] Taking the transitive entropy algorithm as an example, it calculates the time delay of the cause-end device at different times. The conditional mutual information between the result-end devices is considered. When the transfer entropy reaches its peak, it indicates that the flow of information from the cause-end device to the result-end device is at its maximum. The time delay step corresponding to this peak is the causal delay time within this window. The significance is determined by a permutation test. If the test passes, the causal delay time is marked as a significant causal delay time value; if the test fails, no significant causal delay time value is generated in this window.
[0077] S15, determine whether the generation ratio of significant causal delay time values meets the preset conditions; if it does, calculate the median of all significant causal delay time values as the final causal delay time of the device node pair in the causal direction, and generate a time-series causal directed edge from the cause device to the result device with the final causal delay time as the attribute; if it does not meet the conditions, it is not a time-series causal directed edge for the device node pair in the causal direction.
[0078] Furthermore, preset conditions are used to determine whether the obtained significant causal delay time value is sufficient to support the generation of a reliable causal directed edge. The specific conditions are set according to different requirements for the sensitivity and reliability of causal inference.
[0079] In one specific embodiment, to maximize the discovery of potential causal links, a preset condition is that at least one significant causal delay time value exists. This condition has the highest sensitivity and is suitable for exploratory analysis or scenarios with high requirements for the integrity of causal links.
[0080] In another specific embodiment, to enhance the robustness of temporal causal attributes and filter out occasional noise interference, a preset condition is established: the proportion of sliding windows generating significant causal delay values to the total number of sliding windows reaches a preset threshold. This condition, by requiring the causal relationship to be stably reproduced across multiple windows, reduces the risk of spurious causality caused by a single window accidentally passing the significance test, and is suitable for scenarios with high requirements for the reliability of causal inference. The preset threshold is set according to the strictness of the application scenario, for example, set to 30% to 50%.
[0081] S16, connect all generated temporal causal directed edges and their corresponding device nodes to form a temporal causal subgraph.
[0082] S2, when a manufacturing quality anomaly is detected, the corresponding anomaly node in the knowledge graph is determined, the environmental parameters associated with the anomaly node are obtained, and the corresponding handling solution node is obtained through the associated edge connected to the anomaly node. The handling process subgraph is obtained from the handling solution node.
[0083] Furthermore, the manufacturing execution system or quality inspection equipment monitors product quality indicators in real time. When a quality characteristic value exceeds a preset specification limit, an event record is generated, containing an anomaly code, the time of occurrence, associated equipment, and batch information. This anomaly code is then matched with pre-defined anomaly nodes in a knowledge graph to locate the unique corresponding anomaly node.
[0084] In the knowledge graph, anomaly nodes are connected to one or more solution nodes via edges carrying decision parameters. A solution node represents a known handling strategy for the anomaly; it is not an atomic action but a subgraph containing a structured process. Traversing along the edges to the corresponding solution node, the solution process subgraph is parsed from that node. When an anomaly node is connected to multiple edges, the solution process subgraphs in all solution nodes pointed to by these edges are obtained. Candidate handling steps in each subgraph undergo logical constraints and causal timing filtering in subsequent step S3, and are then optimized and selected based on the decision parameters in step S4.
[0085] Environmental parameters include at least one of the following: equipment operating status parameters, process parameters, environmental condition parameters, and the trigger time of operation events for equipment nodes.
[0086] Furthermore, environmental parameters associated with the abnormal node are collected through the sensor network and control system interface.
[0087] Equipment operating status parameters are continuous numerical data that reflect the equipment's own operating conditions, such as equipment vibration amplitude, spindle speed, motor current, and temperature controller output percentage.
[0088] Process parameters reflect the current settings and execution status of the processing, such as the current gas flow rate setting, chamber pressure, feed rate, and cleaning fluid ratio.
[0089] Environmental condition parameters reflect data about the macroscopic environment of the workshop, such as the cleanliness level of the workshop, ambient temperature and humidity, and cooling water inlet temperature.
[0090] Obtaining discrete-time data, such as the trigger time of the operation event, provides a time reference for the pruning based on the temporal causal subgraph in the subsequent step S3. Only by clearly identifying the precise trigger time of the operation on the cause-end device can we accurately determine, in conjunction with the causal delay time, whether the result-end device has been affected by the operation at the current moment, thereby avoiding taking ineffective or harmful candidate measures during the delay period.
[0091] The trigger time of an operation event for a device node is a key parameter for constructing the available action space. This includes not only the trigger time of operation events for the device node directly associated with the malfunctioning node, but also the trigger times of operation events for the causal device nodes connected to the malfunctioning node via directed edges in the temporal causal subgraph. Specifically, the trigger time of an operation event refers to the timestamp of the most recent action performed by the corresponding device, such as the most recent valve opening, the most recent recipe change completion, or the most recent device restart. This timestamp originates from the operation logs of the manufacturing execution system or the historical records of the device's programmable logic controller.
[0092] The processing flow subgraph includes processing step nodes and constraint edges connecting these nodes. The logical constraints of the constraint edges include timing constraints, dependency constraints, and mutual exclusion constraints. Timing constraints indicate the execution order of two processing steps, dependency constraints indicate that the execution of one processing step depends on the completion of the other, and mutual exclusion constraints indicate that two processing steps cannot be executed in parallel within the same processing cycle.
[0093] Furthermore, the process flow sub-graph is a structured abstraction of manual handling experience, with each handling step node representing the smallest operational unit, such as starting a backup vacuum pump or issuing a cleaning work order. To prevent the generated action sequence from being unexecutable in engineering or causing safety accidents, strict logical constraints are defined through constraint edges.
[0094] Taking abnormal chamber pressure in a semiconductor etching process as an example, the handling process flow diagram includes the following handling steps: Step a1 is to reduce the process gas flow rate; Step a2 is to start the backup vacuum pump; Step a3 is to perform chamber cleaning; and Step a4 is to restore the process parameters to their rated values. There is a timing constraint between Step a1 and Step a2, with a1 required to be executed before a2; there is a dependency constraint between Step a3 and Step a4, with a4 requiring the completion of a3; and there is a mutual exclusion constraint between Step a1 and Step a3, meaning they cannot be executed concurrently within the same handling cycle to avoid process instability caused by simultaneous gas flow rate adjustment and cleaning operations.
[0095] By using the above methods, not only are real-time multi-source heterogeneous data incorporated into the decision state space, but more importantly, the stringent process logic and safety regulations of the industrial site are hard-coded in the form of constraint edges.
[0096] S3, based on the logical constraints and time-series causal subgraph of the constraint edges, performs logical constraint elimination and causal readiness determination to construct the available action space in the current state.
[0097] like Figure 3 As shown, the available action space in the current state is constructed, including:
[0098] S31, traverse all candidate processing steps in the processing flow subgraph and eliminate candidate processing steps that violate logical constraints.
[0099] Furthermore, for each candidate disposal step in the disposal process subgraph, the constraint edges involved are checked one by one. If the execution of a candidate disposal step would violate any of the timing constraints, dependency constraints, or mutual exclusion constraints, then the step is removed from the candidate set. Specifically, for timing constraints, if the preceding step has not yet been executed, the subsequent step is removed; for dependency constraints, if the dependent step has not yet been completed, the dependent step is removed; for mutual exclusion constraints, if one of two mutually exclusive steps has been selected for execution in the current disposal cycle, the other is removed.
[0100] S32, for the remaining candidate handling steps, query the temporal causal directed edges in the temporal causal subgraph that have the device node associated with the candidate handling step as the result end; if the current time is less than the sum of the causal delay time of the edge and the most recent operation trigger time of the corresponding cause end device node, then remove the candidate handling step from the available action space.
[0101] Furthermore, this step is based on a temporal causal subgraph to determine causal readiness. Only when the operation of the cause device has passed through the causal delay time and the corresponding impact has been transmitted to the result device, is the execution of the candidate disposal step on the result device effective. If the candidate disposal step is executed on the result device during the delay period, the result device has not yet been affected by the operation of the cause device. At this time, the candidate disposal step may be ineffective or even produce adverse consequences.
[0102] For each remaining candidate action step, firstly, the associated device node is determined. Then, all temporally causal directed edges with that device node as the result end are queried in the temporal causal subgraph. For each found temporally causal directed edge, the corresponding causal delay time attribute and the most recent operation trigger time of the cause device node are obtained, and their sum is calculated as the causal readiness time. The current time is compared with this causal readiness time; if the current time is less than the causal readiness time, the impact of the cause device operation has not yet been transmitted to the result device, and the execution conditions of the candidate action step are not yet mature, so it is eliminated; if the current time is greater than or equal to the causal readiness time, the impact of the cause device operation has been transmitted to the result device, and the candidate action step has passed the causal readiness test.
[0103] S33, the remaining candidate processing steps after elimination constitute the available action space.
[0104] Furthermore, each candidate action step in the available action space simultaneously satisfies both logical constraints and causal readiness conditions, ensuring that the set of candidate action steps optimized by the subsequent sequential decision model is both engineering-executable and causally and temporally reasonable.
[0105] Taking abnormal chamber pressure in semiconductor etching process as an example, the handling process sub-diagram includes step a1 of reducing process gas flow rate, step a2 of starting standby vacuum pump, step a3 of performing chamber cleaning, and step a4 of restoring process parameters to rated values, with the constraint relationship as described in S2.
[0106] Assume the current processing cycle has just begun and no steps have been executed yet. Because there is a timing constraint between a1 and a2, and a1 must precede a2, and the prerequisite step a1 for a2 has not yet been executed, a2 is eliminated. Because there is a dependency constraint between a3 and a4, and a4 depends on a3 for completion, and a3 has not yet been executed, a4 is eliminated. Because there is a mutual exclusion constraint between a1 and a3, and neither has been selected, the mutual exclusion constraint does not trigger elimination for now. After eliminating logical constraints, the remaining candidate processing steps are a1 and a3.
[0107] Assume the device node associated with step a1 is the mass flow controller, and the device node associated with step a3 is the chamber. In the temporal causal subgraph, there exists a temporal causal directed edge with the gas pipeline regulating valve as the cause and the mass flow controller as the result, with a causal delay of 2 seconds; there also exists a temporal causal directed edge with the mass flow controller as the cause and the chamber as the result, with a causal delay of 5 seconds.
[0108] The gas pipeline regulating valve was last triggered at 10:00:03, the mass flow controller was last triggered at 10:00:08, and the current time is 10:00:10.
[0109] For step a1, the associated mass flow controller is the result end, and the corresponding cause end is the gas pipeline regulating valve. The causal readiness time is 10:00:03 + 2 seconds, i.e., 10:00:05. The current time 10:00:10 is greater than 10:00:05, the causal effect has arrived, and step a1 passes the causal readiness test.
[0110] For step a3, the associated chamber is the result end, and the corresponding cause end is the mass flow controller. The causal readiness time is 10:00:08 + 5 seconds = 10:00:13. Since the current time 10:00:10 is less than 10:00:13, the effect of the most recent operation of the mass flow controller has not yet been transmitted to the chamber, and the execution conditions for step a3 are not yet mature; therefore, it is eliminated.
[0111] Ultimately, the available action space contains only step a1. In subsequent step S4, the optimal action will be selected from this available action space based on the sequential decision model. After step a1 is completed, the system will re-enter S3 to update the available action space. At this point, the temporal constraints of a2 will have been satisfied, and the causal readiness of a3 may also have been satisfied, thus expanding the available action space accordingly.
[0112] S4, based on the sequential decision model, optimizes and solves the candidate disposal steps in the available action space to generate a sequence of disposal instructions.
[0113] The optimization objective of the sequential decision-making model is determined based on a reward function that includes a penalty term for equipment safety deviation and a reward term for handling efficiency.
[0114] Furthermore, the sequential decision model uses the real-time operating status and process parameter set of the equipment nodes associated with the abnormal nodes in the knowledge graph as the state space, and the candidate disposal steps in the available action space as the action space. That is, the actions in the sequential decision model are the candidate disposal steps, and the decision parameters carried on the associated edges are used as the basis for the action selection strategy. In each decision step, the action with the highest long-term expected benefit in the current state is selected, and the disposal instruction sequence is gradually generated.
[0115] The reward function is constructed as follows:
[0116] The current state is defined by the real-time operating status and process parameter set of the equipment nodes associated with the abnormal node, and the current action is defined by the candidate handling steps selected from the available action space.
[0117] The reward function is composed of a weighted average of a penalty term for equipment safety deviation and a reward term for handling efficiency, with the weight coefficient of the penalty term for equipment safety deviation being greater than that of the reward term for handling efficiency.
[0118] Furthermore, the negative value of the equipment safety deviation penalty term represents the suppression of unsafe states, while the positive value of the handling efficiency reward term represents the encouragement of efficient handling. The weight coefficient of the equipment safety deviation penalty term is greater than that of the handling efficiency reward term, ensuring that equipment safety always takes precedence over handling efficiency. By embedding the temporal causal relationship into the action space constraints, the reward function can focus on evaluating safety and efficiency, avoiding reward signal bias caused by causal temporal misalignment.
[0119] In one specific embodiment, the reward function satisfies:
[0120]
[0121] in, To perform the action Then from the current state Transition to state The reward value at that time; The penalty for equipment safety deviation is a negative value. The efficiency bonus item is positive. The weighting coefficient for the equipment safety deviation penalty item; The weighting coefficient for efficiency reward items, and , .
[0122] Equipment safety deviation penalty items Based only on the state after the transfer The independent variable is safety deviation, because safety deviation is an objective assessment of the equipment's state after the action is performed, and it is entirely determined by the deviation between the measured values and rated values of each safety operating parameter in the post-transfer state, regardless of the action performed; while the disposal efficiency bonus item... Simultaneously depends on the current state and actions This is because yield statistics are related to the current state, while execution time is directly related to the specific actions.
[0123] The equipment safety deviation penalty is determined based on the deviation between the measured value and the rated value of the equipment's safe operating parameters after the current action is performed. A segmented penalty mechanism is adopted: when the deviation does not exceed the safety allowable deviation threshold, the equipment safety deviation penalty is zero; when the deviation exceeds the safety allowable deviation threshold but does not exceed the critical fault deviation threshold, the equipment safety deviation penalty increases monotonically with the deviation; when the deviation reaches or exceeds the critical fault deviation threshold, the equipment safety deviation penalty takes a saturation value.
[0124] In one specific embodiment, the device safety deviation penalty item The calculation formula is:
[0125]
[0126]
[0127] in, The total number of safe operating parameters for device nodes; The state after the transition The Middle Measured values of each safe operating parameter; For the first A normalized penalty function for each safe operating parameter; For the first The rated values of each safe operating parameter; For the first The safety allowable deviation threshold for each safe operating parameter; For the first Critical fault deviation threshold for each safe operating parameter; , , and Having the same dimensions, and , .
[0128] when When the equipment is within its safe operating range, no penalty is required; when At this point, the equipment has entered the critical fault region. The penalty is no longer increased to avoid the vanishing gradient of the reward function, which could cause the sequential decision-making model to fail to converge. A fixed saturation value of -1 is taken to ensure that the maximum penalty magnitude of a single parameter is symmetrical to the maximum value of the disposal efficiency reward term in terms of scale, facilitating the weighting coefficients. and Unified adjustment. If the absolute value of the saturation value is too large, it will cause the equipment safety deviation penalty item to excessively suppress the handling efficiency reward item after weighting, making the decision too conservative; if the absolute value of the saturation value is too small, it will not be able to fully suppress the critical dangerous state.
[0129] when When the value is 0, it means that the current device node has no safety operation parameter monitoring items and no device safety deviation penalty items. The default value of 0 is used, which means there is no device security deviation penalty.
[0130] Equal-weighted averages are used instead of weighted averages because the rated values and deviation thresholds of each safety operating parameter have been set separately in the normalized penalty function, and the differences in the safety impact of different parameters have been addressed through their respective... and This means that there is no need to introduce additional weighting coefficients to increase the complexity of parameter tuning.
[0131] Safety tolerance threshold and critical fault deviation threshold The specifications are determined based on the safe operating parameters of each equipment node. Specifically, This is the upper limit of the normal allowable fluctuation range of this parameter as specified in the equipment manufacturer's technical specifications; The parameter deviation threshold that triggers equipment safety interlocks or automatic shutdown protection.
[0132] The efficiency reward is determined based on the yield statistics and execution time of the current action. It is composed of yield and timeliness indicators weighted by a quality and timeliness balance coefficient. The yield indicator is determined based on the pass rate within the sliding statistical window, and the timeliness indicator is determined based on the ratio of actual execution time to the standard allowable processing time.
[0133] In one specific embodiment, the processing efficiency reward item The calculation formula is:
[0134]
[0135] in, Current state The quality and timeliness balance coefficient is as follows. ; Current state The total number of workpieces that have completed quality inspection within the sliding statistics window; Current state The number of defective products confirmed by quality inspection is displayed in this sliding statistics window; For the current action The actual execution time; Current state The standard allowable processing time for the anomaly type corresponding to the anomaly node in the middle is and With the same time dimension, this standard allows the handling time to be determined based on the average handling time statistics for this type of abnormality in the standard operating procedure, or set by the process engineer according to the production line cycle requirements.
[0136] when When it is 0, the yield index Take the preset initial value of 1; when When it is 0, the timeliness indicator The preset upper limit value of 1 indicates that the action is completed instantly, with optimal efficiency.
[0137] The weight ratio of the equipment safety deviation penalty item to the handling efficiency reward item is dynamically adjusted according to the severity of the current abnormal node. The severity is determined by the ratio of the number of time-series causal directed edges with the equipment node associated with the current abnormal node as the cause to the number of time-series causal directed edges with the device node as the result. The greater the severity, the greater the weight ratio.
[0138] Furthermore, let A be the number of time-series causal directed edges with the device node associated with the current abnormal node as the cause, and B be the number of time-series causal directed edges with the device node as the result, and let the severity be... When the severity is high, the anomaly of the device node is likely to propagate downstream. In this case, the weight of the safety penalty should be increased to select the response action more prudently. Therefore, the weight ratio of the device safety deviation penalty item to the response efficiency reward item increases with the severity.
[0139] In one specific embodiment, when B is not 0, the formula for calculating the ratio of the weighting coefficients of the equipment safety deviation penalty item and the handling efficiency reward item is as follows:
[0140]
[0141] in, As to the severity, ; The baseline weight ratio, ranging from 2 to 5, is used to ensure that the weight coefficient of the equipment safety deviation penalty item is still greater than the weight coefficient of the handling efficiency reward item even at the lowest severity level. To adjust the sensitivity coefficient, the value ranges from 0.5 to 2.0. In one specific embodiment, during the semiconductor etching process, The value is 3. The value is 1.0.
[0142] When the severity is close to zero, the device node will hardly propagate the anomaly downstream. At this time, the weight ratio of the device safety deviation penalty item to the handling efficiency reward item is close to the benchmark value, and the weight of the handling efficiency reward item is greater. When the severity is much greater than 1, the device node becomes a key causal propagation hub, the weight ratio of the device safety deviation penalty item to the handling efficiency reward item increases significantly, and safety considerations take the lead.
[0143] When B is 0 and A>0, the device node associated with the current abnormal node is not used as the result end of any temporal causal directed edge. The weight coefficient ratio of the device safety deviation penalty item and the handling efficiency reward item is taken as a preset upper limit value, which ranges from 10 to 20.
[0144] When B=0 and A=0, meaning the device node associated with the current abnormal node is an isolated node in the temporal causal subgraph and does not participate in any causal propagation, the severity is set to the preset default value of 0. The weight ratio of the device safety deviation penalty item to the handling efficiency reward item is equal to the baseline weight ratio. The preferred baseline weight ratio is 1:1, indicating that in the absence of causal structure information, safety deviation and handling efficiency are given equal importance.
[0145] The quality and timeliness balance coefficient in the processing efficiency reward item is dynamically adjusted based on the total number of workpieces in the sliding statistical window; the smaller the total number of workpieces, the more the quality and timeliness balance coefficient tends to prioritize ensuring yield; the larger the total number of workpieces, the more the quality and timeliness balance coefficient tends to prioritize ensuring processing time.
[0146] Furthermore, when the total number of workpieces that have completed quality inspection within the sliding statistical window is small, the statistical results fluctuate greatly, and the impact of a single defective product on the yield index is amplified. In this case, priority should be given to ensuring the yield to avoid statistical bias from misleading decision-making, and the balance coefficient between quality and timeliness tends to be larger. When the total number of workpieces that have completed quality inspection within the sliding statistical window is large, the yield statistics have tended to be stable. In this case, the focus should be on improving the processing time to reduce production line downtime losses, and the balance coefficient between quality and timeliness tends to be smaller.
[0147] In one specific embodiment, the formula for calculating the quality-time balance coefficient is as follows:
[0148]
[0149] in, This is an adjustment coefficient, ranging from 0.01 to 0.05. In one specific embodiment, in the semiconductor etching process, the sliding statistical window is set to the most recent 50 wafers. The value is 0.02.
[0150] Its function is to adjust the relative importance between yield and timeliness indicators. When When the yield is small, the statistical fluctuation of yield is large. A tendency towards larger values gives yield metrics greater weight, thus reducing the interference of statistical noise on decision-making; when When the yield is large, the yield statistics are reliable. Lower values give greater weight to timeliness indicators, thus reducing production line downtime losses. Therefore, The design is based on The independent variable is not the yield level.
[0151] When multiple quality anomalies are detected simultaneously and there are action conflicts in multiple sequences of handling instructions due to mutual exclusion constraints, conflict resolution is performed on the multiple sequences of handling instructions:
[0152] Using the safety threshold of the device node corresponding to each anomaly as a hard constraint, the first objective is to minimize the sum of the device safety deviation penalty terms corresponding to each anomaly, and the second objective is to maximize the sum of the handling efficiency reward terms corresponding to each anomaly. The Pareto optimal solution set is then solved.
[0153] Sort the abnormal node pairs involved in the conflict in the knowledge graph from smallest to largest topological distance, select the solution corresponding to the abnormal node pair with the smallest topological distance from the Pareto optimal solution set for joint optimization, and output the conflict-free action sequence as the disposal instruction sequence.
[0154] Furthermore, the safety threshold is the aforementioned critical fault deviation threshold, meaning that no action plan should cause any safe operating parameter of any equipment node to reach or exceed the critical fault deviation threshold.
[0155] In scenarios with multiple concurrent anomalies, the independently generated sequences of handling instructions for each anomaly may contain mutually exclusive steps, and direct execution would violate constraints. Conflict resolution takes the hard satisfaction of safety constraints as a prerequisite. Under this premise, Pareto optimization is performed with the dual objectives of minimizing the sum of equipment safety deviation penalties and maximizing the sum of handling efficiency rewards to obtain a set of non-dominated solutions. Then, the topological structure information of the knowledge graph is used to select solutions; the smaller the topological distance between anomaly node pairs, the higher the degree of physical or technological coupling between the associated devices, and the stronger the necessity for joint optimization. Therefore, the Pareto solution corresponding to the joint optimization of the anomaly node pair with the smallest topological distance is preferentially selected to ensure that anomalies with tight physical coupling are handled in a coordinated manner.
[0156] Taking a semiconductor etching line as an example, assume that two quality anomalies are detected simultaneously: abnormal chamber pressure and abnormal wafer temperature. The instruction sequence for handling abnormal chamber pressure includes step a1 of reducing process gas flow, while the instruction sequence for handling abnormal wafer temperature includes step b1 of increasing heater power. However, a1 and b1 are mutually exclusive because reducing gas flow alters the chamber's thermal conductivity, which is incompatible with simultaneously increasing heater power. Using the safety operating parameter thresholds of the equipment nodes corresponding to the two anomalies as hard constraints, all feasible combinations of a1 and b1 are enumerated. In the solution space satisfying the safety hard constraints, the Pareto optimal solution set is solved with the first objective of minimizing the sum of the chamber pressure deviation penalty and the wafer temperature deviation penalty, and the second objective of maximizing the sum of the handling efficiency rewards for the two anomalies. In the knowledge graph, the topological distance between the chamber node and the heater node is 1, while the topological distance between the chamber node and the cooling water node is 2. Joint optimization solutions involving chamber and heater node pairs are preferentially selected, and conflict-free action sequences are output.
[0157] S5 converts the sequence of processing instructions into executable instructions and sends them to the manufacturing execution equipment.
[0158] like Figure 4 As shown, the sequence of disposal instructions is converted into executable instructions, including:
[0159] S51, based on the action type of the action step node in the action instruction sequence, matches the corresponding instruction template from the predefined action template library. The action template library includes control instruction templates, work order generation templates, and notification templates.
[0160] Furthermore, the action template library predefines instruction templates corresponding to various handling steps. Each template defines the instruction format and parameter placeholders required for that type of action. Control instruction templates are used to generate direct operation instructions for the equipment's programmable logic controller, such as setting parameter values and starting / stopping equipment. Work order generation templates are used to generate equipment maintenance or cleaning work orders, including fields such as work order type, priority, and target equipment identifier. Notification templates are used to generate alarm and notification messages for on-site operators or the dispatch system, including fields such as anomaly description, suggested operations, and deadline.
[0161] S52 utilizes the attribute information of the device nodes associated with the processing step nodes in the knowledge graph to fill in the parameters in the instruction template and generate atomic instructions as executable instructions.
[0162] Furthermore, each processing step node is associated with a specific device node in the knowledge graph. The device node carries attribute information such as device identifier, communication address, current parameter settings, and safety parameter ranges. After matching the corresponding instruction template based on the action type of the processing step node, the required parameter values are extracted from the associated device node attributes and filled into the parameter placeholders in the template, generating a complete atomic instruction. An atomic instruction is the smallest instruction unit that the manufacturing execution device can directly parse and execute.
[0163] S53, for processing step nodes with corresponding temporal causal directed edges, the causal delay time of the temporal causal directed edge is written as the execution pre-wait time into the corresponding atomic instruction.
[0164] Furthermore, when a device node associated with a certain processing step node has an incoming edge in the temporal causal subgraph, it indicates that the state of that device is affected by the delay of the upstream device operation. To ensure that the atomic instruction is executed only after the causal premise is satisfied, the causal delay time of the directed edge of the temporal causal sequence is written into the pre-wait field of the atomic instruction. After receiving the instruction, the manufacturing execution device waits for a specified time before executing it, thereby ensuring the correctness of the causal sequence at the instruction level.
[0165] S6, based on the feedback results returned by the manufacturing execution equipment, update the decision parameters on the associated edges.
[0166] like Figure 5 As shown, based on the feedback results returned by the manufacturing execution equipment, the decision parameters on the associated edges are updated, including:
[0167] S611, Obtain the feedback results returned by the manufacturing execution equipment. The feedback results include the measured values of the safe operating parameters of the equipment nodes after the execution action, the number of defective products and the total number of workpieces in the sliding statistics window, the actual execution time, and a flag indicating whether secondary anomalies were caused.
[0168] Furthermore, after executing each atomic instruction, the manufacturing execution equipment encapsulates the execution result into a feedback message and returns it to the decision-making system. The measured values of the equipment's safe operating parameters after the execution action are used to assess the degree of safety deviation; the number of defective products and the total number of workpieces in the current sliding statistical window are used to update the yield index; the actual execution time is used to assess the timeliness index; and the flag indicating whether secondary anomalies have been triggered is used to trigger online correction of causal delay time.
[0169] S612, based on the feedback results, uses a reward function to calculate the actual reward value.
[0170] Furthermore, this actual reward value reflects the actual effect of the selected action in a real-world execution environment, rather than the model's prediction.
[0171] S613, based on the actual reward value, update the Q value on the associated edge corresponding to the action using the Bellman equation. Here, the Q value is the specific form of the decision parameter, and the associated edges are classified according to the action type.
[0172] Furthermore, through this update, the Q-value carried on the associated edges gradually approaches the true long-term expected return, thereby continuously improving the accuracy of action selection in subsequent decisions.
[0173] In one specific embodiment, the update formula for the Bellman equation is:
[0174]
[0175] in, For the current moment The device abnormal state vector, In the current state The next action to be taken, here and The state and action definitions are consistent with those in the aforementioned reward function calculation; In order to carry out the disposal action Then, the device environment transitions to the next state in the next moment; For the state at the next moment Choose one action from the set of all available actions; the right side of the formula Current state before update Next action of The value, i.e., the current decision parameter carried on the associated edge, represents the long-term expected reward of choosing this action in this state. The left side of the formula... For the updated value; In order to carry out the disposal action And transition to the state of the next time step. The actual reward value obtained is calculated using the reward function described above. State at the next moment Among all available actions The maximum value is used to estimate the maximum long-term expected return that can be obtained starting from the next state; The learning rate is used to control the learning rate in a single update. The adjustment step size for the learning rate ranges from 0.01 to 0.3. A larger learning rate results in a higher step size. The faster the value converges, the more oscillations may occur; the smaller the learning rate, the more stable the update, but the slower the convergence speed. This is a discount factor used to control future rewards in the current context. The weights in the value estimation range from 0.9 to 0.99. The closer the discount factor is to 1, the more the decision focuses on long-term returns, while the smaller the discount factor, the more short-sighted the decision is.
[0176] This Bellman equation is deeply integrated with a causal knowledge graph. First, the Q-value is carried on the association edges between anomalous nodes and solution nodes in the knowledge graph. The storage location of the Q-value is bound to the topological structure of the knowledge graph, making the updating and migration of the Q-value constrained by the knowledge graph structure. Second, the state space consists of the real-time operating status and process parameters of the equipment nodes associated with the anomalous node, while the action space is provided by the available action space after causal readiness judgment. State transitions are constrained by the causal delay time in the temporal causal subgraph. Therefore, the state, action, and state transition triples in the Bellman equation all embed causal temporal information, rather than a general Markov decision process. Third, the actual reward value is calculated by the safety and efficiency dual-driven reward function constructed in this invention. This reward function itself integrates dynamic weights of causal severity and a causal delay time correction mechanism, thus giving the input signal of the Bellman equation causal perception capabilities.
[0177] S614 When the secondary anomaly flag is set, the causal delay time of the time-series causal directed edge between the device node associated with the secondary anomaly and the device node associated with the current anomaly node is corrected to a sliding window statistical value based on the difference between the occurrence time of the secondary anomaly and the trigger time of the cause.
[0178] Furthermore, secondary anomalies refer to new quality anomalies triggered after the execution of the current action. A flag being set indicates a potential problem with the causal timing of the current action, or a discrepancy between the causal delay time recorded in the temporal causal subgraph and the actual causal delay. In this case, the occurrence time of the secondary anomaly and the trigger time of the cause device are recorded, and the difference between the two is calculated as the actual causal delay observed. This actual causal delay is then included in the sliding window statistics, and the median of all observations within the window is updated as the corrected causal delay time. This correction mechanism allows the temporal causal subgraph to evolve online based on actual execution feedback, rather than permanently maintaining the static parameters from the initial construction.
[0179] S615, based on the corrected causal delay time, recalculate the reward function for subsequent decisions.
[0180] Furthermore, the correction of the causal delay time will affect the result of the causal readiness determination in subsequent step S3, thereby changing the composition of the available action space; at the same time, it will also affect the pre-wait time of atomic instructions in step S5. Therefore, after the causal delay time is corrected, the system re-enters step S3 to construct the updated available action space, and in step S4, it recalculates the reward function value based on the new state and available action space to ensure that subsequent decisions are based on the corrected causal timing information.
[0181] like Figure 6 As shown, when the Q-value on the associated edge is updated, the update result is migrated to the associated edges of other abnormal nodes with similar structures to the current abnormal node, including:
[0182] S621, determine the set of anomalous nodes in the knowledge graph that have structural similarity with the current anomalous node; structural similarity is measured by the Jaccard similarity coefficient between the current anomalous node and the set of first-order neighbor node types of any anomalous node in the set of anomalous nodes.
[0183] Furthermore, for the current anomalous node in the knowledge graph, the set of its first-order neighbor node types is obtained; other anomalous nodes in the knowledge graph are traversed, and the set of their respective first-order neighbor node types is obtained; the Jaccard similarity coefficient between the two sets is calculated, which measures the degree of similarity between the two anomalous nodes in the local structure of the knowledge graph. Anomalous nodes with a Jaccard similarity coefficient exceeding a preset similarity threshold are included in the set of structurally similar anomalous nodes, with the preset similarity threshold ranging from 0.5 to 0.8. The lower the preset similarity threshold, the more similar nodes are included, resulting in a wider range of benefits from migration, but this may introduce noise; the higher the preset similarity threshold, the stricter the similarity requirement, resulting in more accurate migration, but a narrower range of benefits.
[0184] S622, for any abnormal node in the abnormal node set, the action type is... For the associated edges, update the Q value using the following formula:
[0185]
[0186] in, Abnormal nodes before the update Upper Action-related edges of value; For abnormal nodes after the update Upper Action-related edges of value; The target abnormal node to be dealt with at present Upper Action-related edges The current value; For transfer learning rate, ; The target abnormal node to be dealt with at present with abnormal nodes Structural similarity between them ; Number the action type.
[0187] Furthermore, the target abnormal node What I have already learned Value information, distributed to similar nodes according to structural similarity ratio. Migration. When When the value is 0, the two nodes are completely dissimilar in structure and no migration occurs. Remain unchanged; when When the value is 1, the two nodes have identical structures and the migration range is the largest. of Value direction of Value direction adjustment. Transfer learning rate. The value ranges from 0.1 to 0.5 and is used to control the overall migration step size. The larger the value, the greater the adjustment range in a single migration, but this may introduce instability. The smaller the value, the smoother the migration but the slower the convergence.
[0188] The Q-value transfer mechanism based on Jaccard structural similarity solves the cold start problem caused by insufficient historical processing data for low-frequency anomalies. It utilizes the topological structure information of the knowledge graph to achieve knowledge transfer across anomaly types, avoiding the inefficiency of learning from scratch. At the same time, it prevents erroneous knowledge transfer between unrelated anomalies through structural similarity constraints.
[0189] In summary, this invention proposes a manufacturing quality anomaly handling method based on a temporal causal knowledge graph. This method constructs a heterogeneous knowledge graph integrating static topology and dynamic temporal causality, using a temporal causal subgraph as a hub to integrate causal reasoning throughout the entire process of anomaly handling decision-making and physical execution. First, at the knowledge graph level, causal delay times between devices are mined from historical equipment operation data using Granger causality testing or transitive entropy algorithms. This explicitly represents the delays in the operation of the causal device affecting the result device in the form of temporal causal directed edges, constructing a temporal causal subgraph. This overcomes the limitations of static qualitative mapping and achieves dynamic temporal modeling under non-stationary operating conditions. Simultaneously, a multi-scale adaptive sliding window and median aggregation strategy are employed to overcome the non-stationarity interference of the manufacturing process, ensuring the robustness of causal delay attributes under complex operating conditions. Secondly, at the action space construction level, a causal readiness determination mechanism is proposed. Based on logical constraint elimination, the causal delay time in the temporal causal subgraph is summed with the most recent operation trigger time of the cause-end device to obtain the causal readiness time, which is then compared with the current time to eliminate processing steps where the causal premise has not yet been met. Furthermore, at the instruction issuance level, the causal delay time is written into the atomic instruction as an execution pre-wait time to ensure the temporal safety of physical execution. This mechanism works from decision generation to device execution, fundamentally avoiding the risk of executing invalid or harmful actions during the delay period and causing secondary anomalies. At the sequential decision optimization level, a dynamic adaptive reward mechanism driven by both safety and efficiency is constructed. The segmented penalty mechanism has zero penalty within the safety allowable range and takes a saturation value above the critical fault, taking into account both safety priority and Q-value convergence stability. The weight coefficient ratio is dynamically adjusted according to the severity measured by the ratio of the causal out-degree to in-degree of the abnormal node, making the handling of causal propagation hub nodes more prudent. The quality and timeliness balance coefficient adaptively switches the focus according to the total number of workpieces in the sliding window. When multiple anomalies concurrently cause mutual exclusion conflicts, a bi-objective Pareto optimization is performed with a safety threshold as a hard constraint. A joint optimization solution is selected based on the knowledge graph's topological distance preference, prioritizing the coordinated handling of physically coupled anomaly pairs. Finally, at the decision evolution level, dual online updates of the knowledge graph and decision parameters are implemented. Based on execution feedback, the Q-values on associated edges are updated using the Bellman equation. When secondary anomalies are triggered, the causal delay time is corrected online, enabling the temporal causal subgraph to self-evolve. Furthermore, based on Jaccard structural similarity, learned Q-values are proportionally transferred to structurally similar low-frequency anomaly nodes to address the cold-start problem. Through these five progressive design levels, a complete closed loop is achieved, from causal discovery, causal constraint decision-making, dynamic balance of safety and efficiency, to decision self-evolution. This enables manufacturing quality anomaly handling to possess accurate causal temporal perception capabilities, dynamic safety and efficiency trade-off capabilities, and continuous cross-scenario learning capabilities.
[0190] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for handling manufacturing quality anomalies based on temporal causal knowledge graphs, characterized in that, include: A knowledge graph for handling manufacturing quality anomalies is constructed. The knowledge graph includes anomaly nodes, handling plan nodes, equipment nodes, and associated edges and time-series causal subgraphs connecting the anomaly nodes and handling plan nodes. The associated edges carry decision parameters. The time-series causal subgraph is composed of time-series causal directed edges connecting equipment nodes, and the time-series causal directed edges have causal delay time as an attribute. When a manufacturing quality anomaly is detected, the corresponding anomaly node in the knowledge graph is determined, the environmental parameters associated with the anomaly node are obtained, and the corresponding handling plan node is obtained through the associated edges connected to the anomaly node. The handling process subgraph is obtained from the handling plan node; the handling process subgraph includes handling step nodes and constraint edges connecting the handling step nodes. Based on the logical constraints and temporal causal subgraph of the constraint edges, logical constraint elimination and causal readiness determination are performed to construct the available action space in the current state. Based on the sequential decision model, the candidate disposal steps in the available action space are optimized and solved to generate a sequence of disposal instructions; the optimization objective of the sequential decision model is determined based on a reward function that includes a penalty term for equipment safety deviation and a reward term for disposal efficiency. The sequence of processing instructions is converted into executable instructions and sent to the manufacturing execution equipment; The decision parameters on the associated edges are updated based on the feedback returned by the manufacturing execution equipment.
2. The manufacturing quality anomaly handling method based on temporal causal knowledge graph according to claim 1, characterized in that, The method for constructing the temporal causal subgraph includes: Based on the topological connection relationship between device nodes as a prior constraint, device node pairs with topological connections are selected, and for each device node pair, at least one causal direction to be tested is specified, with one device as the cause-end device and the other device as the result-end device. Acquire historical operational data for device node pairs; the historical operational data includes continuous numerical data and discrete event-type operation records. The size of the sliding window is determined, and multiple sliding windows are generated based on a preset sliding step size. When there is continuous numerical data in the device node pair, the spectrum analysis of the continuous numerical data of the result device is performed first. If there is no continuous numerical data in the result device, the spectrum analysis of the continuous numerical data of the cause device is performed. The size of the sliding window is determined according to the period of the dominant frequency component of the spectrum. And / or, the size of the sliding window is adjusted according to the business stage label. Within each sliding window, the causal direction to be tested is subjected to Granger causality test and / or transitive entropy algorithm to perform causality test. The causal delay time between device node pairs within the sliding window is calculated and a significance test is performed. The causal delay time that passes the significance test is marked as a significant causal delay time value. Determine whether the proportion of significant causal delay time values generated meets the preset conditions; if it does, calculate the median of all significant causal delay time values as the final causal delay time of the device node pair in the causal direction, and generate a time-series causal directed edge from the cause device to the result device with the final causal delay time as the attribute; if it does not meet the conditions, it is not a time-series causal directed edge for the device node pair in the causal direction. Connect all generated temporal causal directed edges and their corresponding device nodes to form a temporal causal subgraph.
3. The manufacturing quality anomaly handling method based on temporal causal knowledge graph according to claim 1, characterized in that, The environmental parameters include at least one of the following: equipment operating status parameters, process parameters, environmental condition parameters, and the trigger time of operation events of equipment nodes; The logical constraints of the constraint edges include timing constraints, dependency constraints, and mutual exclusion constraints; the timing constraints indicate the execution order of two processing steps, the dependency constraints indicate that the execution of one processing step is contingent upon the completion of another processing step, and the mutual exclusion constraints indicate that two processing steps cannot be executed in parallel within the same processing cycle.
4. The manufacturing quality anomaly handling method based on temporal causal knowledge graph according to claim 3, characterized in that, The construction of the available action space in the current state includes: Traverse all candidate processing steps in the processing flow subgraph and eliminate candidate processing steps that violate logical constraints; For each remaining candidate action step, query the temporal causal directed edge in the temporal causal subgraph that has the device node associated with the candidate action step as the result end; if the current time is less than the sum of the causal delay time of the edge and the most recent operation trigger time of the corresponding cause end device node, then remove the candidate action step from the available action space. The remaining candidate processing steps after elimination constitute the available action space.
5. The manufacturing quality anomaly handling method based on temporal causal knowledge graph according to claim 1, characterized in that, The reward function is constructed as follows: The current state is defined by the real-time operating status and process parameter set of the equipment nodes associated with the abnormal node, and the current action is defined by the candidate handling steps selected from the available action space. The reward function is composed of a weighted average of a device safety deviation penalty term and a handling efficiency reward term, with the weight coefficient of the device safety deviation penalty term being greater than the weight coefficient of the handling efficiency reward term. The equipment safety deviation penalty item is determined based on the deviation between the measured value and the rated value of the equipment safety operating parameters after the current action is performed, and a segmented penalty mechanism is adopted: when the deviation does not exceed the safety allowable deviation threshold, the equipment safety deviation penalty item is zero; When the deviation exceeds the safety allowable deviation threshold but does not exceed the critical fault deviation threshold, the equipment safety deviation penalty term increases monotonically with the deviation; when the deviation reaches or exceeds the critical fault deviation threshold, the equipment safety deviation penalty term takes a saturation value. The efficiency reward is determined based on the yield statistics and execution time of the current action, and is composed of yield and timeliness indicators weighted by a quality and timeliness balance coefficient. The yield indicator is determined based on the pass rate within the sliding statistical window. The timeliness indicator is determined based on the ratio of actual execution time to the standard allowable processing time.
6. The manufacturing quality anomaly handling method based on temporal causal knowledge graph according to claim 5, characterized in that, The weighting ratio of the equipment safety deviation penalty item to the handling efficiency reward item is dynamically adjusted according to the severity of the current abnormal node; the severity is determined by the ratio of the number of time-series causal directed edges with the equipment node associated with the current abnormal node as the cause to the number of time-series causal directed edges with the device node as the result; wherein, the greater the severity, the greater the weighting ratio. The quality and timeliness balance coefficient in the processing efficiency reward item is dynamically adjusted according to the total number of workpieces in the sliding statistical window; wherein, the smaller the total number of workpieces, the more the quality and timeliness balance coefficient tends to prioritize ensuring yield; the larger the total number of workpieces, the more the quality and timeliness balance coefficient tends to prioritize ensuring processing timeliness.
7. The manufacturing quality anomaly handling method based on temporal causal knowledge graph according to claim 5, characterized in that, When multiple quality anomalies are detected simultaneously and there are action conflicts in multiple sequences of handling instructions due to mutual exclusion constraints, conflict resolution is performed on the multiple sequences of handling instructions: Using the safety threshold of the device node corresponding to each anomaly as a hard constraint, minimizing the sum of the device safety deviation penalty terms corresponding to each anomaly as the first objective, and maximizing the sum of the handling efficiency reward terms corresponding to each anomaly as the second objective, we solve for the Pareto optimal solution set. Sort the abnormal node pairs involved in the conflict in the knowledge graph from smallest to largest topological distance, select the solution corresponding to the abnormal node pair with the smallest topological distance from the Pareto optimal solution set for joint optimization, and output the conflict-free action sequence as the disposal instruction sequence.
8. The manufacturing quality anomaly handling method based on temporal causal knowledge graph according to claim 1, characterized in that, The step of converting the sequence of disposal instructions into executable instructions includes: Based on the action type of the action step node in the action instruction sequence, the corresponding instruction template is matched from a predefined action template library; the action template library includes control instruction templates, work order generation templates, and notification templates. By utilizing the attribute information of the device nodes associated with the processing step nodes in the knowledge graph, the parameters in the instruction template are filled in to generate atomic instructions as executable instructions; For processing step nodes with corresponding temporal causal directed edges, the causal delay time of the temporal causal directed edge is written as the execution pre-wait time into the corresponding atomic instruction.
9. A method for handling manufacturing quality anomalies based on a temporal causal knowledge graph according to claim 5, characterized in that, The process of updating the decision parameters on the associated edges based on the feedback results returned by the manufacturing execution equipment includes: Obtain feedback results returned by the manufacturing execution equipment; the feedback results include the measured values of the safe operating parameters of the equipment nodes after the execution action, the number of defective products and the total number of workpieces in the sliding statistics window, the actual execution time, and a flag indicating whether secondary anomalies have been triggered; Based on the feedback results, the actual reward value is calculated using a reward function; Based on the actual reward value, the Q value on the associated edge corresponding to the action is updated using the Bellman equation; the Q value is a specific form of the decision parameter, and the associated edges are classified according to the action type. When the secondary anomaly flag is set, the causal delay time of the time-series causal directed edge between the device node associated with the secondary anomaly and the device node associated with the current anomaly node is corrected to a sliding window statistical value based on the difference between the occurrence time of the secondary anomaly and the trigger time of the cause. Based on the corrected causal delay time, the reward function for subsequent decisions is recalculated.
10. A method for handling manufacturing quality anomalies based on a temporal causal knowledge graph according to claim 9, characterized in that, When the Q-value on an associated edge is updated, the update result is transferred to the associated edges of other abnormal nodes with similar structures to the current abnormal node, including: Determine the set of anomalous nodes in the knowledge graph that have structural similarity to the current anomalous node; the structural similarity is measured by the Jaccard similarity coefficient between the current anomalous node and the set of first-order neighbor node types of any anomalous node in the set of anomalous nodes. For any abnormal node in the abnormal node set, the action type is For the associated edges, update the Q value using the following formula: ; in, Abnormal nodes before the update Upper Action-related edges of value; For abnormal nodes after the update The edge associated with the k-th type of action of value; The target abnormal node to be dealt with at present Upper Action-related edges The current value; For transfer learning rate, ; The target abnormal node to be dealt with at present with abnormal nodes Structural similarity between them ; Number the action type.