A glass production process system bias identification method based on servo trigger measurement

By constructing a dynamic Bayesian probabilistic attribution model in the glass production process and integrating the dynamic response of servo-triggered measuring equipment and production topology data, the problem of the inability to identify the health status and systematic fluctuations of measuring equipment in existing technologies is solved, achieving high-precision deviation root cause diagnosis and improving production stability and efficiency.

CN122260981APending Publication Date: 2026-06-23SHANDONG FULE AUTOMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG FULE AUTOMATION TECH CO LTD
Filing Date
2026-03-31
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing monitoring systems cannot effectively identify measurement errors caused by wear or electrical drift in the glass production process, and they also have difficulty distinguishing between internal equipment faults and deviations caused by systemic environmental fluctuations, leading to errors in judgment and location.

Method used

By acquiring transient dynamic response data from servo-triggered measuring devices and output data from associated devices in the production topology, a dynamic Bayesian probabilistic attribution model is constructed. This model integrates characteristic parameters of the device's inherent health status and systemic fluctuation status to achieve high-precision diagnosis of the root causes of deviations.

Benefits of technology

It enables adaptive and high-precision diagnosis of the root causes of deviations in the glass production process, improving production stability and efficiency and reducing the generation of defective products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122260981A_ABST
    Figure CN122260981A_ABST
Patent Text Reader

Abstract

The application provides a glass production process system deviation identification method based on servo trigger measurement, relates to the technical field of industrial monitoring and control, is applied to a glass production equipment based on a servo driving system, when a deviation root cause is determined as an internal cause of the equipment, hierarchical accurate maintenance instructions can be generated, the instructions not only evaluate the emergency degree of the fault, but also can provide specific operation suggestions for maintenance personnel based on the analysis of short-term fluctuations, medium-term trends or long-term drifts of the deterioration mode of the equipment health state. When the deviation root cause is determined as an external cause of the environment, systematic risk tracing instructions are generated, while actively suppressing the group and informationless single-point alarm caused by the external cause, a systematic risk tracing atlas is generated. The systematic risk tracing atlas can actively identify and highlight the most suspicious shared resource nodes by analyzing the topological relationship of the affected equipment cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial monitoring and control technology, specifically to a method for identifying system deviations in a glass production process based on servo-triggered measurement. Background Technology

[0002] In modern industrial automated production, especially in fields requiring high-precision continuous operation such as glass manufacturing, ensuring product quality through online measurement and process control systems has become a technological consensus. As production systems become increasingly complex and interconnected, the focus of technological development has shifted from simply monitoring finished product quality to conducting deeper, data-driven analysis and optimization of the production process itself. This aims to identify and intervene in deviations at an early stage, which is crucial for improving production stability and efficiency. In glass production processes employing advanced technologies such as servo-triggered measurement, existing monitoring methods still face challenges. Specifically:

[0003] 1. Existing monitoring systems, such as the solution with publication number CN117873009A, determine anomalies by analyzing the morphology, color difference, and other status values ​​of the final product. However, such methods have an inherent limitation: they typically assume that the measuring equipment itself is completely reliable. They lack the ability to diagnose the physical health of the measuring equipment (such as servo drive components) online and cannot identify the source of measurement errors introduced by mechanical wear or electrical drift. This may lead the entire monitoring system to make incorrect judgments based on unreliable data.

[0004] 2. When production deviations are detected, existing technologies struggle to effectively distinguish whether the deviation is caused by an inherent fault in a single piece of equipment (internal cause) or by systemic environmental fluctuations affecting multiple pieces of equipment (external cause) in complex production environments. Their diagnostic logic largely relies on threshold judgments of isolated data points, lacking a comprehensive attribution analysis mechanism that integrates the individual state of equipment with the collective behavior of the system. Summary of the Invention

[0005] The purpose of this invention is to provide a system deviation identification method for glass production process based on servo-triggered measurement, so as to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A method for identifying system deviations in a glass production process based on servo-triggered measurement, comprising the following steps:

[0008] S1: For the target measuring device in the production process, based on the transient dynamic response data of the target measuring device in the servo-triggered measurement process, obtain the first dynamic characteristic parameter characterizing the physical health status of the target measuring device itself;

[0009] S2: Obtain the production process output data of the target measuring device and at least one associated device that is associated with the target measuring device in the production topology, and based on the statistical characteristics of the production process output data and the production topology, obtain a second associated feature parameter characterizing the systematic fluctuation state of the process environment in which the target measuring device is located.

[0010] S3: When a production process deviation of the target measuring equipment is detected, the first dynamic feature parameter and the second associated feature parameter are fused in the probabilistic attribution model to generate a deviation root cause attribute indication signal that characterizes whether the root cause of the production process deviation lies in the internal factors of the equipment or the external factors of the environment.

[0011] Compared with the prior art, the beneficial effects of the present invention are:

[0012] By conducting in-depth analysis of the transient dynamic response of the servo drive component during the physical action of triggering measurement, a first dynamic characteristic parameter that can characterize the evolution trend of the device's intrinsic physical health state over time is extracted. The introduction of this parameter makes it possible to perform microscopic self-inspection and early fault warning for individual measurement devices, thereby filling the gap in existing technologies for monitoring the reliability of measurement data sources.

[0013] This invention transcends the limitations of single-point monitoring. By analyzing the complexity of the output data sequences of the target measuring device and its associated equipment clusters within the production topology (characterized by statistical properties such as information entropy), and combining this with the logical connections between devices, a second correlation feature parameter is constructed that can quantify the propagation intensity of external environmental disturbances within the system. This parameter provides a macroscopic insight into the collective risk of the entire production system. A dynamic Bayesian probabilistic attribution model is constructed. This model nonlinearly fuses the first dynamic feature parameter, which characterizes the microscopic aspects of the equipment, with the second correlation feature parameter, which characterizes the macroscopic insights of the system. In this way, the model can, in probabilistic form—that is, outputting a deviation root cause attribute indication signal containing the posterior probabilities of internal and external factors—to accurately infer whether the root cause of any production deviation originates from the equipment itself or from the external environment. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of the overall method and technical route of the present invention;

[0015] Figure 2 This is a technical roadmap of steps S1 to S3 of the present invention;

[0016] Figure 3 This is a technical roadmap for steps S4 to S5 of the present invention;

[0017] Figure 4This diagram illustrates the differences between the present invention and existing technologies in terms of diagnostic procedures and results.

[0018] Figure 5 This is a schematic diagram showing the states of the glass front end contacting the L-shaped trigger rod (a) and the glass rear end contacting the L-shaped trigger rod (b). Detailed Implementation

[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0020] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0021] Example 1:

[0022] Please see Figures 1 to 5 The present invention provides a technical solution:

[0023] A method for identifying system deviations in a glass production process based on servo-triggered measurement, applied to glass production equipment based on a servo-driven system, wherein the servo-driven system includes a servo motor equipped with an encoder, and includes the following steps:

[0024] S1: For the target measuring equipment in the production process, based on the transient dynamic response data of the target measuring equipment in the servo-triggered measurement process, obtain the first dynamic characteristic parameter characterizing the physical health status of the target measuring equipment itself;

[0025] S2: Obtain the production process output data of the target measuring device and at least one associated device that is associated with the target measuring device in the production topology, and based on the statistical characteristics of the production process output data and the production topology, obtain a second associated characteristic parameter that characterizes the systematic fluctuation state of the process environment in which the target measuring device is located.

[0026] S3: When a deviation in the production process of the target measuring equipment is detected, the first dynamic characteristic parameter and the second associated characteristic parameter are fused in the probabilistic attribution model to generate a deviation root cause attribute indication signal that characterizes whether the root cause of the production process deviation lies in the internal factors of the equipment or the external factors of the environment.

[0027] A core technical feature of this invention lies in extracting a first dynamic characteristic parameter that characterizes the evolution trend of the device's internal physical health state over time by deeply analyzing the transient dynamic response of the servo drive component during the physical process of triggering measurement. Simultaneously, by analyzing the information entropy represented by the complexity of the output data sequence of the target measuring device and its associated device cluster in the production topology, and combining this with the logical connections between devices, a second correlation characteristic parameter is constructed that quantifies the propagation intensity of external environmental disturbances within the system.

[0028] The dynamic Bayesian attribution model achieves adaptive and high-precision diagnosis of the root causes of production deviations by designing nonlinear continuous evidence strength functions for the first dynamic feature parameter and the second associated feature parameter, and using the inference result of the deviation root cause attribute indication signal containing the posterior probabilities of internal and external factors as the final output.

[0029] To further explain, S1 includes:

[0030] S11: Collect the angular position sequence of the servo drive component in the target measurement device within a preset time window as transient dynamic response data;

[0031] S12: In the preset dynamic system model, the diagonal position sequence is processed using the system identification algorithm, and the output transfer function parameter is used as the first dynamic feature parameter.

[0032] Further explanation: The first dynamic characteristic parameter is the dynamic health degradation index; step S12 further includes:

[0033] S121: Form a time series of transfer function parameters over multiple consecutive measurement periods;

[0034] S122: The dynamic health degradation index is calculated based on the short-term volatility, medium-term trend, and long-term drift of the parameter time series.

[0035] To further explain, S2 includes:

[0036] S21: Obtain the length data sequence generated by the target measuring device and associated devices within their respective measurement cycles as production process output data;

[0037] S22: Calculate the information entropy of each data sequence of length within the sliding time window to generate an entropy value sequence;

[0038] S221): Methods for obtaining length data sequences include:

[0039] a) When the leading edge of the object under test pushes the trigger mechanism connected to the servo drive component, causing the trigger mechanism to deflect from the initial position to the preset trigger threshold, the first time signal is acquired;

[0040] b) After the tail of the object under test has completely passed the trigger mechanism, control the servo drive component to perform a preset reset and cleaning action to clean the motion path of the trigger mechanism.

[0041] c) After the reset cleaning action is completed, when the triggering mechanism returns and passes the preset trigger threshold, a second time signal is acquired;

[0042] d) Based on the first time signal, the second time signal, and the conveying speed of the object under test, the length value of the object under test is calculated to form a length data sequence;

[0043] The reset and cleaning action in step b) is: controlling the servo drive component to rotate a preset angle greater than 90 degrees and less than 270 degrees in the opposite direction to the movement of the object under test;

[0044] After the reset cleaning action is completed, the triggering mechanism swings back from the end position of the reset cleaning to the initial position under the action of the restoring force.

[0045] It should be noted that the core of the glass length measurement in this embodiment is the precise coordination between the transmission servo module and the rotation servo measurement module in time and space.

[0046] Drive Servo Module: This module drives the conveyor belt or rollers to transport the object under test at a precisely calibrated, constant linear speed V1. This module serves as the benchmark for the entire time-distance conversion calculation.

[0047] Rotary servo measurement module: This module is the core actuator for measurement. It includes a rotary servo drive component and a swing arm trigger mechanism fixed to the output shaft of the servo drive component. In its natural state, the trigger mechanism is stabilized at a mechanically limited initial position by the restoring force applied by a pre-tensioned torsion spring.

[0048] The method for obtaining the length data sequence, and its execution steps arranged in time sequence, are as follows:

[0049] Phase 1: Measurement Execution;

[0050] Acquire the first time signal (T1) - Leading edge detection: When the leading edge of the object under test driven by the transmission servo module first contacts and pushes the swing arm trigger mechanism, the trigger mechanism begins to deflect around the axis of the rotation servo drive component.

[0051] The control system continuously monitors the encoder readings of the rotary servo. When the reading shows that the deflection angle of the trigger mechanism first reaches and exceeds the preset trigger threshold, the control system immediately captures and records the current high-precision timestamp as the first time signal T1.

[0052] Acquire the second time signal (T2) - trailing edge detection: The object under test continues to move forward under the drive of the transmission servo module. During this period, the swing arm trigger mechanism is always pressed against the side of the object and kept in a deflection state greater than the trigger threshold.

[0053] Once the tail of the object under test has completely passed the trigger mechanism, the support for the trigger mechanism disappears. Under the restoring force applied by the torsion spring, the trigger mechanism quickly and automatically swings back from its maximum deflection position to its initial position without any driving action.

[0054] During this return process, the control system continuously monitors the encoder readings of the rotary servo. When the trigger mechanism, during its return journey, passes through and its deflection angle again falls below the aforementioned trigger threshold, the control system immediately captures and records the current high-precision timestamp as the second time signal T2.

[0055] Calculate the length value L: After obtaining T1 and T2, calculate the time difference between them, ΔT = T2 - T1. This time difference ΔT precisely corresponds to the time it takes for the object under test to pass through the measurement line where the trigger threshold is located.

[0056] Based on the aforementioned constant conveying speed V1 guaranteed by the transmission servo module, the length L of the object to be measured is obtained through the following calculation logic:

[0057] Obtain the value of the first time signal T1. Obtain the value of the second time signal T2. Perform a subtraction operation on T2 and T1 to obtain the time difference ΔT.

[0058] Obtain a preset, constant conveying speed V1. Perform a multiplication operation on the time difference ΔT and the conveying speed V1 to obtain the final length value L. Store the calculated length value L as a data point in a first-in-first-out data queue, thereby gradually forming a length data sequence.

[0059] The second stage: active reset and self-cleaning; the actions in this stage are performed after a complete length measurement is completed (i.e. after T2 is acquired), and its purpose is not for the measurement itself, but to ensure the long-term stability and reliability of the measurement system.

[0060] Execute the preset reset and cleaning action: After the swing arm trigger mechanism automatically returns to the vicinity of the initial position under the action of restoring force, the rotary servo drive component is activated.

[0061] The control system controls the servo drive component to drive the trigger mechanism to perform a rapid, large-angle swing in the opposite direction to the movement of the object under test. This swing angle is set to a preset value greater than 90 degrees and less than 270 degrees; in this embodiment, it is 180 degrees. This action has a dual key function. First, it actively and forcefully removes any debris, dust, or sticky substances that may adhere to the trigger mechanism or its movement path, acting as a self-cleaning agent and preventing measurement errors caused by foreign object interference. Second, it forcibly moves the trigger mechanism away from the measurement area and back, verifying whether its movement is smooth and free from jamming, thus serving a self-diagnostic function.

[0062] After completing the reverse cleaning swing, the torque of the rotary servo drive component is released. The swing arm trigger mechanism then automatically swings back from the end position of the cleaning action to the initial position under the action of the restoring force.

[0063] The control system verifies whether the triggering mechanism can stably return to its initial mechanical limit position. If it fails to return within a preset time, or if the returned position deviates too much from the calibrated initial position, an equipment maintenance warning can be generated.

[0064] The following example illustrates the application of the above: After each length measurement is completed, i.e., after obtaining the final length value L, a structured data packet, i.e., a "digital twin" of the glass, is immediately created in memory. This data packet contains at least the following key information:

[0065] Unique identifier: generated based on high-precision timestamps or serial numbers to ensure the traceability of each piece of glass.

[0066] Precise Length Measurement: The length value L calculated in the previous steps. Entry Timestamp: The first time signal T1 acquired in the previous steps. Reference Transport Speed: The constant speed V1 set by the drive servo module.

[0067] The digital twin data packet is broadcast in real time via Industrial Ethernet (using the OPC-UA protocol) to the Production Execution System (PES) on the production line and all relevant downstream processing units.

[0068] This embodiment takes a "fully automatic glue applicator" on a production line as an example, and its adaptive control process is as follows:

[0069] The control system of the glue applicator receives the digital twin data packet of the glass that is about to arrive in advance.

[0070] The glue applicator control system reads the length value L and compares it with the final target finished product size. Based on this difference, the system recalculates and optimizes the movement trajectory and feed depth of the grinding head in real time to ensure that the final finished product size accurately meets the target regardless of fluctuations in the incoming material size.

[0071] The control system uses the entry timestamp and V1, along with the known physical distance between the grinding station and the measuring station, to predict the arrival time of the glass leading edge at the grinding station entrance using the formula: Arrival Time = T1 + (Distance / V1). Based on this prediction, the glue applicator can complete all preparatory actions in advance, such as adjusting the width of the V-shaped plate / conveyor belt to accommodate the glass thickness, pre-starting the glue mixing pump system, and moving the glue gun head to the starting diagonal position, achieving "seamless entry" of the glass into the processing area and reducing the waiting time between workpieces.

[0072] By converting measurement data into forward-looking control commands for downstream equipment, the processing accuracy, finished product consistency, and overall production cycle time of the production line are improved.

[0073] S23: Based on each entropy value sequence and the production topology obtained from the manufacturing execution system that defines the shared resource relationships between devices, generate the second associated feature parameter.

[0074] To further clarify, the second correlation characteristic parameter is the systemic risk transmission coefficient; step S23 further includes:

[0075] S231: In the production topology, each device is regarded as a node, and the shared resource relationship between devices is regarded as an edge, and a device association graph is constructed; among them, the target measuring device and all associated devices are respectively mapped as unique nodes in the device association graph;

[0076] S232: Calculate the systematic risk transmission coefficient of the target measurement device node based on the entropy sequence of the target measurement device node and the entropy sequence of its adjacent nodes in the device association graph.

[0077] To further clarify, in S3, the probabilistic attribution model is a dynamic Bayesian attribution model; S3 includes:

[0078] S31: Based on the time change rate of the first dynamic characteristic parameter, calculate the first posterior probability of the deviation caused by the intrinsic factors of the target measuring device;

[0079] S32: Based on the synchronization between the second associated feature parameter and the corresponding parameter of the adjacent device in the production topology, calculate the second posterior probability that characterizes the deviation caused by external environmental factors.

[0080] S33: Output the probability vector containing the first posterior probability and the second posterior probability as the bias root cause attribute indication signal.

[0081] To further explain, steps S31 and S32 further include:

[0082] Before inputting the first dynamic feature parameter and the second associated feature parameter into the dynamic Bayesian attribution model, the values ​​of the first dynamic feature parameter and the second associated feature parameter are mapped to the first evidence value and the second evidence value, respectively, representing the strength of internal and external evidence, through a preset continuous evidence strength function.

[0083] The following is a detailed description of the implementation of the above content:

[0084] Servo-triggered length measurement systems are widely used in modern glass production lines. However, existing monitoring methods have significant bottlenecks: when measuring equipment (bearings, couplings) experiences progressive wear, the resulting minute measurement deviations are often submerged in normal production fluctuations, making them difficult to detect in time and eventually leading to batch defects. On the other hand, when external environmental disturbances occur (including unstable grid voltage and gas pressure fluctuations), multiple devices may experience simultaneous performance degradation, leading to an "alarm storm" for operators and making it difficult to pinpoint the root cause of the problem. Conventional single-point threshold alarm methods cannot provide proactive warnings of equipment "sub-health" states, nor can they distinguish the actual device causing the error in a cluster of anomalies. This is one of the technical problems that this embodiment aims to solve.

[0085] The core idea of ​​this invention is to abstract the physical health state of a single measuring device into a transfer function of a dynamic system, whose parameter changes reflect the evolution of the device's inherent physical characteristics. Simultaneously, the stability of the entire production system is abstracted into the evolution of information entropy of a set of interconnected nodes, whose synchronicity reflects the propagation of systemic risk. Finally, in a dynamic Bayesian network, evidence representing device health and system risk is probabilistically fused to make the most probable inference about the root cause of the deviation.

[0086] The method described in this embodiment is implemented as a program script deployed on an industrial edge computing node. The core algorithm logic and specific application strategies of this program are decoupled. During the initialization phase, the program reads a locally stored, structured spreadsheet file through a data loading module. This file predefines all configurable runtime parameters, which will be detailed later in this embodiment, such as model order, algorithm type, time window size, various thresholds, and weighting coefficients. During program execution, all algorithm steps derive their behavioral control from these configuration parameters loaded into memory.

[0087] In this embodiment, the following two key feature parameters are defined:

[0088] 1) The first dynamic characteristic parameter is represented by the Dynamic Health Degradation Index, denoted as DHI. This parameter is a dimensionless normalized value, with its range limited to 0 to 1. It aims to comprehensively characterize the degree of degradation and instability of the physical health status of the target measuring device at the current moment. A DHI value close to 0 indicates that the device is operating in a highly healthy and stable state; while a DHI value close to 1 indicates that the device has a significant risk of performance degradation or operational instability, which is strong evidence that the deviation is caused by "internal factors of the device". The determination process of the Dynamic Health Degradation Index DHI follows these steps:

[0089] A basic physical feature sequence is obtained. During each measurement cycle, when the front end of the object under test impacts the trigger mechanism, the angular position of the servo drive component is acquired at a high frequency every 1 millisecond for a preset transient time window (50 milliseconds), forming an angular position sequence. This sequence is used as input, and a preset system identification algorithm (recursive least squares method) is used to perform online parameter identification on a preset dynamic system model (second-order vibration system). The output of this step is the key transfer function parameter of the model; in this embodiment, the damping ratio is selected as this parameter. The technical consideration for selecting the damping ratio is that it is highly sensitive to changes in physical characteristics such as friction and clearance. The damping ratio values ​​obtained from multiple consecutive measurement cycles (the most recent 1000 cycles) are used to construct a parameter time series.

[0090] The specific implementation of the "dynamic system model" and "identification algorithm" is as follows: The dynamic response of the servo drive component is modeled as a standard second-order vibration system, whose transfer function in the continuous domain is uniquely determined by two physical parameters: natural frequency and damping ratio. By employing the Z-transform method with a zero-order hold, the continuous transfer function is discretized, yielding a difference equation describing the relationship between the angular position output and the control input. This difference equation takes the form: the angular position of the current period equals the product of the first historical angular position and the first coefficient, plus the product of the second historical angular position and the second coefficient, plus a term related to the historical control input. The recursive least squares (RLS) algorithm is configured to identify the first and second coefficients in the difference equation online. Inverse kinematics is then performed to obtain the damping ratio. The calculation steps are as follows: using the identified first and second coefficients, a set of pre-defined algebraic equations derived from Z-transform theory is used to inversely solve for the natural frequency and damping ratio corresponding to the continuous domain model. The inverse kinematics process is deterministic, ensuring that the conversion path from the output of the RLS algorithm to the physical parameter “damping ratio” is unique and reproducible.

[0091] Calculating a short-term volatility index: A short sliding time window is set on the time series of the above parameters to capture the immediate stability of the equipment operation. In this embodiment, the window size is set to 100 data points. This effectively smooths out random noise in a single measurement period while remaining sensitive to recent state changes. Within this window, the standard deviation of all damping ratio values ​​is calculated to obtain a short-term volatility index.

[0092] Calculating a medium-term trend indicator: A moderately long sliding time window is set over the same parameter time series to identify whether there is a continuous, unidirectional deterioration trend in equipment performance. In this embodiment, the window size is set to 500 data points. This length is sufficient to reveal slow changes such as progressive wear. Within this window, a univariate linear regression analysis is performed on the data points, and the slope of the regression line is taken to obtain a medium-term trend indicator.

[0093] Calculating the long-term drift index: The latest damping ratio value is subtracted from the baseline value representing the initial health state of the equipment to obtain a long-term drift index. The baseline value is the average damping ratio measured under ideal operating conditions for a period of time after the equipment has been commissioned or overhauled; in this embodiment, the operating period is at least 100 cycles. This embodiment chooses these three indicators, rather than other statistics, because they can comprehensively depict the evolution of the equipment's health state from different dimensions and in a complementary manner. Conventional methods focus only on a single indicator, which can easily misjudge normal short-term fluctuations as faults or ignore slow but fatal long-term trends. This embodiment achieves a "layered" diagnosis of the equipment's health state by constructing this three-dimensional feature vector, improving the accuracy and predictability of the diagnosis.

[0094] Normalize the three indicators mentioned above. The normalization process for each indicator is as follows: obtain the preset upper and lower limits for each indicator; calculate the difference between the current indicator value and the lower limit; calculate the difference between the upper and lower limits; divide the previous difference by the next difference. In this way, each indicator is mapped to a decimal between 0 and 1. The specific determination logic is as follows: construct a reference historical dataset, which requires a data acquisition module capable of extracting data from a production history database. Extract the complete time series of the "short-term volatility indicator," "medium-term trend indicator," and "long-term drift indicator" corresponding to the target measuring equipment under the confirmed "healthy state" (a long production cycle after new equipment commissioning or just after major overhaul). Perform Gaussian distribution fitting on each of the obtained indicator series. The technical choice for this step is based on the central limit theorem, that is, under stable production conditions, the fluctuations of most process parameters approximately follow a Gaussian distribution. For each indicator, its "lower limit" is determined by subtracting three standard deviations from the mean of the Gaussian distribution; its "upper limit" is determined by adding three standard deviations to the mean of the Gaussian distribution. The theoretical basis for this method is that for a Gaussian distribution, approximately 99.7% of data points will fall within a range of plus or minus three standard deviations. Therefore, using this range as the "normal range" effectively covers most normal fluctuations while maintaining high sensitivity to outliers exceeding this range, achieving a technical balance based on statistical principles between "robustness" and "sensitivity."

[0095] The steps for generating the Dynamic Health Degradation Index (DHI) through weighted fusion are as follows: The normalized short-term volatility indicator, medium-term trend indicator, and long-term drift indicator are weighted and summed. The weight coefficients for each indicator are 0.3, 0.5, and 0.2, respectively; these weight coefficients are read from the aforementioned external spreadsheet file. The technical consideration for setting these weights is that the medium-term trend (slope) is the strongest indicator of irreversible wear and tear, therefore it is given the highest weight. The final sum is the Dynamic Health Degradation Index (DHI). This embodiment provides the following logic for determining the weight coefficients:

[0096] Configure a data annotation module that can access production and maintenance logs and associate historical deviation events with their root causes ("internal" or "external"). Invoke the data annotation module to construct a dataset containing multiple DHI indicator sequences (short-term volatility, medium-term trend, and long-term drift) and their corresponding deviation event labels ("caused by internal factors" or "caused by non-internal factors").

[0097] The goal of optimization is to find a set of weights such that the DHI value calculated from these weights can best distinguish between biases caused by intrinsic factors and biases caused by extrinsic factors. Technically, this optimization objective is defined as maximizing the area under the receiver operating characteristic curve (AUC) of the DHI value between the two classes. AUC is chosen as the objective function because it can comprehensively evaluate the classifier's performance across all possible thresholds, is unaffected by class imbalance, and is considered the gold standard for measuring the performance of binary classification models.

[0098] A grid search algorithm is used to find the optimal weights. The algorithm works as follows: For each weight coefficient, a search range (from 0 to 1) and a search step size of 0.01 are set, while constraining the sum of all weights to be 1; all possible weight combinations are traversed; for each combination, the DHI value of all samples in the dataset is calculated using that combination, and its corresponding AUC value is also calculated; the weight combination that maximizes the AUC value is selected as the final weight coefficients (0.3, 0.5, 0.2) and fixed into the configuration parameters.

[0099] 2) The second correlation characteristic parameter characterizes the systemic risk transmission coefficient and is denoted as SRTC;

[0100] The Systemic Risk Transmission Coefficient (SRTC) is a dimensionless value normalized to the range of 0 to 1. It aims to quantify the extent to which production fluctuations caused by external environmental factors synchronously affect the target measuring equipment and its associated equipment in the production topology. An SRTC value close to 0 indicates that the fluctuations of the target measuring equipment are isolated and specific; while an SRTC value approaching 1 indicates that the fluctuations of the target measuring equipment exhibit a high degree of synchronicity with surrounding equipment, providing strong evidence that the deviation is caused by "external environmental factors."

[0101] The process of determining the SRTC value follows the calculation model below, with the specific steps as follows:

[0102] Obtain the basic statistical characteristic sequence; specifically, for the target measuring device and all associated devices obtained from the Manufacturing Execution System (MES) that share resources with it in the production topology (sharing the same power supply circuit, using the same batch of raw materials), obtain the length data sequence generated by each device within its respective measurement cycle. Apply a sliding window of 100 data points to the length data sequence of each device, calculate the Shannon entropy of the data within the window, and form their respective entropy value sequences. Shannon entropy is chosen because it effectively measures the disorder or uncertainty of the data sequence and is a good characterizer of the stability of the production process.

[0103] Based on the production topology information obtained from MES, a device association graph is constructed. In this graph, each measuring device is abstracted as a node, and the shared resource relationships between devices are abstracted as edges connecting the corresponding nodes, thus constructing the device association graph.

[0104] For the target measurement device node, traverse all its directly adjacent nodes in the device association graph. For each adjacent node, calculate the Pearson correlation coefficient between its entropy sequence and the target measurement device's entropy sequence within the same time period. Weighted aggregation is used to generate the Systemic Risk Transmission Coefficient (SRTC). A weighted average is then performed on all correlation coefficients calculated in the previous step. In this embodiment, the weight is uniformly set to 1, i.e., an arithmetic average operation is performed. The final average correlation coefficient value is used as the Systemic Risk Transmission Coefficient (SRTC). Since the Pearson correlation coefficient ranges from -1 to 1, it is mapped to the 0-1 interval using a linear transformation function for easier subsequent processing.

[0105] 3) This embodiment has a serious flaw if independent DHI and SRTC thresholds are set and then the "IF-THEN" rule is used for judgment. When the values ​​of DHI and SRTC are both in the "gray area" represented by the medium range, this hard thresholding method will completely fail and cannot provide a meaningful diagnosis.

[0106] This invention designs the following innovative fusion mechanism based on a dynamic Bayesian attribution model. Its internal working mechanism is as follows:

[0107] Setting a Continuous Evidence Strength Function: Before inputting DHI and SRTC as evidence into the Bayesian network, a continuous evidence strength function is introduced. This function is technically chosen as a parameterized Sigmoid function. Its role is to non-linearly map the raw values ​​of DHI and SRTC (between 0 and 1) into probability values ​​representing "intrinsic evidence strength" and "extrinsic evidence strength," respectively. Choosing the Sigmoid function smoothly transforms the input values ​​into a probabilistic interpretation and allows for flexible definition of the width of the "gray area" by adjusting its center point and slope parameters. For DHI, the lower the value, the closer the mapped evidence strength is to 0; the higher the value, the closer the evidence strength is to 1; and in the intermediate region, the evidence strength transitions smoothly. This avoids the judgment cliff caused by hard thresholds. The method for determining the continuous evidence strength function is as follows:

[0108] The center point and slope parameters of the two sigmoid functions (one for DHI and one for SRTC), a total of four parameters, are used as variables to be optimized.

[0109] The optimization objective function is defined as follows: to find a set of parameters such that when the DHI and SRTC values ​​from the training dataset are mapped through the Sigmoid function defined by this set of parameters and then input into the initial Bayesian network, the final diagnostic result (i.e., judged as "internal cause" or "external cause") best matches the true label in the dataset. This optimization objective is defined as: maximizing the overall classification accuracy or F1 score of the entire model on the validation set.

[0110] Similarly, using grid search or more efficient gradient descent optimization algorithms, the search is performed within a reasonable range of parameters to ultimately achieve the optimal parameter combination for the objective function, which is the final configuration of the Sigmoid function.

[0111] This embodiment employs a deterministic grid search algorithm. The specific steps of the algorithm are as follows: A one-dimensional search space is defined for the weight coefficients corresponding to short-term volatility, medium-term trend, and long-term drift, respectively, with each space ranging from 0.0 to 1.0. The search step size is uniformly set to 0.05. A good balance is achieved between computational accuracy and search efficiency. A grid of all possible weight combinations is generated, and combinations where the sum of all coefficients is not equal to 1.0 are eliminated. For each remaining valid weight combination in the grid, the DHI value of all samples in the dataset is calculated using this combination, and its corresponding AUC value is also calculated. The set of weight combinations that maximizes the AUC value is recorded and output as the final fixed weight coefficients.

[0112] For a dynamic Bayesian network structure: the network structure contains a parent node "Root Cause of Deviation," whose state is either "Internal Cause" or "External Cause," and two child nodes "Device Status Anomaly" and "System Environment Fluctuation." The relationship between the "Root Cause of Deviation" and its child nodes is defined by a Conditional Probability Table (CPT). The CPT defines the probability that "Device Status Anomaly" is "Yes" given that the "Root Cause of Deviation" is an "Internal Cause."

[0113] The initial parameters of the dynamic Bayesian network structure of the dynamic Bayesian attribution model are derived based on statistical learning of a labeled optimized dataset. The specific process is as follows:

[0114] Calculate the number of biased events in the dataset that are "caused by internal factors" and divide it by the total number of biased events to obtain the initial prior probability that the "root cause of the bias" is an "internal factor". Subtract this probability from 1 to obtain the initial prior probability that it is an "external factor".

[0115] Each item in the CPT is derived through a similar conditional frequency calculation. It includes the following step to determine the probability that "when the root cause of the deviation is internal, the equipment condition abnormality is strong": all deviation events caused by internal factors are identified; among these events, the proportion of events whose corresponding DHI values, after being mapped through the strength of evidence function, have an evidence strength greater than a preset high confidence threshold of 0.8 is calculated; this proportion is the desired conditional probability. All other items in the CPT are populated using this objective and reproducible frequency statistics method.

[0116] When calculating the conditional probability table, the "preset high confidence threshold" is deterministically derived using the following data-driven optimization method: From the aforementioned "labeled optimized dataset," all biased events labeled as "intrinsic causes" are extracted, their corresponding Dynamic Health Deterioration Index (DHI) is calculated, and mapped using its strength of evidence function to obtain a set of "intrinsic cause evidence strength" values. Similarly, all biased events labeled as "non-intrinsic causes" are extracted, and their corresponding "intrinsic cause evidence strength" values ​​are calculated. Receiver operating characteristic (ROC) analysis is performed. Using "intrinsic cause evidence strength" as the classifier's output score and "intrinsic cause / non-intrinsic cause" as the true label, an ROC curve is generated. This curve depicts the relationship between the classifier's true positive rate and false positive rate at different thresholds. The Euclidean distance from each point on the ROC curve to the coordinate (0,1) point is calculated, and the point that minimizes this distance is determined as the optimal operating point. The "intrinsic cause evidence strength" value corresponding to this optimal operating point is adopted as the "high confidence threshold." In this embodiment, by performing the above method, the preferred value of the preset high confidence threshold is calculated to be 0.785, which is rounded to 0.8.

[0117] 4) Further adjustments to the Conditional Probability Table (CPT). When a diagnosis is confirmed as correct, a small, positive update is made to the relevant probability values ​​in the CPT based on the strength of evidence for that case. The core idea of ​​this mechanism is to learn from historical experience and gradually adapt to the unique failure modes of a specific production line, achieving a self-optimization effect.

[0118] The dynamic adjustment mechanism of the conditional probability table is implemented through a clear, parameterized update algorithm, with the following steps: A key parameter named "learning rate" is preset, which controls the magnitude of each update. In this embodiment, the learning rate is set to 0.01. This value is chosen to ensure that the model can learn effectively from new samples while avoiding overfitting due to a single sample, achieving a technical balance. When a manually confirmed true cause label for a deviation event is received, the update process is triggered. The conditional probability entry that needs to be updated is located. If the true cause is "internal cause" and the "internal cause evidence strength" of this event is "strong", then the conditional probability "P(abnormal device status = strong | root cause of deviation = internal cause)" needs to be strengthened. The calculation logic for its update is as follows: obtain the current value of the conditional probability; obtain the learning rate; calculate the product of the learning rate and "1 minus the current value" to obtain an increment; add the increment to the current value to obtain the updated probability value. Normalization is performed. After updating the above conditional probabilities, the system will automatically adjust its complementary probabilities, i.e., "P(abnormal equipment status = weak | root cause of deviation = internal cause)", to ensure that the sum of the probabilities of all possible outcomes under the same conditions is always equal to 1.

[0119] By employing this fusion mechanism, the present invention can: 1) provide probability-based, quantitative diagnostic opinions in cases of ambiguous evidence, rather than simple "yes / no" judgments; 2) enhance the response to strong evidence through nonlinear mapping, while maintaining caution in the face of weak evidence; and 3) enable the model to have adaptive and self-learning capabilities through dynamic adjustment, thereby improving the diagnostic accuracy for long-term applications.

[0120] 5) The overall calculation process of the method disclosed in this invention during a complete deviation diagnosis cycle is as follows:

[0121] Execute two data acquisition and primary processing tasks continuously and in parallel:

[0122] Task 1: For the target measuring device, capture the servo transient response in each measurement cycle, calculate the latest damping ratio, and update its parameter time series.

[0123] Task 2: Obtain the latest length measurement values ​​from the database of the target measuring device and all associated devices, update their respective length data sequences, calculate the latest entropy values, and update their respective entropy value sequences.

[0124] Based on the updated time series, the latest DHI and SRTC values ​​are recalculated in real time and periodically.

[0125] Monitor the length data sequence of the target measuring device. When a deviation event is detected (example includes N consecutive measurements outside the tolerance range), the core attribution diagnostic process is triggered.

[0126] Obtain the DHI and SRTC values ​​at the moment the deviation event occurs. Input these two values ​​into their respective continuous evidence strength functions to calculate the corresponding probability values ​​for "intrinsic evidence strength" and "extrinsic evidence strength".

[0127] The two evidence strength probability values ​​mentioned above are used as observed evidence and input into the dynamic Bayesian attribution model. The model performs inference calculations and finally outputs the posterior probability that the "root cause of bias" is an "internal cause" and the posterior probability that it is an "external cause".

[0128] The vector formed by these two posterior probabilities is output as a root cause attribute indicator signal of the deviation. Subsequent control systems or monitoring platforms can make more advanced decisions based on the value of this signal. For example, when the probability of "internal cause" is significantly higher than the probability of "external cause", a precise maintenance work order for the equipment can be automatically generated.

[0129] To decouple the core algorithm of this invention from specific application strategies and to ensure the configurability and ease of debugging of the technical solution, in the specific implementation path of this invention, all configurable operating parameters (i.e., Class B parameters), such as the aforementioned "transient time window size," "sliding window size," and "high confidence threshold," are read through a standardized "configuration interface." The data source for this interface is a "data storage module" (e.g., a non-transient computer-readable storage medium storing an XML-formatted configuration file or a configuration table in a database), configured to store configuration data in key-value pair format. The above is further elaborated below:

[0130] The root cause indicator signal of the deviation output by this invention is essentially a probability vector containing two elements, in the form of [Pinternal, Pexternal]. Here, Pinternal represents the posterior probability, after model inference, that the root cause of the current production process deviation lies in "internal equipment factors"; Pexternal represents the posterior probability that the root cause lies in "external environmental factors." Since "internal factors" and "external factors" are defined in this model as mutually exclusive and complete sets of events, the sum of these two probability values ​​is always equal to 1.

[0131] When the value of the deviation root cause attribute indicator signal is closer to the vector [1,0], that is, when Pinternal is closer to 1, it indicates that the probabilistic attribution model of the present invention judges with a higher confidence that the currently monitored production deviation is caused by the deterioration of the physical health of the target measuring equipment itself (e.g., component wear, performance drift). This output state provides clear and executable instructions for the subsequent decision-making system, that is, a precise maintenance work order for the specific equipment should be generated immediately, rather than an ineffective investigation of the entire production environment.

[0132] When the value of the deviation root cause attribute indicator signal approaches the vector [0,1], i.e., Pexternal approaches 1, it indicates that the probabilistic attribution model of this invention judges with a higher confidence that the current deviation is not caused by a problem with the target device itself, but rather stems from a systemic, global factor affecting multiple devices, including the target device (e.g., unstable grid voltage, shared gas source pressure fluctuations). This output state can effectively avoid incorrect maintenance operations on healthy devices and guide maintenance personnel to focus their attention on troubleshooting shared, upstream public facilities or environmental factors, thereby solving the problem at its root and avoiding diagnostic chaos under an "alarm storm".

[0133] The core of this invention lies in the effective fusion of two key input parameters with completely different properties: a first dynamic feature parameter and a second associated feature parameter.

[0134] The Dynamic Health Degradation Index (DHI) exhibits a strict positive correlation with the posterior probability (Pinternal) of "internal cause" in the final output vector. DHI is a quantitative indicator that integrates the short-term volatility, medium-term trend, and long-term drift of the equipment's dynamic response parameters. A higher DHI value directly reflects the more severe the degradation of the physical performance (including damping characteristics) of the equipment's servo drive components and the more unstable its operation. The probabilistic attribution model of this invention maps an increase in DHI to an enhancement of "internal cause evidence strength" through a continuous evidence strength function. Within the Bayesian inference framework, when other evidence remains constant, stronger "internal cause evidence" inevitably leads to a monotonically increasing posterior probability (Pinternal) of the "internal cause" hypothesis. This positive correlation design deterministically links a measurable health indicator characterizing the physical entity of the equipment with an abstract logical judgment probability characterizing "fault attribution to internal factors."

[0135] The Systemic Risk Transmission Coefficient (SRTC) exhibits a strong positive correlation with the posterior probability (Pexternal) of the "external factors" in the final output vector. SRTC is an indicator based on production topology that quantifies the synchronicity of the production output data (entropy sequence) of a target device with other related devices. A high SRTC value indicates that the production fluctuations of the target device are not isolated events, but rather exhibit a collective behavior of "rising and falling together" with its "neighbors" in the production network. In real-world production environments, the probability of multiple independent devices experiencing coincidental, unrelated internal failures within the same timeframe is extremely low. Observed highly synchronized fluctuations provide strong evidence of a common, external disturbance source (such as the power grid, gas supply, or raw material batches) acting on the entire device cluster. The probabilistic attribution model maps an increase in SRTC to an increase in the "strength of external factor evidence," leading to a monotonically increasing posterior probability (Pexternal) of the "external factor" hypothesis.

[0136] Another core technical feature of this invention lies in its ability to not only accurately identify the root cause of deviations but also to transform this probabilistic judgment into deterministic intelligent diagnostic instructions with hierarchical and in-depth information. When the cause is determined to be internal to the equipment, a tiered diagnostic and recommendation instruction is generated. This instruction, based on in-depth analysis of the equipment's health status deterioration patterns (abrupt, gradual, or long-term drift), provides maintenance personnel with a severity assessment of the fault and specific operational suggestions. When the cause is determined to be external to the environment, while suppressing invalid alarms, a systemic risk tracing map is generated. By analyzing the topological relationships of the affected equipment cluster, the most suspicious shared resource nodes are proactively identified and highlighted, thereby guiding the focus of fault diagnosis from the symptoms to the root cause.

[0137] Further explanation: The method also includes:

[0138] S4: Compare the first posterior probability and the second posterior probability in the deviation root cause attribute indication signal; if the first posterior probability is greater than the first preset threshold, generate a first diagnostic instruction to indicate precise maintenance of the target measuring equipment.

[0139] Further explanation: The first diagnostic instruction is a graded diagnostic and recommendation instruction; the steps for generating the first diagnostic instruction further include:

[0140] S41: Determine the fault severity level, which characterizes the urgency of the fault, based on the value of the first posterior probability.

[0141] S42: Based on the contribution of each component of the dynamic health degradation index, determine at least one recommended maintenance action, wherein the components include short-term volatility, medium-term trend and long-term drift.

[0142] Further explanation: The method also includes:

[0143] S5: Compare the first posterior probability and the second posterior probability in the deviation root cause attribute indication signal; if the second posterior probability is greater than the second preset threshold, generate a second diagnostic instruction to suppress single-point alarms for the target measuring device and highlight the affected device cluster in the production topology.

[0144] Further explanation: The steps for generating the second diagnostic instruction further include:

[0145] Based on the production topology, identify at least one shared resource node that has a shared resource relationship with the devices in the affected device cluster;

[0146] For at least one shared resource node, a shared resource correlation score is determined by calculating its correlation with devices in the affected device cluster; and,

[0147] Based on the shared resource correlation score, a systemic risk source map is generated as part of the second diagnostic instruction. The systemic risk source map is used to indicate the likelihood that each shared resource node is a source of deviation.

[0148] Further explanation: The systemic risk tracing map includes a sorted list, which arranges at least one shared resource node in descending order of its shared resource correlation score.

[0149] The following is a detailed implementation description of the above content: After obtaining a probabilistic judgment of the root cause of the deviation, if the process stops at displaying a probability value or triggering a general alarm, it presents two major challenges for subsequent processing: 1) For internal faults, after maintenance personnel see the 75% probability of equipment failure, they still need to analyze whether it is an urgent sudden failure or a performance degradation that can be temporarily deferred, lacking guidance on action priorities. 2) For external faults, after operators see three devices (A5, B5, and C5) alarming simultaneously, facing the complex pipeline and circuit layout of the factory, troubleshooting their common cause is like finding a needle in a haystack, resulting in low efficiency. Steps S4 and S5 described in this invention are used to solve this information gap problem after diagnosis but before action.

[0150] After outputting the root cause attribute indication signal of the deviation in step S3, the decision branching mechanism is entered. All behavioral control parameters of this mechanism, such as the first preset threshold, the second preset threshold, and various hierarchical thresholds and rules that will be detailed later, are obtained from the same external spreadsheet file loaded during program initialization.

[0151] It should be noted that the process for determining the first and second preset thresholds is as follows:

[0152] A reference diagnostic dataset is constructed. The data originates from a historical fault data module, which is configured to store deviation events that occurred during historical production processes and whose true root causes (internal or external) have been manually identified and labeled. The dataset output by this module must meet the following key attributes: it contains a balanced number of identified internal and external fault cases to avoid statistical bias caused by imbalanced samples. Each case includes a complete deviation root cause attribute indication signal calculated by the preceding steps (S1-S3) of this invention at the time of its occurrence, i.e., the posterior probability vector [Pinternal, Pexternal].

[0153] The following bias events are induced and calibrated to generate data samples with clearly defined root cause labels through controlled experiments.

[0154] For the generation of intrinsic events: Select a target measuring device in a healthy state. To simulate progressive wear, the virtual damping coefficient in the servo drive component control algorithm is gradually increased in preset, small increments (0.1% of its original value per hour) via a software interface, while continuously running measurements. When the production process deviation is detected to exceed the normal range for the first time, the data calculated and recorded by steps S1-S3 of this invention within the event window is marked as an intrinsic progressive wear type.

[0155] Furthermore, to simulate the physical effects of sluggish system response and increased energy dissipation caused by progressive wear, the control parameters used to suppress the system's dynamic response in the servo drive component control algorithm are adjusted via a software interface. This aims to increase the system's equivalent damping. In one embodiment, this parameter is the proportional gain parameter in the speed loop control loop; by gradually decreasing this proportional gain in small steps, the effect of a slower system response is simulated. In another embodiment, if the controller provides direct vibration suppression functionality, the parameter is the intensity level parameter of this function; by gradually increasing this intensity level, the effect of increased system energy dissipation is simulated. This ensures that this simulation method is not dependent on any specific brand or model of controller.

[0156] To simulate a sudden failure, during normal equipment operation, a very short-duration (50 milliseconds) pulse noise with controllable amplitude is injected into the encoder signal path of the servo drive component via a signal generator. The data recorded in this event window is labeled as intrinsic-mutation type.

[0157] For the generation of external events: Select one or more target measuring devices and related devices that are in a healthy state. To simulate power grid fluctuations, connect a programmable power supply in series to the power supply circuits to which these devices are connected, and control this power supply to generate a brief 100-millisecond, 10% voltage drop. All data recorded by the affected devices within this event window are labeled as external factor - power grid fluctuation.

[0158] To simulate batch issues with raw materials, a batch of test objects, pre-confirmed by high-precision measuring equipment and exhibiting uniform, minute length deviations (0.5% longer than the nominal value), was used in production. All deviation data recorded by all equipment during this batch's production was labeled as external factor-raw material deviation.

[0159] All data samples collected and labeled through the above process are aggregated. The dataset is then examined, and random undersampling is used to ensure that the total number of internal cause cases is approximately equal to the total number of external cause cases, forming the final reference diagnostic dataset.

[0160] Regarding the determination of the first preset threshold: from the reference diagnostic dataset, all cases where the true root cause is internal are selected to form an internal cause case set; at the same time, all cases where the true root cause is external are selected to form an external cause case set.

[0161] Set a candidate threshold variable, starting at 0.01 and iterating up to 0.99 in increments of 0.01. At each candidate threshold, perform the following calculations: Calculate the internal true positive rate: In the internal case set, count the number of cases whose Pinternal value is greater than the current candidate threshold, then divide this number by the total number of cases in the internal case set. Calculate the internal false positive rate: In the external case set, count the number of cases whose Pinternal value is greater than the current candidate threshold, then divide this number by the total number of cases in the external case set.

[0162] For each candidate threshold, a comprehensive benefit score is calculated to evaluate its overall diagnostic effectiveness. The core objective of this comprehensive benefit score is to achieve a satisfactory balance between the intrinsic true positive rate (representing benefits) and the intrinsic false positive rate (representing costs).

[0163] In a preferred embodiment, the calculation logic for the comprehensive benefit score is designed as a linear utility function: obtain the intrinsic true positive rate and a preset benefit weight; obtain the intrinsic false positive rate and a preset cost weight; calculate the product of the first two to obtain a weighted benefit value; then calculate the product of the latter two to obtain a weighted cost value; subtract the weighted cost value from the weighted benefit value to obtain the final comprehensive benefit score.

[0164] The quantitative basis for setting the benefit weight and cost weight in the comprehensive benefit score is as follows: the production management department estimates the economic loss caused by a typical false alarm shutdown (mistaking an external cause for an internal cause), denoted as CFP; secondly, it estimates the average potential loss that may be caused by a missed report delay (failure to detect the internal cause in a timely manner), denoted as CFN. The ratio of cost weight to benefit weight is set to be equal to the ratio of CFP to CFN, thereby directly linking the scoring mechanism to actual production economic benefits.

[0165] After traversing all candidate thresholds, the candidate threshold that maximizes the overall benefit score is selected as the final first preset threshold.

[0166] The process for determining the second preset threshold in this embodiment is completely symmetrical to that of the first preset threshold. It only requires replacing Pinternal with Pexternal in the calculation and swapping the roles of the internal case set and the external case set. No further details are provided.

[0167] The severity level of a fault, denoted as GYz, is a discrete state identifier characterizing the urgency of an internal equipment fault. In this embodiment, it is defined as three levels: Warning, Important, and Urgent. It aims to provide a priority basis for the allocation of maintenance resources.

[0168] The Shared Resource Correlation Score, denoted as SRCS, is a dimensionless numerical value used to quantify the likelihood that a shared resource (including power supply circuits and raw material batches) is the root cause of the current systemic fluctuations. A higher score indicates a stronger correlation between the shared resource and all devices experiencing synchronized fluctuations, and a higher suspicion that it is the source.

[0169] The generation of hierarchical diagnostic and suggestion instructions in step S4 is explained: a two-stage instruction generation logic is designed.

[0170] Phase 1: Determine the severity level GYz of the fault. The calculation logic for this phase is as follows:

[0171] Obtain the first posterior probability Pinternal output from step S3. Read two classification threshold values ​​from the external configuration file: a preset importance threshold of 0.75; and a preset emergency threshold of 0.90.

[0172] The steps for determining the grading threshold are as follows: In the internal case set of the reference diagnostic dataset, based on historical maintenance records, each case is further labeled with its actual urgency at that time, and labeled as a routine internal cause, an important internal cause, or an urgent internal cause based on whether it leads to unplanned downtime.

[0173] The steps for determining the emergency level threshold are as follows: Extract all cases marked as emergency internal causes from the segmented dataset and obtain the corresponding Pinternal values. Perform statistical analysis on these Pinternal values ​​to calculate specific quantiles of their distribution. In this embodiment, the 10th percentile is chosen. The calculated 10th percentile is used as the emergency level threshold. The technical consideration for choosing a lower quantile of 10% as the threshold is to establish a highly sensitive emergency judgment standard, that is, to escalate a few important faults to emergency levels while minimizing the chance of missing any truly emergency faults.

[0174] Determine the importance level threshold: Extract all cases marked as important internal factors from the segmented dataset, and calculate the 10th percentile of their Pinternal value distribution. Use this calculation result as the importance level threshold.

[0175] If Pinternal is greater than the emergency level threshold, then GYz will be assigned the value of emergency.

[0176] If the previous condition is not met, the comparison continues: if Pinternal is greater than the importance level threshold, then GYz is assigned the value of importance.

[0177] If none of the above conditions are met, but the first preset threshold of 0.6 for entering S4 has been met, then GYz will be assigned a warning value.

[0178] Phase Two: Determining Recommended Maintenance Actions. This phase is based on an analysis of the internal composition of DHI, and its calculation logic is as follows:

[0179] Obtain the three normalized components used to calculate the DHI: short-term volatility indicator, medium-term trend indicator, and long-term drift indicator. Compare the values ​​of these three indicators and identify the one with the largest value as the main contributing component.

[0180] Perform the following rule-based mapping:

[0181] If the primary contributing factor is short-term volatility, the recommended maintenance actions are: check the electrical connections and mechanical fasteners of the servo drive components, and identify sources of vibration. This rule is based on the fact that short-term, sharp fluctuations are usually related to poor contact or external shocks.

[0182] If the primary contributing factor is a medium-term trend indicator, the recommended maintenance action is to plan for wear inspection or preventative replacement of the servo drive components. This rule is based on the fact that a continuous unidirectional trend is a typical characteristic of progressive wear.

[0183] If the primary contributing factor is the long-term drift, the recommended maintenance action is to perform a comprehensive accuracy calibration and reference reset on the equipment. This rule is based on the fact that large, long-term drift typically indicates that the equipment's reference has become invalid.

[0184] Through this mechanism, the first diagnostic instruction output is a complete information package containing the internal cause, the severity of the GYz, and suggested actions, which improves the practical value and maintenance efficiency of the diagnostic results.

[0185] It should be further noted that in the aforementioned embodiments, the recommended maintenance actions are determined through rule-based mapping logic, specifically as follows:

[0186] This mapping rule base was derived through analysis and summarization of historical maintenance knowledge datasets. The construction and analysis process of this dataset is as follows:

[0187] The data originates from the maintenance log module, which records every maintenance activity for the target equipment in history. This log must include: the DHI (Discretionary Hierarchy Index) and its three components (short-term volatility, medium-term trend, and long-term drift) calculated by this invention before maintenance; the specific operations performed by maintenance personnel (e.g., tightening terminals, replacing bearings, recalibrating the origin); and confirmation criteria for whether the operation successfully resolved the problem. Data mining methods are used to analyze the strong correlations between significant increases in different components of DHI and specific successful maintenance actions. These strong correlations are refined into deterministic rules of the IF-THEN form described in the preceding embodiments, forming an executable knowledge base. If statistics show that over 90% of cases with significantly increased short-term volatility indicators are ultimately resolved by checking electrical connections and mechanical tightening, this is solidified as a rule.

[0188] Regarding the steps for generating a systemic risk tracing map: when multiple devices alarm, presenting the user with a flat device list is equivalent to throwing a complex topology analysis problem to the user, without truly completing the tracing task. This invention sets up a map generation algorithm that includes topology traversal and scoring sorting, referred to as the topology scoring and sorting algorithm. The specific logic is as follows:

[0189] Identifying shared resource nodes follows the logic as follows:

[0190] Obtain the list of IDs of the affected device cluster that triggered S5; access the complete production topology data obtained from the MES system. This data structure is a graph containing device nodes and resource nodes representing shared resources (including example power supply loop A and gas supply pipeline 3). For each device ID in the affected device cluster, traverse all its adjacent nodes in the topology graph. Collect all traversed nodes of type resource node to form a candidate shared resource list, and remove duplicates.

[0191] Calculate and rank the Shared Resource Relationship Score (SRCS); the logic for this step is as follows:

[0192] For each resource node in the candidate shared resource list generated in the previous step, initialize a counter to zero. Then, iterate through the list of IDs of the affected device cluster again. For each device, determine whether it is directly connected to the currently calculated resource node in the production topology graph.

[0193] If the determination is yes, the counter value is incremented by one. After traversing all affected devices, an initial correlation count value is obtained. Finally, this initial correlation count value is divided by the total number of devices in the affected device cluster to calculate the normalized shared resource correlation score (SRCS), with a value range of [0,1]. After calculating the SRCS for all candidate resource nodes, they are sorted from highest to lowest SRCS value. The second diagnostic instruction output is a structured systemic risk tracing map. This systemic risk tracing map is visualized as a subgraph on the user interface, where affected device nodes are highlighted, and the highest-ranked shared resource nodes are identified with different colors or sizes. This allows maintenance personnel to see that power supply circuit A is the most suspicious because it connects to 90% of the abnormal devices, while gas supply line 3 is less suspicious because it only connects to 20%. This shift from listing symptoms to ranking suspicions is a key step in intelligent diagnosis, reducing fault diagnosis time from hours to minutes.

[0194] Combining S4 and S5 of this embodiment with the preceding S1-S3, the complete deviation diagnosis and instruction generation cycle is as follows:

[0195] It runs continuously in the background, calculating the DHI and SRTC values ​​of each measuring device in real time.

[0196] When a production deviation is detected in any device, the DHI and SRTC values ​​of that device at the time the deviation occurred are immediately retrieved. The root cause attribute indication signal of the deviation, i.e., the posterior probability vector [Pinternal, Pexternal], is calculated through a dynamic Bayesian attribution model.

[0197] Path 1 (Internal Factor): Obtain the Pinternal value and compare it with the first preset threshold. If it is greater than the first preset threshold, the execution process of S4 is triggered to generate a first diagnostic instruction containing GYz and suggested maintenance actions, and send it to the equipment maintenance system (CMMS) as a structured message (JSON format).

[0198] Path Two (External Factors): Obtain the Pexternal value and compare it with the second preset threshold. If it is greater than the second preset threshold, trigger the execution process of S5. Send a flag to the monitoring interface (SCADA) to suppress single-point alarms of the device; execute the source map generation algorithm, and send the second diagnostic command containing the sorted shared resource nodes to the monitoring interface for visualization of the global situation.

[0199] Path 3 (Ambiguous / No Action): If neither Pinternal nor Pexternal exceeds its respective threshold, the current evidence is deemed insufficient to form a definitive conclusion. No diagnostic instructions are generated, and the ambiguous event is simply recorded in the log for further analysis.

[0200] The following further elaboration addresses the above: The diagnostic output of this invention is not a simple alarm signal, but a structured instruction with rich technical content. The changing trend of its key output values ​​is closely coupled with the evolution of the system state, specifically:

[0201] Regarding the fault severity level (driven by the first posterior probability): As the first posterior probability Pinternal changes from a low value (close to the first preset threshold) to a high value (approaching 1), the fault severity level will correspondingly jump from the warning level, through the important level, and finally to the emergency level. This trend represents the increasing confidence of the invention in determining that the deviation originates from within the device. A Pinternal value approaching 1 means that the system is almost completely certain that the fault originates from the device itself, and therefore must be handled with the highest priority. This level classification accurately maps this decision-making logic.

[0202] Regarding the Shared Resource Relationship Score (SRCS): The closer the SRCS value of a shared resource node is to 1, the higher its connectivity with all devices in the cluster experiencing synchronization anomalies. In the extreme case of SRCS=1, it means that the shared resource is the only common upstream node for all anomalous devices, thus making it highly likely to be the root cause. Conversely, when the SRCS value is close to 0, it indicates very little connection between the node and the cluster of anomalous devices, making it low likely to be the root cause.

[0203] Regarding the first diagnostic instruction, the impact of the first posterior probability Pinternal: an increase in Pinternal will monotonically and non-linearly increase the severity level of the fault. Pinternal is a direct quantification of the system's confidence in intrinsic faults. The diagnostic system must assign a higher response priority to judgments with higher confidence. This positive correlation design accurately maps the practical logic of risk management.

[0204] The impact of the Dynamic Health Degradation Index (DHI) components: The three components of the DHI (short-term volatility, medium-term trend, and long-term drift) determine the recommended maintenance actions through a pattern-matching relationship, rather than a simple numerical increase or decrease. Even if the total DHI value is not high, as long as the short-term volatility component becomes the dominant component, a recommendation to check electrical connections will be generated.

[0205] Regarding the second diagnostic instruction: the impact of the affected device cluster size, the cluster size is positively correlated with the maximum possible value of SRCS. The larger the cluster, the more likely the SRCS value of the truly rooted shared resource nodes is to reach a higher value, thus making them more distinguishable from other non-rooted nodes.

[0206] Impact of Production Topology: The connection complexity of the topology directly affects the distribution of SRCS (Search Response Codes). In a sparsely connected topology, the SRCS values ​​of different shared resources are more likely to polarize; while in a highly redundant, mesh-connected complex topology, the distribution of SRCS values ​​is more gradual, making the sorted list particularly important. The SRCS algorithm is designed to locate key nodes in such complex topological relationships by quantifying the correlation degree, a core indicator.

[0207] Furthermore, to quantitatively verify the effectiveness and advancement of the diagnostic logic described in this invention, a discrete event dynamic system model based on MATLAB was constructed. This model simulates a production line containing five processing units (U1 to U5). These units share two core resources: a main power supply circuit and a central cooling system. U4 and U5 also additionally share an auxiliary air path.

[0208] The control group used an independent KPI threshold alarm method: each processing unit independently monitored its single-piece processing time, and when the processing time exceeded three standard deviations of the normal average, the unit independently triggered a red alarm. This invention group fully implemented all logic including S1 to S5. See Table 1 below for details:

[0209] Table 1: Comparison of diagnostic performance of the present invention and existing technologies under different fault scenarios

[0210] Parameters\Scenario Scenario 1: Wear and tear on the internal motor of U2 Scenario 2: Transient interference in U3 encoder Scenario 3: Voltage drop in the main power supply circuit Scenario 4: Decreased efficiency of the central cooling system Scenario 5: Normal Operation Scenario 6: U1 and U5 experience independent internal faults. Scenario 7: Unstable auxiliary air pressure - Test 1 / Test 2 Test 1 / Test 2 Test 1 / Test 2 Test 1 / Test 2 Test 1 / Test 2 Test 1 / Test 2 Test 1 / Test 2 First posterior probability 0.92 / 0.94 0.89 / 0.91 0.03 / 0.02 0.18 / 0.19 0.11 / 0.12 U1: 0.93 / 0.92, U5: 0.90 / 0.91 0.15 / 0.14 Second posterior probability 0.08 / 0.06 0.11 / 0.09 0.97 / 0.98 0.82 / 0.81 0.13 / 0.11 U1: 0.07 / 0.08, U5: 0.10 / 0.09 0.85 / 0.86 Dominant DHI component Medium-term trend Short-term volatility (not applicable) (not applicable) (not applicable) U1: Long-term drift, U5: Medium-term trend (not applicable) Affected device clusters {U2} {U3} {U1,U2,U3,U4,U5} {U2,U3} {} {U1},{U5} {U4,U5} Number of existing technical alarms 1 1 5 2 0 2 2 This invention - Fault Severity Level urgent urgent (not applicable) (not applicable) none All are emergency (not applicable) This invention - Recommended maintenance actions Planned wear inspection Check electrical connections (not applicable) (not applicable) none U1: Full Calibration, U5: Scheduled Wear Check (not applicable) This invention - the root cause of ranking first (not applicable) (not applicable) Main power supply circuit Central cooling system none (not applicable) Auxiliary air path This invention - the first ranked SRCS (not applicable) (not applicable) 1.00 / 1.00 1.00 / 1.00 0 (not applicable) 1.00 / 1.00 This invention - the root cause of the second ranking (not applicable) (not applicable) (none) (none) none (not applicable) Main power supply circuit This invention - SRCS ranked second (not applicable) (not applicable) (none) (none) 0 (not applicable) 0.40 / 0.40 Diagnostic efficiency index 1.00 / 1.00 1.00 / 1.00 1.00 / 1.00 1.00 / 1.00 (not applicable) 1.00 / 1.00 1.00 / 1.00 Existing technology - diagnostic efficiency index 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00 0.00 / 0.00 (not applicable) 0.00 / 0.00 0.00 / 0.00

[0211] The Diagnostic Efficiency Index, denoted as DEI, is calculated as: DEI = (Number of instructions correctly identifying and pointing to the root cause) / (Total number of alarms or instructions generated by the system). For the existing technology represented by the control group, it can only generate alarms and cannot pinpoint the cause, so the numerator is always 0. For this invention, one internal cause diagnosis is counted as one instruction, pointing to one device cause; one external cause diagnosis is counted as one instruction, pointing to one shared resource cause. The Diagnostic Efficiency Index is used to quantify the signal-to-noise ratio of the diagnostic system. A higher DEI indicates a more refined system output, higher information value density, and better ability to avoid alarm storms and directly address the core problem.

[0212] Comparing Scenario 1 and 2 (depth of internal cause diagnosis): In both Scenario 1 and Scenario 2, both existing technologies and the present invention correctly pinpoint the problem to a single device (number of alarms / commands is 1). Existing technologies can only provide a general U2 or U3 alarm. The present invention not only confirms the fault (Pinternal > 0.89) and assesses it as urgent, but also provides completely different and highly targeted maintenance recommendations based on the different dominant DHI components (comparing medium-term trends to short-term fluctuations) (comparing planned wear checks to electrical connection checks). This demonstrates the significant advancement in diagnostic depth achieved by step S42 of the present invention, realizing a leap from fault location to decision support.

[0213] Comparison with Scenario 3 (Global External Cause Diagnosis Efficiency): In this scenario, the shortcomings of existing technologies are glaringly obvious. It generates five independent alarms, forming a typical alarm storm with a DEI of 0, leaving operators unable to determine the root cause amidst a screen full of red alarms. In contrast, this invention, through the logic of S5, correctly identifies the fault as originating externally (Pexternal > 0.97), and then generates only one diagnostic instruction. This instruction, calculated by SRCS (with a value of 1.00), accurately identifies the main power supply circuit as the primary root cause. The DEI of this invention is 1.00, a breakthrough from 0 to 1 compared to the 0 of existing technologies, demonstrating the revolutionary effect of this invention in suppressing alarm storms and improving the efficiency of systemic fault diagnosis.

[0214] Comparison Scenario 7 (Sequencing and Diagnostic Capabilities for Ambiguous External Factors): This scenario is crucial for verifying the value of the sequencing function in step S5 of this invention. When the auxiliary gas path pressure is unstable, only U4 and U5 are affected. Existing technologies generate two isolated alarms, and their DEI remains 0. This invention correctly identifies it as an external event (Pexternal > 0.85) and calculates SRCS for all shared resources:

[0215] Auxiliary gas path: It only connects to U4 and U5, and connects to both of the two affected equipment clusters. Its SRCS is calculated as 2 divided by 2, with a result of 1.00.

[0216] The main power supply circuit connects all five devices, including two in the two affected device clusters. Its SRCS calculation is 2 divided by 5, resulting in 0.40. The output ranking list of this invention is: 1. Auxiliary gas path (SRCS=1.00), 2. Main power supply circuit (SRCS=0.40). This ranking result allows maintenance personnel to immediately focus their attention on the auxiliary gas path with the highest score, a perspective completely unavailable from two isolated alarms in existing technologies. This scenario irrefutably demonstrates the core value of SRCS ranking in effectively distinguishing primary and secondary faults and guiding rapid localization when facing multiple potential fault sources.

[0217] Furthermore, this section uses the Shared Resource Relevance Score (SRCS) as an example to implement the Objective Threshold Derivation (OTD) methodology, demonstrating that the present invention can not only provide diagnostic results, but also guide hierarchical and intelligent response strategies.

[0218] This invention selects Mean Time To Trace (MTTL) as the VEI. This metric directly quantifies the time required from receiving an alarm to maintenance personnel confirming the root cause. MTTL is a core indicator for measuring the economic value and practicality of a diagnostic system. Compared to the indirect number of alarms, it is more directly related to the speed of production recovery, and therefore is the most appropriate indicator for evaluating the effectiveness of step S5 of this invention.

[0219] Experiments were conducted on scenarios involving faulty clusters of varying topologies and sizes, and the relationship curve between SRCS (Search Response Center) and VEIMTTL (Vehicle Interruption Time Tolerance) was plotted. Analysis shows a significant negative correlation and exponential decay. Specifically: in the high SRCS region (SRCS > 0.8), MTTL is low and tends to a stable minimum (close to 0, indicating immediate location); in the median SRCS region, MTTL begins to increase significantly as SRCS decreases; in the low SRCS region (SRCS < 0.4), the increase in MTTL tends to level off, but remains at a relatively high level, indicating a longer investigation time is required.

[0220] The first threshold point is defined as the critical point where the curvature of the exponential decay curve first changes from significantly negative (rapid decline) to near zero. This point is determined to be SRCS≈0.85 through numerical differentiation of the curve.

[0221] The second threshold point is defined as the point where the absolute value of the curvature of the exponential decay curve is the largest, i.e., the inflection point of the curve (the point where the MTTL decreases the fastest). By calculating the second derivative of the curve, this point is determined to be SRCS≈0.50.

[0222] A threshold of 0.85 marks the objective boundary at which a diagnostic conclusion transitions from high confidence to certainty. Values ​​above this threshold are almost equivalent to those requiring manual review and can be directly accepted.

[0223] A threshold of 0.50 marks the objective boundary at which a diagnostic conclusion shifts from vague speculation to high confidence. Above this value, even a small increase in SRCS will result in a gain in MTTL.

[0224] Define Level 1 Intervention: Automatically generate high-priority work orders, directly assigning them to senior maintenance engineers. The work order clearly indicates that the first-ranked shared resource is the primary target for investigation. Cost: Medium (occupies senior human resources). Utility: High (significantly shortens the investigation path).

[0225] Define a secondary intervention: Send a systemic risk recommendation to the central control room, list the top three potential shared resources in the alarm, and suggest that the on-duty engineer investigate according to the production plan. Cost: Low (occupies regular human resources). Utility: Moderate (provides direction for investigation).

[0226] The primary intervention is deployed in the SRCS > 0.85 range because a high SRCS value implies a low risk of misdiagnosis. At this level, utilizing high-cost, senior human resources to resolve the issue as quickly as possible yields benefits far outweighing the costs. The secondary intervention is deployed in the SRCS range of (0.50, 0.85) because the diagnostic conclusion at this level has high confidence, providing a clear direction for investigation, but a slight uncertainty remains. Coordination by the control room achieves the best balance between cost control and problem-solving. For cases with SRCS ≤ 0.50, only log entries are made for post-analysis to avoid issuing interfering information at low confidence levels.

[0227] For appendix Figure 1 The following supplementary explanations are provided: (Attached) Figure 1 The complete logical flow of the method described in this invention is illustrated schematically. The physical entity on the left side of the figure depicts a key measurement step in a laminated glass production line scenario. The figure shows a large, vertically placed glass panel, supported and transported at its bottom by a conveyor system consisting of a series of rollers. The core of the illustration is the "servo-triggered measurement" mechanism located between the conveyor rollers. This mechanism is specifically an L-shaped trigger rod fixed to a rotating output shaft below the glass conveying trajectory. This L-shaped trigger rod is configured such that its short horizontal side extends into the glass conveying path.

[0228] The physical action shown in the attached diagram—the instant when the bottom front edge of the glass contacts the horizontal short side of the L-shaped trigger rod, causing it to rotate in a flipping motion—is the technical starting point for the two core functions of this invention:

[0229] As the starting point for length measurement (corresponding to step a): when the triggering mechanism is pushed from the initial position and its deflection angle reaches a "preset trigger threshold", the system will capture a "first time signal".

[0230] As a data source for condition diagnosis, transient dynamic response data (including torque, velocity, and acceleration curves) is collected throughout the entire flipping action described above, used for in-depth diagnosis of the equipment's health status. Therefore, the L-shaped trigger lever and its associated rotary shaft system together constitute the defined "target measurement device." For clarity, this attached figure only shows this initial triggering action and does not illustrate the subsequent "reset and cleaning action" or the "returning action under restoring force."

[0231] The first flowchart, "Multidimensional Feature Parameter Acquisition," corresponds to steps S1 and S2. It represents the data acquisition stage of this invention, namely: on the one hand, by analyzing the transient dynamic response data of the servo measurement device itself, the "first dynamic feature parameter" (corresponding to S1) characterizing its physical health status is acquired; on the other hand, by analyzing the production output data and production topology of the target device and the associated device cluster, the "second associated feature parameter" (corresponding to S2) characterizing the systematic fluctuations of the external environment is acquired.

[0232] The second flowchart, "Probabilistic Attribution Model Fusion Analysis," corresponds to step S3. This flowchart represents the core analysis stage of the present invention, which uses a probabilistic attribution model to fuse the two feature parameters obtained in the previous stage, representing "microscopic introspection" and "macroscopic insight," respectively, and finally generates a "deviation root cause attribute indication signal" that can quantify the root cause of deviation in probabilistic form.

[0233] The third flowchart, "In-depth Diagnosis and Decision of Internal Faults," corresponds to step S4. This flowchart represents the first diagnostic decision branch of the present invention. When the indication signal generated in step S3 indicates that the deviation is highly likely to originate from "internal equipment factors" (i.e., the first posterior probability is greater than a preset threshold), the system will initiate this process to perform in-depth, graded diagnosis of the equipment and generate a "first diagnostic instruction" containing specific maintenance recommendations.

[0234] The fourth flowchart, "Systematic Source Tracing and Decision-Making of External Risks," corresponds to step S5. This flowchart represents the second diagnostic decision-making branch of the present invention. When the indicator signal suggests that the deviation is highly likely to originate from "external environmental factors" (i.e., the second posterior probability is greater than a preset threshold), the system will initiate this process, suppress single-point alarms, and instead analyze the entire production topology. Through calculation and sorting, a "second diagnostic instruction" that can highlight the root cause of systemic risks will be generated.

[0235] Figure 4The flowchart illustrates the complete process from receiving the same "multi-device concurrent alarm" input to generating a final diagnostic conclusion. The simulation results reveal the fundamental breakthrough of this invention in diagnostic intelligence. In the "Primary Technology" path on the left, systemic fault phenomena are treated merely as a disordered, flat list of alarm devices. This essentially throws the complex topology tracing problem directly at the operators, resulting in extensive manual investigation and a diagnostic conclusion of "inefficiency." In stark contrast, in the "Invention" path on the right, the same alarm input is fed into the core "Topology Scoring and Ranking Algorithm" module. This module analyzes and quantifies the topological relationships between devices, directly outputting a structured "Systemic Risk Source Map." In this map, the most suspicious shared resource nodes (such as "power supply circuits") are prominently identified by their high "shared resource correlation score." This indicates that the method proposed in this invention can theoretically automatically complete the reasoning from phenomenon to root cause, providing direct and visual evidence for the technical effect of "guiding the focus of fault investigation from phenomenon to root cause, improving diagnostic efficiency."

[0236] It should be noted that all calculation formulas in this application employ regression analysis, including but not limited to machine learning algorithms, to deeply analyze the collected parameters and identify their natural trends and interrelationships. Specialized software, such as Python's Scikit-learn library or the R language, is used to automatically generate mathematical models that match the data. Then, cross-validation and other methods are used to objectively evaluate the model performance, and continuous feedback and optimization are combined to ensure that the created formulas truly reflect the inherent laws of the data, thereby guaranteeing their effectiveness and accuracy. In all calculation formulas in this application, the parameters in each formula undergo dimensionless processing within a consistent range to ensure that different physical quantities are compared on the same scale; dimensionless processing techniques include, but are not limited to, min-max-normalization and Z-score standardization.

[0237] The algorithm of this invention is implemented as a Python script. Before executing the core logic, the program first executes a data loading module (e.g., using the widely used pandas library in Python) configured to read the aforementioned spreadsheet file and load its contents into the program's working memory (e.g., a DataFrame data structure). Subsequent algorithm steps will directly query and retrieve the required configuration parameters from this in-memory data structure.

[0238] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for identifying system deviations in a glass production process based on servo-triggered measurement, characterized in that, The specific steps include: S1: For the target measuring device in the production process, based on the transient dynamic response data of the target measuring device in the servo-triggered measurement process, obtain the first dynamic characteristic parameter characterizing the physical health status of the target measuring device itself; S2: Obtain the production process output data of the target measuring device and at least one associated device that is associated with the target measuring device in the production topology, and based on the statistical characteristics of the production process output data and the production topology, obtain a second associated feature parameter characterizing the systematic fluctuation state of the process environment in which the target measuring device is located. S3: When a production process deviation of the target measuring equipment is detected, the first dynamic feature parameter and the second associated feature parameter are fused in the probabilistic attribution model to generate a deviation root cause attribute indication signal that characterizes whether the root cause of the production process deviation lies in the internal factors of the equipment or the external factors of the environment.

2. The method for identifying system deviations in glass production processes based on servo-triggered measurement according to claim 1, characterized in that: S1 includes: S11: Collect the angular position sequence of the servo drive component in the target measuring device within a preset time window as the transient dynamic response data; S12: In the preset dynamic system model, the angular position sequence is processed using a system identification algorithm to output the transfer function parameters as the first dynamic feature parameters.

3. The method for identifying system deviations in glass production processes based on servo-triggered measurement according to claim 2, characterized in that: The first dynamic feature parameter is the dynamic health degradation index; step S12 further includes: S121: The transfer function parameters are formed into a parameter time series over multiple consecutive measurement periods; S122: The dynamic health degradation index is calculated based on the short-term volatility, medium-term trend, and long-term drift of the parameter time series.

4. The method for identifying system deviations in glass production processes based on servo-triggered measurement according to claim 3, characterized in that: S2 includes: S21: Obtain the length data sequence generated by the target measuring device and associated devices within their respective measurement cycles as the production process output data; S22: Calculate the information entropy of each data sequence of length within the sliding time window to generate an entropy value sequence; S23: Based on each entropy value sequence and the production topology obtained from the manufacturing execution system that defines the shared resource relationships between devices, the second associated feature parameter is generated.

5. The method for identifying system deviations in glass production processes based on servo-triggered measurement according to claim 4, characterized in that: The second associated characteristic parameter is the systemic risk transmission coefficient; step S23 further includes: S231: In the production topology, each device is regarded as a node, and the shared resource relationship between devices is regarded as an edge, and a device association graph is constructed; wherein, the target measuring device and all associated devices are respectively mapped as unique nodes in the device association graph; S232: Calculate the systematic risk transmission coefficient of the target measurement device node based on the entropy sequence of the target measurement device node and the entropy sequence of its adjacent nodes in the device association graph.

6. The method for identifying system deviations in glass production processes based on servo-triggered measurement according to claim 5, characterized in that: In S3, the probabilistic attribution model is a dynamic Bayesian attribution model; S3 includes: S31: Based on the time change rate of the first dynamic feature parameter, calculate the first posterior probability that characterizes the deviation caused by the intrinsic factors of the target measuring device; S32: Based on the synchronization between the second associated feature parameter and the corresponding parameter of the adjacent device in the production topology, calculate the second posterior probability characterizing the deviation caused by external environmental factors; S33: Output the probability vector containing the first posterior probability and the second posterior probability as the bias root cause attribute indication signal; Before inputting the first dynamic feature parameter and the second associated feature parameter into the dynamic Bayesian attribution model, the values ​​of the first dynamic feature parameter and the second associated feature parameter are mapped to a first evidence value and a second evidence value representing the strength of internal and external evidence, respectively, through a preset continuous evidence strength function.

7. The method for identifying system deviations in glass production processes based on servo-triggered measurement according to claim 6, characterized in that: The method also includes: S4: Compare the first posterior probability and the second posterior probability in the deviation root cause attribute indication signal; if the first posterior probability is greater than a first preset threshold, generate a first diagnostic instruction to indicate precise maintenance of the target measuring device.

8. The method for identifying system deviations in glass production processes based on servo-triggered measurement according to claim 7, characterized in that: The first diagnostic instruction is a graded diagnostic and recommendation instruction; the step of generating the first diagnostic instruction further includes: S41: Based on the value of the first posterior probability, determine the fault severity level, which characterizes the urgency of the fault; S42: Based on the contribution of each component of the dynamic health degradation index, determine at least one recommended maintenance action, wherein the components include short-term volatility, medium-term trend and long-term drift.

9. A method for identifying system deviations in a glass production process based on servo-triggered measurement according to claim 8, characterized in that: The method also includes: S5: Compare the first posterior probability and the second posterior probability in the deviation root cause attribute indication signal; if the second posterior probability is greater than the second preset threshold, generate a second diagnostic instruction to suppress single-point alarms for the target measuring device and highlight the affected device cluster in the production topology.

10. A method for identifying system deviations in a glass production process based on servo-triggered measurement according to claim 9, characterized in that: The step of generating the second diagnostic instruction further includes: Based on the production topology, at least one shared resource node is identified that has a shared resource relationship with the devices in the affected equipment cluster; For at least one shared resource node, a shared resource correlation score is determined by calculating its correlation with devices in the affected device cluster; and, Based on the shared resource correlation score, a systemic risk source map is generated as part of the second diagnostic instruction. The systemic risk source map is used to indicate the likelihood that each shared resource node is a source of deviation. The systemic risk tracing map includes a sorted list, which arranges at least one shared resource node in descending order of its shared resource correlation score.

Citation Information

Patent Citations

  • Monitoring system based on glass production process

    CN117873009A