Digital factory intelligent control method based on AI
By using AI to identify critical points in equipment health and predict failure times, and combining equipment dependency graphs and graph neural networks, the propagation of risks is dynamically blocked, solving the reliability and stability problems of equipment under traditional maintenance models, and achieving efficient equipment maintenance and stable production operation.
Patent Information
- Application Number
- CN202511045561.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-29
AI Technical Summary
Traditional equipment maintenance methods are difficult to meet the high reliability and stability requirements of modern digital factories. Regular maintenance can lead to resource waste or frequent failures, and post-incident repairs can cause production interruptions and economic losses.
Through AI-based intelligent control methods for digital factories, we can accurately identify critical points in equipment health, predict the time of failure in advance, reasonably judge the reversibility of equipment degradation, formulate targeted and cost-effective intervention strategies, combine equipment dependency maps and graph neural networks to identify risk propagation paths, and implement dynamic blocking of risk propagation.
It enables precise monitoring of equipment health status and graded judgment of degradation status, formulates scientific and reasonable intervention strategies, improves production quality and efficiency, reduces maintenance costs, prevents risk rebound, and ensures smooth production.
Smart Images

Figure CN120848422A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence control technology, and in particular to an AI-based intelligent control method for digital factories. Background Art
[0002] In modern industrial production, digital factories have become an important means to improve production efficiency, ensure product quality, and reduce operating costs. With the deepening promotion of smart manufacturing and Industry 4.0 concepts, factory equipment is gradually developing towards high automation and intelligence. However, with the increasing complexity of equipment and the diversification of operating environments, equipment failures and performance degradation are becoming increasingly prominent, bringing many challenges to production.
[0003] Traditional equipment maintenance primarily relies on periodic inspections and reactive repairs. While periodic inspections can prevent equipment failures to some extent, they often result in over-maintenance or under-maintenance, leading to resource waste or frequent breakdowns. Reactive repairs, on the other hand, involve fixing equipment after it has failed. This approach not only causes production interruptions but can also trigger cascading failures, leading to even greater economic losses. Therefore, traditional maintenance models are ill-suited to meet the high reliability and stability requirements of modern digital factories. Summary of the Invention
[0004] This application provides an AI-based intelligent control method for digital factories, which accurately identifies critical points in equipment health status, predicts the time of failure in advance, reasonably judges the reversibility of equipment degradation, and formulates targeted and cost-effective intervention strategies.
[0005] This application provides an AI-based intelligent control method for digital factories, including: S101, Collect data, extract data features, calculate health index based on data features, and issue early warnings to the device based on health index. The data includes device operating data and device basic data. S102, collect the operating data of the equipment in a healthy state, establish a baseline based on the operating data, compare the real-time collected feature values with the baseline, identify the critical point of state change, and track the deterioration trajectory based on the critical point of state change; S103, Identify the degradation state based on the critical point of state change, calculate the reversibility probability based on the degradation state of the equipment, obtain the reversible range and irreversible point of equipment degradation based on the reversibility probability, and implement corresponding intervention strategies according to the reversibility situation. S104, based on the characteristics of the data, predict the time of failure occurrence, calculate the optimal intervention time based on the predicted failure occurrence time, and match the means according to the means matching rule base based on the failure characteristics.
[0006] Preferably, the formula for calculating the health index is: ,in, For health index, For energy spectrum entropy, For current imbalance, Classified by voiceprint abnormalities, The slope of the pass rate The weights for the energy spectral entropy are derived based on the degree of mechanical wear and vibration signals. The weighting for current imbalance is derived from electrical faults. The weighting for voiceprint anomalies is derived based on voiceprint characteristics. The weights for the slope of the pass rate are derived based on the degree of deterioration in production quality. .
[0007] Preferably, the baseline is a range that is adjusted in real time according to the equipment's operating status, environmental changes, and equipment aging factors. The data features include energy spectrum entropy, current imbalance, abnormal sound signature, and pass rate slope. The baseline range is the average value of the data features over a period of time plus or minus two standard deviations.
[0008] Preferably, the state change threshold includes a first-level threshold and a second-level threshold. The first-level threshold is the point where the health index rises and continues to rise over a period of time, but does not reach a preset risk threshold. The second-level threshold is the point where the health index rises and exceeds a preset risk threshold.
[0009] Preferably, the formula for calculating the reversibility probability is: ,in, The probability of reversibility is represented by n, where n is the number of reversible evidences and N is the total number of evidences. Reversibility is determined based on the calculated probability of reversibility, and a probability threshold range is set, which includes a maximum threshold and a minimum threshold. When the probability of reversibility is greater than the maximum threshold, the device degradation is considered to be highly reversible. When the probability of reversibility is less than or equal to the maximum threshold but greater than the minimum threshold, the device degradation is considered to be critically reversible. When the probability of reversibility is less than or equal to the minimum threshold, the device degradation is considered to be irreversible.
[0010] Preferably, S201: Obtain the technological relationship of the equipment based on the collected basic data of the equipment, calculate the dependency strength, and construct the dependency relationship of the production chain based on the dependency strength and equipment information; S202 collects historical fault data, builds a propagation model based on the historical fault data, monitors risk source nodes in real time, and generates blocking strategies based on the risk source nodes.
[0011] Preferably, the risk source nodes are divided into three risk levels: single-point high-risk, multi-node medium-risk, and full-chain high-risk. The single-point high-risk means there is only one risk source node, the multi-node medium-risk means there are multiple risk source nodes, and the full-chain high-risk means there are multiple risk source nodes that are interconnected and interact with each other, resulting in a complete shutdown.
[0012] Preferably, S301: Obtain equipment degradation entropy value data based on equipment operating data and failure frequency; calculate risk pressure value based on equipment degradation entropy value data; calculate risk transmission intensity based on risk pressure value; identify risk flow direction; and generate risk flow diagram based on risk flow direction. S302 monitors the intensity of risk flows in real time, generates digital blocking barriers based on the intensity of risk flows, identifies risk flow loops according to algorithms, and adjusts strategies according to the risk flow loop level.
[0013] Preferably, the formula for calculating the intensity of risk transmission is: ,in, For the intensity of risk transmission, The risk pressure value for device i. This provides risk transmission characteristic data between device i and device j. As the weight of the risk stress value, Weights for risk transmission characteristic data. + =1.
[0014] Preferably, the risk flow direction refers to the transmission from high-voltage equipment to low-voltage equipment.
[0015] One or more technical solutions provided in this application have at least the following technical effects or advantages: by accurately identifying the critical point of equipment health status, predicting the time of failure in advance, reasonably judging the reversibility of equipment degradation status, and formulating targeted and cost-effective intervention strategies, the application achieves accurate monitoring of equipment health status, critical point identification, degradation status classification judgment, and reversibility determination, and formulates scientific and reasonable intervention strategies. Through effect verification and continuous optimization, the application ensures stable equipment operation, improves production quality and efficiency, and reduces maintenance costs. By using equipment dependency graphs and graph neural networks, it is possible to accurately identify the dependencies between equipment and the risk propagation path, suppress the chain failure of related equipment caused by the irreversible decline of a single piece of equipment in the production line, achieve dynamic blocking of risk propagation, minimize the scope of production stoppage and losses, and improve the stability and economy of production line operation. By combining technologies such as digital blocking barriers and smart circuit breakers, it is possible to more accurately block risk transmission paths, prevent risk rebound, effectively protect healthy equipment, improve the overall reliability of equipment operation, realize dynamic topology analysis of equipment risks, accurately predict and effectively block risk transmission through risk flow simulation and digital twin blocking technology, prevent risk rebound, improve the stability and reliability of equipment operation, and ensure the smooth operation of production. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating an AI-based intelligent control method for digital factories according to the present invention. Figure 2 A schematic diagram illustrating the process of constructing the production chain dependency relationship for this invention; Figure 3 This is a schematic diagram of the process for generating a risk flow diagram for this invention. DETAILED DESCRIPTION
[0017] To facilitate understanding of the present invention, a more complete description of this application will be given below with reference to the accompanying drawings, which illustrate preferred embodiments of the invention. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to enable a more thorough and complete understanding of the disclosure of the present invention.
[0018] It should be noted that the terms "vertical," "horizontal," "up," "down," "left," "right," and similar expressions used in this article are for illustrative purposes only and do not represent the only possible implementation.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to limit the invention; the term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0020] Example 1: Figure 1 This is a flowchart illustrating an AI-based intelligent control method for a digital factory according to an embodiment of the present invention, including: S101: Collect data, extract data features, calculate health index based on data features, and issue early warnings to devices based on health index; Furthermore, temperature sensors, voltage sensors, vibration sensors, acoustic sensors, and current clamps are installed on the equipment to monitor the equipment's temperature, voltage, vibration signals, mechanical noise, and three-phase current. The vibration sensor is installed on the equipment's spindle to monitor the spindle's vibration in real time, including amplitude, frequency, and phase. The acoustic sensor is installed on the equipment's housing to capture abnormal mechanical noise during operation. The current clamp is installed at the motor input or distribution cabinet output to monitor fluctuations in the motor's three-phase current. The same type of component is evenly distributed to multiple devices for production, and the number of qualified and unqualified components for each device during the production cycle is collected in real time. The component qualification rate is calculated based on the collected number of qualified and unqualified components.
[0021] The acquired vibration signal is subjected to a Fast Fourier Transform (FFT) to obtain the energy distribution in the 0-5kHz frequency domain. The 0-5kHz frequency domain is then divided into multiple sub-bands of equal or non-equal width to obtain the energy of each band. The formula for calculating the energy spectral entropy is as follows: ,in, The energy spectral entropy quantifies the degree of disorder in the frequency domain energy distribution of the vibration signal. N is the total number of frequency bands (dividing the 0-5kHz frequency domain into multiple sub-bands), and i is the index of the frequency band, indicating the current frequency band number being calculated. The energy proportion of the i-th frequency band reflects the concentration of energy distribution in the frequency domain. Low energy spectral entropy indicates that energy is concentrated in a few frequency bands (e.g., the vibration energy of normal equipment is mainly distributed in the low frequency band), indicating an orderly energy distribution. High energy spectral entropy indicates that energy is dispersed in multiple frequency bands (e.g., a fault leads to an increase in high-frequency components), indicating a disordered energy distribution and an increased risk of mechanical wear or faults. The formula for calculating the current imbalance based on the monitored three-phase current is: ,in, This refers to the current imbalance, ranging from 0 to 1. It is the maximum value among the three-phase currents. It is the minimum value among the three-phase currents. The average value of the three-phase current, when When the ratio is 0, it indicates that the three-phase currents are completely balanced, which is an ideal state. A larger value indicates a more severe current imbalance and a higher risk of electrical faults. The current imbalance value at the current moment is obtained and a timestamp is recorded. The spectral resolution of the acquired mechanical signal is calculated based on the signal length and sampling rate. A single-sided amplitude spectrum, i.e., the frequency domain spectrum, is generated based on the spectral resolution. The measured spectrum is compared with the theoretical fault frequency database. The formula for calculating the voiceprint anomaly score is as follows: ,in, For abnormal voiceprint readings, A represents the fault frequency amplitude, B represents the noise floor, B represents the background noise level of the spectrum under normal operating conditions, and T represents the significance threshold. This represents the upper limit of the dynamic range; the slope of the pass rate is calculated based on the obtained component pass rate. The formula for calculating the slope of the pass rate is: ,in, The slope of the pass rate The current pass rate, The pass rate at the previous moment. The time interval is defined as follows: a positive slope indicates an increase in the pass rate, while a negative slope indicates a decrease. A larger absolute value of the pass rate slope indicates a more pronounced trend of quality deterioration. The health index is calculated based on the calculated energy spectral entropy, current imbalance, abnormal voiceprint score, and pass rate slope. The formula for calculating the health index is: ,in, For health index, For energy spectrum entropy, For current imbalance, Classified by voiceprint abnormalities, The slope of the pass rate The weights for the energy spectral entropy are derived based on the degree of mechanical wear and vibration signals. The weighting for current imbalance is derived from electrical faults. The weighting for voiceprint anomalies is derived based on voiceprint characteristics. The weights for the slope of the pass rate are derived based on the degree of deterioration in production quality. ; The system sets a first-level warning threshold based on the calculated health index. When the health index exceeds the preset first-level warning threshold, the system automatically triggers a first-level warning, indicating that the equipment has a high probability of failure risk. Warning information is pushed through SMS, email or industrial APP, and shutdown maintenance is required.
[0022] S102, collect the operating data of the equipment in a healthy state, establish a dynamic baseline based on the operating data, compare the real-time collected feature values with the dynamic baseline, identify the critical point of state change, and track the deterioration trajectory based on the critical point; Specifically, the dynamic baseline is a range that is adjusted in real time according to factors such as equipment operating status, environmental changes, and equipment aging. It collects operating data of the equipment in a healthy state, which includes the state after new equipment installation and commissioning, major overhaul, or when production enters a stable period. Data acquisition is initiated in this state, and features such as temperature, voltage, vibration signal, mechanical noise, and three-phase current are extracted from the collected data. Feature values such as energy spectrum entropy, current imbalance, acoustic anomaly score, and pass rate slope are calculated. The average and standard deviation of the energy spectrum entropy, current imbalance, acoustic anomaly score, and pass rate slope over a period of time are calculated. Based on the statistical process control (SPC) principle, the initial normal range for each feature value is set to the average value plus or minus twice the standard deviation. For example, the baseline range for vibration energy spectrum entropy is 0.1 - 0.3. Simultaneously, a rolling window averaging method is used to calculate statistics to avoid the influence of extreme values on the threshold, ensuring that the initial baseline range can accurately reflect the characteristics of normal equipment operation. The threshold is automatically recalculated every quarter, incorporating newly collected health data into the calculation range. When the equipment is maintained or key components are replaced, the operating characteristics of the equipment may change, requiring immediate baseline reset. The baseline is dynamically adjusted based on the pass rate. If sensor characteristics are within the baseline range, but the pass rate shows an abnormal decline, a decision tree model is used in conjunction with the pass rate signal to rank the importance of features, recalibrate feature weights, prioritize retaining features highly correlated with the pass rate, and adjust their threshold range, enabling the baseline to more accurately determine the health status of the equipment.
[0023] The real-time collected feature values are compared with the established dynamic baseline. For each feature value, it is determined whether it exceeds the baseline range. The number of times each feature exceeds the limit and the duration are recorded. Features with more than three instances of exceeding the limit are marked as abnormal events. When at least two features exceed the limit three times consecutively, and the pass rate slope is negative, this situation is marked as potential degradation. For features exceeding the limit, the deterioration rate is quantified using a simple moving average method. At the same time, a health index is calculated, and a normal range and risk threshold for the health index are set. When the health index starts to rise from the normal range and continues to rise for a period of time, but does not reach the risk threshold, and the pass rate slope is slightly negative, the equipment is determined to be at the first-level critical point. When the health index rises continuously and exceeds the set risk threshold, the equipment is considered to be at the first-level critical point. When the pass rate slope drops sharply, the equipment is judged to be at the second-level critical point. From single-feature anomaly to multiple consecutive over-limits, the rise in health index, and the pass rate slope from a slightly negative value to a sharp drop, the complete deterioration trajectory of the equipment from initial anomaly to gradual decline to critical failure is tracked. When the first-level critical point is identified, a first-level warning is triggered. At this time, the equipment load is automatically reduced by 20% to alleviate the operating pressure of the equipment and slow down the rate of equipment deterioration. At the same time, maintenance personnel are notified to inspect the equipment. When the second-level critical point is identified, a critical warning is immediately triggered. At this time, the backup equipment is immediately switched to ensure the continuity of production. At the same time, a detailed diagnostic report is generated, which includes abnormal feature combinations, deterioration trajectory diagrams, and possible cause analysis.
[0024] S103, Identify the degradation state based on the critical point of state change, obtain the reversible range and irreversible point of equipment degradation based on the degradation of the equipment, and implement corresponding intervention strategies based on the degradation level and reversibility. Furthermore, other data during equipment operation are collected, including mechanical response characteristics (such as the rate of vibration entropy fallback after lubrication), material deformation characteristics (metal fatigue obtained through current waveform analysis), and recovery potential characteristics (success rate of historical maintenance of the same type). When the health index continues to rise and abnormalities occur, or when the pass rate slope shows a significant negative change, it is determined that the equipment may be entering a deterioration state. The deterioration state is determined by combining mechanical response characteristics, material deformation characteristics, and recovery potential characteristics. If the rate of vibration entropy fallback after lubrication is lower than the normal level, it indicates that the mechanical response of the equipment is abnormal. If the cumulative metal fatigue index obtained through current waveform analysis exceeds the low-risk threshold, it indicates that the material deformation has affected the normal operation of the equipment. If the success rate of historical maintenance of the same type is lower than a certain level, it means that the equipment may be difficult to recover through routine maintenance in the current state. Based on the above, when multiple characteristics show abnormalities at the same time, it is determined that the equipment is in a deterioration state.
[0025] To determine whether a device is reversible, reversible and irreversible thresholds are set based on equipment characteristics. The real-time monitored vibration entropy fall rate is compared with these thresholds. If the fall rate is greater than the reversible threshold, the device is reversible; if it is less than the irreversible threshold, it is irreversible. Historical maintenance records of the same type are queried, and the historical maintenance success rate is calculated. If the success rate is >80%, the device is reversible; if it is <30%, it is irreversible. The metal fatigue accumulation index is obtained from current waveform analysis. If the index is <0.5, the device shows a tendency towards reversibility; if it is >0.8, the device shows a tendency towards irreversibility. The vibration entropy fall rate, historical maintenance success rate, and fatigue accumulation index are all evidence for determining reversibility. The formula for calculating the probability of reversibility based on these evidence items is as follows: ,in, Let n represent the reversibility probability, n be the number of reversible pieces of evidence, and N be the total number of pieces of evidence. Reversibility is determined based on the calculated reversibility probability. If... If the value is >0.7, the equipment degradation is determined to be highly reversible (corresponding to degradation level 0-1); if 0.4 < If the value is ≤0.7, the equipment degradation is determined to be critically reversible (corresponding to degradation level 1); if ≤0.4 indicates irreversible degradation of the equipment (corresponding to degradation level 2-3). Degradation levels 0 and 1 fall within the reversible range. Degradation level 0 indicates the equipment is in an elastic degradation state, such as a short-term temperature exceedance. In this case, the equipment has the ability to recover automatically, which is within the reversible range. Degradation level 1 indicates the equipment is experiencing plastic degradation, such as insufficient lubrication. With external intervention, such as automatic lubrication or load adjustment, the equipment can recover its baseline performance within 72 hours, which is also within the reversible range. Degradation levels 2 and 3 are irreversible points. Degradation level 2 indicates partial failure of the equipment, such as a bearing crack. At this point, the equipment has suffered relatively serious damage and requires component replacement to restore normal operation. If not handled in time, the equipment condition will further deteriorate, which is close to an irreversible state. Degradation level 3 indicates complete equipment failure, such as a broken spindle. The equipment can no longer operate and can only be decommissioned, which is a clear irreversible point.
[0026] Based on the intervention strategies corresponding to the identified degradation levels, when the degradation level is 0-1 and highly reversible, for degradation caused by insufficient lubrication, the system automatically triggers the lubrication device to add an appropriate amount of lubricating oil to the equipment, improving the lubrication condition. According to the equipment's performance and operating status, the system automatically adjusts the equipment load to avoid overload operation, ensuring the equipment operates within a reasonable load range and promoting performance recovery. For degradation level 1 and critically reversible, preventative replacement of easily damaged parts is performed, such as replacing belts and filter elements nearing the end of their service life, to prevent further deterioration of the equipment. Production process parameters are adjusted to compensate for equipment operation and reduce performance fluctuations caused by degradation, for example, adjusting processing accuracy and speed parameters to ensure product quality. For degradation levels 2-3 and irreversible, a decommissioning warning signal is sent to factory management personnel, reminding them to promptly arrange equipment decommissioning plans. The intelligent allocation module is automatically invoked to transfer production tasks from irreversibly deteriorating equipment to first-level performance equipment, ensuring production continuity. Simultaneously, a decommissioning schedule is generated to rationally arrange equipment decommissioning and upgrade work, triggering spare parts procurement requests (through the MES system) to promptly procure the spare parts needed for equipment upgrades; at the same time, an equipment upgrade budget application is submitted to the ERP system to provide financial support for equipment upgrades.
[0027] S104, Based on the characteristics of the data, predict the time of failure occurrence, calculate the optimal intervention time based on the predicted failure occurrence time, and match the means according to the means matching rule base based on the failure characteristics. Specifically, based on the calculated data characteristics, when the daily increase in energy spectrum entropy is greater than 0.1 and the slope of the pass rate decreases by less than -0.15 / min, it indicates that the equipment fault propagation speed is relatively fast, and a fault is expected to occur within 12 hours. When the current imbalance is greater than 15% and the pass rate falls below the preset minimum pass rate threshold, it means that there is a serious problem with the equipment's electrical system, and a shutdown is expected to occur within 2 hours. The formula for calculating the optimal intervention time is: ,in, For the optimal time to intervene, The estimated time of failure is derived from data characteristics. The production gap cutoff time is determined based on the current production status of the equipment. To ensure sufficient preparation and execution time for intervention operations and prevent equipment failure due to untimely intervention, a safety margin is set by default to 1 hour. Fault characteristic combinations are obtained based on data features, including vibration dominance, voltage fluctuations and a sharp drop in pass rate, and multimodal anomalies and health indices exceeding risk thresholds. Means matching is performed from a means matching rule base. The steps are as follows: For vibration-dominant fault characteristics, load reduction and automatic lubrication are implemented. Reducing the load can decrease mechanical stress on the equipment and inhibit further mechanical wear; automatic lubrication can improve the lubrication condition of the equipment and reduce friction. For fault characteristics such as voltage fluctuations and a sharp drop in the pass rate, switching to a regulated power supply and reprocessing defective products is recommended. Switching to a regulated power supply ensures the electrical stability of the equipment and prevents damage caused by voltage fluctuations. Reprocessing defective products can improve the pass rate and reduce losses. For fault characteristics such as multimodal anomalies and health indices exceeding the risk threshold, switching to backup equipment and triggering a maintenance work order is recommended. When equipment exhibits multimodal anomalies and a high health index, it indicates that the equipment may have a serious fault. Switching to backup equipment can prevent production line interruptions and ensure production continuity. Triggering a maintenance work order allows for timely arrangement of professional personnel to repair the faulty equipment.
[0028] The technical solutions in the above-described embodiments of this application have at least the following technical effects or advantages: by accurately identifying the critical point of equipment health status, predicting the time of failure in advance, reasonably judging the reversibility of equipment degradation status, and formulating targeted and cost-effective intervention strategies, the system achieves accurate monitoring of equipment health status, critical point identification, degradation status classification judgment, and reversibility determination, and formulates scientific and reasonable intervention strategies. Through effect verification and continuous optimization, the system ensures stable equipment operation, improves production quality and efficiency, and reduces maintenance costs.
[0029] Example 2: Based on the irreversible equipment failures in Example 1, and addressing the problem of cascading failures of related equipment caused by irreversible degradation of a single piece of equipment in a production line (such as a spindle breakage leading to gearbox overload), this example solves the limitations of traditional single-point maintenance by constructing a networked cascading risk control system. Figure 2 As shown.
[0030] S201: Obtain the technological relationship of the equipment based on the collected basic data of the equipment, calculate the dependency strength, and construct the dependency relationship of the production chain based on the dependency strength and equipment information; Furthermore, the collected equipment includes equipment models, serial numbers, and functions. Through the production site, the technological processes and connections between various pieces of equipment in the production chain are understood. For example, in the automotive manufacturing production chain, the sequence and collaboration methods between stamping equipment, welding equipment, painting equipment, and final assembly equipment are clarified. The input and output information of each piece of equipment during the production process is recorded, forming a complete list of equipment technological relationships. Through theoretical analysis, experimental testing, or actual production data monitoring, the load transfer coefficients between equipment in the production chain are obtained. Historical fault data is collected, including the time of fault occurrence, equipment name, fault type, and fault propagation information. Based on the historical fault data, the number of historical fault propagations between each pair of equipment is counted. The formula for calculating the dependency strength using the number of historical fault propagations is: ,in, Dependency strength measures the extent to which a failure in one device affects another. Historical fault propagation count refers to the number of times, within a certain period of time, a failure in one device caused a failure or abnormality in another device. The current load ratio refers to the relative proportion of load between two devices under the current production state. A judgment threshold is set based on the dependency strength. When the calculated dependency strength is greater than the preset judgment threshold, the connection of that device is marked as a red warning edge. Using the visualization tool Gephi, the devices in the production chain are displayed as nodes, and the connections between devices are represented by edges. Different colors and styles are set for the edges according to the calculated dependency strength. For example, high dependency relationships (red warning edges) are represented by eye-catching red lines, and other dependencies are distinguished by lines of different thicknesses or colors. At the same time, the device name is marked on the node, and key information such as dependency strength is displayed on the edge, making the dependencies between devices clearly visible. When irreversible decay device information is obtained, it is marked as a risk source node in the device dependency graph; when reversible decay device information is obtained, it is marked as a potential transmission node.
[0031] S202: Collect historical fault data, build a propagation model based on the historical fault data, monitor risk source nodes in real time, and generate blocking strategies based on risk source nodes; Specifically, representative cases are selected from fault records based on the significant characteristics of cascading faults, such as the chain reaction of faults and causing large-scale production line shutdowns. For these selected cases, basic information such as the time, location, equipment model, and fault symptoms are recorded. Based on this basic information, equipment degradation status is identified, dependency strength is calculated, and the degradation probability of related equipment is identified. For the degradation probability of related equipment, the likelihood of related equipment in the case experiencing degradation after the failure of the risk source equipment is identified. The probability value is determined through a combination of expert evaluation and data analysis. The compiled and labeled case information is stored in a unified format to construct a structured training dataset. Using the graph neural network framework PyTorch Geometric, the equipment degradation status and dependency strength data are input into the input layer, organizing the data into a graph structure. Equipment is used as nodes in the graph, dependency strength is used as the weight of the edges between nodes, and equipment degradation status is used as the attribute of the nodes. The equipment degradation status and dependency strength data are then populated into the PyTorch Geometric graph. In the Geometric data structure, the input layer serves as the input to the output layer, which outputs the predicted probability of device degradation and the time of failure. Two independent output branches are used: one outputs the degradation probability (ranging from 0 to 1), and the other outputs the failure time. The degradation probability output branch uses a Sigmoid activation function to limit the output value to 0-1; the failure time output branch uses a linear activation function to directly output the predicted time value. The constructed training dataset is divided into training, validation, and test sets. The training set is used for model training, employing the Adam optimization algorithm to adjust the model parameters to minimize the loss function. The number of training epochs and batch size are set, and the model parameters are continuously updated through iterative training to gradually bring the model's predictions closer to the actual data. The validation set is used to adjust the model parameters and evaluate its performance during training. The test set is used to ultimately evaluate the model's generalization ability. Real-time monitoring of risk source nodes includes monitoring equipment operating parameters (such as temperature, vibration, and current) and fault alarm information.Real-time status data of equipment is acquired through sensors and data acquisition devices. Risk warning thresholds are set. When the status data of a risk source node exceeds the warning threshold, a risk simulation process is triggered. Based on the structure of the equipment dependency graph, first-order related devices directly connected to the risk source node are extracted. Using an adjacency list traversal algorithm, the adjacent nodes of the risk source node are traversed and extracted as first-order related devices. The decay entropy value of the risk source node is obtained. The decay entropy value can be calculated based on the equipment's operating parameters and historical data and is used to measure the degree of equipment decay and failure risk. The dependency strength between the risk source node and the first-order related devices is obtained from the equipment dependency graph. The propagation decay probability = source node decay entropy value × dependency strength. Based on actual production needs and risk tolerance, a propagation decay probability threshold is set. When the calculated propagation decay probability is greater than the threshold, second-order related devices directly connected to the first-order related devices are scanned. The adjacent nodes of the first-order related devices are traversed to obtain the information of the second-order related devices.
[0032] Risk levels are categorized based on the number, scope, and severity of risk source nodes. A single high-risk point is defined as one node whose impact is limited to a few closely related devices, with minimal disruption to the overall production process. Multiple nodes exist, but if they haven't yet created a chain reaction leading to a complete production line shutdown, it's a medium-risk multi-node situation. A severe situation with multiple interconnected and interacting risk source nodes, causing a complete line shutdown, is defined as a high-risk end-to-end situation. For single-point high-risk situations, the risk source device is shut down, redundant equipment is activated, the emergency stop button is pressed, and the power is turned off. For multi-node situations… For medium-risk situations, a strategy of slowing down production and implementing process compensation is adopted. By adjusting the control parameters of production equipment (such as motor speed, conveyor belt speed, etc.), the production speed is reduced to 50% of the original speed. Based on the production situation after the speed reduction, the production process is adjusted accordingly, such as increasing processing time and adjusting processing parameters, to ensure that product quality is not affected. When a high-risk situation occurs throughout the entire chain, a strategy of isolating the faulty segment and dynamically reorganizing the production flow is implemented. By setting up physical isolation devices (such as isolation barriers, valves, etc.) or logical isolation measures (such as the locking function in the software control system), the faulty segment with multiple risk source nodes is isolated from other normal production segments to prevent the risk from spreading further.
[0033] The technical solutions described in the above embodiments of this application have at least the following technical effects or advantages: by using equipment dependency graphs and graph neural networks, it is possible to accurately identify the dependency relationships and risk propagation paths between equipment, suppress the chain failure problem of related equipment caused by the irreversible decline of a single equipment in the production line, realize dynamic blocking of risk propagation, minimize the scope of production stoppage and losses, and improve the stability and economy of production line operation.
[0034] Example 3: In the above examples, equipment risk association is mainly based on static topology relationships (such as process connections). This solution will introduce risk flow to simulate the transmission path of fault energy in the equipment network and achieve virtual blocking, such as... Figure 3 As shown.
[0035] S301: Obtain equipment degradation entropy value data based on equipment operating data and failure frequency, calculate risk pressure value based on equipment degradation entropy value data, identify risk flow direction, and generate risk flow diagram based on risk flow direction; Furthermore, all equipment on the production line is transformed into energy nodes. Based on collected historical operating data and equipment fault records—the historical operating data recording the equipment's operating status and performance over different time periods to understand past operating conditions and performance trends, and the equipment fault records documenting past problems, frequency, and severity—a weighted average is calculated based on the historical operating data, fault records, and expert evaluation results to obtain the degradation entropy value for each piece of equipment. A risk pressure value is then calculated based on this degradation entropy value: Risk Pressure Value = Equipment Degradation Entropy Value × Process Dependence Coefficient. The equipment degradation entropy value reflects the degree of degradation of the equipment; a higher value indicates the equipment is closer to a fault state, and the greater the risk of failure. The process dependence coefficient reflects the interrelationship and influence of equipment within the production line process. In the production line, equipment does not operate in isolation; each piece of equipment... The operating status of a device directly or indirectly affects other equipment. A higher process dependency coefficient indicates a closer process correlation between devices, meaning a failure or performance degradation in one device will have a more significant impact on others. For example, after in-depth analysis of historical spindle operating data, detailed review of fault records, and comprehensive consideration of expert evaluation, the spindle's degradation entropy value was determined to be 0.9. This indicates that the spindle has shown significant signs of degradation and a relatively high probability of failure. Furthermore, analysis of the process relationships between devices revealed a close process connection between the spindle and the gearbox. As a key component of power transmission, the spindle's operating status significantly impacts the normal operation of the gearbox. Evaluation determined that the gearbox's process dependency coefficient on the spindle is 0.7, meaning that the gearbox's operation largely depends on the stable operation of the spindle. Any failure or performance fluctuation of the spindle could have a significant impact on the gearbox. Based on the risk pressure value calculation formula, the spindle's degradation entropy value is set at 0.9. Multiplying this by the process dependence coefficient of 0.7, we get 0.9 × 0.7 = 0.63. This calculated result of 0.63 is the risk pressure value of the spindle. The spindle faces a certain level of risk pressure in the current production line environment. The higher the risk pressure value, the more severe the risk situation of the spindle, and the more attention and importance it needs to attract.
[0036] Passive RFID tags are installed in the gaps between equipment. These tags can sense subtle signal changes generated during risk transmission. For example, when equipment A malfunctions, its abnormal vibrations or electromagnetic interference will be transmitted to surrounding equipment through the gaps. The passive RFID tags can promptly detect these changes and collect risk transmission characteristics, including vibration wave frequency shift and electromagnetic interference changes. The vibration wave frequency shift reflects abnormal changes in the equipment's operating state. When equipment malfunctions or its performance deteriorates, its vibration frequency and amplitude change, resulting in a change in vibration wave frequency. The electromagnetic interference change is caused by internal electrical faults in the equipment or interference from the external electromagnetic environment. The collected risk transmission characteristics are transmitted to the data processing center. The risk flow direction is from high-voltage equipment (i.e., the fault source, equipment with a high risk pressure value) to low-voltage equipment (equipment with a low risk pressure value). A risk flow diagram is generated based on the risk transmission characteristic data, risk pressure value, and risk flow direction. Specifically, based on the physical layout and process flow of the production line equipment, the connection relationship between the equipment is determined. For each piece of equipment, the risk transmission intensity to surrounding equipment is calculated based on its risk pressure value and risk transmission characteristic data. The formula is: ,in, For the intensity of risk transmission, The risk pressure value for device i. This provides risk transmission characteristic data between device i and device j. As the weight of the risk stress value, Weights for risk transmission characteristic data. + =1, inputting risk flow information (transmission from high-pressure equipment to low-pressure equipment) and the connection relationship between equipment into the Dijkstra algorithm. The Dijkstra algorithm outputs the shortest path of risk propagation. A two-dimensional graphical interface is used to display the production line equipment on a plane in a graphical way. Each equipment is represented by a specific graphic symbol. The size and color of the graphic are initialized according to the risk pressure value of the equipment. Based on the obtained risk propagation path, lines with arrows are drawn on the risk flow diagram to indicate the direction of risk transmission. The thickness and color of the lines are set according to the intensity of risk transmission, thus generating the risk flow diagram.
[0037] S302 monitors the intensity of risk flows in real time, generates digital blocking barriers based on the intensity of risk flows, identifies risk flow loops according to algorithms, and adjusts the strategy according to the risk flow loop level; Specifically, the calculation of risk flow intensity is achieved by integrating real-time equipment operating status data collected by various pre-deployed sensors (such as vibration, temperature, and pressure sensors). This data is pre-processed through filtering and amplification by the data acquisition system before being input into the risk flow intensity calculation model for analysis. This model analyzes the connection methods between equipment based on the production line topology (such as mechanical transmission, pipeline transport, or electrical interconnection) to establish a physical relationship network reflecting the risk propagation path. Secondly, key sensor parameters (such as flow velocity / pressure difference in pipeline connections, current harmonics / temperature rise in electrical circuits) are extracted for different connection types. Multi-source data fusion technology is used to eliminate measurement noise and standardize the data. Then, combining historical fault data and mechanism analysis, statistical modeling or machine learning methods (such as Bayesian networks and random forests) are used to quantify the contribution weight of each parameter to risk transmission, constructing parameter coupling equations. Finally, dynamic correlation indicators (such as...) are calculated in real-time. The system outputs a comprehensive strength value reflecting the probability and degree of risk propagation between devices (energy transfer efficiency, abnormal fluctuation covariance), and continuously iterates and optimizes model parameters with sensor data streams. Ultimately, it outputs an index characterizing the intensity of risk propagation between devices. A risk flow intensity threshold is set based on historical operating data. When the real-time monitoring shows a risk flow intensity exceeding the set threshold, a trigger signal is issued. Based on the trigger signal, a digital blocking barrier is automatically generated. The digital twin serves as a virtual mapping of the production line equipment, synchronizing the equipment's operating status and parameters in real time. This barrier is dynamically constructed in the digital twin environment through software programming and algorithms. Its core is an embedded phase cancellation array—an anti-frequency vibration system composed of multiple independently adjustable vibration units. When the barrier is activated, the phase cancellation array analyzes the frequency and phase characteristics of the risk flow in real time and precisely releases a reverse anti-frequency vibration wave. The released anti-frequency vibration wave superimposes with the vibration wave in the risk flow, resulting in phase cancellation. By precisely controlling the frequency and phase of the anti-frequency vibration wave, the energy of the vibration wave in the risk flow is weakened or completely canceled, thereby blocking the transmission of the risk flow from the risk source equipment to healthy equipment.
[0038] The Depth-First Search (DFS) algorithm is used to monitor the risk transmission paths between devices in real time. Starting from each node, the algorithm recursively traverses the downstream paths, maintaining an access stack to record the current path. When a node is found to have been visited and exists in the current path stack, a risk loop is determined to have formed (e.g., A→B→C→A), and the total risk intensity of the loop path is calculated (by summing the weights of each edge). If the algorithm finds that a risk flow forms a loop (e.g., A→B→C→A), the loop is marked as a high-risk loop. When the algorithm detects a risk loop (e.g., D→E→F→D) but its cumulative risk intensity is lower than a preset high-risk threshold, the system marks the loop as a medium-risk loop. For high-risk loops, critical nodes in the loop are first identified through risk analysis. Critical nodes refer to devices or connection points in the loop that have a significant impact on risk transmission. Smart circuit breakers are then installed at the identified critical nodes. The smart circuit breaker has the function of automatically detecting and disconnecting the circuit, and it can monitor parameters such as current and voltage in the loop in real time. When a risk indicator in a circuit exceeds a set value, the intelligent circuit breaker automatically disconnects the circuit, cutting off the connection of the weakest device, thus breaking the risk circuit. For medium-risk circuits, the critical nodes in the circuit are first identified, and a phase offset module is inserted at the critical node. The phase offset module can change the phase of risk transmission. By adjusting the phase relationship of the risk wave, the risk undergoes a phase change during transmission, reducing the possibility of risk rebound. For example, in a vibration transmission circuit, after inserting a phase offset module, the phase of the vibration wave will change, thereby reducing the reflection and superposition of vibration in the circuit and reducing the impact of risk on the equipment.
[0039] The technical solutions described in the above embodiments of this application have at least the following technical effects or advantages: by combining technologies such as digital blocking barriers and smart circuit breakers, it is possible to more accurately block risk transmission paths, prevent risk rebound, effectively protect healthy equipment, improve the overall reliability of equipment operation, realize dynamic topology analysis of equipment risks, accurately predict and effectively block risk transmission through risk flow simulation and digital twin blocking technology, prevent risk rebound, improve the stability and reliability of equipment operation, and ensure the smooth progress of production.
[0040] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. An AI-based intelligent control method for digital factories, characterized in that, include: S101, Collect data, extract data features, calculate health index based on data features, and issue early warnings to the device based on health index. The data includes device operating data and device basic data. S102, collect the operating data of the equipment in a healthy state, establish a baseline based on the operating data, compare the real-time collected feature values with the baseline, identify the critical point of state change, and track the deterioration trajectory based on the critical point of state change; S103, Identify the degradation state based on the critical point of state change, calculate the reversibility probability based on the degradation state of the equipment, obtain the reversible range and irreversible point of equipment degradation based on the reversibility probability, and implement corresponding intervention strategies according to the reversibility situation. S104, based on the characteristics of the data, predict the time of failure occurrence, calculate the optimal intervention time based on the predicted failure occurrence time, and match the means according to the means matching rule base based on the failure characteristics.
2. The AI-based intelligent control method for digital factories as described in claim 1, characterized in that, The formula for calculating the health index is: ,in, For health index, For energy spectrum entropy, For current imbalance, Classified by voiceprint abnormalities, The slope of the pass rate The weights for the energy spectral entropy are derived based on the degree of mechanical wear and vibration signals. The weighting for current imbalance is derived from electrical faults. The weighting for voiceprint anomalies is derived based on voiceprint characteristics. The weights for the slope of the pass rate are derived based on the degree of deterioration in production quality. .
3. The AI-based intelligent control method for digital factories as described in claim 1, characterized in that, The baseline is a range that is adjusted in real time according to the equipment's operating status, environmental changes, and equipment aging factors. The data features include energy spectrum entropy, current imbalance, abnormal sound signature, and pass rate slope. The baseline range is the average value of the data features over a period of time plus or minus two standard deviations.
4. The AI-based intelligent control method for digital factories as described in claim 1, characterized in that, The state change thresholds include a first-level threshold and a second-level threshold. The first-level threshold is when the health index rises and continues to rise over a period of time, but does not reach a preset risk threshold. The second-level threshold is when the health index rises and exceeds a preset risk threshold.
5. The AI-based intelligent control method for a digital factory as described in claim 1, characterized in that, The formula for calculating the probability of reversibility is: ,in, The probability of reversibility is represented by n, where n is the number of reversible evidences and N is the total number of evidences. Reversibility is determined based on the calculated probability of reversibility, and a probability threshold range is set, which includes a maximum threshold and a minimum threshold. When the probability of reversibility is greater than the maximum threshold, the device degradation is considered to be highly reversible. When the probability of reversibility is less than or equal to the maximum threshold but greater than the minimum threshold, the device degradation is considered to be critically reversible. When the probability of reversibility is less than or equal to the minimum threshold, the device degradation is considered to be irreversible.
6. The AI-based intelligent control method for a digital factory as described in claim 1, characterized in that, S201: Obtain the technological relationship of the equipment based on the collected basic data of the equipment, calculate the dependency strength, and construct the dependency relationship of the production chain based on the dependency strength and equipment information; S202 collects historical fault data, builds a propagation model based on the historical fault data, monitors risk source nodes in real time, and generates blocking strategies based on the risk source nodes.
7. The AI-based intelligent control method for a digital factory as described in claim 6, characterized in that, Risk source nodes are classified into three risk levels: single-point high-risk, multi-node medium-risk, and full-chain high-risk. The single-point high-risk refers to a situation where there is only one risk source node, the multi-node medium-risk refers to a situation where there are multiple risk source nodes, and the full-chain high-risk refers to a situation where there are multiple risk source nodes, and these multiple risk source nodes are interconnected and interact with each other, resulting in a complete shutdown.
8. The AI-based intelligent control method for a digital factory as described in claim 1, characterized in that, S301: Obtain equipment degradation entropy value data based on equipment operating data and failure frequency; calculate risk pressure value based on equipment degradation entropy value data; calculate risk transmission intensity based on risk pressure value; identify risk flow direction; and generate risk flow diagram based on risk flow direction. S302 monitors the intensity of risk flows in real time, generates digital blocking barriers based on the intensity of risk flows, identifies risk flow loops according to algorithms, and adjusts strategies according to the risk flow loop level.
9. The AI-based intelligent control method for a digital factory as described in claim 8, characterized in that, The formula for calculating the intensity of risk transmission is: ,in, For the intensity of risk transmission, The risk pressure value for device i. This provides risk transmission characteristic data between device i and device j. As the weight of the risk stress value, Weights for risk transmission characteristic data. + =1.
10. The AI-based intelligent control method for a digital factory as described in claim 8, characterized in that, The risk flow refers to the transmission from high-voltage equipment to low-voltage equipment.
Citation Information
Patent Citations
PEMFC decline fusion prediction method suitable for dynamic working conditions
CN118052133A
IT equipment asset maintenance intelligent management method and system
CN119420662A
Monitoring and maintaining method and system for forging and pressing equipment
CN119609032A
Solar power generation integrated management system, device and method based on AI intelligence
CN120090561A
State detection method for water energy storage unit
CN120296556A
Cited By
Blood detection equipment performance evaluation method and system based on big data
CN121117644A