An industrial Internet fault diagnosis method based on digital twin and edge computing

By adopting fault diagnosis methods of edge computing and digital twin technology in the industrial Internet, fault diagnosis delay problems caused by cloud computing are solved, more efficient and real-time fault detection and decision-making are achieved, and the reliability and safety of the manufacturing process are improved.

CN116016146BActive Publication Date: 2025-06-06HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211630943.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-19
Publication Date
2025-06-06
Estimated Expiration
2042-12-19

AI Technical Summary

Technical Problem

Among the existing industrial Internet fault diagnosis methods, cloud computing is highly dependent on the network environment. When network blockage occurs, it will cause a large delay, resulting in high delay in industrial Internet fault diagnosis, affecting the reliability and security of the manufacturing process.

Method used

The industrial Internet fault diagnosis method based on edge computing and digital twin technology is adopted. By deploying edge embedded devices near sensors, data is collected and fault diagnosis is carried out in real time, and failure detection and decision-making are used in the cloud to improve the real-time and visualization of the system.

Benefits of technology

It effectively reduces the workload of cloud servers, reduces the amount of data during fault detection, reduces communication costs, solves the delay problem of industrial fault diagnosis, and improves the human-computer interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116016146B_ABST
    Figure CN116016146B_ABST
Patent Text Reader

Abstract

The present invention discloses an industrial Internet fault diagnosis method based on digital twins and edge computing, and belongs to the field of industrial Internet and fault diagnosis. First, the present invention uses the collected industrial historical data to train a machine learning model, so as to find the key variables corresponding to each fault and the machine learning model for fault diagnosis. Secondly, for a single fault, find a set of edge embedded devices that can cover all key variables of the fault, and find a minimum set of edge embedded devices that can diagnose all types of faults. Finally, data is collected in real time in the industrial Internet scenario, the machine learning model obtained by input is output, the diagnosis result is reflected in the digital twin system interface. The present invention solves the delay problem of industrial fault diagnosis, reduces communication costs, and better realizes human-computer interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of industrial Internet and fault diagnosis, and in particular relates to an industrial Internet fault detection method based on digital twins and edge computing. Background Art

[0002] Fault diagnosis is a basic requirement of the Industrial Internet of Things and is very important for manufacturing companies. As smart manufacturing becomes more automated, digital, and intelligent, people are paying more and more attention to the reliability and safety of the manufacturing process. Minor faults in the production process may cause irreparable damage.

[0003] There are many methods for fault diagnosis in the industrial Internet of Things. The most commonly used method is based on prior process knowledge. This method can be roughly divided into two categories: physical model-based and data-driven. Model-based methods require accurate process descriptions, while data-driven methods do not require understanding the principles and structures of machine equipment and can give results based solely on intermediate data. Because physical model-based methods are difficult to design and have poor results, they have been gradually eliminated in recent years.

[0004] Another type of data-driven method is mainly divided into three categories. The first type is statistical methods, including single-variable statistical methods and multivariate statistical methods. They mainly characterize and utilize the correlation between variables. They are suitable for fault detection and diagnosis of high-dimensional systems. However, they require data to be independent in time series and are not suitable for industrial Internet data, because most industrial data is sampled based on time. The second type is signal processing methods, which are mainly divided into spectral analysis and wavelet transform methods. By using a large amount of rich expert experience and knowledge and state information during system operation for analysis and processing, a comprehensive evaluation of system operation status and fault status is obtained. However, this type of method has high data requirements and is complex and time-consuming to process, so it is not suitable for industrial Internet. The third type of method is based on quantitative machine learning methods, which mainly use machine learning such as artificial intelligence technology, combined with the powerful computing power of cloud computing, and can simulate real systems and give relatively accurate prediction results. This type of method can be well applied to industrial Internet scenarios. However, machine learning requires a lot of computing resources for data processing and model training. The mainstream method is to upload all data to cloud servers for task processing, including historical data and real-time data. However, cloud computing is highly dependent on the network environment. When network congestion occurs, it will cause large delays. In the industrial Internet, such problems will have very serious consequences.

[0005] The use of edge computing architecture can effectively solve the above problems. By performing necessary computing tasks near the sensors of the machine equipment that generates data, the workload of the cloud server can be effectively reduced.

[0006] At the same time, with the rapid development and application of information technology, digital twin technology has received more and more attention in various fields such as product design and manufacturing, medical analysis, engineering construction, process optimization, and workshop scheduling. In the field of intelligent manufacturing, digital twin technology has also played a pivotal role and has become a powerful weapon to promote the development of intelligent manufacturing. It completes the mapping in the virtual space and reflects the entire life cycle of the corresponding physical equipment. Digital twin technology can not only reduce design and maintenance costs, but also improve manufacturing efficiency and quality. Summary of the invention

[0007] In order to overcome the shortcomings of the prior art, the present invention provides an industrial Internet fault diagnosis method based on edge computing and digital twins to solve the problem of high processing mode delay in traditional data centers and cloud centers, while achieving overhead optimization and improving the overall real-time performance of the system. The digital twin technology is combined to make fault diagnosis more specific. A digital twin model is deployed in the cloud and combined with fault detection, which can better provide stronger protection and more effective decision-making for the production and manufacturing process of the smart factory, and has a higher visualization for fault diagnosis, enhancing the human-computer interaction experience.

[0008] An industrial Internet fault diagnosis method based on digital twin and edge computing includes the following steps:

[0009] S1: Extract key variables: Use the collected industrial historical data to train a machine learning model, and use the characteristics of the machine learning model to find the characteristic variables with high importance, so as to achieve the purpose of variable screening, thereby finding the key variables corresponding to each fault and obtaining a machine learning model for fault diagnosis.

[0010] S2: Find a set of edge embedded devices that cover a single fault: Since edge embedded devices can only collect data on a few characteristic variables, it is necessary to find a set of edge embedded devices that can cover all key variables of a single fault.

[0011] S3: Find the minimum set of edge embedded devices that can cover all faults: After obtaining all sets of edge embedded devices that cover a single fault, combine them together to form a list to find a minimum set of edge embedded devices that can diagnose all types of faults using only some edge embedded devices.

[0012] S4: Fault diagnosis and interaction: Based on the minimum set of edge embedded devices obtained in S3, real-time data is collected in the industrial Internet scenario, input into the machine learning model obtained in S1, output the fault diagnosis results, and reflect the diagnosis results on the digital twin system interface.

[0013] The step S1 of extracting key variables comprises the following steps:

[0014] S1.1: After screening and cleaning the industrial historical data set, all fault-related data, i.e., feature variables, are input into the machine learning model for training, and the fault diagnosis results and the weights of the corresponding feature variables are output.

[0015] S1.2: Sort all characteristic variables in each type of fault from high to low according to their weight values, and retain the first n characteristic variables.

[0016] S1.3: Count the number of times each feature variable is retained in each type of fault, select the m feature variables with the highest number of occurrences as a new training set, retrain a machine learning model, repeat steps S1.1 and S1.2, the retained feature variables are the key variables, and save the machine learning model obtained through training, so that the data requirements for fault detection are smaller and detection can be performed faster.

[0017] In the step S2, the set of all edge embedded devices that can provide key variables is found for a single fault. Based on the existing set of edge embedded devices, a search tree is constructed, in which each node represents a device. Assuming a single fault f1, starting from the root node, all nodes on the path are collected according to the deep traversal method, and these nodes are combined into a new set. Calculate the perceived quality value of the new set: if the perceived quality value is greater than or equal to the average threshold of the corresponding fault, this set is retained, and returns to the previous node to continue traversing the set of edge embedded devices. After the traversal is completed, the retained set corresponds to the set of all edge embedded devices that can provide key variables for fault f1.

[0018] In the step S3, the number of times the edge embedded device appears in each fault corresponding set is recorded, and the edge embedded device number with the least number of appearances is found. If the edge embedded device appears in more than two sets, it is determined whether deleting the edge embedded device will cause the set in the corresponding list to be missing. If so, the edge embedded device is retained, otherwise the edge embedded device is deleted, and the device with the second least number of appearances is found instead. Repeat the above operation until the edge embedded device traversal is completed, and the minimum edge embedded device set that can cover all faults is found.

[0019] In step S4, the digital twin system includes a service layer, a model layer, and a data layer. After the fault detection result comes out, it is transmitted to the data layer of the digital twin system by a wired or wireless network. After receiving the data, the data layer packages the detection result data and sends it to the model layer. The model layer matches the detection result one by one through key variables and sends the result to the service layer. The result is reflected in the fault visualization management system and monitoring platform of the service layer, which is convenient for operators to check at any time and understand the overall status of the machine equipment.

[0020] Compared with the prior art, the present invention has the following beneficial effects: the present invention uses edge embedded devices deployed near sensors to collect data. Then, the trained model is used to find multiple key features related to a single fault, and then the minimum set of edge embedded devices covering all key variables is found, and the edge embedded devices are selected to collect necessary data for processing and upload for diagnosis, and finally the detection results are reflected in the digital twin system of the machine equipment. Compared with the prior art, the present invention adopts an edge computing architecture, which can effectively reduce the workload of the cloud server by performing necessary computing tasks near the sensors of the machine equipment that generates data, and reduces the amount of data required for fault detection, solves the delay problem of industrial fault diagnosis, reduces communication costs, completes the fault result mapping of the machine equipment in the digital twin system, and reflects it on the display platform, which is convenient for operators to view and better realizes human-computer interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic diagram of the process of the present invention;

[0022] Figure 2 A flow chart for finding the set of edge embedded devices corresponding to the fault;

[0023] Figure 3 Schematic diagram of the search tree in the algorithm for finding the edge embedded device set corresponding to the fault;

[0024] Figure 4 Flowchart for finding the minimum covering set;

[0025] Figure 5 This is the architecture diagram of the digital twin platform;

[0026] Figure 6 Schematic diagram of the service layer management system and monitoring platform interface Figure 1 ;

[0027] Figure 7 Schematic diagram of the service layer management system and monitoring platform interface Figure 2 . DETAILED DESCRIPTION

[0028] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0029] The embodiment tests it on a data set of actual industrial Internet machines and equipment, including normal state and 21 different faults. Each state is divided into a training set and a test set, with 660 sets of data and 1510 sets of data respectively. The entire plant consists of five main operating units: a reactor, a condenser, a circulating compressor, a gas-liquid separator, and a product analysis tower. And data monitoring is performed for aspects such as whether the inlet is blocked, the speed of adding inlet materials, machine pressure, temperature, discharge flow, and control variables, with a total of 49 variables. Each data sampling interval is set to 3 minutes. 660 sets of data samples are generated after the entire chemical process runs for 54 hours. Among them, a fault is introduced after 9 hours of stable operation. There will be 110 sets of data samples in a fault sample without fault, and the remaining 550 sets of data samples are the state of generating faults. There are 13 edge embedded devices in the set experimental scenario, and the machine equipment has 21 kinds of faults and 52 characteristic variables. These 13 edge embedded devices can each be connected to several fixed sensors. Device a can collect variables 31, 29, 22, 23, 30, and 32, and device b can collect variables 1, 4, 2, 3, 5, and 49, as shown in Table 1.

[0030] Table 1

[0031]

[0032] like Figure 1 As shown, an industrial Internet fault diagnosis method based on digital twins and edge computing mainly includes an edge layer, and the method includes the following steps:

[0033] S1: Extract key variables: Use the collected industrial historical data to train a machine learning model, and use the characteristics of the machine learning model to find the characteristic variables with high importance, so as to achieve the purpose of variable screening, thereby finding the key variables corresponding to each fault and obtaining a machine learning model for fault diagnosis.

[0034] S2: Find a set of edge embedded devices that cover a single fault: Since edge embedded devices can only collect data on a few characteristic variables, it is necessary to find a set of edge embedded devices that can cover all key variables of a single fault.

[0035] S3: Find the minimum set of edge embedded devices that can cover all faults: After obtaining all sets of edge embedded devices that cover a single fault, combine them together to form a list to find a minimum set of edge embedded devices that can diagnose all types of faults using only some edge embedded devices.

[0036] S4: Fault diagnosis and interaction: Based on the minimum set of edge embedded devices obtained in S3, data is collected in real time in the industrial Internet scenario, input into the machine learning model obtained in S1, output the fault diagnosis results, and reflect the diagnosis results in the digital twin system interface for operators to view and analyze.

[0037] In step S1, after screening and cleaning the industrial historical data set, all data related to the fault, that is, feature variables, are brought into the random forest model for training. After the training is completed, the results of fault diagnosis and the weights of the corresponding feature variables are output. All feature variables are sorted according to the value of the weight. The higher the value, the more suitable the variable is for the diagnosis of this type of fault. The higher the importance of the fault, the more the feature variable should be retained. According to the specific circumstances of the experiment, several fixed key variables are screened out, and these variables are used as new training sets to retrain a random forest model. After such operations, the key variables are successfully screened out, and the model is updated and saved, so that the data requirements for fault detection are smaller and detection can be performed faster.

[0038] In step S2, all edge embedded devices that can provide characteristic variables are found for a single fault. The specific steps are to set a fault threshold T, which is the average value of the key variables of each fault. If the edge embedded device u 1 Can collect two variables v 1 and v 2 , another edge embedded device u 2 Able to collect variable v 2 and v 3 If there is a fault f 1 The key variable is v 1 and v 3 , then we need a 1 and u 2 The edge set of can complete the detection of the fault. Then for f 1 This fault 1 The perceived quality at the edge is q 11 ,u 2 The perceived quality at the edge is q 12 , the perceived quality refers to the number of key variables required to detect a certain fault when an edge embedded device detects the fault. After deduplication, if the total perceived quality is greater than or equal to the fault threshold T, then it can be said that u 1 and u 2 The edge segments combined together are able to diagnose the fault f1. In order to quantify the above process into a formula, assume that there is a fault f t , and there are edge embedded devices u v ,uz ,u x ,…,u c , the judgment process can be expressed as follows:

[0039] q vt +q zt +q xt +…+q ct ≥T

[0040] where q vt For fault f t , through edge embedded devices u v The perceived quality of the acquisition variable, q zt ,q xt ,q ct Same reason.

[0041] Find collections such as Figure 2 As shown in the figure, the specific process is as follows: Based on the existing set of edge embedded devices, a search tree is constructed, such as Figure 3 As shown in the figure, each node represents an edge embedded device. If there is a device set {D1, D2, D3}, all child nodes are different from the parent node or ancestor node, and the child nodes of the parent node are all nodes except the ancestor node. Assuming there is a fault f1, starting from the root node, according to the method of deep traversal, collect all nodes on the path, and combine these nodes into a new set. Assuming that the set {D1, D2} is obtained, because it is greater than or equal to the threshold of the corresponding fault, there is no need to traverse downwards, and return directly in advance. Calculate the collection quality value of the new set: If the collection quality value is greater than or equal to the threshold of the corresponding fault, this set is retained, and return to the previous node to continue traversing the set of edge embedded devices. After the traversal, the set obtained is {{D1, D2, D3}, {D1, D3}, {D3, D1}}, which corresponds to the set of all edge embedded devices that can provide key variables for the single fault f1.

[0042] In step S3, a suitable subset is selected from the edge embedded device set that can detect each type of fault, and these subsets form a theoretically minimum edge embedded device set. All faults can be diagnosed using only the edge embedded devices in this set, which can greatly reduce the communication and computing costs of fault detection.

[0043] The steps to find the minimum device coverage set are as follows: Figure 4As shown, record the number of times the edge embedded device appears in each fault corresponding set, find the edge embedded device number with the least number of appearances, if the edge embedded device appears in more than two sets, determine whether deleting the edge embedded device will cause the set in the corresponding list to be missing, if so, keep the edge embedded device, otherwise delete the edge embedded device; instead, find the device with the second least number of appearances and repeat the above operation until the device traversal is completed, and the minimum coverage set can be obtained. The final result is shown in Table 2:

[0044] Table 2

[0045]

[0046] Table 3

[0047]

[0048] In step S4, according to the minimum set of edge embedded devices obtained in S3, real-time data is collected in the industrial Internet scenario, the random forest model saved in S1 is input, and the diagnosis results are output. The fault detection rate results are shown in Table 3. The digital twin system includes a service layer, a model layer, and a data layer. The service layer includes a visual fault management system and a monitoring platform, such as Figure 5 After the fault detection and processing results come out, they are transmitted to the digital twin system data layer by wired or wireless network. After receiving the data, the data layer packages the detection result data and sends it to the model layer. The model layer matches the detection results one by one through key variables and sends them to the service layer. The results are reflected in the management system and monitoring platform of the service layer, which is convenient for operators to check at any time and understand the overall status of the machine equipment. Figure 6 and Figure 7 shown.

[0049] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Technical personnel in the relevant field should understand that on the basis of the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.

Claims

1. An industrial Internet fault diagnosis method based on digital twins and edge computing, It is characterized in that The following steps are involved: S1: Extract key variables: Use the collected industrial historical data to train a machine learning model, find the key variables corresponding to each fault, and obtain a machine learning model for fault diagnosis; the details are as follows: S1.1: After screening and cleaning the industrial historical data set, all fault-related data, i.e., feature variables, are input into the machine learning model for training, and the fault diagnosis results and the weights of the corresponding feature variables are output; S1.2, sort all characteristic variables in each type of fault from high to low according to their weight values, and retain the first n characteristic variables; S1.3: Count the number of times each feature variable is retained in each type of fault, select the m feature variables with the highest number of occurrences as the new training set, retrain a machine learning model, repeat steps S1.1 and S1.2, the retained feature variables are the key variables, and save the machine learning model obtained by the training; S2: Find a set of edge embedded devices that cover a single fault: For a single fault, find a set of edge embedded devices that can cover all key variables of the fault; the details are as follows: According to the existing set of edge embedded devices, a search tree is constructed, in which each node represents an edge embedded device; For a single fault f1, starting from the root node, all nodes on the path are collected according to the deep traversal method, and these nodes are combined into a new set. The perceived quality value of the new set is calculated: if the perceived quality value is greater than or equal to the average threshold of the corresponding fault, this set is retained and the traversal is continued back to the previous node. After the traversal is completed, the retained set corresponds to the set of all edge embedded devices that can provide its key variables for the fault f1. In the search tree, all child nodes are not identical to the parent node or ancestor node, and the parent node has child nodes other than the ancestor node; The average threshold is the average value of the key variables of each fault; S3: Find the minimum set of edge embedded devices that can cover all faults: Combine the sets of edge embedded devices that cover all single faults to form a list, and find a minimum set of edge embedded devices that can diagnose all types of faults using only some edge embedded devices; The specific process of finding a minimum set of edge embedded devices that can diagnose all types of faults using only some edge embedded devices is as follows: S3.

1. Find a method that can diagnose all types of faults using only some edge embedded devices, specifically: Record the number of times the edge embedded device appears in each fault corresponding set, and find the edge embedded device with the least number of appearances; If the edge embedded device appears in more than two sets, determine whether deleting the edge embedded device will cause the set in the corresponding list to be missing, if so, retain the edge embedded device, otherwise delete the edge embedded device; S3.

2. Find the edge embedded device with the second smallest number of occurrences and repeat S3.1 until all edge embedded devices are traversed and the minimum set of edge embedded devices that can cover all faults is obtained; S4: Fault diagnosis and interaction: Based on the minimum set of edge embedded devices obtained in S3, real-time data is collected in the industrial Internet scenario, input into the machine learning model obtained in S1, output the fault diagnosis results, and reflect the diagnosis results on the digital twin system interface; The digital twin system includes a service layer, a model layer and a data layer; The method of reflecting the diagnosis results on the digital twin system interface includes: transmitting to the digital twin system data layer by a wired or wireless network, and after receiving the data, the data layer packages the detection result data and sends it to the model layer, and the model layer corresponds the detection results one-to-one through key variables and sends them to the service layer, which is reflected in the fault visualization management system and monitoring platform of the service layer.

Citation Information

Patent Citations

  • Intelligent production system and method based on edge calculation and digital twinning

    CN111857065A

  • Federal learning method and system based on edge digital twin association

    CN113419857A