Production equipment fault cloud edge collaborative diagnosis method and system based on transfer learning

CN118277828BActive Publication Date: 2026-09-29CHONGQING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410380776.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-31
Publication Date
2026-09-29
Estimated Expiration
2044-03-31

AI Technical Summary

Technical Problem

然而,目前基于云边协同的故障诊断方法还存在以下问题:1)将故障诊断模型部署到云平台上时,由于云平台和边缘平台巨大的数据传输和可能存在的网络交互延迟,将直接导致故障诊断的延迟高、实时性不好

Benefits of technology

[0063]本发明通过设备平台获取待测生产设备上关键部件的实时采集(传感器)数据,然后通过边缘平台进行预处理和工况判断并将结果上传至云平台,最终接收云平台下发的故障诊断模型并进行故障诊断,进而输出对应的故障诊断结果。其中边缘平台设置于实际生产环境附近,在边缘平台部署故障诊断模型有效解决了传统云端范式框架的高延迟响应问题,能够提高整个系统的灵活性和可扩展性,从而提高生产设备关键部件故障诊断的实时性。同时,本发明在边缘平台部署的是经过云端预先训练好或迁移学习得到的故障诊断模型,其能够在保留模型诊断效果的情况下,降低对计算资源和数据资源的要求,拥有更快的收敛速度和节约更多的个性化训练时间,能够进一步提高生产设备关键部件状态诊断的准确率和实时性,特别是在网络拥堵和个性化训练样本不足的情况下,本发明的诊断效果更加显著。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118277828B_ABST
    Figure CN118277828B_ABST
Patent Text Reader

Abstract

The application discloses a kind of production equipment fault cloud edge collaborative diagnosis method and system based on transfer learning.The method includes: the real-time running data after pre-processing is judged to working condition: if known historical working condition, the information of known historical working condition and real-time running data are packed into same working condition task;If it is unknown working condition, the information of several similar historical working conditions closest to the unknown working condition and real-time running data are packed into different working condition task;Cloud platform: if same working condition task is accepted, the trained fault diagnosis model is issued to edge platform;If different working condition task is accepted, pre-training fault diagnosis model is trained by transfer learning, and the trained fault diagnosis model is issued to edge platform to realize fault diagnosis.The trained fault diagnosis model is deployed in edge platform to realize more efficient fault diagnosis, and the fault diagnosis model corresponding to each working condition is matched to realize individualized fault diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of the Internet and big data, specifically to a cloud-edge collaborative diagnosis method and system for production equipment faults based on transfer learning. Background Technology

[0002] Driven by next-generation information and computing technologies such as artificial intelligence, big data, cloud computing, and edge computing, intelligent manufacturing has been integrated into all aspects of the manufacturing industry. Big data provides a rich knowledge base for intelligent manufacturing, while cloud computing and edge computing provide a powerful computing foundation. Through real-time calculation and analysis of data, real-time allocation and scheduling of manufacturing resources are achieved, thereby realizing an intelligent and automated manufacturing process.

[0003] In real-world manufacturing scenarios, critical components such as rolling bearings are prone to various defects (such as wear, spalling, and corrosion) under complex environments involving high speed, heavy loads, and prolonged impacts. This can lead to decreased equipment performance and even safety accidents. Therefore, to avoid these issues and reduce maintenance costs, it is essential to employ accurate methods for identifying operational faults in critical components of production equipment. This is crucial for stable equipment operation and efficient equipment maintenance.

[0004] In modern manufacturing, the data generated by manufacturing systems is experiencing explosive growth, originating and being collected from every stage of industrial manufacturing. Systematic computational analysis of manufacturing data to form more informed decisions improves the effectiveness of intelligent manufacturing. In other words, data-driven manufacturing can be seen as a necessary condition for intelligent manufacturing. In intelligent manufacturing systems, monitoring equipment generates a large amount of industrial data during the production process, including signals used for equipment fault diagnosis. With the development of big data and artificial intelligence algorithms, data-driven intelligent fault diagnosis methods have been extensively researched. Deep learning-based methods, with their powerful advantage of automatically extracting fault feature information from raw signals, have been widely applied in equipment fault diagnosis.

[0005] Cloud computing, as a scalable computing platform for processing and analyzing large-scale historical data, provides ample computing and storage resources for data-driven prediction algorithms. To fully leverage the data processing advantages of cloud computing and the rapid response advantages of edge computing, cloud-edge collaboration has become a new research hotspot in intelligent manufacturing. However, current fault diagnosis methods based on cloud-edge collaboration still suffer from the following problems: 1) When deploying fault diagnosis models to cloud platforms, the massive data transmission and potential network interaction delays between the cloud and edge platforms directly lead to high latency and poor real-time performance in fault diagnosis. 2) Fault signals under various operating conditions often exhibit different amplitude characteristics, time-domain characteristics, or frequency-domain characteristics due to differences in fault type, equipment status, and working environment. This makes it difficult for general models trained on fault datasets from all or a specific historical operating condition to effectively cope with various operating conditions, and even results in low diagnostic accuracy under most conditions. Therefore, improving the real-time performance and accuracy of cloud-edge collaborative fault diagnosis for production equipment is a pressing technical problem that needs to be solved. Summary of the Invention

[0006] To address the shortcomings of the existing technologies, the technical problem to be solved by this invention is: how to provide a cloud-edge collaborative diagnosis method and system for production equipment faults based on transfer learning, which deploys the fault diagnosis model trained on the cloud platform on the edge platform to achieve more efficient fault diagnosis, and at the same time can match the corresponding fault diagnosis model for fault signals under various working conditions to achieve personalized fault diagnosis, thereby improving the real-time performance and accuracy of cloud-edge collaborative diagnosis of production equipment faults.

[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0008] A cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning includes:

[0009] S1: Obtain real-time operating data of the corresponding part of the production equipment under test through the equipment platform and upload it to the edge platform;

[0010] S2: The edge platform receives real-time running data and preprocesses it, using the preprocessed data as the feature information to be tested;

[0011] S3: The edge platform performs operating condition judgment on the feature information to be tested: if it is a known historical operating condition, the information of the known historical operating condition and the real-time running data are packaged into a same operating condition task and uploaded to the cloud platform; if it is an unknown operating condition, the information of several similar historical operating conditions closest to the unknown operating condition and the real-time running data are packaged into a different operating condition task and uploaded to the cloud platform.

[0012] S4: The cloud platform receives tasks under the same or different operating conditions uploaded by the edge platform. If a task under the same operating condition is received, the fault diagnosis model trained based on the fault dataset of the corresponding known historical operating conditions is sent to the edge platform. If a task under different operating conditions is received, the fault datasets of each similar historical operating condition and the real-time running data are used together as training data. The pre-trained fault diagnosis model trained from the fault datasets of all historical operating conditions is trained by transfer learning, and finally the trained fault diagnosis model is sent to the edge platform.

[0013] S5: The edge platform receives the fault diagnosis model sent out and uses the feature information to be tested as the model input, and outputs the corresponding fault diagnosis result through the fault diagnosis model.

[0014] Preferably, in step S2, the preprocessing of real-time running data includes data cleaning and data transformation;

[0015] Data cleaning includes deleting or completing missing and abnormal data, as well as deleting redundant data;

[0016] Data transformation refers to the process of converting one-dimensional real-time running data into a two-dimensional signal time-frequency diagram through wavelet transform, which can then be used as the feature information to be measured.

[0017] Preferably, in step S3, the edge platform is deployed with a working condition discrimination model library, which stores working condition discrimination models trained from fault datasets for each historical working condition;

[0018] When determining the operating condition of the feature information to be tested: First, the feature information to be tested is input into the operating condition discrimination model corresponding to each historical operating condition, and the corresponding reconstructed data is output by each operating condition discrimination model; then, the reconstruction error between the feature information to be tested and the reconstructed data output by each operating condition discrimination model is calculated; it is determined whether there is a reconstruction error less than or equal to the error threshold: if so, the historical operating condition to which the operating condition discrimination model corresponding to the reconstructed data of the reconstruction error belongs is taken as the known historical operating condition of the current feature information to be tested; otherwise, the operating condition of the current feature information to be tested is an unknown operating condition.

[0019] Preferably, in step S3, the working condition discrimination model is the LAE model, which is based on the VAE network model but removes the intermediate layer to map the features to a distribution. That is, the features generated by the encoder in the VAE network model are directly input into the decoder to generate the corresponding reconstructed data.

[0020] Preferably, in step S4, the fault diagnosis model includes a feature extraction module, a domain fusion module, and a classification module; wherein the domain fusion module is a variational autoencoder.

[0021] 1) During fault diagnosis, the working logic of the fault diagnosis model is as follows:

[0022] Use the feature information to be tested as input to the fault diagnosis model;

[0023] The feature extraction module performs multi-scale feature extraction on the feature information to be tested, generating fault features containing multi-scale information.

[0024] The encoder in the domain fusion module extracts features from the fault characteristics and generates potential features;

[0025] The intermediate layer of the domain fusion module transforms latent features into a feature distribution and samples from the feature distribution to generate sampled features;

[0026] The classification module classifies based on the sampled features and generates corresponding fault diagnosis results;

[0027] 2) During transfer learning training, the working logic of the fault diagnosis model is as follows:

[0028] Use source domain data and target domain data as input to the fault diagnosis model;

[0029] The feature extraction module performs multi-scale feature extraction on the source domain data and the target domain data respectively, generating fault features containing multi-scale information;

[0030] The encoder of the domain fusion module learns from the fault characteristics of the source domain data and the target domain data to generate domain-invariant features of the source domain and the target domain.

[0031] The intermediate layer of the domain fusion module transforms the domain-invariant features into feature distributions for each operating condition, and samples from the feature distributions for each operating condition to generate sampled features for each operating condition.

[0032] The decoder in the domain fusion module generates corresponding reconstructed data based on the sampling features under various operating conditions;

[0033] The classification module classifies samples based on their characteristics under various operating conditions and generates fault diagnosis results for each condition.

[0034] Preferably, the feature extraction module includes two first depthwise convolutional layers, a first pooling layer, a second depthwise convolutional layer, a second pooling layer, two third depthwise convolutional layers, and a third pooling layer connected end to end in sequence; wherein the input of the first first convolutional layer is the input of the fault diagnosis model, and the output of the second third depthwise convolutional layer is the input of the domain fusion module;

[0035] The first deep convolutional layer includes three branches, the inputs of which are all inputs to the first deep convolutional layer, and the outputs of the three branches are concatenated to serve as the output of the first deep convolutional layer. The first branch includes a 1×1 basic convolutional layer, the second branch includes a 1×1 basic convolutional layer and a 3×3 basic convolutional layer connected end to end in sequence, and the third branch includes a 1×1 basic convolutional layer and a 5×5 basic convolutional layer connected end to end in sequence.

[0036] The first pooling layer includes two branches, the input of which is the output of the second first convolutional layer. The outputs of the two branches are connected and used as the input of the 2×2 max pooling layer, and the output of the 2×2 max pooling layer is used as the output of the first pooling layer. The first branch includes a 1×1 basic convolutional layer, and the second branch includes a 1×1 basic convolutional layer and a 3×3 basic convolutional layer connected end to end.

[0037] The second deep convolutional layer includes three branches. The inputs of the three branches are the outputs of the first pooling layer. The outputs of the three branches are concatenated to serve as the output of the second deep convolutional layer. The first branch includes a 1×1 basic convolutional layer, the second branch includes a 1×1 basic convolutional layer and a 5×5 basic convolutional layer connected end to end in sequence, and the third branch includes a 1×1 basic convolutional layer and a 7×7 basic convolutional layer connected end to end in sequence.

[0038] The second pooling layer includes two branches, the input of which is the output of the second depth convolutional layer. The outputs of the two branches are connected and used as the input of the 2×2 max pooling layer, and the output of the 2×2 max pooling layer is used as the output of the second pooling layer. The first branch includes a 1×1 basic convolutional layer, and the second branch includes a 1×1 basic convolutional layer and a 5×5 basic convolutional layer connected end to end.

[0039] The third deep convolutional layer includes three branches, the input of which is the output of the second pooling layer. The outputs of the three branches are concatenated to serve as the output of the third deep convolutional layer. The first branch includes a 1×1 basic convolutional layer, the second branch includes a 1×1 basic convolutional layer and a 7×7 basic convolutional layer connected end to end, and the third branch includes a 1×1 basic convolutional layer and a 9×9 basic convolutional layer connected end to end.

[0040] The structure of the third pooling layer is the same as that of the second pooling layer;

[0041] The x×x basic convolutional layer consists of x×x convolutional layers, batch normalization layers, and activation function layers connected end to end in sequence.

[0042] Preferably, the domain fusion module includes an encoder and a decoder, wherein the output of the encoder is used as a domain-invariant feature;

[0043] The encoder contains four 3×3 convolutional layers connected end-to-end in sequence; the decoder contains three 3×3 deconvolutional layers, a 5×5 deconvolutional layer, and a 3×3 deconvolutional layer connected end-to-end in sequence.

[0044] Preferably, the classification module includes two fully connected layers connected first and second in sequence and a softmax classifier, wherein neurons between the two fully connected layers are randomly frozen by a Dropout layer to prevent overfitting.

[0045] Preferably, in step S4, the pre-trained fault diagnosis model is trained through transfer learning through the following steps:

[0046] S401: Use the fault dataset with similar historical operating conditions as the source domain data and the real-time operating data as the target domain data;

[0047] S402: After preprocessing the source domain data and target domain data, they are used as input to the pre-trained fault diagnosis model;

[0048] S403: Initialize the weight values ​​of each layer of the pre-trained fault diagnosis model using a normal distribution;

[0049] S404: First, the feature extraction module extracts features from the preprocessed source and target domain data to generate fault features. Then, the encoder of the domain fusion module learns the domain-invariant features of the source and target domains. Next, the intermediate layer of the domain fusion module transforms the domain-invariant features into feature distributions for each operating condition, and uses KL divergence to align the feature distributions with a 0-1 normal distribution. Sampling is then performed on the feature distributions to generate sampled features for various operating conditions. Furthermore, the decoder of the domain fusion module generates corresponding reconstructed data based on the sampled features for various operating conditions. Finally, the encoder and decoder of the feature extraction module and the domain fusion module are trained by minimizing the reconstruction error between the source and target domain data and the corresponding reconstructed data.

[0050] S405: The classification module classifies the samples based on the sampling features under various working conditions, generates fault diagnosis results for various working conditions, and finally optimizes the parameters of the classification module through the loss function of the classification module.

[0051] S406: Repeat steps S404 to S405 until the model converges, and the trained fault diagnosis model is obtained.

[0052] This invention also discloses a cloud-edge collaborative diagnosis system for production equipment faults based on transfer learning. The implementation of the cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning in this invention includes:

[0053] The equipment platform is used to acquire real-time operating data of the corresponding parts of the production equipment under test;

[0054] The edge platform receives and preprocesses real-time operating data uploaded from the device platform, using the preprocessed data as the feature information to be tested. Then, it performs operating condition judgment on the feature information to be tested: if it is a known historical operating condition, the information of the known historical operating condition and the real-time operating data are packaged into a same operating condition task and uploaded to the cloud platform; if it is an unknown operating condition, the information of several similar historical operating conditions closest to the unknown operating condition and the real-time operating data are packaged into a different operating condition task and uploaded to the cloud platform.

[0055] The cloud platform is used to receive tasks under the same or different operating conditions uploaded by the edge platform. If a task under the same operating condition is received, the fault diagnosis model trained based on the fault dataset of the corresponding known historical operating conditions will be sent to the edge platform. If a task under different operating conditions is received, the fault datasets of each similar historical operating condition and the real-time running data will be used as training data. The pre-trained fault diagnosis model trained from the fault datasets of all historical operating conditions will be trained by transfer learning, and finally the trained fault diagnosis model will be sent to the edge platform.

[0056] The edge platform is also used to receive the fault diagnosis model and take the feature information to be tested as the model input, and output the corresponding fault diagnosis result through the fault diagnosis model.

[0057] This invention also discloses a cloud-edge collaborative diagnosis system for production equipment faults based on transfer learning. The implementation of the cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning according to this invention includes:

[0058] The equipment platform is used to acquire real-time operating data of the corresponding parts of the production equipment under test;

[0059] The edge platform receives and preprocesses real-time operating data uploaded from the device platform, using the preprocessed data as the feature information to be tested. Then, it performs operating condition judgment on the feature information to be tested: if it is a known historical operating condition, the information of the known historical operating condition and the real-time operating data are packaged into a same operating condition task and uploaded to the cloud platform; if it is an unknown operating condition, the information of several similar historical operating conditions closest to the unknown operating condition and the real-time operating data are packaged into a different operating condition task and uploaded to the cloud platform.

[0060] The cloud platform is used to receive tasks under the same or different operating conditions uploaded by the edge platform. If a task under the same operating condition is received, the fault diagnosis model trained based on the fault dataset of the corresponding known historical operating conditions will be sent to the edge platform. If a task under different operating conditions is received, the fault datasets of each similar historical operating condition and the real-time running data will be used as training data. The pre-trained fault diagnosis model trained from the fault datasets of all historical operating conditions will be trained by transfer learning, and finally the trained fault diagnosis model will be sent to the edge platform.

[0061] The edge platform is also used to receive the fault diagnosis model and take the feature information to be tested as the model input, and output the corresponding fault diagnosis result through the fault diagnosis model.

[0062] Compared with existing technologies, the cloud-edge collaborative diagnosis method and system for production equipment faults based on transfer learning in this invention has the following advantages:

[0063] This invention acquires real-time (sensor) data from key components of the production equipment under test through a device platform. This data is then preprocessed and assessed for operational status via an edge platform, and the results are uploaded to a cloud platform. Finally, the cloud platform receives a fault diagnosis model and performs fault diagnosis, outputting the corresponding results. The edge platform is located near the actual production environment. Deploying the fault diagnosis model on the edge platform effectively solves the high latency response problem of traditional cloud-based paradigms, improving the flexibility and scalability of the entire system and thus enhancing the real-time performance of fault diagnosis for key components of the production equipment. Furthermore, the fault diagnosis model deployed on the edge platform is pre-trained or obtained through transfer learning in the cloud. This model reduces the requirements for computing and data resources while preserving the diagnostic effectiveness of the model, resulting in faster convergence and less time for personalized training. This further improves the accuracy and real-time performance of diagnosing the condition of key components of the production equipment, especially under conditions of network congestion and insufficient personalized training samples.

[0064] This invention uses an edge platform to determine the operating conditions of the feature information under test and uploads the results (for tasks with the same or different operating conditions) to a cloud platform. The cloud platform then distributes existing fault diagnosis models for known operating conditions and trains transfer learning models for unknown operating conditions for both tasks with the same and different operating conditions. This allows for the matching of personalized fault diagnosis models to fault signals under different operating conditions. Firstly, the characteristics of fault signals differ under different operating conditions. By matching corresponding diagnostic models, the fault characteristics of the fault signals can be captured more accurately, and the fault type and location can be identified more precisely. This not only reduces false alarms and missed alarms but also shortens fault diagnosis time and improves overall work efficiency. Secondly, matching sensor data for different operating conditions makes the fault diagnosis model more comprehensive and flexible, thus adapting to more practical application scenarios and improving the generalization ability of the fault diagnosis model, enabling it to maintain high diagnostic accuracy even when facing unknown or complex operating conditions. Finally, for unknown operating conditions, this invention trains a pre-trained fault diagnosis model using transfer learning. On one hand, transfer learning allows the fault diagnosis model to better generalize to new tasks. There may be shared characteristics or patterns between the source and target domains, and transfer learning can capture these characteristics, making the model perform better on new tasks. On the other hand, transfer learning utilizes a pre-trained fault diagnosis model trained on fault datasets from all historical operating conditions. Therefore, the computational resources required for unknown operating conditions can be significantly reduced, thereby further improving the efficiency of diagnosing the condition of key components in production equipment. Attached Figure Description

[0065] To make the objectives, technical solutions, and advantages of the invention clearer, the invention will now be described in further detail with reference to the accompanying drawings, wherein:

[0066] Figure 1 Here is a logical block diagram of a cloud-edge collaborative diagnostic method and system for production equipment faults based on transfer learning;

[0067] Figure 2 A flowchart for performing operational condition assessment on an edge platform;

[0068] Figure 3 This is a network structure diagram of the fault diagnosis model;

[0069] Figure 4 Network structure diagram for Domain Fusion Module (VAE);

[0070] Figure 5 To migrate fault diagnosis results;

[0071] Figure 6 Visualizing the results of t-sne

[0072] Figure 7 Transfer analysis of small sample labeled data

[0073] Figure 8 The comparison results are with and without a pre-trained model. Detailed Implementation

[0074] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but only to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0075] It should be noted that similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the figures, or the orientation or positional relationship commonly used when the product is in use. They are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance. In addition, the terms "horizontal," "vertical," etc., do not mean that the component is required to be absolutely horizontal or suspended, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical," and does not mean that the structure must be completely horizontal, but can be slightly tilted. In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0076] The following detailed explanation illustrates the specific implementation methods:

[0077] Example 1:

[0078] This embodiment discloses a cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning.

[0079] like Figure 1 As shown, the cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning includes:

[0080] S1: Obtain real-time operating data of the corresponding part of the production equipment under test through the equipment platform and upload it to the edge platform;

[0081] S2: The edge platform receives real-time running data and preprocesses it, using the preprocessed data as the feature information to be tested;

[0082] S3: Combination Figure 2 As shown, the edge platform performs operating condition judgment on the feature information to be tested: if it is a known historical operating condition, the information of the known historical operating condition and the real-time running data are packaged into a same operating condition task and uploaded to the cloud platform; if it is an unknown operating condition, the information of several similar historical operating conditions closest to the unknown operating condition and the real-time running data are packaged into a different operating condition task and uploaded to the cloud platform.

[0083] S4: The cloud platform receives tasks under the same or different operating conditions uploaded by the edge platform. If a task under the same operating condition is received, the fault diagnosis model trained based on the fault dataset of the corresponding known historical operating conditions is sent to the edge platform. If a task under different operating conditions is received, the fault datasets of each similar historical operating condition and the real-time running data are used together as training data. The pre-trained fault diagnosis model trained from the fault datasets of all historical operating conditions is trained by transfer learning, and finally the trained fault diagnosis model is sent to the edge platform.

[0084] S5: The edge platform receives the fault diagnosis model sent from the network and uses the feature information to be tested as model input. It then outputs the corresponding fault diagnosis result through this model. Taking a bearing as a key component of the production equipment under test as an example, the fault points are distributed on the inner ring (IR), outer ring (OR), and balls (BA) of the bearing. There are three diameters for the faults: 0.007 mil, 0.014 mil, and 0.021 mil, resulting in a total of 9 fault states. Adding the bearing's normal operating state (NC), bearing faults are divided into 10 types, and the final diagnosis result is one of these fault types.

[0085] In this embodiment, when task t is executed on the device platform, its task completion latency is mainly the local computation latency, i.e. p t For task volume, c device For computing power; when a task is executed on an edge platform, the task completion latency is expressed as For data transmission rate; considering the powerful computing capabilities of cloud platforms, this invention ignores the computing latency of cloud platforms and only considers the data transmission latency from the device platform to the edge platform and then to the cloud platform.

[0086] This invention acquires real-time (sensor) data from key components of the production equipment under test through a device platform. This data is then preprocessed and assessed for operational status via an edge platform, and the results are uploaded to a cloud platform. Finally, the cloud platform receives a fault diagnosis model and performs fault diagnosis, outputting the corresponding results. The edge platform is located near the actual production environment. Deploying the fault diagnosis model on the edge platform effectively solves the high latency response problem of traditional cloud-based paradigms, improving the flexibility and scalability of the entire system and thus enhancing the real-time performance of fault diagnosis for key components of the production equipment. Furthermore, the fault diagnosis model deployed on the edge platform is pre-trained or obtained through transfer learning in the cloud. This model reduces the requirements for computing and data resources while preserving the diagnostic effectiveness of the model, resulting in faster convergence and less time for personalized training. This further improves the accuracy and real-time performance of diagnosing the condition of key components of the production equipment, especially under conditions of network congestion and insufficient personalized training samples.

[0087] This invention uses an edge platform to determine the operating conditions of the feature information under test and uploads the results (for tasks with the same or different operating conditions) to a cloud platform. The cloud platform then distributes existing fault diagnosis models for known operating conditions and trains transfer learning models for unknown operating conditions for both tasks with the same and different operating conditions. This allows for the matching of personalized fault diagnosis models to fault signals under different operating conditions. Firstly, the characteristics of fault signals differ under different operating conditions. By matching corresponding diagnostic models, the fault characteristics of the fault signals can be captured more accurately, and the fault type and location can be identified more precisely. This not only reduces false alarms and missed alarms but also shortens fault diagnosis time and improves overall work efficiency. Secondly, matching sensor data for different operating conditions makes the fault diagnosis model more comprehensive and flexible, thus adapting to more practical application scenarios and improving the generalization ability of the fault diagnosis model, enabling it to maintain high diagnostic accuracy even when facing unknown or complex operating conditions. Finally, for unknown operating conditions, this invention trains a pre-trained fault diagnosis model using transfer learning. On one hand, transfer learning allows the fault diagnosis model to better generalize to new tasks. There may be shared characteristics or patterns between the source and target domains, and transfer learning can capture these characteristics, making the model perform better on new tasks. On the other hand, transfer learning utilizes a pre-trained fault diagnosis model trained on fault datasets from all historical operating conditions. Therefore, the computational resources required for unknown operating conditions can be significantly reduced, thereby further improving the efficiency of diagnosing the condition of key components in production equipment.

[0088] In practice, vibration sensors, acceleration sensors, and other signal acquisition components are installed on key parts of the production equipment under test (such as bearings). Analog signals are converted into digital signals using a digital acquisition card, and the acquired signals are aggregated and transmitted to a computing-capable edge platform via a networked device (such as a bidirectional communication gateway device embedded with a 5G module or a PLC controller). In this embodiment, the equipment platform, as the platform for data perception and action execution, is the lowest-level physical entity of the fault diagnosis system and also the main body for the optimized operation and health maintenance of the entire system.

[0089] Taking the acquisition of sensor data from bearings on production equipment under test as an example, several data acquisition cards and controllers are installed on the equipment with the bearing under test. Accelerometers are installed in the horizontal and vertical directions of the equipment and connected to the data acquisition cards to acquire the vibration signals of the bearing under test, i.e., the sensor data. The sensor data during equipment operation is uploaded to the edge platform via DDS for further processing. At the same time, the controller installed on the equipment can receive feedback control information (including control commands and warning signals) returned by the edge platform and control the equipment to take corresponding measures.

[0090] In the specific implementation process, the edge platform preprocesses real-time running data, including data cleaning and data transformation.

[0091] Data cleaning includes deleting or completing missing and abnormal data, as well as deleting redundant data;

[0092] Data transformation refers to converting one-dimensional real-time operational data into a two-dimensional time-frequency signal graph using wavelet transform, which serves as the feature information to be measured. In this invention, the two-dimensional time-frequency graph obtained by wavelet transform can extract features strongly correlated with faults, thereby reducing the complexity of model training and minimizing overfitting.

[0093] In the specific implementation process, the edge platform is deployed with a working condition discrimination model library, which stores working condition discrimination models trained from fault datasets of each historical working condition;

[0094] When determining the operating condition of the feature information to be tested: First, the feature information to be tested is input into the operating condition discrimination model corresponding to each historical operating condition, and the corresponding reconstructed data is output by each operating condition discrimination model; then, the reconstruction error between the feature information to be tested and the reconstructed data output by each operating condition discrimination model is calculated; it is determined whether there is a reconstruction error less than or equal to the error threshold: if so, the historical operating condition to which the operating condition discrimination model corresponding to the reconstructed data of the reconstruction error belongs is taken as the known historical operating condition of the current feature information to be tested; otherwise, the operating condition of the current feature information to be tested is an unknown operating condition.

[0095] In practice, the operating condition discrimination model is the LAE model, a lightweight autoencoder model. The overall model structure is the same as the VAE in the domain fusion module, but the intermediate layer feature mapping to a distribution is omitted. Instead, the features generated by the encoder in the VAE network model are directly input into the decoder to generate the corresponding reconstructed data. This is because VAE has stronger feature generalization capabilities, and calculating the mean and variance through fully connected layers during distribution mapping significantly increases computational cost. LAE, on the other hand, directly inputs the encoder-generated features into the decoder to reconstruct the data, focusing more on extracting the unique features of fault data under each operating condition. Therefore, using LAE not only lightens the model and reduces the burden on the edge platform, but also more clearly distinguishes the relationship between the current task's operating condition and historical operating conditions.

[0096] Combination Figure 3 As shown, the fault diagnosis model includes a feature extraction module, a domain fusion module, and a classification module; the domain fusion module is a variational autoencoder.

[0097] 1) During fault diagnosis, the working logic of the fault diagnosis model is as follows:

[0098] Use the feature information to be tested as input to the fault diagnosis model;

[0099] The feature extraction module performs multi-scale feature extraction on the feature information to be tested, generating fault features containing multi-scale information.

[0100] The encoder in the domain fusion module extracts features from the fault characteristics and generates potential features;

[0101] The intermediate layer of the domain fusion module transforms latent features into a feature distribution and samples from the feature distribution to generate sampled features;

[0102] During the fault diagnosis phase, the decoder of the domain fusion module is not working.

[0103] The classification module classifies based on the sampled features and generates corresponding fault diagnosis results;

[0104] 2) During transfer learning training, the working logic of the fault diagnosis model is as follows:

[0105] Use the (preprocessed) source domain data and target domain data as input to the fault diagnosis model;

[0106] The feature extraction module performs multi-scale feature extraction on the source domain data and the target domain data respectively, generating fault features containing multi-scale information;

[0107] The encoder of the domain fusion module learns from the fault characteristics of the source domain data and the target domain data to generate domain-invariant features of the source domain and the target domain.

[0108] The intermediate layer of the domain fusion module transforms the domain-invariant features into feature distributions for each operating condition, and samples from the feature distributions for each operating condition to generate sampled features for each operating condition.

[0109] The decoder in the domain fusion module generates corresponding reconstructed data based on the sampling features under various operating conditions;

[0110] The classification module classifies samples based on their characteristics under various operating conditions and generates fault diagnosis results for each condition.

[0111] 1) Feature extraction module

[0112] To address the problem that traditional convolutional neural networks using fixed-size convolutional kernels cannot fully extract feature information from fault signals under different operating conditions, this invention proposes a multi-scale convolutional kernel feature extraction module (MSKNet) by setting multiple sets of convolutional kernels of different sizes. This improves the network's ability to extract domain-invariant and domain-transferable features. Figure 3As shown, the feature extraction module includes two first depth convolutional layers, a first pooling layer, a second depth convolutional layer, a second pooling layer, two third depth convolutional layers, and a third pooling layer connected end to end in sequence; wherein the input of the first first convolutional layer is the input of the fault diagnosis model, and the output of the second third depth convolutional layer is the input of the domain fusion module;

[0113] The first deep convolutional layer includes three branches, the input of which is the input of the first deep convolutional layer, and the output of which is concatenated to serve as the output of the first deep convolutional layer. The first branch includes a 1×1 basic convolutional layer (BC), the second branch includes a 1×1 basic convolutional layer and a 3×3 basic convolutional layer connected end to end in sequence, and the third branch includes a 1×1 basic convolutional layer and a 5×5 basic convolutional layer connected end to end in sequence.

[0114] To avoid overfitting, DropBlock regularization is used after each basic convolutional layer. Therefore, the calculation formula for the first depthwise convolutional layer is expressed as:

[0115] out1=DropBlock(BasicConv(X,1×1));

[0116] out2=Dropblock(BasicConv(BasicConv(X,1×1),3×3));

[0117] out3=DropBlock(BasicConv(BasicConv(X,1×1),5×5));

[0118] out_A=concat(out1,out2,out3);

[0119] The first pooling layer includes two branches, the input of which is the output of the second first convolutional layer. The outputs of the two branches are connected and used as the input of the 2×2 max pooling layer, and the output of the 2×2 max pooling layer is used as the output of the first pooling layer. The first branch includes a 1×1 basic convolutional layer, and the second branch includes a 1×1 basic convolutional layer and a 3×3 basic convolutional layer connected end to end.

[0120] The x×x basic convolutional layer consists of a series of x×x convolutional layers connected end-to-end, a batch normalization (BN) layer, and an activation function layer. Therefore, the pooling layer is calculated as follows:

[0121]

[0122] PoolLayer(x)=MaxPool(BC(x));

[0123] Where x is the input data, w k This represents a convolution kernel of size k*k. Indicates the convolution operation, b k The bias is denoted by , ReLU is the activation function, and batch normalization (BN) is used between the convolutional layer and the activation function.

[0124] The second deep convolutional layer includes three branches. The inputs of the three branches are the outputs of the first pooling layer. The outputs of the three branches are concatenated to serve as the output of the second deep convolutional layer. The first branch includes a 1×1 basic convolutional layer, the second branch includes a 1×1 basic convolutional layer and a 5×5 basic convolutional layer connected end to end in sequence, and the third branch includes a 1×1 basic convolutional layer and a 7×7 basic convolutional layer connected end to end in sequence.

[0125] The second pooling layer includes two branches, the input of which is the output of the second depth convolutional layer. The outputs of the two branches are connected and used as the input of the 2×2 max pooling layer, and the output of the 2×2 max pooling layer is used as the output of the second pooling layer. The first branch includes a 1×1 basic convolutional layer, and the second branch includes a 1×1 basic convolutional layer and a 5×5 basic convolutional layer connected end to end.

[0126] The third deep convolutional layer includes three branches, the input of which is the output of the second pooling layer. The outputs of the three branches are concatenated to serve as the output of the third deep convolutional layer. The first branch includes a 1×1 basic convolutional layer, the second branch includes a 1×1 basic convolutional layer and a 7×7 basic convolutional layer connected end to end, and the third branch includes a 1×1 basic convolutional layer and a 9×9 basic convolutional layer connected end to end.

[0127] The structure of the third pooling layer is the same as that of the second pooling layer.

[0128] Since fault feature information under different operating conditions will not remain on the same scale, a single-structure convolutional kernel cannot fully extract fault information under different operating conditions. Therefore, the feature extraction module of this invention uses the feature that the larger the convolutional kernel, the larger the receptive field to design convolutional layers of different sizes. Multi-scale features can express richer original data information, which is conducive to the model learning more domain-invariant features.

[0129] 2) Domain Fusion Module (VAE)

[0130] Combination Figure 3 and Figure 4As shown, the domain fusion module includes an encoder, an intermediate layer, and a decoder, wherein the output of the encoder is used as a domain-invariant feature; the encoder contains four 3×3 convolutional layers connected end to end in sequence; the decoder contains three 3×3 deconvolutional layers, a 5×5 deconvolutional layer, and a 3×3 deconvolutional layer connected end to end in sequence.

[0131] The domain fusion module (variational autoencoder, VAE) is a variant of the autoencoder (AE) that incorporates variational Bayesian theory. It generates new data by introducing a latent variable z and training a model x = g(z). This model maps the original data to a probability distribution and resamples to reconstruct the original data. Noise is introduced during the feature distribution mapping process to expand the feature encoding region. When data from different operating conditions are input into the VAE, the sampling probability of domain-invariant features between these operating conditions will continuously increase due to the shared encoder and decoder, while the difference in feature distribution will continuously decrease. This structurally greatly improves the model's generalization ability. The specific implementation process of the VAE is as follows:

[0132] Suppose we have a batch of data samples {x1, x2, ..., x} n Given x ∈ x, we hope that according to {x1, x2, ..., x...} n We obtain the distribution p(x) of x, and then sample based on p(x), thus obtaining all possible x (including {x1, x2, ..., x...}). n (excluding}). However, in practical applications, this ideal generative model is difficult to achieve

[22] . By introducing the latent variable z, the generative model can be expressed in the form of marginal likelihood integral:

[0133] p θ (x)=∫p θ (z)p θ (x|z)dz (1)

[0134] Since the latent variable z is a random variable, the integral in the above equation cannot be calculated. Taking the logarithm of both sides of the equation, the marginal likelihood integral consists of the sum of the marginal likelihoods of individual data points. Therefore, the above equation can be rewritten as:

[0135] log p θ (x)=D KL (q φ (z|x)||p θ (z|x))+L(θ,φ;x) (2)

[0136] In the formula, q φ (z|x) and p θ (z|x) represent the variational posterior distribution and the true posterior distribution, respectively, and D KL(·) is used to calculate the Kullback-Leibler (KL) divergence between the two; φ and θ are the parameters of the learned recognition model and the generated model, respectively; L(θ,φ;x) represents the lower bound of the marginal likelihood (variational) of the data x, which can also be written in the following form:

[0137] L(θ,φ;x)=-D KL (qφ(z|x)||p θ (z))+E qφ(z|x) [log p θ (x|z)] (3)

[0138] The first term in the equation represents minimizing the variational posterior distribution q. φ (z|x) and prior distribution p θ The distribution of latent variables is regularized by the KL divergence between (z); the second term is obtained by maximizing log p. θ (x|z) ensures that the generated data has the same distribution as the real data when the model converges.

[0139] In the VAE model (intermediate layer), it is assumed that the variational posterior distribution q φ (z|x) = follows a normal distribution and p θ (z) follows a standard normal distribution, and the mean and variance of the variational posterior distribution are calculated using a neural network, i.e.:

[0140] log q φ (z|x)=log N(z;μ,σ 2 I) (4)

[0141] μ=f1(x) (5)

[0142] log σ 2 =f2(x) (6)

[0143] When calculating the variance in equation (6), the fitted logσ is selected. 2 Instead of directly fitting σ 2 Because of σ 2 It is always non-negative and needs to be processed using an activation function, while logσ 2 It can be positive or negative, and no activation function is needed.

[0144] After the above processing, we obtain the variational posterior distribution that x follows. Then, we combine it with the random vector e sampled from the standard normal distribution to generate the latent variable z. After passing through the generator, we obtain x = g(z). This generator can restore the input x.

[0145] z=μ+e×σ, (e~N(0,1)).

[0146] 3) Classification Module

[0147] Combination Figure 3 As shown, the output of the VAE encoder undergoes a distribution transformation in an intermediate layer and is then resampled to obtain fault features, which serve as the input to the classification module. The classification module contains two fully connected layers, using Dropout to randomly freeze neurons to prevent overfitting, and finally, a softmax classifier is used to obtain the classification result.

[0148] In the specific implementation process, the pre-trained fault diagnosis model is trained through transfer learning through the following steps:

[0149] S401: Use the fault dataset with similar historical operating conditions as the source domain data and the real-time operating data as the target domain data;

[0150] S402: After preprocessing the source domain data and target domain data, they are used as input to the pre-trained fault diagnosis model; where preprocessing refers to transforming one-dimensional real-time running data into a two-dimensional signal time-frequency diagram through wavelet transform;

[0151] S403: Initialize the weight values ​​of each layer of the pre-trained fault diagnosis model using a normal distribution;

[0152] The formula is expressed as:

[0153]

[0154] Where: n in and n out ω represents the dimensions of the input and output, respectively; i Indicates the initial weights; N represents a normal distribution;

[0155] S404: First, the feature extraction module extracts features from the preprocessed source and target domain data to generate fault features. Then, the encoder of the domain fusion module learns the domain-invariant features of the source and target domains. Next, the intermediate layer of the domain fusion module transforms the domain-invariant features into feature distributions for each operating condition, and uses KL divergence to align the feature distributions with a 0-1 normal distribution (i.e., clustering the fault feature distributions under different operating conditions towards a standard normal distribution, continuously increasing the sampling probability of invariant features and decreasing the sampling probability of domain-specific features). Then, sampling is performed in the feature distribution to generate sampled features for various operating conditions. Further, the decoder of the domain fusion module generates corresponding reconstructed data based on the sampled features under various operating conditions. Finally, the encoder and decoder of the feature extraction module and the domain fusion module are trained by minimizing the reconstruction error between the source and target domain data and the corresponding reconstructed data.

[0156] The loss function of the domain fusion module is expressed as:

[0157]

[0158]

[0159] Where: N s and N t represents the amount of data in the source domain and the target domain, respectively; x represents the actual data in the source domain and the target domain. Indicates reconstructed data; ω lm λ represents the weights of each layer in the pre-trained fault diagnosis model; λ is a hyperparameter; M represents the number of weight parameters in each layer; μ and σ 2 These represent the expectation and variance, which are key parameters that determine the characteristic distribution.

[0160] S405: The classification module classifies the samples based on the sampling features under various working conditions, generates fault diagnosis results for various working conditions, and finally optimizes the parameters of the classification module through the loss function of the classification module.

[0161] 1) The calculation formula for the classification module is expressed as follows:

[0162]

[0163] In the formula: n represents the number of neurons in the output layer; y k a represents the output of the k-th neuron; k This represents the output value of a specific neuron; a i This represents the output value of each neuron;

[0164] 2) The loss function for the classification module is expressed as:

[0165] L cs =MSE(f(x) i ),y i );

[0166] In the formula: MSE represents the notation for the cross-entropy loss function; f(x) i ) represents the output of the model; y i Indicates the true result;

[0167] S406: Repeat steps S404 to S405 until the model converges, and the trained fault diagnosis model is obtained.

[0168] Example 2:

[0169] This embodiment discloses a cloud-edge collaborative diagnosis system for production equipment faults based on transfer learning, which is implemented based on the cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning in Embodiment 1.

[0170] A cloud-edge collaborative diagnostic system for production equipment faults based on transfer learning, comprising:

[0171] The equipment platform is used to acquire real-time operating data of the corresponding parts of the production equipment under test;

[0172] Vibration sensors, acceleration sensors, and other signal acquisition components are installed on the corresponding parts of the production equipment under test. The analog signals are converted into digital signals by a digital acquisition card, and the acquired signals are aggregated and transmitted to an edge platform with computing capabilities through a network device (such as a two-way communication gateway device with an embedded 5G module or a PLC controller).

[0173] In this embodiment, the device platform, as the platform for data perception and action execution, is the lowest-level physical entity of the fault diagnosis system, and also the main body for the optimization of the entire system's operation and health maintenance.

[0174] Taking the acquisition of sensor data from bearings on production equipment under test as an example, several data acquisition cards and controllers are installed on the equipment with the bearing under test. Accelerometers are installed in the horizontal and vertical directions of the equipment and connected to the data acquisition cards to acquire the vibration signals of the bearing under test, i.e., the sensor data. The sensor data during equipment operation is uploaded to the edge platform via DDS for further processing. At the same time, the controller installed on the equipment can receive feedback control information (including control commands and warning signals) returned by the edge platform and control the equipment to take corresponding measures.

[0175] The edge platform receives and preprocesses real-time operating data uploaded from the device platform, using the preprocessed data as the feature information to be tested. Then, it performs operating condition judgment on the feature information to be tested: if it is a known historical operating condition, the information of the known historical operating condition and the real-time operating data are packaged into a same operating condition task and uploaded to the cloud platform; if it is an unknown operating condition, the information of several similar historical operating conditions closest to the unknown operating condition and the real-time operating data are packaged into a different operating condition task and uploaded to the cloud platform.

[0176] The edge platform's preprocessing of real-time running data includes data cleaning and data transformation. Data cleaning includes deleting or completing missing and abnormal data, as well as deleting redundant data. Data transformation refers to converting one-dimensional real-time running data into a two-dimensional signal time-frequency diagram through wavelet transform, which can be used as the feature information to be measured.

[0177] The cloud platform is used to receive tasks under the same or different operating conditions uploaded by the edge platform. If a task under the same operating condition is received, the fault diagnosis model trained based on the fault dataset of the corresponding known historical operating conditions will be sent to the edge platform. If a task under different operating conditions is received, the fault datasets of each similar historical operating condition and the real-time running data will be used as training data. The pre-trained fault diagnosis model trained from the fault datasets of all historical operating conditions will be trained by transfer learning, and finally the trained fault diagnosis model will be sent to the edge platform.

[0178] The edge platform is also used to receive the fault diagnosis model and take the feature information to be tested as the model input, and output the corresponding fault diagnosis result through the fault diagnosis model.

[0179] This invention acquires real-time (sensor) data from key components of the production equipment under test through a device platform. This data is then preprocessed and assessed for operational status via an edge platform, and the results are uploaded to a cloud platform. Finally, the cloud platform receives a fault diagnosis model and performs fault diagnosis, outputting the corresponding results. The edge platform is located near the actual production environment. Deploying the fault diagnosis model on the edge platform effectively solves the high latency response problem of traditional cloud-based paradigms, improving the flexibility and scalability of the entire system and thus enhancing the real-time performance of fault diagnosis for key components of the production equipment. Furthermore, the fault diagnosis model deployed on the edge platform is pre-trained or obtained through transfer learning in the cloud. This model reduces the requirements for computing and data resources while preserving the diagnostic effectiveness of the model, resulting in faster convergence and less time for personalized training. This further improves the accuracy and real-time performance of diagnosing the condition of key components of the production equipment, especially under conditions of network congestion and insufficient personalized training samples.

[0180] This invention uses an edge platform to determine the operating conditions of the feature information under test and uploads the results (for tasks with the same or different operating conditions) to a cloud platform. The cloud platform then distributes existing fault diagnosis models for known operating conditions and trains transfer learning models for unknown operating conditions for both tasks with the same and different operating conditions. This allows for the matching of personalized fault diagnosis models to fault signals under different operating conditions. Firstly, the characteristics of fault signals differ under different operating conditions. By matching corresponding diagnostic models, the fault characteristics of the fault signals can be captured more accurately, and the fault type and location can be identified more precisely. This not only reduces false alarms and missed alarms but also shortens fault diagnosis time and improves overall work efficiency. Secondly, matching sensor data for different operating conditions makes the fault diagnosis model more comprehensive and flexible, thus adapting to more practical application scenarios and improving the generalization ability of the fault diagnosis model, enabling it to maintain high diagnostic accuracy even when facing unknown or complex operating conditions. Finally, for unknown operating conditions, this invention trains a pre-trained fault diagnosis model using transfer learning. On one hand, transfer learning allows the fault diagnosis model to better generalize to new tasks. There may be shared characteristics or patterns between the source and target domains, and transfer learning can capture these characteristics, making the model perform better on new tasks. On the other hand, transfer learning utilizes a pre-trained fault diagnosis model trained on fault datasets from all historical operating conditions. Therefore, the computational resources required for unknown operating conditions can be significantly reduced, thereby further improving the efficiency of diagnosing the condition of key components in production equipment.

[0181] To better illustrate the advantages of the technical solution of the present invention, the following experiment is disclosed in this embodiment.

[0182] This experiment uses the public dataset for motor bearings provided by Case Western Reserve University (CWRU) to test the fault diagnosis model proposed in this invention (hereinafter also referred to as MSDF-VAE).

[0183] 1. Data description and comparison algorithms

[0184] 1.1 Data Description

[0185] The CWRU bearing dataset was developed using vibration data collected from a 2-horsepower Reliance Electric motor via accelerometers located near and far from the bearing. Fault points were distributed across the inner ring (IR), outer ring (OR), and balls (BA) of the bearing. Three fault diameters were identified: 0.007 mil, 0.014 mil, and 0.021 mil, resulting in nine fault states. Combined with the normal bearing state (NC), bearing faults were categorized into ten types, each with 1000 samples and each sample being 1200 characters long.

[0186] 1.2 Comparison Algorithm

[0187] 1) CNN-AE: A traditional convolutional neural network that uses AE to extract domain-invariant features.

[0188] 2) CNN-VAE: Traditional convolutional neural networks use VAE to extract domain-invariant features.

[0189] 3) MSDF-AE: The MSDF module (feature extraction module) proposed in this invention uses the traditional AE to extract domain-invariant features.

[0190] 1.3 Experimental Environment

[0191] The experimental setup consisted of Windows 11 and a GPU 1060-super. All models were compiled using the PyTorch 2.3 framework and Python 3.9 under Anaconda. Based on the four load conditions in the CWRU dataset, data collected at 0HP, 1HP, 2HP, and 3HP loads were designated as datasets A, B, C, and D, respectively. Each dataset contained 1024 data points as a single sample. The samples were preprocessed using CMOR wavelets to generate time-frequency graphs, each saved as a 224*224 image. Each dataset ultimately contained 10,000 samples, which were then divided into training, test, and validation sets in a 7:2:1 ratio. The batch size was 16, and the Adam optimizer was used to optimize the loss function with an initial learning rate of 0.001.

[0192] 2. Experimental Analysis

[0193] 2.1 MSDF-VAE Algorithm Test Results

[0194] The bearing fault data is divided into 9 categories according to different loads and the size of the fault diameter. In addition to the normal data, there are a total of 10 categories. The labels for each category are 0-9, as shown in Table 1.

[0195] Table 1. Label settings for the dataset

[0196]

[0197] Figure 5 The test results of the data set A→B under the model algorithm proposed in this invention are shown. A is the source domain data and B is the target domain data. Only 30% of the labeled data in the target domain is used for classification training. Figure 5 (a) shows the decreasing curves of KL divergence and reconstruction error of the domain fusion module during the training process. KL divergence can converge quickly after several iterations, indicating that the domain fusion module can quickly extract the shared features between the source domain and the target domain. Figure 5 (b) Shows the training loss, test loss, and accuracy of the classification module. After 25 iterations, the accuracy of the classification model reaches 99.28%. Figure 5 In (c), only one sample in the 1000 samples of dataset B had a classification error, that is, the data with label 8 was classified as label 2.

[0198] 2.2 Comparison of MSDF-VAE with other algorithms

[0199] In this experiment, we first conducted an ablation study on the feature extraction module and domain fusion module of MSDF-VAE to verify the effectiveness of the two modules. The platform and framework on which the comparison algorithms ran were the same as in Experiment 2.1. To compare the effectiveness of the experiments, the depth of the feature extraction network was the same as the model proposed in this invention, and multiple sets of comparison experiments were added to prevent experimental randomness. The comparison results are shown in Table 2. Six sets of experiments were conducted for each algorithm, with each set of experiments repeated 5 times and the average accuracy was taken.

[0200] Table 2. Validity Tests of MSDF-VAE

[0201]

[0202] As shown in Table 2, traditional CNN and AE as feature extraction and domain fusion modules yielded the lowest fault diagnosis results. Using VAE improved the transfer learning performance by nearly 8%, which is attributed to VAE's stronger feature generalization ability. MSDF-AE achieved an accuracy nearly 10% higher than CNN-AE and nearly 2% higher than CNN-VAE, indicating that the multi-scale feature extraction of this invention can obtain more abstract features from both the source and target domains, and these abstract features are shared by both. Finally, combining MSDF and VAE resulted in a transfer learning fault diagnosis accuracy of 99.3%.

[0203] Figure 6 The t-SNE visualization method is used to visualize the classification results of the above four algorithms. Figure 6 The classification effect is clearly visible. The best classification effect is when each category clusters into a single area with no overlap between areas. In CNN-AE, categories 3, 4, 7, and 9 are relatively dispersed, and the classification blocks for categories 4, 7, and 9 overlap. In MSDF-AE and CNN-VAE, the classification blocks for categories 0, 1, 4, and 7 are very clear, but the classification blocks for the other categories are dispersed and overlap. In the classification effect of MSDF-VAE, it is obvious that each category has its own distinct area without overlap, proving the effectiveness of the model proposed in this invention.

[0204] This experiment also compared the accuracy of MSDF-VAE with several transfer learning algorithms listed in Table 3 in fault diagnosis. DBN extracts features through stacked restricted Boltzmann machines; CNN uses MMD to optimize the difference between the source and target domains; DAFD uses AE as a single-layer representation model of the deep structure, reduces distribution differences through MMD, and uses weight regularization to prevent data loss; DANN extracts domain-invariant features through adversarial training of the feature extractor and the domain discriminator; DIBDN and SAE-CSDF are two state-of-the-art DTL methods. The former improves upon DBN by reducing the distribution differences of hidden units to extract domain-invariant features, while the latter uses stacked AE to extract domain-invariant features and achieves unsupervised transfer learning through class separation and domain fusion.

[0205] Table 3 Comparison of algorithms with related works

[0206]

[0207] Table 3 shows that traditional DBN and CNN algorithms perform poorly in bearing fault transfer learning, with accuracies below 70%. DAFD and DANN improve transfer learning accuracy by using regularization techniques, achieving accuracies above 90%. DIBDN and SAE-CSDF both reach 97% accuracy. Compared to the above methods, MSDF-VAE does not use MMD to optimize the distribution error between the source and target domains during training. Instead, it uses the intermediate layer of VAE to align the distribution. In our experiments, we compared the efficiency of VAE's alignment distribution with MMD. VAE converged after 60 iterations in 134 seconds, while MMD required 200 iterations to converge in 887 seconds. This shows that VAE is about 7 times more efficient at extracting domain-invariant features than MMD. In conclusion, the MSDF-VAE of this invention can not only efficiently extract domain-invariant features between the source and target domains but also achieve high-precision transfer fault diagnosis.

[0208] 2.3 Small Sample Label Data Transfer Analysis

[0209] In experiments 2.1 and 2.2, the amount of labeled data used accounted for 30% of the target domain data. In this experiment, the classification model was trained by reducing the amount of labeled data in the target domain data. Comparative experiments were conducted using labeled data amounts of 20%, 10%, 5%, 3%, 2%, and 0% of the target domain data. The experimental results are as follows: Figure 7 As shown, when the amount of labeled data accounts for more than 3% of the target domain data, the migration accuracy is above 96%. When the amount of labeled data is less than 3%, the accuracy shows a rapid downward trend. However, when the amount of labeled data is 0, the accuracy can still reach 86.92%.

[0210] 2.4 Effectiveness Analysis of Cloud Platform Migration Fault Diagnosis Strategies

[0211] To validate the cloud-edge collaboration strategy in the absence of labeled data in the target domain, the previous three experiments used only datasets A, B, and C. This experiment adds dataset D. When one dataset is used as the target domain data to publish a fault diagnosis task on the device platform, the other three datasets are stored as historical fault datasets on the cloud platform. First, a correlation analysis experiment is conducted on the four datasets A, B, C, and D to compare the reconstruction error of one dataset under one model. Subscripts d and m are used to distinguish between the dataset and the model. For example, A... d →B m This represents the reconstruction error of dataset A under model B (trained from dataset B). The experimental results are shown in Table 5.

[0212] Table 5 Reconstruction errors of the dataset under each model (*10) -3 )

[0213]

[0214] The load on the four datasets A, B, C, and D increases sequentially. As shown in Table 5, the greater the difference in operating conditions, the greater the reconstruction error of the data under the same model. Two model offloading strategies are defined in the cloud platform based on the difference in reconstruction error. When the reconstruction error of the new operating condition data is the same as that of the historical data, the high-precision historical fault diagnosis model is directly offloaded to the edge platform. When the reconstruction errors are different, n historical operating condition datasets similar to the new operating condition are selected to adjust the global model. Table 7 shows the fault diagnosis accuracy when n takes different values.

[0215] Table 7 shows the fault diagnosis accuracy when n takes different values.

[0216]

[0217] As shown in Table 7, even without labeled data in the target domain, the accuracy of transfer fault diagnosis can reach over 95% after using the pre-trained model. Compared to training the transfer learning model directly using historical data, the accuracy of bearing fault diagnosis is improved by about 10%. Further analysis of the experimental results reveals that as the value of n increases, the accuracy of fault diagnosis does not necessarily increase further. In experiments with target domains A and B, the accuracy decreases as n increases. In the experiment with target domain D, the accuracy of fault diagnosis increases when n=2 compared to n=1, but decreases slightly when n=3. This indicates that when training the transfer learning model, historical operating condition data similar to the new operating condition should be selected. If too much data with significantly different operating conditions is selected, this data may be treated as noise, interfering with model training and ultimately reducing the model's fault diagnosis accuracy. Further experiments show that the fault diagnosis strategy in the cloud platform not only improves the accuracy of fault diagnosis but also accelerates the model's convergence speed, thereby reducing the model's training time. Figure 8 The results of the B→A experiment are compared with and without a pre-trained model. The pre-trained model was trained on three historical datasets: B, C, and D.

[0218] from Figure 8As can be seen, when training a fault diagnosis model suitable for dataset A, loading the parameters from the pre-trained model significantly improves the model's convergence speed, allowing it to converge after only 10 iterations. In contrast, training the model directly using the default initial parameters requires approximately 40 iterations. Furthermore, without a training model to validate the loss, the loss value increases after 40 iterations, indicating that further training after convergence will lead to overfitting and instability in the model training process. Moreover, with a pre-trained model, the model's loss value is smaller, which also means that the model of this invention will have higher fault diagnosis accuracy.

[0219] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit the technical solutions. Those skilled in the art should understand that any modifications or equivalent substitutions to the technical solutions of the present invention without departing from the spirit and scope of the present invention should be covered within the scope of the claims of the present invention.

Claims

1. A cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning, characterized in that, include: S1: Obtain real-time operating data of the corresponding part of the production equipment under test through the equipment platform and upload it to the edge platform; S2: The edge platform receives real-time running data and preprocesses it, using the preprocessed data as the feature information to be tested; S3: The edge platform performs operating condition judgment on the feature information to be tested: if it is a known historical operating condition, the information of the known historical operating condition and the real-time running data are packaged into a same operating condition task and uploaded to the cloud platform; if it is an unknown operating condition, the information of several similar historical operating conditions closest to the unknown operating condition and the real-time running data are packaged into a different operating condition task and uploaded to the cloud platform. In step S3, the edge platform deploys a working condition discrimination model library, which stores working condition discrimination models trained from fault datasets for each historical working condition. When determining the operating condition of the feature information to be tested: First, the feature information to be tested is input into the operating condition discrimination model corresponding to each historical operating condition, and the corresponding reconstructed data is output by each operating condition discrimination model; then, the reconstruction error between the feature information to be tested and the reconstructed data output by each operating condition discrimination model is calculated; it is determined whether there is a reconstruction error less than or equal to the error threshold: if so, the historical operating condition to which the operating condition discrimination model corresponding to the reconstructed data of the reconstruction error belongs is taken as the known historical operating condition of the current feature information to be tested; otherwise, the operating condition of the current feature information to be tested is an unknown operating condition. S4: The cloud platform receives tasks under the same or different operating conditions uploaded by the edge platform. If a task under the same operating condition is received, the fault diagnosis model trained based on the fault dataset of the corresponding known historical operating conditions is sent to the edge platform. If a task under different operating conditions is received, the fault datasets of each similar historical operating condition and the real-time running data are used together as training data. The pre-trained fault diagnosis model trained from the fault datasets of all historical operating conditions is trained by transfer learning, and finally the trained fault diagnosis model is sent to the edge platform. S5: The edge platform receives the fault diagnosis model sent out and uses the feature information to be tested as the model input, and outputs the corresponding fault diagnosis result through the fault diagnosis model.

2. The cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning as described in claim 1, characterized in that: In step S2, the preprocessing of real-time running data includes data cleaning and data transformation; Data cleaning includes deleting or completing missing and abnormal data, as well as deleting redundant data; Data transformation refers to the process of converting one-dimensional real-time running data into a two-dimensional signal time-frequency diagram through wavelet transform, which can then be used as the feature information to be measured.

3. The cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning as described in claim 1, characterized in that: In step S3, the working condition discrimination model is the LAE model, which is based on the VAE network model but removes the intermediate layer to map the features to a distribution. That is, the features generated by the encoder in the VAE network model are directly input into the decoder to generate the corresponding reconstructed data.

4. The cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning as described in claim 1, characterized in that: In step S4, the fault diagnosis model includes a feature extraction module, a domain fusion module, and a classification module; wherein the domain fusion module is a variational autoencoder. 1) During fault diagnosis, the working logic of the fault diagnosis model is as follows: Use the feature information to be tested as input to the fault diagnosis model; The feature extraction module performs multi-scale feature extraction on the feature information to be tested, generating fault features containing multi-scale information. The encoder in the domain fusion module extracts features from the fault characteristics and generates potential features; The intermediate layer of the domain fusion module transforms latent features into a feature distribution and samples from the feature distribution to generate sampled features; The classification module classifies based on the sampled features and generates corresponding fault diagnosis results; 2) During transfer learning training, the working logic of the fault diagnosis model is as follows: Use source domain data and target domain data as input to the fault diagnosis model; The feature extraction module performs multi-scale feature extraction on the source domain data and the target domain data respectively, generating fault features containing multi-scale information; The encoder of the domain fusion module learns from the fault characteristics of the source domain data and the target domain data to generate domain-invariant features of the source domain and the target domain. The intermediate layer of the domain fusion module transforms the domain-invariant features into feature distributions for each operating condition, and samples from the feature distributions for each operating condition to generate sampled features for each operating condition. The decoder in the domain fusion module generates corresponding reconstructed data based on the sampling features under various operating conditions; The classification module classifies samples based on their characteristics under various operating conditions and generates fault diagnosis results for each condition.

5. The cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning as described in claim 4, characterized in that: The feature extraction module includes two first depth convolutional layers, a first pooling layer, a second depth convolutional layer, a second pooling layer, two third depth convolutional layers, and a third pooling layer connected end to end in sequence; wherein the input of the first depth convolutional layer is the input of the fault diagnosis model, and the output of the second third depth convolutional layer is the input of the domain fusion module; The first depthwise convolutional layer consists of three branches. The inputs of the three branches are all inputs to the first depthwise convolutional layer, and the outputs of the three branches are concatenated to serve as the output of the first depthwise convolutional layer. The first branch consists of a 1×1 basic convolutional layer, the second branch consists of a 1×1 basic convolutional layer and a 3×3 basic convolutional layer connected end to end in sequence, and the third branch consists of a 1×1 basic convolutional layer and a 5×5 basic convolutional layer connected end to end in sequence. The first pooling layer includes two branches. The inputs of both branches are the outputs of the second first-depth convolutional layer. The outputs of the two branches are concatenated and used as the input of the 2×2 max pooling layer. The output of the 2×2 max pooling layer is used as the output of the first pooling layer. The first branch includes a 1×1 basic convolutional layer, and the second branch includes a 1×1 basic convolutional layer and a 3×3 basic convolutional layer connected end to end. The second deep convolutional layer consists of three branches. The inputs of the three branches are the outputs of the first pooling layer. The outputs of the three branches are concatenated to serve as the output of the second deep convolutional layer. The first branch consists of a 1×1 basic convolutional layer, the second branch consists of a 1×1 basic convolutional layer and a 5×5 basic convolutional layer connected end to end in sequence, and the third branch consists of a 1×1 basic convolutional layer and a 7×7 basic convolutional layer connected end to end in sequence. The second pooling layer includes two branches, the input of which is the output of the second depth convolutional layer. The outputs of the two branches are connected and used as the input of the 2×2 max pooling layer, and the output of the 2×2 max pooling layer is used as the output of the second pooling layer. The first branch includes a 1×1 basic convolutional layer, and the second branch includes a 1×1 basic convolutional layer and a 5×5 basic convolutional layer connected end to end. The third deep convolutional layer consists of three branches. The inputs of the three branches are the outputs of the second pooling layer. The outputs of the three branches are concatenated to serve as the output of the third deep convolutional layer. The first branch consists of a 1×1 basic convolutional layer, the second branch consists of a 1×1 basic convolutional layer and a 7×7 basic convolutional layer connected end to end, and the third branch consists of a 1×1 basic convolutional layer and a 9×9 basic convolutional layer connected end to end. The structure of the third pooling layer is the same as that of the second pooling layer; The basic convolutional layer consists of a 1×1 convolutional layer, a batch normalization layer, and an activation function layer connected sequentially; a 3×3 basic convolutional layer consists of a 3×3 convolutional layer, a batch normalization layer, and an activation function layer connected sequentially; a 5×5 basic convolutional layer consists of a 5×5 convolutional layer, a batch normalization layer, and an activation function layer connected sequentially; a 7×7 basic convolutional layer consists of a 7×7 convolutional layer, a batch normalization layer, and an activation function layer connected sequentially; and a 9×9 basic convolutional layer consists of a 9×9 convolutional layer, a batch normalization layer, and an activation function layer connected sequentially.

6. The cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning as described in claim 4, characterized in that: The domain fusion module includes an encoder and a decoder, where the output of the encoder is used as a domain-invariant feature; The encoder contains four 3×3 convolutional layers connected end-to-end in sequence; the decoder contains three 3×3 deconvolutional layers, a 5×5 deconvolutional layer, and a 3×3 deconvolutional layer connected end-to-end in sequence.

7. The cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning as described in claim 4, characterized in that: The classification module consists of two fully connected layers connected end-to-end and a softmax classifier. The neurons between the two fully connected layers are randomly frozen by a Dropout layer to prevent overfitting.

8. The cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning as described in claim 7, characterized in that: In step S4, the pre-trained fault diagnosis model is trained through transfer learning using the following steps: S401: Use the fault dataset with similar historical operating conditions as the source domain data and the real-time operating data as the target domain data; S402: After preprocessing the source domain data and target domain data, they are used as input to the pre-trained fault diagnosis model; S403: Initialize the weight values ​​of each layer of the pre-trained fault diagnosis model using a normal distribution; S404: First, the feature extraction module extracts features from the preprocessed source domain data and target domain data to generate fault features; then, the encoder of the domain fusion module learns the domain-invariant features of the source domain and target domain. The domain-invariant features are then transformed into feature distributions for each operating condition through the intermediate layer of the domain fusion module. KL divergence is used to align the feature distributions with a 0-1 normal distribution, and sampling is performed on the feature distributions to generate sampled features for various operating conditions. The decoder of the domain fusion module then generates corresponding reconstructed data based on the sampled features for various operating conditions. Finally, the encoder and decoder of the feature extraction module and the domain fusion module are trained by minimizing the reconstruction error between the source domain data, the target domain data, and the corresponding reconstructed data. S405: The classification module classifies the samples based on the sampling features under various working conditions, generates fault diagnosis results for various working conditions, and finally optimizes the parameters of the classification module through the loss function of the classification module. S406: Repeat steps S404 to S405 until the model converges, and the trained fault diagnosis model is obtained.

9. A cloud-edge collaborative diagnostic system for production equipment faults based on transfer learning, characterized in that: The implementation of the cloud-edge collaborative diagnosis method for production equipment faults based on transfer learning as described in claim 1 includes: The equipment platform is used to acquire real-time operating data of the corresponding parts of the production equipment under test; The edge platform receives and preprocesses real-time operating data uploaded from the device platform, using the preprocessed data as the feature information to be tested. Then, it performs operating condition judgment on the feature information to be tested: if it is a known historical operating condition, the information of the known historical operating condition and the real-time operating data are packaged into a same operating condition task and uploaded to the cloud platform; if it is an unknown operating condition, the information of several similar historical operating conditions closest to the unknown operating condition and the real-time operating data are packaged into a different operating condition task and uploaded to the cloud platform. The edge platform is equipped with a working condition discrimination model library, which stores working condition discrimination models trained from fault datasets for each historical working condition. When determining the operating condition of the feature information to be tested: First, the feature information to be tested is input into the operating condition discrimination model corresponding to each historical operating condition, and the corresponding reconstructed data is output by each operating condition discrimination model; then, the reconstruction error between the feature information to be tested and the reconstructed data output by each operating condition discrimination model is calculated; it is determined whether there is a reconstruction error less than or equal to the error threshold: if so, the historical operating condition to which the operating condition discrimination model corresponding to the reconstructed data of the reconstruction error belongs is taken as the known historical operating condition of the current feature information to be tested; otherwise, the operating condition of the current feature information to be tested is an unknown operating condition. The cloud platform is used to receive tasks under the same or different operating conditions uploaded by the edge platform. If a task under the same operating condition is received, the fault diagnosis model trained based on the fault dataset of the corresponding known historical operating conditions will be sent to the edge platform. If a task under different operating conditions is received, the fault datasets of each similar historical operating condition and the real-time running data will be used as training data. The pre-trained fault diagnosis model trained from the fault datasets of all historical operating conditions will be trained by transfer learning, and finally the trained fault diagnosis model will be sent to the edge platform. The edge platform is also used to receive the fault diagnosis model and take the feature information to be tested as the model input, and output the corresponding fault diagnosis result through the fault diagnosis model.

Citation Information

Patent Citations

  • Gearbox fault diagnosis method based on cloud edge collaboration and deep convolution transfer learning

    CN115293234A