Industrial Equipment Fault Diagnosis System Based on Multi-Class Time-Series Data Analysis
Through an industrial equipment fault diagnosis system based on multi-category time sequence data analysis, the problems of low fault diagnosis efficiency and low accuracy in the existing technology are solved, and efficient and automated fault diagnosis is achieved, which is suitable for a variety of industrial equipment and data characteristics.
Patent Information
- Application Number
- CN202510329544.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-20
AI Technical Summary
The prior art is inefficient and has low accuracy in the diagnosis of industrial equipment faults. In particular, small and medium-sized enterprises need to manually label abnormal data for each device, and the data characteristics of different sensors are different, making it difficult to adapt.
An industrial equipment fault diagnosis system based on multi-class timing data analysis is adopted, which includes a data acquisition unit, a model training unit and a fault diagnosis unit. By preprocessing a variety of runtime sequence data to be trained, analyzing data types and generating target fault diagnosis models, it can adapt to different types of data characteristics.
It improves the efficiency and accuracy of industrial equipment fault diagnosis, reduces manual intervention, reduces dependence on a large amount of manual annotation data, can quickly adapt to different equipment types and data characteristics, significantly improving the reliability of diagnostic results.
Smart Images

Figure CN119828674B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial equipment monitoring and management, and particularly relates to an industrial equipment fault diagnosis system based on multi-class time series data analysis. Background Art
[0002] With the development of industrial technology, the types of industrial equipment are becoming increasingly rich. In an industrial scenario, a factory usually includes various industrial equipment, such as wind turbines, port cranes, chemical equipment, and power transmission equipment. As the operation time of industrial equipment increases, the equipment performance will gradually decline, and the probability of failure will also gradually increase. Moreover, the occurrence of faults in industrial equipment often triggers a chain reaction. If not discovered and eliminated in time, it will affect the operation of industrial production and even cause serious economic losses and casualties. Therefore, the fault diagnosis of industrial equipment is very important.
[0003] Currently, the fault diagnosis of industrial equipment mainly involves manually annotating data and then performing anomaly detection and fault prediction through methods such as machine learning. However, there are a large number of types of industrial equipment. For small and medium-sized enterprises, a large amount of abnormal data needs to be annotated for each type of equipment, and the generality of the fault diagnosis system is insufficient, which will lead to low efficiency of equipment fault diagnosis. On the other hand, the fault modes of equipment are usually diverse and not fixed, and the sampling data of different sensors have different characteristics. At this time, if the same fault diagnosis standard is adopted, the accuracy of the diagnosis results will decrease.
[0004] Therefore, how to improve the efficiency and accuracy of industrial equipment fault diagnosis is an urgent problem to be solved at present. Summary of the Invention
[0005] In order to solve the technical problem of how to improve the efficiency and accuracy of industrial equipment fault diagnosis, the purpose of the present invention is to provide an industrial equipment fault diagnosis system based on multi-class time series data analysis. The specific technical solution adopted is as follows:
[0006] In an embodiment of the present application, an industrial equipment fault diagnosis system based on multi-class time series data analysis is provided. The system includes:
[0007] A data acquisition unit for acquiring the operation data of multi-class industrial equipment, where the operation data includes multiple types of operation time series data to be trained and multiple types of operation time series data to be diagnosed;
[0008] A model training unit for analyzing multiple data types corresponding to the multiple types of operation time series data to be trained, and transmitting the operation time series data to be trained corresponding to each data type to a pre-constructed initial fault diagnosis model for training to obtain a target fault diagnosis model corresponding to each data type;
[0009] A fault diagnosis unit, configured to analyze multiple data types corresponding to the multiple types of operation timing data to be diagnosed, and transmit the operation timing data to be diagnosed to the corresponding target fault diagnosis model according to the data types, so as to obtain a fault diagnosis result.
[0010] In an embodiment of the present application, the system further includes:
[0011] A preprocessing unit, configured to perform windowing processing and differencing processing on the multiple types of operation timing data to be trained, so as to obtain preprocessed data to be trained, wherein the preprocessed data to be trained is used to be transmitted to the model training unit to analyze and obtain the multiple data types.
[0012] In an embodiment of the present application, the preprocessing unit includes:
[0013] A frequency domain conversion subunit, configured to perform frequency domain conversion on the multiple types of operation timing data to be trained, so as to obtain multiple types of operation frequency domain data to be trained;
[0014] A differencing operation subunit, configured to analyze the periodicity performance of the multiple types of operation frequency domain data to be trained, determine the periodicity significance level, determine periodic data and aperiodic data according to the periodicity significance level, perform windowing processing on the periodic data and the aperiodic data, and perform a differencing operation on the windowed periodic data and aperiodic data;
[0015] A normalization subunit, configured to perform normalization processing on the data after the differencing operation, so as to obtain the preprocessed data to be trained.
[0016] In an embodiment of the present application, the analyzing the periodicity performance of the multiple types of operation frequency domain data to be trained and determining the periodicity significance level includes:
[0017] Obtaining the extreme values of the multiple types of operation frequency domain data to be trained;
[0018] Determining an envelope line according to the extreme values;
[0019] Determining the periodicity significance level according to the frequency domain distribution of the multiple types of operation frequency domain data to be trained and the distance distribution of the extreme values corresponding to the envelope line.
[0020] In an embodiment of the present application, the determining periodic data and aperiodic data according to the periodicity significance level and performing windowing processing on the periodic data and the aperiodic data includes:
[0021] Take the data with a periodicity significance greater than a preset periodicity significance threshold as the periodic data, and take the data with a periodicity significance less than or equal to the preset periodicity significance threshold as the aperiodic data;
[0022] Take the average distance of all extreme values on the envelope line as the first window size, and perform windowing processing on the periodic data based on the first window size;
[0023] Perform clustering processing on the aperiodic data to obtain multiple clusters, divide the aperiodic data into data corresponding to multiple states according to the distances between data points between the multiple clusters, and perform windowing processing on the data corresponding to each state.
[0024] In an embodiment of the present application, performing a difference operation on the windowed aperiodic data includes:
[0025] For multiple data segments in the data corresponding to any one of the states, when performing forward difference on the data segments and the amount of data in the data segments is greater than the amount of data in the adjacent previous data segment, divide the data segments into data segments of the same length as the adjacent previous data segment for difference operation, and perform alignment difference on the remaining parts.
[0026] In an embodiment of the present application, the initial fault diagnosis model includes a long short-term memory network model, and the model training unit includes:
[0027] A data classification subunit, configured to analyze the noise manifestation and fault manifestation of the multiple types of to-be-trained operation time series data, and determine the multiple data types according to the analysis results;
[0028] A model training subunit, configured to transmit the to-be-trained operation time series data corresponding to each data type to the pre-constructed long short-term memory network model for training to obtain a target fault diagnosis model corresponding to each data type.
[0029] In an embodiment of the present application, analyzing the noise manifestation and fault manifestation of the multiple types of to-be-trained operation time series data, and determining the multiple data types according to the analysis results includes:
[0030] For any one of the to-be-trained operation time series data, intercept a first classification window according to a preset second window size, move the first classification window, obtain a second classification window based on new data, analyze the data changes in the first classification window and the second classification window, and determine the data change consistency;
[0031] Based on the data change consistency, analyze the data changes in the first classification window to determine the degree of fault manifestation;
[0032] Compare the data in the window with the greatest degree of fault manifestation in the two pieces of to-be-trained operation time series data to determine the combined training index;
[0033] Combine the multiple pieces of to-be-trained operation time series data according to the magnitude of the combined training index to determine the multiple data types.
[0034] In an embodiment of the present application, the step of transmitting the to-be-trained operation time series data corresponding to each data type to the pre-constructed long short-term memory network model for training to obtain the target fault diagnosis model corresponding to each data type includes:
[0035] Extract a training data set from the to-be-trained operation time series data corresponding to each data type;
[0036] Transmit the training data set to the pre-constructed long short-term memory network model, output the possibility of the fault state, and use the fault state with the greatest possibility of the fault state as the fault state corresponding to the training data set, where the fault state includes the normal state and the fault states corresponding to multiple fault types.
[0037] In an embodiment of the present application, the step of analyzing the multiple data types corresponding to the multiple pieces of to-be-diagnosed operation time series data includes:
[0038] Perform windowing processing and differencing processing on the multiple pieces of to-be-diagnosed operation time series data to obtain the pre-processed to-be-diagnosed data;
[0039] For any one of the to-be-diagnosed data, determine the data change consistency, analyze the data change based on the data change consistency to determine the degree of fault manifestation, compare the data in the window with the greatest degree of fault manifestation in the two pieces of to-be-diagnosed data to determine the combined training index;
[0040] Combine the multiple pieces of to-be-diagnosed data according to the magnitude of the combined training index to determine the multiple data types.
[0041] The present invention has the following beneficial effects:
[0042] First, the data acquisition unit acquires the operation data of multiple types of industrial equipment, where the operation data includes multiple types of operation time series data to be trained and multiple types of operation time series data to be diagnosed. Then, the model training unit analyzes the multiple data types corresponding to the multiple types of operation time series data to be trained, and transmits the operation time series data corresponding to each data type to the pre-constructed initial fault diagnosis model for training, obtaining the target fault diagnosis model corresponding to each data type. Finally, the fault diagnosis unit analyzes the multiple data types corresponding to the multiple types of operation time series data to be diagnosed, and transmits the operation time series data to be diagnosed to the corresponding target fault diagnosis model according to the data type, obtaining the fault diagnosis result. In the present invention, the data acquisition unit acquires the operation data of multiple types of industrial equipment, and can receive the operation data from different types of industrial equipment, not limited to a specific type of equipment only, making the system have wide applicability and being applicable to various industrial scenarios such as wind turbines, port cranes, chemical equipment, power transmission equipment, etc. By the model training unit analyzing the multiple data types corresponding to the multiple types of operation time series data to be trained, it can identify and process different types of data (such as vibration data, temperature data, current data, etc.), thus adapting to the data characteristics collected by different sensors, further enhancing the versatility of the system and avoiding the need for separate modeling for each device. The model training unit will automatically analyze the data types and transmit the corresponding data to the initial fault diagnosis model for training, reducing manual intervention, significantly improving the automation level of the fault diagnosis system, and reducing the dependence of small and medium-sized enterprises on a large amount of manually labeled data. By generating the corresponding target fault diagnosis model for each data type, it can quickly adapt to different device types and data characteristics without the need to redesign or adjust the entire diagnosis process, thus improving the diagnosis efficiency. Since the system generates the corresponding target fault diagnosis model according to different data types, each target fault diagnosis model can better capture the characteristics and fault patterns of specific types of data. Compared with the method of using a single model to process all data, this method can significantly improve the accuracy of the diagnosis result. The fault diagnosis unit will select the corresponding target fault diagnosis model according to the data type of the operation time series data to be diagnosed, ensuring that the diagnosis process is always based on the model most suitable for this data type, thus avoiding misjudgment problems caused by using a unified standard. By integrating the operation data of multiple types of industrial equipment, the system can analyze the device status from multiple dimensions, further improving the reliability of the diagnosis result. Brief Description of the Drawings
[0043] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.
[0044] Figure 1 Schematic diagram of the implementation environment of an industrial equipment fault diagnosis system based on multi-class time series data provided by an embodiment of the present invention;
[0045] Figure 2 Schematic diagram of the structure of an industrial equipment fault diagnosis system based on multi-class time series data provided by an embodiment of the present invention;
[0046] Figure 3 Schematic diagram of a piece of time series data to be trained and run provided by an embodiment of the present invention;
[0047] Figure 4 Schematic diagram of another piece of time series data to be trained and run provided by an embodiment of the present invention;
[0048] Figure 5 Schematic diagram of the structure of a long short-term memory network provided by an embodiment of the present invention;
[0049] Figure 6 Schematic diagram of the flow of a method for industrial equipment fault diagnosis based on multi-class time series data provided by an embodiment of the present invention. Detailed implementation manners
[0050] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in combination with the accompanying drawings and preferred embodiments, details the specific implementation manners, structures, features and effects of an industrial equipment fault diagnosis system based on multi-class time series data proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0051] It should be noted that the terms "first", "second", etc. in the specification of this application and the above-mentioned accompanying drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs.
[0053] The following specifically describes the specific solution of an industrial equipment fault diagnosis system provided by the present invention based on multi-class time series data analysis with reference to the accompanying drawings.
[0054] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the implementation environment of an industrial equipment fault diagnosis system provided by an embodiment of the present invention based on multi-class time series data analysis. As shown in Figure 1 , the implementation environment includes a fault diagnosis terminal 101, a data acquisition terminal 102, and an industrial equipment 103. The fault diagnosis terminal 101 can be a terminal device configured with a target client, including but not limited to a laptop computer, a tablet computer, a personal digital assistant, a PAD, a desktop computer, etc. with local computing capabilities. The target client can be a client installed with an industrial equipment fault diagnosis system based on multi-class time series data analysis, such as a video client, an instant messaging client, a browser client, etc. The fault diagnosis terminal 101 can communicate with the data acquisition terminal 102 through a network, which can include but not limited to: a wired network, a wireless network. Among them, the wired network includes: a local area network, a metropolitan area network, and a wide area network, and the wireless network includes: Bluetooth, WIFI, and other networks that implement wireless communication. The above-mentioned fault diagnosis terminal 101 can include but not limited to a human-computer interaction screen, a processor, and a memory. The above-mentioned human-computer interaction screen can be used to display the fault diagnosis result. The above-mentioned processor can be used to respond to human-computer interaction operations, execute corresponding operations, or generate corresponding instructions.
[0055] As an optional method, the operation data of the industrial equipment 103 can be collected by the data acquisition terminal 102. The data acquisition terminal 102 can include, for example: sensors, data acquisition modules, and spectrum analyzers. Among them, the sensors can include: vibration sensors, which are used to measure the vibration signals during equipment operation and are applicable to equipment such as wind turbines and port cranes; temperature sensors, which are used to monitor the temperature changes of key parts of the equipment and are applicable to scenarios such as chemical equipment and power transmission equipment; current / voltage sensors, which are used to collect the current and voltage data of electrical equipment and are applicable to equipment such as motors and transformers; pressure sensors, which are used to measure the pressure values of hydraulic or pneumatic systems and are applicable to equipment such as port cranes and chemical equipment; acceleration sensors, which are used to capture the acceleration changes of the equipment and are applicable to high-speed rotating mechanical equipment; acoustic sensors.
[0056] As an optional method, the industrial equipment 103 can be, for example: wind turbines, port cranes, chemical equipment, power transmission equipment, manufacturing production lines, elevator systems.
[0057] As an alternative, the above-mentioned fault diagnosis terminal 101 can also be a server, which can be a single server, a server cluster composed of multiple servers, or a cloud server. The above is only an example, and this embodiment does not make any limitation thereto.
[0058] As an alternative, an industrial equipment fault diagnosis system based on multi-class time series data analysis can be installed on the fault diagnosis terminal 101. The structure of the industrial equipment fault diagnosis system based on multi-class time series data analysis is specifically as follows:
[0059] A data acquisition unit, configured to acquire the operation data of multiple types of industrial equipment, where the operation data includes multiple types of to-be-trained operation time series data and multiple types of to-be-diagnosed operation time series data;
[0060] A model training unit, configured to analyze the multiple data types corresponding to the multiple types of to-be-trained operation time series data, and transmit the to-be-trained operation time series data corresponding to each data type to a pre-constructed initial fault diagnosis model for training to obtain a target fault diagnosis model corresponding to each data type;
[0061] A fault diagnosis unit, configured to analyze the multiple data types corresponding to the multiple types of to-be-diagnosed operation time series data, and transmit the to-be-diagnosed operation time series data to the corresponding target fault diagnosis model according to the data type to obtain a fault diagnosis result.
[0062] Based on the above method, the data acquisition unit acquires the operation data of multiple types of industrial equipment, and can receive the operation data from different types of industrial equipment, not limited to a specific type of equipment, making the system have wide applicability and being applicable to various industrial scenarios such as wind turbines, port cranes, chemical equipment, and power transmission equipment. By analyzing multiple data types corresponding to multiple operation time series data to be trained through the model training unit, it can identify and process different types of data (such as vibration data, temperature data, current data, etc.), thus adapting to the data characteristics collected by different sensors, further enhancing the versatility of the system and avoiding the need for separate modeling for each device. The model training unit will automatically analyze the data types and transmit the corresponding data to the initial fault diagnosis model for training, reducing manual intervention, significantly improving the automation level of the fault diagnosis system, and reducing the dependence of small and medium-sized enterprises on a large amount of manually labeled data. By generating a corresponding target fault diagnosis model for each data type, it can quickly adapt to different device types and data characteristics without the need to redesign or adjust the entire diagnosis process, thereby improving the diagnosis efficiency. Since the system generates a corresponding target fault diagnosis model according to different data types, each target fault diagnosis model can better capture the characteristics and fault patterns of specific types of data. Compared with the method of using a single model to process all data, this method can significantly improve the accuracy of the diagnosis results. The fault diagnosis unit will select the corresponding target fault diagnosis model according to the data type of the operation time series data to be diagnosed, ensuring that the diagnosis process is always based on the model most suitable for this data type, thus avoiding misjudgment problems caused by using a unified standard. By integrating the operation data of multiple types of industrial equipment, the system can analyze the equipment status from multiple dimensions, further improving the reliability of the diagnosis results.
[0063] As an optional example, in this embodiment, the host for the industrial equipment fault diagnosis system based on multi-class time series data analysis is not limited. The industrial equipment fault diagnosis system based on multi-class time series data analysis can be installed on the fault diagnosis terminal 101. For example, when the fault diagnosis terminal 101 is a desktop computer, some or all components of the industrial equipment fault diagnosis system based on multi-class time series data analysis can be installed on the desktop computer.
[0064] In an embodiment of the present application, an industrial equipment fault diagnosis system based on multi-class time series data analysis is provided. Figure 2 The following is a schematic structural diagram of an industrial equipment fault diagnosis system based on multi-class time series data analysis provided by an embodiment of the present invention. Refer to Figure 2 The industrial equipment fault diagnosis system based on multi-class time series data analysis includes the units described in 210 to 230 as follows:
[0065] A data acquisition unit 210, configured to acquire the operation data of multiple types of industrial equipment, where the operation data includes multiple types of operation time series data to be trained and multiple types of operation time series data to be diagnosed.
[0066] Among them, multiple types of industrial equipment refer to that the system can adapt to multiple types of industrial equipment, such as wind turbines, port cranes, chemical equipment, power transmission equipment, etc. The operation data of these equipment has different characteristics and patterns, so a general method is needed to process them.
[0067] Among them, multiple types of operation time series data to be trained are operation status information collected from multiple types of industrial equipment, usually including data in the form of time series, such as vibration signals, temperature changes, current fluctuations, etc. These data are used to train the fault diagnosis model.
[0068] Among them, multiple types of operation time series data to be diagnosed come from multiple types of industrial equipment, but need to be diagnosed in actual applications. Similar to the training data, they also contain operation status information in the form of time series.
[0069] A model training unit 220, configured to analyze multiple data types corresponding to the multiple types of operation time series data to be trained, and transmit the operation time series data to be trained corresponding to each data type to a pre-constructed initial fault diagnosis model for training, so as to obtain a target fault diagnosis model corresponding to each data type.
[0070] Among them, the data type refers to the result of classifying the operation data according to the characteristics of the data (such as source, physical meaning, statistical characteristics, etc.). For example, vibration data, temperature data, and current data can be regarded as different data types.
[0071] Among them, the initial fault diagnosis model is a pre-constructed basic model, configured to receive training data of a specific data type and generate a more accurate target fault diagnosis model through training. It can be a model based on deep learning or other machine learning models.
[0072] Among them, when analyzing multiple data types corresponding to the multiple types of operation time series data to be trained, the system will analyze the input operation time series data to be trained and identify different data types therein. For example, distinguish vibration data, temperature data, and current data from a set of data.
[0073] Among them, according to the analysis result, the operation time series data to be trained of each data type are respectively input into the initial fault diagnosis model for training. After training, a target fault diagnosis model for each data type is generated. For example, a model specifically for vibration data analysis is trained from vibration data.
[0074] A fault diagnosis unit 230 is configured to analyze multiple data types corresponding to the multiple to-be-diagnosed operation timing data, and transmit the to-be-diagnosed operation timing data to the corresponding target fault diagnosis model according to the data types, so as to obtain a fault diagnosis result.
[0075] Among them, in the actual diagnosis process, the system will analyze the received to-be-diagnosed operation timing data and identify different data types therein. According to the analysis result, the to-be-diagnosed operation timing data of each data type is respectively input into the corresponding target fault diagnosis model, and finally a fault diagnosis result is generated.
[0076] Exemplarily, assume there is a factory including the following three industrial devices: a wind turbine, which mainly collects vibration data and temperature data; a port crane, which mainly collects current data and pressure data; and a power transmission device, which mainly collects voltage data and current data. The data acquisition unit collects operation data from the above devices to obtain the following two types of data: to-be-trained operation timing data: data for training the model; to-be-diagnosed operation timing data: data for actual diagnosis. The model training unit analyzes the to-be-trained operation timing data and identifies the following data types: wind turbine: vibration data, temperature data; port crane: current data, pressure data; power transmission device: voltage data, current data. The to-be-trained operation timing data of each data type is input into the initial fault diagnosis model for training to generate the following target fault diagnosis models: vibration data model, temperature data model, current data model, pressure data model, voltage data model. The fault diagnosis unit analyzes the received to-be-diagnosed operation timing data and identifies the data type. According to the data type, the to-be-diagnosed operation timing data is input into the corresponding target fault diagnosis model to obtain a fault diagnosis result. For example: if the vibration data of the wind turbine is received, it is input into the vibration data model; if the current data of the port crane is received, it is input into the current data model.
[0077] In an embodiment of the present application, the system further includes:
[0078] A preprocessing unit is configured to perform windowing processing and differencing processing on the multiple to-be-trained operation timing data to obtain preprocessed to-be-trained data, where the preprocessed to-be-trained data is used to be transmitted to the model training unit to analyze and obtain the multiple data types.
[0079] Among them, taking the operation data of a wind power generation unit as an example, it includes wind speed, generator speed, power output, temperature, current, frequency, pressure, unit vibration, etc. Refer to Figure 3 , Figure 3 which is a schematic diagram of to-be-trained operation timing data provided by an embodiment of the present invention. The to-be-trained operation timing data in this schematic diagram is unit vibration timing data; refer toFigure 4 , Figure 4 Another schematic diagram of the to-be-trained operation time series data provided by an embodiment of the present invention. The to-be-trained operation time series data in this schematic diagram is wind speed time series data. As can be seen from Figure 3 and Figure 4 , affected by the external environment, mechanical noise, and mechanical operation cycle, the sampling frequencies of different data are different, and the periodic performance and change characteristics of the data are different. Therefore, if the original data is input into the model for learning, the detection accuracy will be insufficient. Therefore, preprocessing is required. The differential operation can reflect the change characteristics of the data, but in data with high noise and weak regularity, the use effect of the differential operation is poor. By windowing, part of the continuous data is regarded as a whole, and then differential processing is performed, which can ensure the processing effect of the data.
[0080] Among them, windowing processing refers to dividing the time series data into windows (or segments) of a fixed length for analyzing local data. The data within each window can be regarded as an independent data segment, which is convenient for subsequent processing and modeling. For example, for a vibration signal, the entire signal can be divided into multiple windows by setting the window size to 100 sampling points and the step size to 10 sampling points.
[0081] Among them, differential processing is a common time series data analysis method used to eliminate the trend or periodic components in the data and extract the change characteristics of the data. Specifically, differential processing generates a new sequence by calculating the differences between adjacent data points. For example, for the time series data x1, x2, x3, the first-order differential results are x2 - x1, x3 - x2.
[0082] Among them, when performing windowing processing and differential processing on the multiple to-be-trained operation time series data, first perform windowing processing on the multiple to-be-trained operation time series data to divide the data into multiple windows of a fixed length; then perform differential processing on the data within each window to extract the change characteristics of the data. This preprocessing method can reduce noise interference, highlight the data change pattern, and provide more meaningful input data for subsequent model training.
[0083] Exemplarily, assume that a vibration signal is collected from a wind turbine, which contains the following time-series data: [1.0, 1.2, 1.5, 1.8, 2.0, 2.2, 2.5, 2.8, 3.0, 3.2]. Set the window size to 4 and the step size to 2, then the original data is segmented into the following windows: Window 1, [1.0, 1.2, 1.5, 1.8]; Window 2, [1.8, 2.0, 2.2, 2.5]; Window 3, [2.5, 2.8, 3.0, 3.2]. Perform first-order difference processing on the data in each window: The difference result of Window 1 is: [1.2 - 1.0, 1.5 - 1.2, 1.8 - 1.5] = [0.2, 0.3, 0.3]; The difference result of Window 2 is: [2.0 - 1.8, 2.2 - 2.0, 2.5 - 2.2] = [0.2, 0.2, 0.3]; The difference result of Window 3 is: [2.8 - 2.5, 3.0 - 2.8, 3.2 - 3.0] = [0.3, 0.2, 0.2]. Take the data (i.e., the difference result) after windowing and difference processing as the preprocessed data to be trained, and transmit it to the model training unit for analyzing the data type and generating the target fault diagnosis model.
[0084] In the above embodiment, the windowing process can segment long time-series data into short windows, which is convenient for focusing on local features and reducing the influence of noise on the overall data. The difference processing can eliminate the trend or periodic components in the data and highlight the change characteristics of the data, thereby improving the sensitivity of the model to abnormal patterns. The preprocessed data is more stable and representative, which can effectively reduce noise interference and data redundancy, and enhance the robustness and generalization ability of the model. By performing windowing and difference processing on the data, the system can more accurately capture the change patterns of the device operating state, thereby improving the accuracy of fault diagnosis. The preprocessed data structure is more regular, which is convenient for subsequent classification, feature extraction and model training, and significantly simplifies the entire fault diagnosis process.
[0085] In an embodiment of the present application, the preprocessing unit includes:
[0086] A frequency-domain conversion sub-unit for performing frequency-domain conversion on the multiple pieces of training operation time-series data to obtain multiple pieces of training operation frequency-domain data;
[0087] A difference operation sub-unit for analyzing the periodic performance of the multiple pieces of training operation frequency-domain data, determining the degree of periodic significance, determining periodic data and non-periodic data according to the degree of periodic significance, performing windowing processing on the periodic data and the non-periodic data, and performing difference operations on the windowed periodic data and non-periodic data;
[0088] A normalization sub-unit for performing normalization processing on the data after difference operations to obtain the preprocessed data to be trained.
[0089] Among them, frequency-domain conversion is the process of converting time-series data from the time domain to the frequency domain, usually achieved through Fourier Transform or Wavelet Transform. Frequency-domain conversion can reveal the frequency components and periodic characteristics in the data, helping to analyze the oscillation patterns of signals.
[0090] Among them, periodic manifestation refers to the periodic characteristics exhibited by time-series data in the frequency domain. For example, during the operation of certain devices, there may be vibrations or fluctuations at a fixed frequency, and these characteristics will be manifested as energy concentration at specific frequencies in the frequency domain.
[0091] Among them, the degree of periodic significance is used to quantify the intensity of the periodic characteristics of the data. The degree of periodic significance can be determined by analyzing the energy distribution, envelope shape or other statistical indicators of the frequency-domain data. For example, if the energy proportion of a certain frequency is very high, it indicates that the data has strong periodicity.
[0092] Among them, periodic data refers to data with obvious periodic characteristics, usually manifested as energy concentration at specific frequencies in the frequency domain. For example, the rotating components of a wind turbine may generate periodic vibration signals.
[0093] Among them, aperiodic data refers to data without obvious periodic characteristics, and its frequency-domain representation is usually continuous or dispersed. For example, random noise or irregular fluctuations caused by equipment failures.
[0094] Among them, normalization is a data preprocessing method used to scale the data to a fixed range (such as [0, 1] or [-1, 1]) to eliminate the influence of dimension differences and numerical ranges. Normalization can improve the convergence speed and stability of model training, and at the same time avoid misjudgment caused by data scale differences.
[0095] Exemplarily, obtain n kinds of operating data of several industrial devices ; among them, the sampling frequency of each kind of data is ; the sampling quantity of each kind of data is . Periodicity is a basic characteristic of the data. The window size can be determined according to the periodic manifestation of the operating data. By drawing the upper and lower envelope lines, the extreme value changes of the data can be obtained, and then the period can be obtained. Due to the existence of extreme noise, the periodic characteristics reflected by the envelope lines will be affected, but the extreme noise data is small in quantity, and the influence of extreme noise data can be reduced through frequency-domain conversion. Perform differential operations on all the data, and then use the sigmoid function (an activation function) for normalization processing to convert all the data to the [0, 1] range, completing the data preprocessing.
[0096] In the above embodiments, the frequency-domain conversion can convert time-series data into a frequency-domain representation, thus more clearly revealing the periodic characteristics in the data. This is particularly important for analyzing the operating state of industrial equipment because many equipment failures will exhibit specific periodic changes. The analysis of the periodicity significance further quantifies the periodic characteristics of the data, enabling the system to distinguish between periodic data and aperiodic data and providing a basis for subsequent processing. By performing windowing processing and differential operations on periodic data and aperiodic data respectively, the system can adopt different processing strategies for different types of data, thereby improving the accuracy of data classification. The normalization processing eliminates the influence of dimensional differences and numerical ranges between data, making the model training process more stable and efficient. The normalized data is easier to be learned by models such as neural networks, thus accelerating convergence and improving model performance.
[0097] In an embodiment of the present application, the analysis of the periodic performance of the multiple pieces of to-be-trained operating frequency-domain data to determine the periodicity significance includes:
[0098] Obtain the extreme values of the multiple pieces of to-be-trained operating frequency-domain data;
[0099] Determine the envelope line according to the extreme values;
[0100] Determine the periodicity significance according to the frequency-domain distribution of the multiple pieces of to-be-trained operating frequency-domain data and the distance distribution of the extreme values corresponding to the envelope line.
[0101] Among them, the extreme value refers to the position where the energy distribution reaches a local maximum or minimum in the frequency-domain data. These extreme value points usually reflect the important characteristics of the periodic components in the signal. For example, in the frequency-domain graph, the energy peaks at certain frequencies may correspond to the vibration frequencies or harmonic frequencies during the operation of the equipment.
[0102] Among them, the envelope line is a curve used to describe the range of signal amplitude changes and can reflect the periodic or aperiodic characteristics of the signal. In frequency-domain analysis, the envelope line can be obtained by fitting the extreme values of the frequency-domain data and is used to further quantify the energy distribution law of the signal.
[0103] Among them, the frequency-domain distribution describes the energy distribution of the signal at different frequencies. By analyzing the frequency-domain distribution, the main frequency components of the signal and their intensities can be identified. For example, the vibration signal of a wind turbine may show significant energy concentration within a certain specific frequency range.
[0104] Among them, the distance distribution refers to the statistical characteristics of the distances between adjacent extreme points on the envelope line. By analyzing the distribution of these distances, the significance degree of the periodicity of the signal can be judged. If the distances between adjacent extreme points are relatively uniform, it indicates that the signal has strong periodicity; on the contrary, if the distance distribution is chaotic, it indicates that the signal is non-periodic.
[0105] Exemplarily, when determining the significance degree of periodicity, the following steps can be taken: Obtain all the maximum values and minimum values of the data, connect them with a straight line to obtain the upper and lower envelope lines; for the upper and lower envelope lines of any data, obtain all the maximum values and minimum values of the envelope lines. Any two adjacent maximum values or minimum values reflect the period of the data, and the distance distribution between the maximum values or minimum values reflects the significance degree of the periodicity of the data; taking all the maximum values of the upper envelope line of any data as an example, obtain all the data between any two adjacent maximum values, and obtain the corresponding frequency domain distribution ; obtain the frequency domain distribution of the distances between all adjacent maximum values ; based on the law of the frequency domain distribution and the law of the distances between the maximum values of the envelope line, calculate the significance degree of the periodicity of any data; among them, the distance between adjacent maximum values is the interval between two adjacent maximum values, and this interval can be an interval in frequency; among them, the values within the distance between adjacent maximum values are the amplitudes or energy magnitudes corresponding to the frequency components.
[0106] Exemplarily, the representation method of the significance degree of periodicity can be:
[0107]
[0108] Among them, represents the significance degree of the periodicity of any data; represents the exponential function in the mathematical calculation method; represents the number of types of different distances between adjacent maximum values; is the number of types of the r-th type of distance between adjacent maximum values; represents the number of occurrences of all the maximum value distances; represents the number of maximum values; represents the minimum value in the mathematical calculation method; represents the number of numerical distributions within the distance between adjacent maximum values; represents the frequency difference between the j-th value of the c-th maximum value segment (the data segment between the c-th maximum value and the (c + 1)-th maximum value) and the s-th maximum value segment.
[0109] Among them, It represents the ratio of the number of types of different adjacent maximum distances on the envelope line to the number of occurrences of all maximum distances, characterizing the degree of periodicity significance of the data. The larger this value is, the wider the numerical distribution is, and the weaker the periodicity is.
[0110] Among them, It represents the sum of the frequency differences of all numerical values in the c-th maximum segment and the s-th maximum segment.
[0111] Among them, It represents obtaining the sum of the frequency differences with the smallest frequency difference between the c-th maximum segment and all maximum segments, and this value reflects the difference between the c-th maximum segment and the overall data.
[0112] Among them, It represents obtaining the difference between all maximum segments and the overall data. The larger this value is, the weaker the regularity and periodicity of the data are.
[0113] In the above embodiment, by extracting the extreme values of the frequency domain data and fitting the envelope line, the system can more accurately capture the periodic characteristics of the signal. This method is more intuitive and quantitative compared to directly analyzing the frequency domain distribution. The extreme value distance distribution corresponding to the envelope line provides a quantitative index for evaluating the degree of periodicity significance of the signal. This multi-dimensional analysis method can reduce the risk of misjudgment and improve the reliability of the diagnosis results. For the signals of complex industrial equipment, the frequency domain distribution may contain multiple frequency components and noise interference. Through envelope line analysis and distance distribution statistics, the system can quickly screen out the main periodic characteristics and simplify the subsequent processing flow.
[0114] In an embodiment of the present application, determining periodic data and aperiodic data according to the degree of periodicity significance and performing windowing processing on the periodic data and the aperiodic data includes:
[0115] Taking the data with the degree of periodicity significance greater than the preset periodicity significance threshold as the periodic data, and taking the data with the degree of periodicity significance less than or equal to the preset periodicity significance threshold as the aperiodic data;
[0116] Taking the average value of the distances of all extreme values on the envelope line as the first window size, and performing windowing processing on the periodic data based on the first window size;
[0117] Performing clustering processing on the aperiodic data to obtain multiple clusters, dividing the aperiodic data into data corresponding to multiple states according to the distances between the data points among the multiple clusters, and performing windowing processing on the data corresponding to each state.
[0118] Among them, clustering is an unsupervised learning method used to divide data points into several groups (clusters) such that data points within the same group have high similarity while data points between different groups have significant differences. In this solution, clustering is applied to non-periodic data with the aim of dividing it into different states according to the distribution characteristics of the data. Cluster classes are the results of clustering, representing a group of data points with similar characteristics. Each cluster class corresponds to a certain pattern or state in the data. For example, during the operation of industrial equipment, there may be three states: "normal operation", "minor anomaly", and "severe fault", and these states can be obtained as corresponding cluster classes through cluster analysis. Multiple states refer to different data categories divided through clustering, and each state corresponds to a specific situation during equipment operation. For non-periodic data, multiple states can reflect the dynamic change characteristics of the equipment under different conditions.
[0119] Exemplarily, the higher the degree of periodicity significance, the stronger the regularity of the data, and the better the analysis effect of the data through differential operation by increasing the window. Conversely, the worse the effect of the differential operation. The preset threshold for the degree of periodicity significance is 0.7. Data with a degree of periodicity significance greater than 0.7 is recorded as periodic data, and the first window size is set as the average distance between all adjacent maximum values of the upper envelope line. Exemplarily, for non-periodic data, its changes are mostly related to the external environment, human operations, and natural wear and tear of the equipment. Through clustering, the data can be classified. By adding windows to different categories of data and then moving the windows for differential operation, the differential effect can be ensured, and the states of the operating data can also be distinguished.
[0120] Exemplarily, when performing clustering on non-periodic data, the following hierarchical clustering algorithm can be used for clustering: Consider each data point as a separate cluster; for any data point, merge it with the data point with the closest distance to obtain several new cluster classes; and perform iteration and continue merging; until the distance between any cluster class and all the remaining cluster classes is greater than (maximum value - minimum value) or all the data is merged into one cluster class; the distance between two cluster classes is the minimum value of the distances between any two data points among them; obtain several cluster classes, and record data with the minimum value of the distances between any two data points between the cluster classes less than or equal to (maximum value - minimum value) as data in the same state, and obtain several states of non-periodic data; since the time when the data is in the same state is not the same, the window size is not fixed.
[0121] In the above embodiments, by clustering the non-periodic data into multiple clusters, the hidden patterns and state change characteristics in the data can be captured more accurately. This method is more refined than the traditional unified processing method, avoiding misjudgments caused by data diversity. Different windowing strategies are adopted for data in different states, enabling the system to flexibly adapt to various data characteristics. For example, data in a low-load state may require a smaller window to capture rapid changes, while data in a high-load state requires a larger window to smooth out noise interference.
[0122] In an embodiment of the present application, performing a difference operation on the windowed non-periodic data includes:
[0123] For multiple data segments in the data corresponding to any one of the states, when performing a forward difference on the data segments and the amount of data in the data segments is greater than the amount of data in the adjacent previous data segment, the data segments are divided into data segments of the same length as the adjacent previous data segment for difference operation, and the remaining parts are subjected to alignment difference.
[0124] Wherein, a data segment refers to multiple subsequences obtained by windowing the non-periodic data. Each data segment contains a certain number of data points for subsequent difference operation or other analysis.
[0125] Wherein, forward difference is a common difference calculation method used to extract the change trend of data. Forward difference is applied to each data segment to capture the change characteristics of the data.
[0126] Wherein, alignment difference means that in the case of inconsistent data segment lengths, by adjusting the length or position of the data segments, they are aligned with the adjacent data segments before performing the difference operation. Specifically, when the amount of data in a certain data segment is greater than that of the adjacent previous data segment, it is divided into parts of the same length as the previous data segment, and the remaining parts are separately subjected to difference operation.
[0127] Exemplarily, for the data segments of any one state , where represents the v-th data segment of the u-th state, when performing a forward difference on , if the amount of data in is greater than , then first divide into data segments of the same length as for difference, and then the remaining parts continue with alignment difference to complete the difference operation.
[0128] In the above embodiments, through the alignment difference method, the comparability of the difference operations between data segments of different lengths is ensured, and misjudgment caused by inconsistent data segment lengths is avoided. When processing longer data segments, the information of all data points is retained through the division and alignment operations, minimizing data loss. The difference operation can effectively extract the change characteristics of the data, and the alignment difference method further improves the accuracy and stability of feature extraction. For example, in the fault diagnosis of industrial equipment, the difference results can reflect the change trend of the equipment operation state and help identify potential abnormalities.
[0129] In an embodiment of the present application, the initial fault diagnosis model includes a long short-term memory network model, and the model training unit includes:
[0130] A data classification subunit, configured to analyze the noise manifestations and fault manifestations of the multiple types of to-be-trained operation time series data, and determine the multiple data types according to the analysis results;
[0131] A model training subunit, configured to transmit the to-be-trained operation time series data corresponding to each data type to the pre-constructed long short-term memory network model for training, and obtain the target fault diagnosis model corresponding to each data type.
[0132] Among them, due to the diverse changes in data, the data is divided into periodic data and non-periodic data. Moreover, the noise manifestations of the data are different, and the manifestations of the fault data are different. If all data is trained by the same model, it will lead to incorrect fault judgments. If each type of data is trained separately, it will result in insufficient data volume, and still cause enterprises to need to manage multiple data of multiple devices separately. Therefore, the data should also be classified according to the noise manifestations and fault manifestations of the data, and then each type of data is trained separately.
[0133] Among them, the long short-term memory network model is a special recurrent neural network (RNN), designed specifically for processing time series data. It can effectively capture the long-term dependencies in time series data by introducing a cell state and gating mechanisms (such as an input gate, a forget gate, and an output gate). In this solution, the LSTM model is used as the initial fault diagnosis model to analyze the complex patterns of industrial equipment operation data.
[0134] Among them, the noise manifestation refers to the fluctuations or abnormal manifestations in the operation time series data caused by environmental interference, measurement errors, or other non-fault factors. For example, the vibration signal of a wind turbine may be affected by changes in wind speed, and these changes do not necessarily reflect equipment failures.
[0135] Among them, the fault manifestation refers to the abnormal patterns or characteristics in the operation time series data caused by equipment failures. For example, bearing damage may cause harmonic components with specific frequencies to appear in the vibration signal, and this phenomenon can be regarded as a fault manifestation.
[0136] Among them, the data classification subunit analyzes the noise manifestation and fault manifestation of the running timing data, identifies different characteristics in the data, and accordingly divides the data into multiple categories. For example, some data may mainly contain noise components, while other data shows obvious fault characteristics. Through this analysis, the system can more accurately determine the data categories, thereby providing a basis for subsequent modeling.
[0137] In the above embodiments, the noise is randomly distributed and will repeat in the timing data. The fault data is special and is different from the range and variation of normal data. Therefore, the manifestation of the fault data can be obtained by calculating the repeatability of the change in any data segment in the overall data. Through the analysis of the noise manifestation and fault manifestation, the system can more accurately distinguish data with different characteristics, avoiding misjudging noise as a fault or ignoring potential fault signals. Training target fault diagnosis models for different data categories respectively enables each model to focus on specific types of data characteristics, thereby improving the accuracy of the diagnosis results.
[0138] In an embodiment of the present application, analyzing the noise manifestation and fault manifestation of the multiple types of to-be-trained running timing data, and determining the multiple data categories according to the analysis results includes:
[0139] For any one of the to-be-trained running timing data, according to a preset second window size, intercept a first classification window, move the first classification window, and based on the new data, obtain a second classification window, and analyze the data changes in the first classification window and the second classification window to determine the data change consistency;
[0140] Based on the data change consistency, analyze the data changes in the first classification window to determine the degree of fault manifestation;
[0141] Compare the data in the window with the maximum degree of fault manifestation in two types of to-be-trained running timing data to determine the combined training index;
[0142] According to the magnitude of the combined training index, combine the multiple types of to-be-trained running timing data to determine the multiple data categories.
[0143] Among them, by setting a preset window size (second window size), intercept a first classification window from the to-be-trained running timing data, and move the window to obtain new data to form a second classification window. Subsequently, analyze the data changes within these two windows to capture the change patterns of local data.
[0144] Among them, data change consistency is used to measure the similarity of data changes between the first classification window and the second classification window. If the data change trends of the two windows are consistent, they are considered to have high data change consistency; otherwise, it is low. For example, if the data fluctuation amplitudes, frequencies, or directions of the two windows are similar, their consistency is high.
[0145] Among them, the degree of fault manifestation is based on the result of data change consistency analysis, and further evaluates whether there are potential fault characteristics in the first classification window. Usually, statistical indicators (such as mean, variance, frequency domain distribution, etc.) are used to quantify the significance of fault characteristics.
[0146] Among them, similarity comparison refers to comparing the data in the window with the largest degree of fault manifestation in two kinds of to-be-trained operation time series data, and calculating the similarity between them. Common methods include Euclidean distance, cosine similarity, or mutual information, etc.
[0147] Among them, the combined training index is a quantitative index used to evaluate whether two kinds of to-be-trained operation time series data can be combined into the same data type for training. If the similarity comparison result shows that the two kinds of data have high similarity, the combined training index is large; otherwise, it is small.
[0148] Among them, according to the size of the combined training index, it is decided whether to combine multiple kinds of to-be-trained operation time series data into the same data type. If the combined training index exceeds the preset threshold, these data are classified into the same category; otherwise, they are kept separately processed.
[0149] Exemplarily, for any one of the to-be-trained operation time series data, a first classification window is intercepted (10 - 100 sampling points, with a step size of 1), for the data segment within the first classification window, the data segment is translated left and right to reach the second classification window , and the first classification window is translated up and down so that as much data as possible within the second classification window appears in the first classification window .
[0150] Exemplarily, the representation method of the data change consistency between the first classification window and the second classification window can be:
[0151]
[0152] Among them, represents the data change consistency between the first classification window and the second classification window ; Represents the exponential function in mathematical calculation methods; Represents the window size; Represents the first classification window The difference between the a-th data and the previous data in the first classification window, where when a equals 1, the difference between the first data and the last data is obtained; Represents the second classification window The difference between the a-th data and the previous data in the second classification window, where when a equals 1, the difference between the first data and the last data is obtained.
[0153] Among them, Represents the first classification window And the second classification window The consistency of data changes at corresponding positions. If the difference is smaller, it indicates that the data trends in the window are consistent, and thus the data change consistency is greater.
[0154] Exemplarily, by performing multiple up-and-down translations and left-and-right translations on the first classification window, multiple reached second classification windows can be obtained. Furthermore, the data change consistency between the first classification window and all the reached second classification windows can be obtained. Sort all the obtained data change consistencies from largest to smallest. In this way, each first classification window corresponds to a sorting result of data change consistency. Among them, if there are more high data change consistencies in the sorting result, it indicates that the data changes in the first classification window appear more frequently in the overall data, and the degree of fault manifestation is smaller, that is, the fault manifestation is weaker. On the contrary, if the data change consistencies in the sorting result are relatively low, it indicates that the degree of fault manifestation of the first classification window is greater, that is, the fault manifestation is higher. Among them, the level of data change consistency can be judged through a preset data change consistency threshold.
[0155] Exemplarily, the way to represent the degree of fault manifestation of the first classification window can be:
[0156]
[0157] Among them, Represents the degree of fault manifestation of the first classification window; Represents the exponential function in mathematical calculation methods; w is the number of second classification windows corresponding to the first w data change consistencies in the sorting result of data change consistency. In this embodiment, w can be 100; Represents the first classification window And the second classification window Of data change consistency; Represents adding the data change consistencies of the first classification window and w second classification windows.
[0158] It should be noted that in the embodiments of the present application, first, a first classification window is intercepted according to the second window size, and then the degree of fault manifestation corresponding to the first classification window is determined using the above steps; then, according to the second window size, a new first classification window is intercepted, and then the degree of fault manifestation corresponding to the new first classification window is determined using the above steps; and so on, the degrees of fault manifestation corresponding to multiple first classification windows can be obtained.
[0159] Exemplarily, for any data, e first classification windows with the largest degree of fault manifestation are obtained. If the data in the windows with the largest degree of fault manifestation of two data are relatively similar, the two data can be trained together. The representation of the combined training index of any two data can be:
[0160]
[0161] Wherein, represents the combined training index of the i-th data and the h-th data; represents the exponential function in the mathematical calculation method; represents the difference between the a-th data of the e-th first classification window with the largest degree of fault manifestation in the i-th data and the previous data. Wherein, when a is equal to 1, the difference between the first data and the last data is obtained; represents the difference between the a-th data of the e-th first classification window with the largest degree of fault manifestation in the h-th data and the previous data. Wherein, when a is equal to 1, the difference between the first data and the last data is obtained; represents the total amount of data within the e-th first classification window with the largest degree of fault manifestation.
[0162] Wherein, represents adding up the changes of all data in the e-th first classification window with the largest degree of fault manifestation in the two data. The smaller the difference, the larger the combined training index of the data. Since the window sizes may be different, the smallest difference part of the two windows can be taken for calculation.
[0163] Then, the two data with the largest combined training index are combined until there is no data with a combined training index greater than 0.5, and several types of data are obtained.
[0164] In the above embodiments, through the consistency analysis of the data changes within the classification window, the system can more accurately identify the noise and fault characteristics in the data, thereby improving the accuracy of data classification. Based on similarity comparison and combined training indices, the system can automatically determine which data can be combined into the same data category for training, reducing the need for manual intervention while improving the training efficiency. The target fault diagnosis model generated after data combination can cover more diverse data characteristics, thereby enhancing the generalization ability of the model and improving its adaptability to unknown data. By combining similar data categories, the system avoids multiple modeling of duplicate data, significantly reducing the consumption of computing resources.
[0165] In one embodiment of the present application, the step of transmitting the to-be-trained operation time-series data corresponding to each data category to the pre-constructed long short-term memory network model for training to obtain the target fault diagnosis model corresponding to each data category includes:
[0166] Extracting a training data set from the to-be-trained operation time-series data corresponding to each data category;
[0167] Transmitting the training data set to the pre-constructed long short-term memory network model, outputting the possibility of the fault state, and taking the fault state with the highest possibility of the fault state as the fault state corresponding to the training data set, where the fault state includes the normal state and the fault states corresponding to multiple fault types.
[0168] Among them, the training data set is a set of samples extracted from the to-be-trained operation time-series data corresponding to each data category and is used for model training. These data are pre-processed (such as windowing, differencing, etc.) and can better reflect the change characteristics of the equipment operation state.
[0169] Among them, the possibility of the fault state is a probability value output by the long short-term memory (LSTM) model, indicating the possibility that the input data belongs to a certain fault state. For example, for a certain input data, the model may output the following results: the possibility of the normal state, 0.8; the possibility of fault type A, 0.15; the possibility of fault type B, 0.05.
[0170] Among them, the fault state refers to the state categories that may occur during the operation of the equipment, including the normal state and the fault states corresponding to multiple fault types. The normal state indicates that the equipment is operating normally without obvious abnormalities; the fault state corresponds to different fault modes of the equipment.
[0171] Among them, the fault states corresponding to multiple fault types refer to various specific fault modes that the equipment may have. For example, a wind turbine may have the following fault types: bearing damage, gearbox wear, generator overheating.
[0172] Exemplarily, refer toFigure 5 , Figure 5 is a schematic structural diagram of a long short - term memory network provided by an embodiment of the present invention. As Figure 5 shown, the long short - term memory network includes three gating mechanisms (the process from left to right represents the forget gate, the input gate, and the output gate respectively), C is the cell state, x is the input, y is the output, and tanh is the activation function. 70% of the data is extracted from the time - series data to be trained and run corresponding to the data types as the training data set. The training data set is input into the LSTM network model, and the output is the possibility of each state of each data. The one with the greatest possibility is recorded as the state of the data. The states include normal and fault types.
[0173] In the above - mentioned embodiment, through the fault state possibility output by the LSTM model, the system can more accurately identify the operating state of the device and avoid misjudgment caused by a single - threshold judgment. The model can not only distinguish between the normal state and the fault state, but also further subdivide into various specific fault types, so as to provide more detailed diagnostic results.
[0174] In an embodiment of the present application, the analysis of the multiple data types corresponding to the multiple time - series data to be diagnosed includes:
[0175] Performing windowing processing and differential processing on the multiple time - series data to be diagnosed to obtain the pre - processed data to be diagnosed;
[0176] For any one of the data to be diagnosed, determining the data change consistency. Based on the data change consistency, analyzing the data change, determining the degree of fault manifestation, comparing the data in the window with the greatest degree of fault manifestation in the two data to be diagnosed, and determining the combined training index;
[0177] Merging the multiple data to be diagnosed according to the size of the combined training index to determine the multiple data types.
[0178] It should be noted that in this embodiment, when analyzing the multiple data types corresponding to the multiple time - series data to be diagnosed, the method used is the same as the method for analyzing the multiple data types corresponding to the multiple time - series data to be trained, that is, first pre - processing the multiple time - series data to be diagnosed, then determining the combined training index, and merging the multiple data to be diagnosed through the combined training index to obtain multiple data types. Specifically, the steps for analyzing the multiple data types corresponding to the multiple time - series data to be diagnosed can be, for example:
[0179] Performing frequency - domain conversion on the multiple time - series data to be diagnosed to obtain multiple frequency - domain data to be diagnosed;
[0180] Analyze the periodicity performance of multiple frequency-domain data to be diagnosed, determine the degree of periodicity significance, determine periodic data and aperiodic data according to the degree of periodicity significance, perform windowing processing on the periodic data and aperiodic data, and perform differential operations on the windowed periodic data and aperiodic data;
[0181] Normalize the data after the differential operation to obtain the data to be diagnosed after preprocessing;
[0182] For any data to be diagnosed after preprocessing, according to the preset second window size, intercept to obtain the first classification window, move the first classification window, obtain the second classification window based on the new data, analyze the data changes in the first classification window and the second classification window, determine the data change consistency, and based on the data change consistency, analyze the data changes in the first classification window to determine the degree of fault manifestation;
[0183] Compare the data in the window with the largest degree of fault manifestation in the two data to be diagnosed, determine the combined training index, and merge the data to be diagnosed after preprocessing according to the size of the combined training index to determine multiple data types.
[0184] Based on the above embodiments, the data acquisition unit acquires the operation data of multiple types of industrial equipment, and can receive the operation data from different types of industrial equipment, not limited to a specific type of equipment, making the system have wide applicability and being applicable to various industrial scenarios such as wind turbines, port cranes, chemical equipment, and power transmission equipment; by analyzing multiple data types corresponding to multiple operation time series data to be trained through the model training unit, it can identify and process different types of data (such as vibration data, temperature data, current data, etc.), thus adapting to the data characteristics collected by different sensors, further enhancing the versatility of the system and avoiding the need to build separate models for each equipment; the model training unit automatically analyzes the data types and transmits the corresponding data to the initial fault diagnosis model for training, reducing manual intervention, significantly improving the automation level of the fault diagnosis system, and reducing the dependence of small and medium-sized enterprises on a large amount of manually labeled data; by generating corresponding target fault diagnosis models for each data type, it can quickly adapt to different equipment types and data characteristics without redesigning or adjusting the entire diagnosis process, thereby improving the diagnosis efficiency; since the system generates corresponding target fault diagnosis models according to different data types, each target fault diagnosis model can better capture the characteristics and fault patterns of specific types of data. Compared with the method of using a single model to process all data, this method can significantly improve the accuracy of the diagnosis results; the fault diagnosis unit selects the corresponding target fault diagnosis model according to the data type of the operation time series data to be diagnosed, ensuring that the diagnosis process is always based on the model most suitable for this data type, thus avoiding misjudgment problems caused by using a unified standard; by integrating the operation data of multiple types of industrial equipment, the system can analyze the equipment status from multiple dimensions, further improving the reliability of the diagnosis results.
[0185] In an embodiment of the present application, a method for fault diagnosis of industrial equipment based on multi-type time series data analysis is further provided. Figure 6 The flowchart of a method for fault diagnosis of industrial equipment based on multi-type time series data analysis provided by an embodiment of the present invention is shown in Figure 6 This method includes the following steps:
[0186] S610. Acquire the operation data of multiple types of industrial equipment, where the operation data includes multiple operation time series data to be trained and multiple operation time series data to be diagnosed;
[0187] S620. Analyze multiple data types corresponding to the multiple operation time series data to be trained, and transmit the operation time series data to be trained corresponding to each data type to a pre-constructed initial fault diagnosis model for training to obtain a target fault diagnosis model corresponding to each data type;
[0188] S630. Analyze the multiple data types corresponding to the multiple to-be-diagnosed operation time series data, and transmit the to-be-diagnosed operation time series data to the corresponding target fault diagnosis model according to the data types to obtain a fault diagnosis result.
[0189] Through the industrial equipment fault diagnosis method based on multi-class time series data analysis in the embodiments of the present application, by obtaining the operation data of multi-class industrial equipment, it can receive the operation data from different types of industrial equipment, not limited to a certain specific type of equipment, and has wide applicability, being applicable to various industrial scenarios such as wind turbines, port cranes, chemical equipment, and power transmission equipment; by analyzing the multiple data types corresponding to the multiple to-be-trained operation time series data, it can identify and process different types of data (such as vibration data, temperature data, current data, etc.), thus adapting to the data characteristics collected by different sensors, further enhancing the generality and avoiding the need for separate modeling for each device; by automatically analyzing the data types and transmitting the corresponding data to the initial fault diagnosis model for training, it reduces manual intervention, significantly improves the degree of automation, and reduces the dependence of small and medium-sized enterprises on a large amount of manually labeled data; by generating a corresponding target fault diagnosis model for each data type, it can quickly adapt to different device types and data characteristics without the need to redesign or adjust the entire diagnosis process, thereby improving the diagnosis efficiency; by generating a corresponding target fault diagnosis model according to different data types, each target fault diagnosis model can better capture the characteristics and fault patterns of specific types of data. Compared with the method of using a single model to process all data, this method can significantly improve the accuracy of the diagnosis result; by selecting the corresponding target fault diagnosis model according to the data type of the to-be-diagnosed operation time series data, it ensures that the diagnosis process is always based on the model most suitable for this data type, thus avoiding misjudgment problems caused by using a unified standard; by integrating the operation data of multi-class industrial equipment, it can analyze the equipment status from multiple dimensions and further improve the reliability of the diagnosis result.
[0190] For the specific embodiments of the industrial equipment fault diagnosis method based on multi-class time series data analysis in the present application, reference can be made to the examples shown in the above industrial equipment fault diagnosis system based on multi-class time series data analysis, and these examples will not be elaborated here.
[0191] It should be noted that: the above sequence of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous.
[0192] Each embodiment in this specification is described in a progressive manner, and the same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. An industrial equipment fault diagnosis system based on multi-class time series data analysis, characterized in that: The system comprises: A data acquisition unit, used to acquire operation data of various types of industrial equipment, wherein the operation data includes various operation time series data to be trained and various operation time series data to be diagnosed; A preprocessing unit is used to perform windowing and differential processing on the multiple types of running time series data to be trained to obtain preprocessed data to be trained, including: a frequency domain conversion subunit, used to perform frequency domain conversion on the multiple types of running time series data to be trained to obtain multiple types of running frequency domain data to be trained; a differential operation subunit, used to analyze the periodic performance of the multiple types of running frequency domain data to be trained, determine the periodic significance, and determine periodic data and non-periodic data according to the periodic significance; use the data whose periodic significance is greater than a preset periodic significance threshold as the periodic data, and use the data whose periodic significance is less than or equal to the preset periodic significance threshold as the non-periodic data. The non-periodic data; taking the distance mean of all extreme values on the envelope as the first window size, and performing windowing processing on the periodic data based on the first window size; performing clustering processing on the non-periodic data to obtain multiple clusters, dividing the non-periodic data into data corresponding to multiple states according to the distances of data points between the multiple clusters, and performing windowing processing on the data corresponding to each state; a model training unit, used to analyze multiple data types corresponding to the multiple types of running time series data to be trained, and transmitting the running time series data to be trained corresponding to each of the data types to a pre-constructed initial fault diagnosis model for training, so as to obtain a target fault diagnosis model corresponding to each of the data types; The fault diagnosis unit is used to analyze multiple data types corresponding to the multiple types of operating sequence data to be diagnosed, and transmit the operating sequence data to be diagnosed to the corresponding target fault diagnosis model according to the data types to obtain a fault diagnosis result.
2. The industrial equipment fault diagnosis system based on multi-type time series data analysis according to claim 1, characterized in that: The differential operation subunit is further used to perform differential operation on the periodic data and the non-periodic data after windowing; The preprocessing unit also includes a normalization subunit, which is used to perform normalization processing on the data after the difference operation to obtain the preprocessed data to be trained.
3. The industrial equipment fault diagnosis system based on multi-type time series data analysis according to claim 1, characterized in that: The step of analyzing the periodic performance of the plurality of frequency domain data to be trained and determining the periodic significance includes: Obtaining extreme values of the plurality of frequency domain data to be trained; Determining an envelope according to the extreme value; The periodic significance degree is determined according to the frequency domain distribution of the plurality of frequency domain data to be trained and the distance distribution of the extreme values corresponding to the envelope.
4. The industrial equipment fault diagnosis system based on multi-type time series data analysis according to claim 1, characterized in that: Performing a differential operation on the non-periodic data after the windowing process, including: For multiple data segments in the data corresponding to any of the states, when forward differencing is performed on the data segments and the data volume of the data segments is greater than the data volume of the adjacent previous data segments, the data segments are divided into data segments with equal lengths to the adjacent previous data segments for differential operations, and the remaining parts are aligned for differential operations.
5. The industrial equipment fault diagnosis system based on multi-type time series data analysis according to claim 1, characterized in that: The initial fault diagnosis model includes a long short-term memory network model, and the model training unit includes: A data classification subunit, configured to analyze the noise manifestation and fault manifestation of the plurality of training running time series data, and determine the plurality of data types according to the analysis results; The model training subunit is used to transmit the to-be-trained running time sequence data corresponding to each of the data types to the pre-built long short-term memory network model for training, so as to obtain the target fault diagnosis model corresponding to each of the data types.
6. The industrial equipment fault diagnosis system based on multi-type time series data analysis according to claim 5, characterized in that: The analyzing the noise manifestation and fault manifestation of the plurality of training running time series data, and determining the plurality of data types according to the analysis results, includes: For any of the training running time series data, a first classification window is obtained according to a preset second window size, the first classification window is moved, a second classification window is obtained based on new data, and data changes in the first classification window and the second classification window are analyzed to determine data change consistency; Based on the consistency of the data changes, analyzing the data changes in the first classification window to determine the degree of fault manifestation; Perform similarity comparison on the data in the window with the largest degree of fault manifestation in the two types of running time series data to be trained, and determine a combined training index; The multiple types of running time series data to be trained are merged according to the size of the merged training index to determine the multiple types of data.
7. The industrial equipment fault diagnosis system based on multi-type time series data analysis according to claim 6, characterized in that: The step of transmitting the to-be-trained running time series data corresponding to each of the data types to the pre-built long short-term memory network model for training to obtain the target fault diagnosis model corresponding to each of the data types includes: Extracting a training data set from the runtime data to be trained corresponding to each of the data types; The training data set is transmitted to the pre-built long short-term memory network model, the possibility of the fault state is output, and the fault state with the greatest possibility of the fault state is taken as the fault state corresponding to the training data set, wherein the fault state includes a normal state and fault states corresponding to multiple fault types.
8. The industrial equipment fault diagnosis system based on multi-type time series data analysis according to claim 6, characterized in that: The analyzing of the multiple data types corresponding to the multiple types of running time series data to be diagnosed includes: Performing windowing and differential processing on the plurality of to-be-diagnosed running time series data to obtain pre-processed to-be-diagnosed data; For any of the data to be diagnosed, determine the consistency of data changes, analyze the data changes based on the consistency of data changes, determine the degree of fault manifestation, perform similarity comparison on the data in the window with the largest degree of fault manifestation in the two data to be diagnosed, and determine the combined training index; The multiple types of data to be diagnosed are merged according to the size of the merged training index to determine the multiple types of data.
Citation Information
Patent Citations
Server multi-performance index anomaly detection method and device based on time sequence
CN115412455A
Abnormality processing method based on time series data, network equipment and readable storage medium
CN115495274A
Moving-device fault detection method and system
WO2024212832A1