Feature Extraction Method and Fault Diagnosis Method for Multi-Source Heterogeneous Data of Large Rotating Machinery
By performing feature extraction and fusion processing on text, tables and timing data associated with large rotating machinery, low-dimensional fusion feature vectors are constructed, which solves the problem of insufficient utilization of multi-source heterogeneous data in the prior art, and improves the accuracy of fault diagnosis and lifetime prediction.
Patent Information
- Application Number
- CN202210771293.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-06-30
AI Technical Summary
The prior art is difficult to effectively integrate text data, tabular data and timing data associated with large rotating machinery, resulting in the failure to fully utilize the operating status information and maintenance value information of these multi-source heterogeneous data in healthy operation and maintenance applications such as fault diagnosis and life prediction.
The multi-source heterogeneous data feature extraction method is adopted. By performing sentence-based and word-partitioning processing on text data, and using the Bert model for word embedding encoding, the table data is subjected to sentence-based and word-partitioning processing is performed for word embedding encoding, the timing data is segmented and the pre-trained autoencoder is used for encoding processing. Finally, the feature representation vectors of different types of data are spliced and fusion and dimensional reduction encoding processing are processed to construct low-dimensional fusion feature vectors.
Fusion data feature mining and extraction of multi-source heterogeneous data of large rotating machinery is realized, so that the extracted features can more fully present the operating status information of the equipment and maintenance value information, thereby improving the accuracy of fault diagnosis and life prediction.
Smart Images

Figure CN115062720B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of engineering applications and industrial big data, and particularly relates to a method for extracting multi-source heterogeneous data features and a fault diagnosis method for large rotating machinery. Background Art
[0002] Large rotating units are core equipment serving in the main battlefield of the national economy and the national defense field, especially aero-engines and gas turbines. Due to long-term operation in extremely harsh environments such as high temperature, high pressure, high speed and strong vibration, they are extremely prone to failures, resulting in significant economic losses and even serious safety accidents. According to relevant data statistics, aero-engine failures occur frequently in the civil aviation field, exceeding 1 / 3 of all aircraft failures; the maintenance support costs of gas compressors almost account for 60% of the total life cycle costs of the units. At the same time, large rotating units have complex structures and numerous fault causes, leading to difficult maintenance decisions. Therefore, developing health operation and maintenance application technologies such as fault diagnosis and life prediction for large rotating mechanical equipment, ensuring the safe operation of the units and reducing the maintenance costs of the units are of great significance to the national economy and national defense security.
[0003] In recent years, driven by the big data environment, deep learning and various neural network models have developed rapidly. In the field of fault diagnosis and prediction of large rotating machinery, intelligent operation and maintenance methods based on deep learning technology have eliminated the dependence on precise physical models and rich signal processing experience, attracting wide attention, and scholars have carried out a large number of studies. For example, Zhang Xiangyang et al. proposed a monitoring method based on the characteristics of casing vibration signals, combined three vibration signal preprocessing methods including matrix graph method, kurtosis graph method and wavelet scale spectrum method, and used a convolutional neural network to adaptively extract fault features to achieve fault monitoring and identification; Memarzadeh et al. proposed a fault monitoring model based on interpretable deep learning and used semi-supervised training to perform multi-class anomaly detection in flight data; Bleu-Laine et al. proposed a multi-fault classifier based on multi-instance learning (MIL) and multi-head convolutional neural network-recurrent neural network (MCNN-RNN) to achieve the prediction of aircraft adverse events and their precursors; Tayarani et al. proposed a gas turbine fault diagnosis based on multi-layer perceptron (MLP), dynamic neural network (DNM), and time-delay neural network (TDMM), and used the residual signals generated by DNM and TDMM as the input of MLP to complete the fault isolation of a twin-spool gas turbine engine; Peng Jun et al. proposed an engine gas path fault diagnosis based on deep belief neural network and used the deep belief network algorithm to solve the fault data of the performance degradation of aircraft engine components generated by simulation software; Shen et al. used a fully convolutional network (FCN) to automatically identify and locate the damage of aircraft engine pipeline mirror images; Mosallam et al. selected sensitive signals from multi-sensor signals using unsupervised information metrics, and then used principal component analysis and empirical mode decomposition to extract the principal component degradation trend from the multi-source signal feature set as a health indicator to predict the remaining life of an aeroengine. Ragab et al. proposed a remaining life prediction method based on Kaplan-Meier survival analysis, using both time data and condition monitoring data, and applied it to the remaining life prediction of aeroengines.
[0004] Driven by the development of new technologies and health management service models, the demand for multi-source heterogeneous data collection and in-depth analysis of massive data is increasing. There are numerous and complex types of data sources available for the monitoring and diagnosis of large rotating equipment, including both unstructured data represented by operation logs and monitoring time series, and structured data represented by work order forms. The above-mentioned deep learning-based operation and maintenance methods in the prior art have achieved certain results, but they mainly use the time series data obtained by monitoring means such as vibration, temperature, and pressure for large rotating machinery as feature data to perform applications such as life prediction and fault diagnosis of large rotating machinery. However, in addition to vibration, temperature, and pressure time series data, large rotating mechanical equipment will also generate a large amount of operation and maintenance historical record texts and historical data tables and other data sources in different dimensions during its service life. These data sources in different dimensions can more fully and detailedly present the operation state information and maintenance value information of large rotating mechanical equipment; however, compared with time series data, there are significant differences in the data structure characteristics of these text data and table data as data bodies, forming multi-source heterogeneous data, which is difficult to directly extract and analyze data features in a unified dimension in data analysis. Therefore, in the application of the operation and maintenance methods of large rotating mechanical equipment in the prior art, the text data and table data generated by large rotating mechanical equipment during its service life have not been fully mined and utilized. Summary of the Invention
[0005] Aiming at the deficiencies of the above prior art, the actual problem to be solved by the present invention is: how to provide a method for extracting multi-source heterogeneous data features of large rotating machinery to better mine and extract the fusion data features of the text data, table data, and time series data associated with large rotating machinery, so that the extracted multi-source heterogeneous data features can more fully present the operation state information and maintenance value information of large rotating mechanical equipment to help more accurately perform health operation and maintenance applications such as fault diagnosis and life prediction of large rotating mechanical equipment.
[0006] To solve the above technical problems, the present invention adopts the following technical solutions:
[0007] A method for extracting multi-source heterogeneous data features of large rotating machinery, comprising the following steps:
[0008] S1: Obtain the multi-source heterogeneous data of large rotating machinery; the multi-source heterogeneous data of the large rotating machinery includes text data, table data, and time series data associated with the large rotating machinery;
[0009] S2: Respectively perform sentence splitting and word segmentation processing on the text information of the text data and the text information of each cell in the table data to obtain the corresponding sentence splitting and word segmentation information;
[0010] S3: Perform word embedding encoding on the clause and word segmentation information of the text data, and use the obtained word encoding vectors of the text data as the feature representation vectors of the text data;
[0011] S4: Perform word embedding encoding on the clause and word segmentation information of each cell in the table data respectively, and splice and fuse the obtained word encoding vectors of each cell in the table data to obtain the encoding vector matrix of the table data as the feature representation vector of the table data;
[0012] S5: Perform segmented cutting on the time series data, and perform encoding processing on each time series data segment obtained by cutting the time series data using a pre-trained autoencoder respectively, and then splice and fuse them to obtain the encoding vector of the time series data as the feature representation vector of the time series data;
[0013] S6: Splice, fuse and perform dimensionality reduction encoding processing on the feature representation vectors of the text data, table data and time series data associated with the large rotating machinery, and use the obtained low-dimensional fusion feature vectors as the multi-source heterogeneous data feature vectors of the large rotating machinery.
[0014] In the above method for extracting multi-source heterogeneous data features of large rotating machinery, as a preferred solution, the step S2 specifically includes:
[0015] S201: Perform clause processing on the text information of the text data and the text information of each cell in the table data respectively to obtain the sentence segments of each text information clause;
[0016] S202: Perform word segmentation processing on each sentence segment of the text information respectively to obtain the feature words included in each sentence segment;
[0017] S203: Use the set of feature words included in each sentence segment of the text information in the text data as the clause and word segmentation information of the text data; use the set of feature words included in each sentence segment of the text information in each cell of the table data as the clause and word segmentation information of the corresponding cell.
[0018] In the above method for extracting multi-source heterogeneous data features of large rotating machinery, as a preferred solution, before performing clause processing on the text information of the text data and the text information of each cell in the table data in the step S201, it further includes:
[0019] Perform text preprocessing on the text information of the text data and the text information of each cell in the table data, and the text preprocessing includes one or more of misspelling correction processing, error symbol correction processing, error grammar correction processing, stop word removal processing, and synonym expression consistency processing of the text information.
[0020] In the above method for extracting multi-source heterogeneous data features of large rotating machinery, as a preferred solution, the step S3 specifically includes:
[0021] S301: For each feature word included in each sentence segment in the sentence segmentation and word segmentation information of the text data, use the Bert model to perform word embedding encoding respectively to obtain a 1×B dimensional word encoding vector for each feature word, where B is the encoding dimension size of the word embedding encoding by the Bert model;
[0022] S302: For a single text data, use the concat method to splice and fuse the word encoding vectors of the feature words included in each sentence segment in the sentence segmentation and word segmentation information of the text data to obtain the dimensional word encoding vector of the text data as the feature representation vector of the text data; where, m w represents the number of sentence segments obtained by sentence segmentation of the text data, and n w,i represents the number of feature words included in the i-th sentence segment of the text data.
[0023] In the above method for extracting features of multi-source heterogeneous data of large rotating machinery, as a preferred solution, the step S4 specifically includes:
[0024] S401: For each feature word included in each sentence segment in the sentence segmentation and word segmentation information of each cell of the table data, use the Bert model to perform word embedding encoding respectively to obtain a 1×B dimensional word encoding vector for each feature word, where B is the encoding dimension size of the word embedding encoding by the Bert model;
[0025] S402: For a single cell in the table data, use the concat method to splice and fuse the word encoding vectors of the feature words included in each sentence segment in the sentence segmentation and word segmentation information of the cell to obtain the dimensional word encoding vector of the cell; where, m c represents the number of sentence segments obtained by sentence segmentation of the text information in a single cell, and n c,i represents the number of feature words included in the i-th sentence segment of the text information in a single cell;
[0026] S403: For each cell of the N tuples × M field attributes included in the table data, first use the concat method to splice and fuse the word encoding vectors of the cells with M different field attributes in the same tuple to obtain the dimensional tuple encoding vector; then, taking the tuple as a unit, splice and fuse the tuple encoding vectors of the N different tuples included in the table data to obtain the dimensional encoding vector matrix of the table data as the feature representation vector of the table data.
[0027] In the above method for extracting features of multi-source heterogeneous data of large rotating machinery, as a preferred solution, the step S5 specifically includes:
[0028] S501: Segment and cut the time series data according to the set segment length to obtain each time series segment of the segmented and cut time series data;
[0029] S502: Use the encoding dimension size B of the word embedding encoding for clause segmentation and word segmentation information as the encoding dimension size of the autoencoder, and use the pre-trained autoencoder to perform encoding processing on each time series segment of the time series data respectively, and obtain the 1×B-dimensional data segment encoding vector of each time series segment;
[0030] S503: For a single time series data, splice and fuse the data segment encoding vectors of each time series segment of the time series data through the concat method to obtain the m t ×B-dimensional encoding vector of the time series data as the feature representation vector of the time series data; where, m t represents the number of time series segments obtained by segmenting and cutting the time series data.
[0031] In the above method for extracting features of multi-source heterogeneous data of large rotating machinery, as an optimal solution, the autoencoder is trained through the following steps:
[0032] Step 5021: Obtain multiple sample time series data of large rotating machinery from the multi-source heterogeneous database;
[0033] Step 5022: Segment and cut each sample time series data according to the set segment length to obtain each time series segment of the segmented and cut sample time series data as the sample time series data set;
[0034] Step 5023: Select training samples and test samples from the sample time series data set according to the set training and test ratio to obtain the training sample set and the test sample set;
[0035] Step 5024: Use the training sample set and the test sample set as the input of the autoencoder, and use minimizing the mean square loss as the training objective to perform unsupervised learning training on the autoencoder;
[0036] The network model of the autoencoder includes an encoding layer and a decoding and verification layer; among them, the encoding layer of the autoencoder contains 5 Linear layers, and a single time series segment of the sample time series data obtains a 1×B-dimensional data segment encoding vector through the encoding layer of the autoencoder; the decoding and verification layer of the autoencoder contains 5 Linear layers, and the 1×B-dimensional data segment encoding vector obtained by the encoding layer is decoded and restored into a time series segment through the decoding and verification layer of the autoencoder for comparison and verification with the original time series segment;
[0037] Step 5025: After completing the unsupervised learning training, obtain the trained autoencoder.
[0038] In the above method for extracting multi-source heterogeneous data features of large rotating machinery, as a preferred solution, step S6 specifically includes:
[0039] S601: Use the concat method to splice and fuse the -dimensional feature representation vectors of text data, the -dimensional feature representation vectors of tabular data, and the m t ×B-dimensional feature representation vectors of time series data to obtain an -dimensional fused feature representation matrix;
[0040] Among them, B represents the encoding dimension size for word embedding encoding; m w represents the number of sentence segments obtained by splitting sentences in text data, n w,i represents the number of feature words in the i-th sentence segment of text data; m c represents the number of sentence segments obtained by splitting sentences in the text information of a single cell in tabular data, n c,i represents the number of feature words in the i-th sentence segment of the text information in a single cell, N represents the number of tuples in tabular data, and M represents the number of field attributes in tabular data; m t represents the number of time series data segments obtained by segmenting and cutting time series data;
[0041] S602: Input the fused feature representation matrix obtained by fusion into a pre-trained dimensionality reduction encoding model, and use the 1×D B -dimensional low-dimensional fused feature vector output by the dimensionality reduction encoding model as the multi-source heterogeneous data feature vector of large rotating machinery; among them, D B represents the dimensionality reduction encoding dimension size of the dimensionality reduction encoding model.
[0042] In the above method for extracting multi-source heterogeneous data features of large rotating machinery, as a preferred solution, the dimensionality reduction encoding model is trained through the following steps:
[0043] Step 6021: Obtain multiple groups of sample text data, sample tabular data, and sample time series data related to large rotating machinery from a multi-source heterogeneous database;
[0044] Step 6022: Process each group of sample text data, sample tabular data, and sample time series data respectively to obtain the -dimensional feature representation vectors of the sample text data in each group, the -dimensional feature representation vectors of the sample tabular data, and the m t ×B-dimensional feature representation vectors of the sample time series data, and splice and fuse them to obtain the corresponding -dimensional fused feature representation matrix as the sample data set;
[0045] Step 6023: Select training samples and test samples from the sample dataset according to the set training and test ratio to obtain a training sample set and a test sample set;
[0046] Step 6024: Use the training sample set and the test sample set as the input of the dimensionality reduction and encoding model, and use minimizing the mean square loss as the training objective to perform unsupervised learning training on the dimensionality reduction and encoding model;
[0047] The dimensionality reduction and encoding model includes an encoding layer and a decoding and verification layer; among them, the encoding layer of the dimensionality reduction and encoding model contains 1 Linear layer, 3 deconvolution operator convolutional layers and 1 residual module, The fusion feature representation matrix of dimension 1×D is obtained through the encoding layer of the dimensionality reduction and encoding model to obtain a low-dimensional fusion feature vector of dimension 1×D; the decoding and verification layer of the dimensionality reduction and encoding model contains 5 Linear layers, and the 1×D dimensional low-dimensional fusion feature vector obtained from the encoding layer is re-decoded and restored into B a fusion feature representation matrix of dimension 1×D, which is used to compare and verify with the original fusion feature representation matrix; B Step 6025: After completing the unsupervised learning training, obtain the trained dimensionality reduction and encoding model. a fusion feature representation matrix of dimension 1×D, which is used to compare and verify with the original fusion feature representation matrix;
[0048] Step 6025: After completing the unsupervised learning training, obtain the trained dimensionality reduction and encoding model.
[0049] Correspondingly, the present invention also provides a fault diagnosis method for large rotating machinery, including the following steps:
[0050] Step A: Obtain multi-source heterogeneous data of the large rotating machinery to be detected, and use the multi-source heterogeneous data feature extraction method described in any one of claims 1 to 8 to perform feature extraction to obtain a multi-source heterogeneous data feature vector of the large rotating machinery to be detected;
[0051] Step B: Input the multi-source heterogeneous data feature vector of the large rotating machinery to be detected into the trained fault classification and recognition model, and output the predicted diagnosis result of the fault category of the large rotating machinery to be detected.
[0052] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0053] 1. The method for extracting multi-source heterogeneous data features of large rotating machinery according to the present invention adopts different data feature encoding methods for the text data, tabular data, and time-series data associated with large rotating machinery. After performing sentence splitting and word segmentation on the text data and tabular data, word embedding encoding is carried out. After segmenting the time-series data, auto-encoding is performed, so that the text data, tabular data, and time-series data are all converted into encoded vector forms with a unified data dimension, serving as their respective feature representation vectors, and better retaining the operating state information and maintenance value information carried by each of the three. Furthermore, the encoded vectors of the three can be further spliced, fused, and dimension-reduced encoded under the unified data dimension, constructing a low-dimensional fusion feature vector that retains the operating state information and maintenance value information carried by the three, serving as the multi-source heterogeneous data feature vector of large rotating machinery, and realizing the extraction of fusion data features of the text data, tabular data, and time-series data associated with large rotating machinery.
[0054] 2. The multi-source heterogeneous data feature vector of large rotating machinery extracted by the method of the present invention is used as the feature data for health operation and maintenance applications such as fault diagnosis and life prediction of large rotating machinery equipment. Since the multi-source heterogeneous data feature vector originates from multiple data source dimensions, it can more fully present the operating state information and maintenance value information of large rotating machinery equipment, and thus can better improve the accuracy of health operation and maintenance applications such as fault diagnosis and life prediction of large rotating machinery equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to make the objectives, technical solutions, and advantages of the invention clearer, the present invention will be further described in detail below with reference to the drawings, where:
[0056] Figure 1 is the flowchart of the method for extracting multi-source heterogeneous data features of large rotating machinery according to the present invention.
[0057] Figure 2 is the flowchart of a detailed process example of the method for extracting multi-source heterogeneous data features of large rotating machinery according to the present invention.
[0058] Figure 3 is the schematic diagram of the principle of extracting sentence data by the Bert model.
[0059] Figure 4 is the structural example diagram of the autoencoder AE model.
[0060] Figure 5 is the schematic diagram of the principle of the deconvolution operator.
[0061] Figure 6 is the schematic diagram of the operation process of generating the deconvolution operator.
[0062] Figure 7It is a structural example diagram of an AE model based on a deconvolution operator. Specific implementation manners
[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention usually described and illustrated in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0064] As Figure 1 shown, the present invention discloses a method for extracting features of multi-source heterogeneous data of large rotating machinery, including the following steps:
[0065] S1: Obtain multi-source heterogeneous data of large rotating machinery; the multi-source heterogeneous data of the large rotating machinery includes text data, table data, and time-series data associated with the large rotating machinery;
[0066] S2: Respectively perform sentence splitting and word segmentation processing on the text information of the text data and the text information of each cell in the table data to obtain corresponding sentence splitting and word segmentation information;
[0067] S3: Perform word embedding encoding on the sentence splitting and word segmentation information of the text data, and use the obtained word encoding vector of the text data as the feature representation vector of the text data;
[0068] S4: Perform word embedding encoding on the sentence splitting and word segmentation information of each cell in the table data respectively, and splice and fuse the obtained word encoding vectors of each cell in the table data to obtain an encoding vector matrix of the table data as the feature representation vector of the table data;
[0069] S5: Perform segmented cutting on the time-series data, and perform encoding processing on each time-series data segment obtained by cutting the time-series data using a pre-trained autoencoder respectively and then splice and fuse them to obtain an encoding vector of the time-series data as the feature representation vector of the time-series data;
[0070] S6: Splice, fuse, and perform dimensionality reduction encoding processing on the feature representation vectors of the text data, table data, and time-series data associated with the large rotating machinery, and use the obtained low-dimensional fusion feature vector as the multi-source heterogeneous data feature vector of the large rotating machinery.
[0071] The multi-source heterogeneous data feature extraction method for large rotating machinery of the present invention adopts different data feature encoding methods for the text data, table data, and time series data associated with large rotating machinery. After performing sentence splitting and word segmentation on the text data and table data, word embedding encoding is carried out. After segmenting the time series data, auto-encoding is performed, so that the text data, table data, and time series data are all converted into encoded vector forms with a unified data dimension, serving as their respective feature representation vectors, and the operation state information and maintenance value information carried by the three are better retained. Furthermore, the encoded vectors of the three can be further spliced, fused, and dimension-reduced encoded under the unified data dimension to construct a low-dimensional fusion feature vector that retains the operation state information and maintenance value information carried by the three, serving as the multi-source heterogeneous data feature vector of large rotating machinery, realizing the extraction of fusion data features of the text data, table data, and time series data associated with large rotating machinery.
[0072] Taking the multi-source heterogeneous data feature vector of large rotating machinery extracted by the method of the present invention as the feature data for health operation and maintenance applications such as fault diagnosis and life prediction of large rotating machinery equipment, since the multi-source heterogeneous data feature vector comes from multiple data source dimensions, it can more fully present the operation state information and maintenance value information of large rotating machinery equipment, and thus can better improve the accuracy of health operation and maintenance applications such as fault diagnosis and life prediction of large rotating machinery equipment.
[0073] Figure 2 The flowchart shows a detailed process example of the multi-source heterogeneous data feature extraction method for large rotating machinery of the present invention; next, taking this as an example, the multi-source heterogeneous data feature extraction method of the present invention will be described in more detail.
[0074] In specific implementation, in step S1, taking the CMS system of a wind farm as an example, time series data such as vibration data and SCADA data of large rotating machinery such as wind turbines in the wind farm CMS system can be collected, as well as text data such as relevant text descriptions and text records, and table data such as table descriptions and table records; for the collected text data, table data, and time series data, through manual processing, annotation can be carried out according to the relevance to large rotating machinery, or the text data, table data, and time series data associated with large rotating machinery can be classified into a folder, and in these ways, the association relationships between the text data, table data, and time series data and large rotating machinery are recorded, facilitating subsequent feature fusion extraction of the associated text-table-time series data.
[0075] In specific implementation, step S2 specifically includes:
[0076] S201: Sentence processing is performed on the text information of the text data and the text information of each cell in the table data to obtain the sentence segments of each text information sentence;
[0077] S202: performing word segmentation processing on each sentence segment of the text information to obtain feature words contained in each sentence segment;
[0078] S203: taking a set of characteristic words contained in each sentence of the text information in the text data as sentence segmentation information of the text data; taking a set of characteristic words contained in each sentence of the text information in each cell of the table data as sentence segmentation information of the corresponding cell.
[0079] In view of the heterogeneity of tabular data, the semantic similarity of data is more important than the character similarity of data. Therefore, the text data and tabular data can be encoded by word embedding through the pre-trained Bert model and converted into vector form; the maximum number of input characters of the Bert model determines the encoding dimension size of word embedding encoding. If the text information (text, characters, data and other types of information) contained in the text data and tabular data is too long, some text information may be lost after word embedding encoding. Therefore, it is necessary to first perform sentence and word segmentation on the text information of the text data and the text information of each cell in the tabular data.
[0080] The sentence and word segmentation processing of text data is relatively mature. Usually, the sentence segmentation is performed with punctuation marks such as ".", "?", "!" and ";" used for sentence segmentation as separators; for the Chinese and English character features in the sentence, the word segmentation method of the prior art can be used for processing. In addition, before the text information of the text data and the text information of each cell in the table data are processed by sentence segmentation, the text information of the text data and the text information of each cell in the table data can also be pre-processed. The text pre-processing includes the correction of typos, the correction of incorrect symbols, the correction of incorrect grammar, the removal of stop words, the consistency of synonym expression, etc., and one or more text pre-processing operations can be selected according to the actual situation and needs of the text information. When segmenting words, Chinese is segmented by characters, and English is segmented by words. By querying the dictionary, typos, incorrect symbols, incorrect grammar, synonym replacement can be performed to make the expression consistent, and redundant and meaningless stop words can be removed. Pre-processing, only feature words with semantically relevant information are retained, providing a relatively pure data space for subsequent feature extraction. For example, the data in data.txt is "The intermediate-stage bearing is worn, the high-speed shaft outputs the large gear with unbalanced load, and the meshing is uneven." If a single character is used as the feature word for word segmentation, the feature word data obtained by the word segmentation are "middle", "intermediate", "level", "shaft", "bearing", "wear", "loss", "high", "speed", "shaft", "output", "large", "tooth", "wheel", "unbalanced", "load", "meshing", "engagement", "uneven", and "uniform".
[0081] Of course, word segmentation processing can also use words, phrases, etc. as word segmentation units.
[0082] Tabular data is usually two-dimensional structured data containing tuple dimensions and field attribute dimensions. Without loss of generality, the header of the tabular data can be recorded as {A_1, A_2, .., A_M}, the tuple is recorded as t_i, and t_i[A_k] represents the value of the i-th tuple on the attribute A_k. In view of the structural characteristics of tabular data, the tabular data is first split into cells, that is, a single tabular data is divided into multiple t_i[A_k], and the cell data t_i[A_k] is regarded as a single text data for subsequent operations. For example, the structure and content of the data.xlsx tabular data are shown in Table 1:
[0083] Table 1
[0084] Fan failure number Failure time Failed component #1 December 25, 2018 High-speed shaft of planetary gearbox #2 December 26, 2019 Fan blade
[0085] The table data can be divided into 9 cell data by tuple and header.
[0086] In specific implementation, step S3 specifically includes:
[0087] S301: For each feature word included in each sentence segment in the sentence segmentation and word segmentation information of the text data, use the Bert model to perform word embedding encoding respectively to obtain a 1×B dimensional word encoding vector for each feature word, where B is the encoding dimension size of the Bert model for word embedding encoding;
[0088] S302: For a single text data, use the concat method to splice and fuse the word encoding vectors of the feature words included in each sentence segment in the sentence segmentation and word segmentation information of the text data to obtain a dimensional word encoding vector of the text data as the feature representation vector of the text data; where, m w represents the number of sentence segments obtained by segmenting the text data, and n w,i represents the number of feature words included in the i-th sentence segment of the text data.
[0089] In specific application implementation, according to the sentence structure of the original data, a sample set can be constructed with the original sentence pattern structure, and the sample set is based on a single text or tabular data as a unit; then set the batch size, and send the sample set into the Bert model in batches for pre-training, and obtain its word vector representation through word embedding query. After training, the obtained Bert model can be applied to the word embedding encoding of text data and tabular data. The concat method is a mature technical method for connecting two or more data vectors, used for splicing and connecting two or more data vectors.
[0090] Taking the encoding dimension size B of the Bert model for word embedding encoding as 768 as an example, the size of the word encoding vector obtained by word embedding encoding is 1×768 dimensions. In this way, for a single text data, the final word encoding vector of the text data is represented as a dimensional matrix vector, where, m w represents the number of sentence segments obtained by segmenting the text data, and n w,i represents the number of feature words included in the i-th sentence segment of the text data.
[0091] For example, as Figure 3 shown, for the text data data.txt "The intermediate bearing is worn, the output large gear of the high-speed shaft is eccentrically loaded, and the meshing is uneven.", a 22×768 dimensional word encoding vector can be finally obtained through the Bert model. This vector, as the feature representation vector of the text data data.txt, not only contains the character information and semantic information of the global text, but also contains the position information of each character, and has sufficient expressive ability.
[0092] In specific implementation, step S4 specifically includes:
[0093] S401: For each feature word included in each sentence segment in the sentence segmentation and word segmentation information of each cell of the tabular data, use the Bert model to perform word embedding encoding respectively to obtain a 1×B dimensional word encoding vector for each feature word, where B is the encoding dimension size of the Bert model for word embedding encoding;
[0094] S402: For a single cell in the tabular data, use the concat method to splice and fuse the word encoding vectors of the feature words included in each sentence segment in the sentence segmentation and word segmentation information of the cell to obtain the dimensional word encoding vector of the cell; where, m c represents the number of sentence segments obtained by sentence segmentation of the text information in a single cell, and n c,i represents the number of feature words included in the i-th sentence segment of the text information in a single cell;
[0095] S403: For each cell of the N tuples × M field attributes included in the tabular data, first use the concat method to splice and fuse the word encoding vectors of the cells with M different field attributes in the same tuple to obtain the dimensional tuple encoding vector; then, taking the tuple as a unit, splice and fuse the tuple encoding vectors of the N different tuples included in the tabular data to obtain the dimensional encoding vector matrix of the tabular data as the feature representation vector of the tabular data.
[0096] In a specific application implementation, taking the encoding dimension size B of the word embedding encoding by the Bert model as 768 for the N×M cell data split from a single tabular data, N×M dimensional word encoding vectors can be obtained through step 3, where M represents the number of field attributes included in the tabular data; m t represents the number of time series data segments obtained by segmenting and cutting the time series data. Then, use the concat method to splice and fuse the word encoding vectors of the cells with M different field attributes in the same tuple to obtain the dimensional tuple encoding vector; then, taking the tuple as a unit, splice and fuse the tuple encoding vectors of the N different tuples included in the tabular data to obtain the dimensional encoding vector matrix of the tabular data as the feature representation vector of the tabular data.
[0097] For example, the data in the data.xlsx table shown in Table 1 is preprocessed and divided into 9 cell data. The word encoding vectors obtained by inputting them into the Bert model are word encoding vectors of 6×768 dimensions, 4×768 dimensions, 4×768 dimensions, 2×768 dimensions, 6×768 dimensions, 8×768 dimensions, 2×768 dimensions, 6×768 dimensions, and 4×768 dimensions respectively; then, according to the field attributes, the word encoding vectors of the cells with different field attributes in the same tuple are fused to obtain tuple encoding vectors of 14×768 dimensions, 16×768 dimensions, and 12×768 respectively; finally, the feature vectors are fused according to different tuples, and a 42×768-dimensional encoding vector matrix is finally obtained. This vector, as the feature representation vector of the data.xlsx table data, contains not only global information but also information about each tuple and field attribute.
[0098] Specifically in implementation, step S5 specifically includes:
[0099] S501: Segment and cut the time series data according to the set segment length to obtain each time series data segment after segment and cut of the time series data;
[0100] S502: Use the encoding dimension size B of the word embedding encoding for clause tokenization information as the encoding dimension size of the autoencoder, and use the pre-trained autoencoder to perform encoding processing on each time series data segment of the time series data respectively to obtain a 1×B-dimensional data segment encoding vector for each time series data segment;
[0101] S503: For a single time series data, splice and fuse the data segment encoding vectors of each time series data segment of the time series data through the concat method to obtain an m t ×B-dimensional encoding vector of the time series data as the feature representation vector of the time series data; where m t represents the number of time series data segments obtained by segment and cut of the time series data.
[0102] Among them, the autoencoder is trained through the following steps:
[0103] Step 5021: Obtain multiple sample time series data of large rotating machinery from a multi-source heterogeneous database;
[0104] Step 5022: Segment and cut each sample time series data according to the set segment length to obtain each time series data segment after segment and cut of each sample time series data as the sample time series data set;
[0105] Step 5023: Select training samples and test samples from the sample time series data set according to the set training and test ratio to obtain a training sample set and a test sample set;
[0106] Step 5024: Use the training sample set and the test sample set as the input of the autoencoder, and use minimizing the mean square loss as the training objective to perform unsupervised learning training on the autoencoder;
[0107] The network model of the autoencoder includes an encoding layer and a decoding verification layer; among them, the encoding layer of the autoencoder contains 5 Linear layers, and a single time series data segment of the sample time series data obtains a 1×B-dimensional data segment encoding vector through the encoding layer of the autoencoder; the decoding verification layer of the autoencoder contains 5 Linear layers, and the 1×B-dimensional data segment encoding vector obtained by the encoding layer is decoded and restored into a time series data segment through the decoding verification layer of the autoencoder for comparison and verification with the original time series data segment;
[0108] Step 5025: After completing the unsupervised learning training, obtain the trained autoencoder.
[0109] Taking the set segmentation length of 4096 data points as an example, for a single time series data, first slice and divide it according to 4096 data points to obtain multiple time series data segments. After all time series data are divided, the obtained several time series data segments are used as the sample set and divided into a training set and a test set at a ratio of 5:1.
[0110] Then construct the autoencoder AE model, and its model structure example is as Figure 4 shown. Among them, the Encoder layer (encoding layer) contains 5 Linear layers, and a single time series data segment finally obtains a 1×768-dimensional feature vector through the Encoder layer; the Decoder layer (decoding verification layer) contains 5 Linear layers, and the 1×768-dimensional feature vector obtained through the Encoder is decoded and restored into a 1×4096-dimensional input vector, and the AE model is trained and optimized by minimizing the MSE loss (mean square loss).
[0111] The calculation formula of the MSE loss function is as follows:
[0112]
[0113] Among them, y i and respectively represent the true value and the predicted value of the i-th sample, and m is the number of samples.
[0114] As a preferred parameter selection, the optimizer of the autoencoder AE model AE neural network is Adam, the learning rate is 1e -4 , weight-declay is 2e -5 , the batch size is 30, the Dropout random inactivation rate is 0.4, and a total of 20 epochs are trained.
[0115] Then, use the partitioned sample set as the input of the AE model, and train the AE model through unsupervised learning.
[0116] After completing the unsupervised learning training, use the trained AE model to extract features from the time series data. A time series data segment composed of a single 4096 points is extracted by the Encoder to obtain a 1×768-dimensional feature vector. The feature vectors of multiple time series data segments are concatenated by the concat method, and finally an m t ×768-dimensional encoded vector is obtained as the feature representation vector of the time series data.
[0117] For example, for the vibration time series data collected by the wind farm CMS system, the number of data points collected in one day is 32×4096 points. Take the collected data for two months, slice and block it to get 60×32 pieces of data as the sample set, and divide the training set and the test set according to the ratio of 5:1; the number of samples in the training set is 1600, and the number of samples in the test set is 320; input the sample set into the constructed AE model for training; finally, input the collected data for one day into the trained AE model respectively to obtain a 32×768-dimensional encoded vector as the feature representation vector of the vibration time series data collected on this day.
[0118] In specific implementation, step S6 specifically includes:
[0119] S601: Concatenate and fuse the -dimensional feature representation vector of the text data, the -dimensional feature representation vector of the tabular data, and the m t ×B-dimensional feature representation vector of the time series data through the concat method to obtain a -dimensional fused feature representation matrix;
[0120] Among them, B represents the encoding dimension size for word embedding encoding; m w represents the number of sentence segments obtained by splitting sentences of the text data, n w,i represents the number of feature words included in the i-th sentence segment of the text data; m c represents the number of sentence segments obtained by splitting sentences of the text information in a single cell of the tabular data, n c,i represents the number of feature words included in the i-th sentence segment of the text information in a single cell; N represents the number of tuples included in the tabular data, and M represents the number of field attributes included in the tabular data; m t represents the number of time series data segments obtained by segmenting and cutting the time series data;
[0121] S602: Input the fused feature representation matrix obtained by fusion into a pre-trained dimensionality reduction encoding model, and the 1×D output by the dimensionality reduction encoding model BThe low-dimensional fusion feature vector of dimension D is used as the feature vector of multi-source heterogeneous data of large rotating machinery; where D B represents the dimension size of the dimensionality reduction encoding of the dimensionality reduction encoding model.
[0122] Among them, the dimensionality reduction encoding model is obtained through the following steps:
[0123] Step 6021: Obtain multiple groups of sample text data, sample table data, and sample time series data associated with large rotating machinery from the multi-source heterogeneous database;
[0124] Step 6022: Process each group of sample text data, sample table data, and sample time series data respectively to obtain the dimensional feature representation vectors of the sample text data in each group, the dimensional feature representation vectors of the sample table data, and the m t ×B dimensional feature representation vectors of the sample time series data, and splice and fuse them to obtain the corresponding dimensional fusion feature representation matrix of each group as the sample data set;
[0125] Step 6023: Select training samples and test samples from the sample data set according to the set training and test ratio to obtain a training sample set and a test sample set;
[0126] Step 6024: Use the training sample set and the test sample set as the input of the dimensionality reduction encoding model, and use minimizing the mean square loss as the training objective to perform unsupervised learning training on the dimensionality reduction encoding model;
[0127] The dimensionality reduction encoding model includes an encoding layer and a decoding and verification layer; among them, the encoding layer of the dimensionality reduction encoding model contains 1 Linear layer, 3 deconvolution operator convolutional layers and 1 residual module, The dimensional fusion feature representation matrix of dimension is input into the encoding layer of the dimensionality reduction encoding model to obtain a 1×D B dimensional low-dimensional fusion feature vector; the decoding and verification layer of the dimensionality reduction encoding model contains 5 Linear layers, and the 1×D B dimensional low-dimensional fusion feature vector obtained from the encoding layer is decoded and restored into dimensional fusion feature representation matrix through the decoding and verification layer of the dimensionality reduction encoding model, which is used to compare and verify with the original fusion feature representation matrix;
[0128] Step 6025: After completing the unsupervised learning training, obtain the trained dimensionality reduction encoding model.
[0129] Similarly, taking the encoding dimension size B of the word embedding encoding as 768 as an example, the feature representation vectors of the associated text, table, and time-series data are obtained through the above steps. First, the 3 individual matrix vectors are concatenated and fused into 1 multi-dimensional matrix vector by the concat method as a single set of associated text-table-time series data samples, and its size is For example, through the above example, its Therefore, the size of the low-dimensional fusion feature vector of a single set of associated text-table-time series data obtained should be 96×768 dimensions.
[0130] Through the above step method, multiple sets of associated text-table-time series data samples can be obtained as a sample set, which is divided into a training set and a test set at a ratio of 5:1. For example, taking two months of collected data, text data, and table data, 1920 feature vectors of 96×768 are obtained through the above steps. Divide the training set and the test set according to the ratio of 5:1; the number of training set samples is 1600, and the number of test set samples is 320.
[0131] Then, construct a deconvolution operator AE model as a dimensionality reduction encoding model. Among them, use the deconvolution operator to construct a convolutional layer, and use the deconvolution operator convolutional layer to design the Encoder (encoding layer); specifically, the Encoder (encoding layer) contains 1 Linear layer, 3 deconvolution operator convolutional layers, and 1 residual module. A single set of associated text-table-time series data finally obtains a 1×1024-dimensional feature vector through the Encoder layer; the Decoder (decoding verification layer) contains 5 Linear layers, and the 1×1024-dimensional feature vector obtained through the Encoder is decoded and restored into a 96×768-dimensional input vector. Train and optimize the deconvolution operator AE model by minimizing the MSE loss (mean square loss).
[0132] Input the sample set into the dimensionality reduction encoding model for unsupervised learning.
[0133] Finally, after completing the unsupervised learning training, use the trained dimensionality reduction encoding model to extract the features of the associated text-table-time series data, and obtain the low-dimensional fusion feature vector of the text-table-time series data as the multi-source heterogeneous data feature vector of the corresponding group of text-table-time series data of the large rotating machinery. For example, the 96×768-dimensional text-table-time series data in the above example is finally obtained as a 1×1024-dimensional low-dimensional fusion feature vector through feature extraction by the dimensionality reduction encoding model.
[0134] The characteristics of the involution kernel are the opposite of convolution, with spatial specificity and channel invariance, that is, sharing the kernel in the channel dimension and using a spatially specific kernel in the spatial dimension for more flexible modeling. The involution kernel s where H×W represents the size of the feature map, K×K represents the size of the kernel, and G represents sharing G kernels for all channels. For a single involution kernel (i,j) is a pixel point The coordinates on the feature map, where C is the number of channels of the feature map. Thus, at different spatial positions, the sizes of the involution kernels are also different, and the formula for generating the involution kernel is as follows:
[0135]
[0136] where, ψ i,j is a set of indices in the neighborhood of the coordinates (i,j), then represents a certain patch on the feature map that contains X i,j . After the involution kernel is generated, deconvolution calculation can be performed. The calculation process of deconvolution (Involution) is as follows:
[0137]
[0138] where,[[]] represents the set of neighborhood offset amounts for convolving the central pixel point, and its expression is:
[0139]
[0140] where, × represents the Cartesian product.[[]]
[0141] To simplify the generation method of the involution kernel, ψ i,j is taken as the singleton set {(i,j)}, that is represents a single pixel point with coordinates (i,j) on the feature map, thus obtaining an instantiation method of the involution kernel:
[0142]
[0143] where,[[]] and are linear transformation matrices, r is the channel reduction ratio, and σ is the intermediate Batch Nomalization layer and the non-linear activation function ReLU layer, etc.[[]]
[0144] The schematic diagram of the principle of the deconvolution operator is as shown in Figure 5As shown in the figure, for the feature vector at a coordinate point of the input feature map, it is first expanded into the shape of the kernel through the φ(FC-BN-ReLU-FC) and reshape(channel-to-space) transformations, so as to obtain the corresponding involution kernel at this coordinate point, and then the feature vectors in the neighborhood of this coordinate point on the input feature map are subjected to Multiply-Add to obtain the finally output feature map. represents the multiplication operation that propagates across C channels, represents the summation operation that aggregates within the spatial neighborhood. The specific operation process and tensor shape changes for generating the inverse convolution operator are as Figure 6 shown. Among them, Ω i,j is the K×K neighborhood near the coordinates (i, j). Specifically, the method for generating the inverse convolution operator is already a relatively mature existing technology, and more details will not be elaborated here.
[0145] An example of the model structure for constructing the AE model of the deconvolution operator is as Figure 7 shown. As a preferred parameter selection, its neural network optimizer is SGD, the initial learning rate is 0.1, the weight-declay is 2e -4 , the momentum is 0.9, the batch size is 16, the Dropout random inactivation rate is 0.5, and it is trained for a total of 60 epochs. The learning rate decay strategy is that the learning rate decays by 99% every 20 epochs.
[0146] The present invention also provides a fault diagnosis method for large rotating machinery, including the following steps:
[0147] Step A: Obtain multi-source heterogeneous data of the large rotating machinery to be detected, and perform feature extraction using the multi-source heterogeneous data feature extraction method of the present invention for the large rotating machinery to obtain the multi-source heterogeneous data feature vectors of the large rotating machinery to be detected;
[0148] Step B: Input the multi-source heterogeneous data feature vectors of the large rotating machinery to be detected into the trained fault classification and recognition model, and output the predicted diagnosis result of the fault category of the large rotating machinery to be detected.
[0149] Among them, the fault classification and recognition model is obtained through the following steps of training:
[0150] Step b1: The multi-source heterogeneous database obtains multiple groups of sample text data, sample table data, and sample time series data associated with the large rotating machinery, and each group of sample text data, sample table data, and sample time series data has been labeled with the corresponding fault category label;
[0151] Step b2: Use the multi-source heterogeneous data feature extraction method of the present invention described above to extract the multi-source heterogeneous data feature vectors of each group of sample text data, sample table data, and sample time series data associated with large rotating machinery, and form a multi-source heterogeneous data sample set;
[0152] Step b3: Select training samples and test samples from the multi-source heterogeneous data sample set to form a training sample set and a test sample set respectively;
[0153] Step b4: Use each training sample in the training sample set as the input of the fault classification and recognition model, and use the fault category label of each training sample in the training sample set as the output verification label to perform fault category classification prediction training on the fault classification and recognition model to adjust the fault category classification parameters of the fault classification and recognition model;
[0154] Step b5: Input the test sample quantity in the test sample set into the fault classification and recognition model for fault category prediction, and use the fault category label of each test sample in the test sample set as the output verification label to compare and verify the fault category prediction result of the fault classification and recognition model, and evaluate the fault category prediction performance of the fault classification and recognition model;
[0155] Step b6: If the fault category prediction performance of the fault classification and recognition model does not reach the preset target, return to execute Step b4; if the fault category prediction performance of the fault classification and recognition model reaches the preset target, the training is completed, and a trained fault classification and recognition model is obtained.
[0156] In specific implementation, the fault category prediction performance indicators for evaluating the fault classification and recognition model include precision rate, precision, recall rate, F value, etc. These are common performance indicators for neural network model training.
[0157] Using this method for fault diagnosis of large rotating machinery, since the multi-source heterogeneous data feature vectors are derived from multiple data source dimensions, they can more fully present the operation state information and maintenance value information of large rotating machinery equipment, and thus can better improve the accuracy of fault diagnosis of large rotating machinery equipment.
[0158] Similarly, the multi-source heterogeneous data feature vectors of large rotating machinery extracted by the method of the present invention can also be used as feature data for applications such as life prediction of large rotating machinery equipment to help improve its prediction accuracy.
[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the technical solutions. Those of ordinary skill in the art should understand that those who modify or equivalently replace the technical solutions of the present invention without departing from the purpose and scope of the present technical solution should be covered by the scope of the claims of the present invention.
Claims
1. Feature extraction method for multi-source heterogeneous data of large rotating machinery, characterized in that, It includes the following steps: S1: Obtain the multi-source heterogeneous data of large rotating machinery; the multi-source heterogeneous data of the large rotating machinery includes text data, tabular data, and time-series data associated with the large rotating machinery; S2: Perform sentence splitting and word segmentation on the text information of the text data and the text information of each cell in the tabular data respectively to obtain the corresponding sentence-split and word-segmented information; Step S2 specifically includes: S201: Perform sentence splitting on the text information of the text data and the text information of each cell in the tabular data respectively to obtain the sentence segments of each text information sentence; S202: Perform word segmentation on each sentence segment of the text information respectively to obtain the feature words included in each sentence segment; S203: Take the set of feature words included in each sentence segment of the text information in the text data as the sentence-split and word-segmented information of the text data; take the set of feature words included in each sentence segment of the text information in each cell of the tabular data as the sentence-split and word-segmented information of the corresponding cell; S3: Perform word embedding encoding on the sentence-split and word-segmented information of the text data, and take the obtained word encoding vector of the text data as the feature representation vector of the text data; Step S3 specifically includes: S301: Perform word embedding encoding on each feature word included in each sentence segment in the sentence-split and word-segmented information of the text data respectively using the Bert model to obtain a 1×B dimensional word encoding vector for each feature word, where B is the encoding dimension size of the word embedding encoding by the Bert model; S302: For a single text data, the word encoding vectors of the feature words contained in each sentence segment in the sentence splitting and word segmentation information of the text data are concatenated and fused through the concat method to obtain the -dimensional word encoding vector, which serves as the feature representation vector of the text data; where m w represents the number of sentence segments obtained by splitting the text data into sentences, and n w,i represents the number of feature words contained in the i-th sentence segment of the text data; S4: Perform word embedding encoding on the sentence-split and word-segmented information of each cell in the tabular data respectively, and splice and fuse the obtained word encoding vectors of each cell in the tabular data to obtain an encoding vector matrix of the tabular data as the feature representation vector of the tabular data; Step S4 specifically includes: S401: Perform word embedding encoding on each feature word included in each sentence segment in the sentence-split and word-segmented information of each cell in the tabular data respectively using the Bert model to obtain a 1×B dimensional word encoding vector for each feature word, where B is the encoding dimension size of the word embedding encoding by the Bert model; S402: For a single cell in the tabular data, the word encoding vectors of the feature words contained in each sentence segment in the sentence segmentation and word segmentation information of the cell are concatenated and fused through the concat method to obtain the -dimensional word encoding vector of the cell; where m c represents the number of sentence segments obtained by sentence segmentation of the text information in a single cell, and n c,i represents the number of feature words contained in the i-th sentence segment of the text information in a single cell; S403: For each cell of the N tuples × M field attributes contained in the tabular data, first use the concat method to splice and fuse the word encoding vectors of the cells with M different field attributes in the same tuple to obtain a -dimensional tuple encoding vector; then, taking the tuple as a unit, splice and fuse the tuple encoding vectors of the N different tuples contained in the tabular data to obtain a -dimensional encoding vector matrix of the tabular data, which serves as the feature representation vector of the tabular data; S5: Perform segmented cutting on the time-series data, and perform encoding processing on each obtained time-series data segment using a pre-trained autoencoder respectively and then splice and fuse them to obtain an encoding vector of the time-series data as the feature representation vector of the time-series data; S6: Perform splicing, fusion, and dimensionality reduction encoding processing on the feature representation vectors of the text data, tabular data, and time-series data associated with the large rotating machinery, and take the obtained low-dimensional fusion feature vector as the multi-source heterogeneous data feature vector of the large rotating machinery.
2. The method for extracting features of multi-source heterogeneous data of large rotating machinery according to claim 1, wherein In the step S201, before performing sentence splitting on the text information of the text data and the text information of each cell in the tabular data, it further includes: Perform text preprocessing on the text information of the text data and the text information of each cell in the tabular data, and the text preprocessing includes one or more of misspelling correction processing, error symbol correction processing, error grammar correction processing, stop word removal processing, and synonym expression consistency processing of the text information.
3. The method for extracting features of multi-source heterogeneous data of large rotating machinery according to claim 1, wherein, The step S5 specifically includes: S501: Segment the time-series data according to the set segment length to obtain each time-series data segment after segmentation of the time-series data; S502: Use the encoding dimension size B of the word embedding encoding for clause segmentation and word segmentation information as the encoding dimension size of the autoencoder, and use the pre-trained autoencoder to perform encoding processing on each time-series data segment of the time-series data respectively, and obtain a 1×B-dimensional data segment encoding vector for each time-series data segment; S503: For a single time-series data, the data segment encoding vectors of each time-series data segment of the time-series data are spliced and fused through the concat method to obtain an m t ×B-dimensional encoding vector, which serves as the feature representation vector of the time-series data; where m t represents the number of time-series data segments obtained by segmenting and cutting the time-series data.
4. The method for extracting features of multi-source heterogeneous data of large rotating machinery according to claim 3, wherein The autoencoder is obtained through training by the following steps: Step 5021: Obtain multiple sample time-series data of large rotating machinery from a multi-source heterogeneous database; Step 5022: Segment and cut each sample time-series data according to the set segment length to obtain each time-series data segment after segmentation of each sample time-series data, as a sample time-series data set; Step 5023: Select training samples and test samples from the sample time-series data set according to the set training-test ratio to obtain a training sample set and a test sample set; Step 5024: Use the training sample set and the test sample set as the input of the autoencoder, and use minimizing the mean square loss as the training objective to perform unsupervised learning training on the autoencoder; The network model of the autoencoder includes an encoding layer and a decoding and verification layer; among them, the encoding layer of the autoencoder contains 5 Linear layers, and a single time-series data segment of the sample time-series data passes through the encoding layer of the autoencoder to obtain a 1×B-dimensional data segment encoding vector; the decoding and verification layer of the autoencoder contains 5 Linear layers, and the 1×B-dimensional data segment encoding vector obtained by the encoding layer is re-decoded and restored to a time-series data segment through the decoding and verification layer of the autoencoder for comparison and verification with the original time-series data segment; Step 5025: After completing the unsupervised learning training, obtain the trained autoencoder.
5. The method for extracting features of multi-source heterogeneous data of large rotating machinery according to claim 1, characterized in that, The specific steps of step S6 include: S601: Concatenate and fuse the -dimensional feature representation vectors of text data, the -dimensional feature representation vectors of tabular data, and the m t ×B-dimensional feature representation vectors of time series data to obtain a -dimensional fused feature representation matrix; Among them, B represents the encoding dimension size for word embedding encoding; m w represents the number of sentence segments obtained by splitting the text data into sentences, n w,i represents the number of feature words contained in the i-th sentence segment of the text data; m c represents the number of sentence segments obtained by splitting the text information in a single cell of the table data into sentences, n c,i represents the number of feature words contained in the i-th sentence segment of the text information in a single cell, N represents the number of tuples contained in the table data, and M represents the number of field attributes contained in the table data; m t represents the number of time series data segments obtained by segmenting and cutting the time series data; S602: Input the fusion feature representation matrix obtained by fusion into a pre-trained dimensionality reduction and encoding model, and use the 1×D B -dimensional low-dimensional fusion feature vector output by the dimensionality reduction and encoding model as the multi-source heterogeneous data feature vector of the large rotating machinery; where D B represents the dimensionality reduction and encoding dimension size of the dimensionality reduction and encoding model.
6. The method for extracting features of multi-source heterogeneous data of large rotating machinery according to claim 5, characterized in that, The dimensionality reduction encoding model is obtained through training by the following steps: Step 6021: Obtain multiple groups of sample text data, sample table data, and sample time-series data associated with large rotating machinery from a multi-source heterogeneous database; Step 6022: Process each group of sample text data, sample table data, and sample time series data respectively to obtain the -dimensional feature representation vectors of the sample text data in each group, the -dimensional feature representation vectors of the sample table data, and the m t ×B-dimensional feature representation vectors of the sample time series data, and perform splicing and fusion to obtain the -dimensional fusion feature representation matrix corresponding to each group as the sample data set; Step 6023: Select training samples and test samples from the sample data set according to the set training-test ratio to obtain a training sample set and a test sample set; Step 6024: Use the training sample set and the test sample set as the input of the dimensionality reduction encoding model, and use minimizing the mean square loss as the training objective to perform unsupervised learning training on the dimensionality reduction encoding model; The dimensionality reduction encoding model includes an encoding layer and a decoding and verification layer; among them, the encoding layer of the dimensionality reduction encoding model contains 1 Linear layer, 3 deconvolution operator convolution layers, and 1 residual module. The fused feature representation matrix of [dimension] is processed by the encoding layer of the dimensionality reduction encoding model to obtain a 1×D B dimensional low-dimensional fused feature vector; the decoding and verification layer of the dimensionality reduction encoding model contains 5 Linear layers, and the 1×D B dimensional low-dimensional fused feature vector obtained from the encoding layer is re-decoded and restored to the fused feature representation matrix of [dimension], which is used for comparison and verification with the original fused feature representation matrix. Step 6025: After completing the unsupervised learning training, obtain the trained dimensionality reduction encoding model.
7. A fault diagnosis method for a large rotating machine, characterized in that, It includes the following steps: Step A: Obtain the multi-source heterogeneous data of the large rotating machinery to be detected, and perform feature extraction using the large rotating machinery multi-source heterogeneous data feature extraction method described in any one of claims 1 to 5 to obtain the multi-source heterogeneous data feature vector of the large rotating machinery to be detected; Step B: Input the multi-source heterogeneous data feature vector of the large rotating machinery to be detected into the trained fault classification and recognition model, and output the fault category prediction and diagnosis result of the large rotating machinery to be detected.
Citation Information
Patent Citations
Platform for promoting intelligent development of industrial internet of things system
CN114424167A
Abnormality detection method, device and equipment for time series data and operation machine
CN114547147A