Industrial equipment anomaly detection method and system based on word embedding coding

Through an industrial equipment abnormality detection system based on word embedding encoding, data processing and autoencoder training are used using Transformer and word2vec models, the problem of high false alarm rate of equipment abnormality detection in the prior art is solved, and effective monitoring and abnormal detection of equipment health status is realized.

CN120353207APending Publication Date: 2025-07-22SHANGHAI ZHUSUAN TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311671273.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the detection of abnormalities of industrial equipment, it is difficult to effectively distinguish the true health status of the equipment based on the "baseline-spatial distance" paradigm, resulting in a high false alarm rate, and traditional methods fail to fully mine the statistical distribution information and joint distribution information of equipment historical data.

Method used

An industrial equipment abnormality detection system based on word embedding encoding is adopted, including data acquisition, quality evaluation, operating condition screening, data symbolization processing, word embedding encoding and baseline training modules. The data is represented by Transformer and word2vec models, and the baseline model is trained through an autoencoder to calculate the reconstruction error to judge the abnormality.

Benefits of technology

It improves the accuracy of equipment abnormality detection, reduces the false alarm rate, and better recognizes changes in the health status of the equipment. Through distributed representation and optimization of the autoencoder, effective monitoring of the equipment status is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353207A_ABST
    Figure CN120353207A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial equipment anomaly detection method and system based on word embedding coding and an industrial equipment anomaly detection system based on word embedding coding. Comprising a data acquisition module, a data quality evaluation module, a working condition screening module, a data symbolization processing module, a word embedding coding module, a baseline training module, a real-time abnormal value calculation module and an alarm output module. A distributed representation technology for processing industrial continuous data or time series data is introduced, and word embedding representation of the continuous data is obtained by learning the continuous data by using Transform or word2vec, so that distributed representation using a fixed-width vector and fusing historical data information can be realized; compared with traditional feature engineering directly using original data or based on the original data, the method is more effective; and an auto-encoder is utilized, so that the defects in a traditional anomaly detection method are optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of anomaly detection, and particularly to an industrial equipment anomaly detection method and system based on word embedding coding. Background Art

[0002] In the industrial field, a lot of sensor measurement data are continuous analog data, and these data are directly applied to functions such as device control or safety warning. The warning based on threshold in the control system usually only focuses on the measured value itself. However, there are many influencing factors for the measuring points of sensors. For example, the temperature sensor of the generator bearing is affected by its working conditions, the ambient temperature, and the temperature of adjacent mechanical equipment. The single threshold detection method is prone to false alarms under normal conditions.

[0003] For industrial equipment, users are more concerned about the real state of the equipment, which means the health condition of the equipment represented by the data over a period of time, rather than the information at a single time point.

[0004] In the above problems, some methods have been proposed to solve such problems, such as technologies improved based on statistics or machine learning and other technologies. However, these technologies often only process the data itself. Considering the situation where device historical data is usually available, the meaning of the data value itself has not been more deeply explored, such as combining its statistical distribution information in historical data, combining its joint distribution information with other variables / data points, etc.

[0005] In addition, the high-level representation of industrial data and time series data is a broad research field, and many high-level representations of time series have been proposed for data mining. Although many symbolic representations of time series have been introduced in the past few decades, they all have defects, including: the dimension of the symbolic representation is the same as that of the original data, and the dimension scalability of almost all data mining algorithms is very poor. Secondly, although distance metrics can be defined on symbolic methods, these distance metrics have little correlation with the distance metrics defined on the original time series.

[0006] In addition, in the technical path of anomaly detection based on either the original data or the secondary representation of the original data, the "baseline - spatial distance" paradigm is one of the most common solutions. This paradigm does not rely on labeled data and is a general technology widely used in anomaly detection. This solution uses normal data to train a "baseline" representing the normal state of the device. This baseline can be a single-dimensional or multi-dimensional spatial hyperplane. After calculating the "baseline" of the normal state of the device through data, the distance metric between the data and the baseline is calculated, and the distance representing the current state of the device from the "baseline" can be obtained. This distance can represent the degree of anomaly of the device. The farther away from the "baseline", the greater the degree of anomaly it represents.

[0007] However, in practical applications, it is very difficult to obtain the healthiest state of the device. Even if the data of the normal state of the device is obtained, it is not necessarily the so-called "most normal" state. At this time, an inevitable situation will occur. If the device is more normal than the "baseline", the calculated spatial distance ("outlier") will be relatively large, while it should be smaller under normal logic. How to optimize the above industrial device anomaly detection technology based on the "baseline - spatial distance" paradigm to better detect anomalies or faults in the device is an urgent problem to be solved. Therefore, an industrial device anomaly detection method and system based on word embedding coding are proposed. Summary of the Invention

[0008] The object of the present invention is to solve the defects existing in the prior art, and to propose an industrial device anomaly detection method and system based on word embedding coding.

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] According to one aspect of the present invention, an industrial device anomaly detection system based on word embedding coding is provided.

[0011] The industrial device anomaly detection system based on word embedding coding includes a data acquisition module, a data quality assessment module, a working condition screening module, a data symbolization processing module, a word embedding coding module, a baseline training module, a real-time outlier calculation module, and an alarm output module;

[0012] Among them:

[0013] The data acquisition module is used to connect to the corresponding control system to collect device data;

[0014] The data quality assessment module is used to evaluate the data quality by setting different data quality assessment logics and algorithms for the data;

[0015] The working condition screening module is used to select specific device working conditions to make the evaluated device states comparable;

[0016] The data symbolization processing module is used to symbolize continuous data in an adapted manner;

[0017] The word embedding coding module is used to, after obtaining the symbolized representations of each data dimension, use a natural language model to train based on all training data to achieve the encoding of the symbols;

[0018] The baseline training module trains all document vectors on all training sets through an autoencoder to obtain the representation of the normal state of the device;

[0019] A real-time outlier calculation module, which is used to calculate the outliers of real-time device data after obtaining the baseline model and the threshold;

[0020] An alarm output module, which is used to judge the reconstruction error output by the real-time outlier calculation module and the set threshold, and comprehensively judge whether to output an alarm according to the historical reconstruction error value.

[0021] Further, the specific step process of the word embedding encoding module is as follows:

[0022] 1: Based on the Transformer model, by means of the masked language model method, randomly mask some words in the input sequence, and then train the model by letting the model predict the masked words. After obtaining the trained model, adopt the document input model, take the CLS Token vector of the output of the model encoding layer, and the obtained vector is the vector of this document;

[0023] 2: Implement the encoding of document vectors under different time windows based on the doc2vec technology of word2vec. The data is cyclically slid with a certain time window size to extract data, and the data symbols in each dimension under each time window form a "document";

[0024] 3: After obtaining the above data for training, use word2vec to train in the training data, and through training, the word vector encoding of each symbol and the document vector encoding of each "document" will be obtained.

[0025] Further, in the baseline training module, a trained autoencoder model, that is, the baseline model, will be obtained through training, and at the same time, the reconstruction error between the document vector of each "document" in the training data and the baseline can be calculated;

[0026] By calculating the distribution of the reconstruction error, appropriate criteria are used to set the outlier threshold for different morphological distributions. Among them: for the reconstruction error of the Gaussian distribution, the criterion of its mean + 3 times the standard deviation is usually used to set the threshold.

[0027] According to another aspect of the present invention, a method based on the above-mentioned industrial equipment anomaly detection system based on word embedding encoding is provided.

[0028] An industrial equipment anomaly detection method based on word embedding encoding includes the following steps:

[0029] S1: Data acquisition: Through communication with the device SCADA system, collect the controller signals of the device, and the acquisition frequency is 1s;

[0030] S2: Working condition screening: Through the working condition screening module, limit the data for the next data analysis and processing to the working condition where the wind turbine starts to be grid-connected and is facing the wind;

[0031] S3: Data symbolization processing: Use different discretized binning methods for data in different dimensions, symbolize the continuous data, and add prefix symbols to the data in each dimension for differentiation;

[0032] S4: Word embedding encoding;

[0033] S5: Baseline training: After obtaining the document vectors of each document in the data, input them into an autoencoder to train the baseline model representing the normal state of the device. By training the baseline model, the distribution of the reconstruction error of each document vector can also be obtained;

[0034] S6: Real-time outlier calculation: When the above steps are completed, implement the data to be detected. Through the same processing, input it into the above baseline model, calculate the reconstruction error of each document, and calculate the average reconstruction error of this batch of data;

[0035] S7: Alarm output: When the obtained average reconstruction error exceeds the calculated threshold, output the alarm.

[0036] Further, in step S1: Collect signals from the device SCADA system for monitoring the vibration and temperature of the wind turbine and related operating condition data. The dimensional data includes the bearing temperature at the drive end of the generator, the temperature at the non-drive end of the generator, the effective value of the vibration speed at the drive end of the generator, the effective value of the vibration speed at the non-drive end of the generator, the generator speed, the active power of the generator, the wind speed, the yaw angle, and the blade angle.

[0037] Further, in step S2, filter the data by setting the condition combination of the blade angle <= 3 degrees and -1 <= yaw angle < 1 degree.

[0038] Further, step S4 specifically further includes the following steps:

[0039] S401: After symbolizing all the data, construct the corpus used for training word vectors. At the same time, extract "documents" from all the data with a non-overlapping sliding window;

[0040] S403: Construct training data: Further, for the convenience of use in Transformer, use the Embedding Layer to convert the symbolized data obtained in the previous step into real number vectors of a fixed dimension. After obtaining the mapped vector data, for each document, randomly select and replace some symbols in its symbol sequence with special "[MASK]" symbols to simulate the masking operation of MLM to construct such training samples.

[0041] S405: Training: Input the constructed training samples into the model. The model needs to predict the masked symbols. Use the cross-entropy loss function to compare the model's predictions with the true masked symbols, and perform backpropagation and parameter updates according to the loss function. Repeat this process until the model converges;

[0042] S407: After training is completed, input all documents into the model. By using the vector representation of the CLS Token in the output sequence of the Transformer Encoder extracted from the last layer of the entire Transformer model, obtain the vector representation of the document, that is, the document vector.

[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0044] Introduce a distributed representation technology for processing industrial continuous data or time series data. By using Transformer or word2vec to learn the word embedding representation of continuous data, it is possible to achieve a distributed representation using fixed-width vectors and integrating historical data information, which is more effective than directly using the original data or feature engineering based on the original data in the traditional way;

[0045] Utilize autoencoders to optimize the defects in traditional anomaly detection methods, which has been proven to be widely effective in practice. Description of the Drawings

[0046] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention.

[0047] Figure 1 It is a schematic flowchart of an industrial equipment anomaly detection system based on word embedding coding proposed by the present invention. Detailed Embodiments

[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0049] In this application, for the convenience of describing the implementation solutions of the present technology, the following further elaborates on the concepts and technical terms that appear in the embodiments:

[0050] 1) Word Embedding

[0051] Word embedding, also known as word vector, refers to a class of technologies that use numerical vectors to represent documents and words. Word embeddings can be obtained using language modeling and feature learning techniques, where words or phrases in the vocabulary are mapped to real-valued vectors.

[0052] 2) Word2vec

[0053] Word2vec is a natural language processing (NLP) technique released in 2013. The word2vec algorithm uses a neural network model to learn the distributed representations of words from a large text corpus, obtaining word vectors, which are a set of related models for generating word embeddings. These models are shallow two-layer neural networks that are trained to reconstruct the linguistic context of words. Word2vec takes a large text corpus as input and generates a vector space (usually with hundreds of dimensions), where each unique word in the corpus is assigned a corresponding vector in this space. Word2vec can use either of two model architectures to generate the distributed representation of words: Continuous Bag of Words (CBOW) or Continuous Skip-gram. In both architectures, word2vec considers individual words and a sliding context window when iterating through the corpus.

[0054] 3) Transformer Model

[0055] The Transformer model is a deep learning architecture first proposed in 2017 that relies on a parallel multi-head attention mechanism. The Transformer model consists of an Encoder and a Decoder. The Encoder converts the input sequence into continuous representations, while the Decoder converts these representations into an output sequence. Each Encoder and Decoder is composed of multiple layers (usually the same number of layers), and each layer contains two sub-layers: the multi-head self-attention mechanism and a feed-forward neural network. The self-attention mechanism is used to establish dependencies between each position in the sequence and other positions, while the feed-forward neural network is used to perform a non-linear transformation on the representation of each position. In the self-attention mechanism, the association degree between each position of the input sequence and other positions is calculated, and a weight is assigned to each position based on these association degrees. In this way, the model can consider the information of other positions in the input sequence when processing each position. By using the multi-head mechanism, the Transformer model can simultaneously learn multiple different attention representations, thus better capturing the information in the sequence.

[0056] 4) Autoencoder

[0057] An autoencoder is a neural network that, when trained, attempts to copy its input to its output. Internally, it has a hidden layer h that describes the code used to represent the input. The network can be seen as consisting of two parts: an encoder function h = f(x) and a decoder that produces a reconstruction r = g(h). If the autoencoder simply learns to set g(f(x)) = x everywhere successfully, then it is not particularly useful. Instead, autoencoders are designed not to learn to copy perfectly. Typically, they are restricted to approximately copying and only copying inputs similar to the training data. Since the model will prioritize copying specific dimensions or aspects, it usually learns useful properties of the data.

[0058] 5) Reconstruction error

[0059] The reconstruction error refers to the difference between the input data and the output of the autoencoder decoder. Mean squared error or cross-entropy is usually used as a measure of the reconstruction error.

[0060] 6) CLS Token

[0061] In BERT and Transformer models, the main purpose of the CLS token is to enable predictions for downstream tasks such as text classification. During pre-training, the model learns the context information of each word through tasks such as MLM (Masked Language Model), and in downstream tasks, the output of the CLS token is used for the representation of the entire text.

[0062] Refer to Figure 1 , an industrial equipment anomaly detection system based on word embedding encoding, including a data acquisition module, a data quality assessment module, a working condition screening module, a data symbolization processing module, a word embedding encoding module, a baseline training module, a real-time outlier calculation module, and an alarm output module;

[0063] Among them:

[0064] The data acquisition module is used to collect device data. The device data can come from its control system, such as PLC, DCS, or the upper system, such as SCADA, the centralized control system, etc., or it can come from sensors installed additionally. The adoption rates of different systems can be different.

[0065] The data quality assessment module is used to evaluate the data quality by setting different data quality assessment logics and algorithms for the data. Among them: retain the data that meets the data quality standards;

[0066] The working condition screening module is used to select specific device working conditions to make the evaluated device states comparable. Among them: different device data collected, different conditions are used for working condition screening;

[0067] A data symbolization processing module is used to symbolize continuous data in an adapted manner. Among them, it is preferably achieved by discretizing the data. For example, for continuous data in some (equal-width or non-equal-width) intervals preset by parameters, it is judged whether the data belongs to the interval, and the preset interval symbol is taken as the symbol representation of the discretization of the value. In multi-dimensional data, in order to distinguish data of different dimensions, a specific prefix setting is usually added before the symbol after the above discretization processing;

[0068] A word embedding encoding module is used to, after obtaining the symbolized representations of each data dimension, use a natural language model to train based on the "corpus", that is, all training data, to achieve the encoding of the symbols;

[0069] It should be further noted that after obtaining the symbolized representations of each data dimension, the original numerical data is processed into symbols, which is equivalent to the concept of "words" in natural language processing. In order to achieve the encoding of the symbols, a natural language model needs to be trained based on the "corpus", that is, all training data. All data in the normal state used for training is symbolized to form the "corpus". The "vocabulary" is composed of all unique words in the corpus.

[0070] A baseline training module trains all document vectors on all training sets through an autoencoder to obtain a representation of the normal state of the device;

[0071] A real-time outlier calculation module is used to calculate the outliers of real-time device data after obtaining the baseline model and the threshold. Among them: the real-time device data is obtained through the same symbolization processing, and the document vectors under each time window of it are calculated according to the same time window. Each document vector is input into the baseline model to calculate its reconstruction error as the outlier;

[0072] An alarm output module is used to judge the reconstruction error output by the real-time outlier calculation module and the set threshold, and comprehensively judge whether to output an alarm according to the historical reconstruction error value. Among them: when the reconstruction error calculated continuously for n times exceeds the threshold, an alarm is output, where n consecutive times is a parameter that can be set.

[0073] In a specific embodiment of the present application, the specific step process of the word embedding encoding module is as follows:

[0074] 1: Based on the Transformer model (transformer model), through the Masked Language Model method, some words are randomly masked in the input sequence, and then the model is trained in the way of letting the model predict the masked words. After obtaining the trained model, the document is input into the model, and the CLS Token vector of the output of the encoding layer of the model is taken, and the obtained vector is the vector of the document;

[0075] 2: Implement document vector encoding for different time windows based on the doc2vec technology of word2vec. Extract data by sliding a window of a certain size over the data (symbols) in a cyclic manner. The data symbols in each dimension under each time window form a "document".

[0076] 3: After obtaining the above data for training, use word2vec to train on the training data. Through training, the word vector encoding of each symbol and the document vector encoding of each "document" will be obtained.

[0077] Furthermore:

[0078] In the baseline training module, a trained autoencoder model, i.e., the baseline model, will be obtained through training, and at the same time, the reconstruction error between the document vector of each "document" in the training data and the baseline can be calculated.

[0079] By calculating the distribution of the reconstruction error, appropriate criteria are used to set the anomaly threshold for different forms of distribution. Among them: for the reconstruction error of a Gaussian distribution, the criterion of its mean + 3 times the standard deviation is usually used to set the threshold.

[0080] According to an embodiment of the present invention, a method for an industrial equipment anomaly detection system based on the above word embedding encoding is also provided.

[0081] An industrial equipment anomaly detection method based on word embedding encoding includes the following steps:

[0082] S1: Data acquisition: Collect the controller signals of the equipment through communication with the equipment SCADA system, and the acquisition frequency is 1 s.

[0083] S2: Operating condition screening: Through the operating condition screening module, limit the data for the next data analysis and processing to the operating condition where the wind turbine starts to grid-connected power generation and is facing the wind.

[0084] S3: Data symbolization processing: Use different discretization binning methods for data in different dimensions to symbolize continuous data, and add a prefix symbol to the data in each dimension for distinction.

[0085] S4: Word embedding encoding;

[0086] S5: Baseline training: After obtaining the document vector of each document in the data, input it into an autoencoder to train the baseline model representing the normal state of the equipment. By training the baseline model, the distribution of the reconstruction error of each document vector can be obtained at the same time.

[0087] S6: Real-time outlier calculation: After the above steps are completed, the data to be detected is implemented. Through the same processing, it is input into the above baseline model to calculate the reconstruction error of each document, and the average reconstruction error of this batch of data is calculated;

[0088] S7: Alarm output: When the obtained average reconstruction error exceeds the calculated threshold, output the alarm.

[0089] To better illustrate the technical solution of this application, the following further explains it in combination with specific embodiments.

[0090] S1: Data collection: Through communication with the device SCADA system, the controller signals of the device are collected, and the collection frequency is 1 s;

[0091] In step S1: Collect signals from the device SCADA system for monitoring the vibration and temperature of the wind turbine and related operating condition data, including the temperature of the generator drive-end bearing, the temperature of the generator non-drive-end, the root mean square (RMS) of the vibration speed of the generator drive-end, the root mean square (RMS) of the vibration speed of the generator non-drive-end, the generator speed, the active power of the generator, the wind speed, the wind alignment angle, the blade angle, etc., a total of 9 dimensions of data.

[0092] S2: Operating condition screening: Through the operating condition screening module, the data used for the next data analysis and processing is limited to the operating condition where the wind turbine starts grid-connected power generation and is facing the wind;

[0093] In step S2, the data is screened by setting the condition combination of the blade angle <= 3 degrees and -1 <= the wind alignment angle < 1 degree;

[0094] S3: Data symbolization processing: Use different discretization binning methods for data of different dimensions to symbolize continuous data, and add a prefix symbol to the data of each dimension for distinction;

[0095] In step S3, a row of data [60.5248, 58.7601, 4.2381, 4.0122, 20.1205, 2681.67, 9.8901, 0, 0.13] composed of the generator drive-end bearing temperature, generator non-drive-end temperature, generator drive-end vibration velocity effective value (RMS), generator non-drive-end vibration velocity effective value (RMS), generator speed, generator active power, wind speed, wind alignment angle, and blade angle is processed by setting equally spaced bins for each dimension, and after processing, it becomes [60.5, 58.8, 4.24, 4.02, 20.1, 2682, 9.9, 0, 0]. Different dimensions are scaled in the same way to remove the decimal points, and then dimension prefix symbols are added. Finally, it is processed into [A605, B588, C424, D402, E201, F2682, G99, H0, I0].

[0096] S4: Word embedding encoding;

[0097] As a preferred embodiment of the present application, step S4 specifically further includes the following steps:

[0098] S401: After symbolizing all the data, a "corpus" used for training word vectors is constructed. At the same time, we will use a non-overlapping sliding window to extract "documents" from all the data.

[0099] Among them: In this embodiment, processing is carried out with a 10s time window. That is, each "document" contains the symbolized data symbols arranged in the order of 10s time and dimension order;

[0100] In this embodiment, a Transformer model is used for training.

[0101] S403: Construct training data: Further, for the convenience of use in Transformer, the symbolized data obtained in the previous step is used with an Embedding Layer to convert symbols into real number vectors of a fixed dimension. After obtaining the mapped vector data, for each document, some symbols in its symbol sequence are randomly selected and replaced with special "[MASK]" symbols to simulate the masking operation of MLM to construct such training samples.

[0102] S405: Training: The constructed training samples are input into the model, and the model needs to predict the masked symbols. The cross-entropy loss function is used to compare the model's predictions with the real masked symbols. Backpropagation and parameter updates are performed according to the loss function. This process is repeated until the model converges.

[0103] S407: After the training is completed, all documents are input into the model. By using the vector representation of the CLS Token in the output sequence of the Transformer Encoder extracted from the last layer of the entire Transformer model, the vector representation of the document, that is, the document vector, is obtained. Among them, in this embodiment, the dimension of the document vector obtained by this method is 768 dimensions.

[0104] S5: Baseline training: After obtaining the document vectors of each document in the data, input them into an autoencoder to train the baseline model representing the normal state of the device. By training the baseline model, the distribution of the reconstruction error of each document vector can be obtained at the same time. Among them: take the mean of the reconstruction error + 3 times the standard deviation as the threshold;

[0105] S6: Real-time outlier calculation: When the above steps are completed, the data to be detected is implemented. Through the same processing, input it into the above baseline model, calculate the reconstruction error of each document, and calculate the average reconstruction error of this batch of data;

[0106] S7: Alarm output: When the obtained average reconstruction error exceeds the calculated threshold, output the alarm.

[0107] As can be seen from the above solution:

[0108] 1. A new industrial data anomaly detection method is proposed, which is widely applicable to various types of industrial data. Through the symbolic processing of continuous data, the numerical meanings that are the same or similar are aggregated, and some redundant information is ignored to obtain more dimensional information, reduce the noise in the data, improve the signal-to-noise ratio, and through the use of neural network (Transformer or word2vec, etc.) technology for the second distributed representation of symbols, so that the re-expressed symbols contain more information;

[0109] 2. By using the autoencoder as the baseline model and utilizing the representation learning ability of the autoencoder, which is also to utilize its defects, the information that accounts for a small proportion and is not needed by us is removed through the training of the autoencoder. The reconstruction error is used to represent the measure of anomaly, which can more effectively describe the deviation of the anomaly from the normal baseline. For those data states that are "healthier" than the normal baseline, they can be normally represented without obtaining large anomaly values or being considered as possibly having anomalies.

[0110] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. An industrial equipment anomaly detection system based on word embedding encoding, characterized in that It includes a data acquisition module, a data quality assessment module, a working condition screening module, a data symbolization processing module, a word embedding encoding module, a baseline training module, a real-time outlier calculation module, and an alarm output module; Among them: The data acquisition module is used to connect to the corresponding control system to collect device data; The data quality assessment module is used to evaluate the data quality by setting different data quality assessment logics and algorithms for the data; The working condition screening module is used to select specific device working conditions to make the evaluated device states comparable; The data symbolization processing module is used to perform symbolization processing on continuous data in an adapted manner; The word embedding encoding module is used to, after obtaining the symbolized representations of each data dimension, use a natural language model to train based on all training data to achieve the encoding of symbols; The baseline training module trains all document vectors on all training sets through an autoencoder to obtain the representation of the normal state of the device; The real-time outlier calculation module is used to calculate the outliers of real-time device data after obtaining the baseline model and the threshold; The alarm output module is used to judge the reconstruction error output by the real-time outlier calculation module and the set threshold, and comprehensively judge whether to output an alarm according to the historical reconstruction error value.

2. The industrial equipment anomaly detection system based on word embedding coding according to claim 1, characterized in that, The specific step process of the word embedding encoding module is as follows: 1: Based on the Transformer model, through the masked language model method, randomly mask some words in the input sequence, and then train the model by letting the model predict the masked words. After obtaining the trained model, adopt the document input model, and take the CLS Token vector of the output of the encoding layer of the model. The obtained vector is the vector of this document; 2: Implement the encoding of document vectors under different time windows based on the doc2vec technology of word2vec. Slide the data in a cyclic manner with a certain time window size, and the data symbols of each dimension under each time window form a "document"; 3: After obtaining the data used for training above, use word2vec to train in the training data. Through training, the word vector encoding of each symbol and the document vector encoding of each "document" will be obtained.

3. The industrial equipment anomaly detection system based on word embedding coding according to claim 2, characterized in that, In the baseline training module, a trained autoencoder model, that is, the baseline model, will be obtained through training, and at the same time, the reconstruction error between the document vector of each "document" in the training data and the baseline can be calculated; By calculating the distribution of the reconstruction error, appropriate criteria are used to set the outlier threshold for different forms of distributions. Among them: for the reconstruction error of the Gaussian distribution, the criterion of its mean + 3 times the standard deviation is usually used to set the threshold.

4. A method for detecting anomalies in industrial equipment based on word embedding encoding, according to the industrial equipment anomaly detection system based on word embedding encoding described in any one of claims 1-3, characterized in that, It includes the following steps: S1: Data acquisition: Through communication with the device SCADA system, collect the controller signals of the device, and its acquisition frequency is 1s; S2: Working condition screening: Through the working condition screening module, limit the data for the next data analysis and processing to the working condition where the wind turbine starts to grid-connected power generation and is facing the wind; S3: Data Symbolization Processing: Use different discretized binning methods for data in different dimensions to symbolize continuous data, and add prefix symbols to the data in each dimension for distinction; S4: Word Embedding Encoding; S5: Baseline Training: After obtaining the document vectors of each document in the data, input them into an auto encoder to train the baseline model representing the normal state of the device. By training the baseline model, the distribution of the reconstruction error of each document vector can also be obtained; S6: Real-time Outlier Calculation: After the above steps are completed, implement the data to be detected. Through the same processing, input it into the above baseline model, calculate the reconstruction error of each document, and calculate the average reconstruction error of this batch of data; S7: Alarm Output: When the obtained average reconstruction error exceeds the calculated threshold, output the alarm.

5. The industrial equipment anomaly detection method based on word embedding coding according to claim 4, wherein In step S1: Collect signals from the device SCADA system for monitoring the vibration and temperature of the wind turbine and related operating condition data. The dimensional data includes the bearing temperature at the drive end of the generator, the temperature at the non-drive end of the generator, the effective value of the vibration speed at the drive end of the generator, the effective value of the vibration speed at the non-drive end of the generator, the generator speed, the active power of the generator, the wind speed, the wind alignment angle, and the blade angle.

6. The industrial equipment anomaly detection method based on word embedding encoding according to claim 5, wherein In step S2, the data is filtered by setting the condition combination of the blade angle <= 3 degrees and -1 <= the wind alignment angle < 1 degree.

7. The industrial equipment anomaly detection method based on word embedding encoding according to claim 6, wherein, Step S4 specifically further includes the following steps: S401: After symbolizing all the data, construct the corpus used for training word vectors. At the same time, use a non-overlapping sliding window to extract documents from all the data; S403: Construct Training Data: Use the embedding layer to convert the symbolized data obtained in the previous step into real number vectors of a fixed dimension. After obtaining the mapped vector data, for each document, randomly select and replace some symbols in its symbol sequence with special symbols to simulate the masking operation of MLM to construct such training samples; S405: Training: Input the constructed training samples into the model. The model needs to predict the masked symbols. Use the cross-entropy loss function to compare the model's prediction with the real masked symbols, and perform backpropagation and parameter update according to the loss function. Repeat this process until the model converges; S407: After training is completed, input all the documents into the model. By using the vector representation of the CLS Token in the output sequence of the Transformer Encoder in the last layer of the entire Transformer model, obtain the vector representation of the document, that is, the document vector.

Citation Information

Cited By

  • AI analysis alarm platform and method for video monitoring

    CN121305824A