Filling equipment fault automatic detection method based on TimeNet and CRF fusion network
Through the method of TimesNet and CRF converged network, the problems of slow computing speed and no hidden cycle information in the prior art are solved, and efficient and accurate automatic detection of filling equipment failures is achieved.
Patent Information
- Application Number
- CN202510404141.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-22
AI Technical Summary
In the existing automatic identification method of time series state variable production line equipment failure, the LSTM calculation speed is slow, the Transform network does not perform Fourier transform to obtain hidden cycle information, and most fault classification networks do not consider the transfer probability distribution between fault types.
Using a method based on TimesNet and CRF fusion network, data is segmented through a fixed-length sliding window, combined with the DataEmbedding layer, fast Fourier variation, Inception_Block layer and CRF layer, hidden period information is obtained and fault type transfer probability is constrained, and fault automatic detection is achieved.
Improves the accuracy of model prediction, avoids gradient explosion and gradient disappearance, improves running speed, and generates the optimal failure type sequence.
Smart Images

Figure CN120354280A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic identification of equipment failures, and particularly to an automatic detection method for filling equipment failures based on a TimesNet and CRF fusion network. Background Art
[0002] The automatic identification of production line equipment failures based on time series state variables refers to analyzing the time series state variables output by sensors to predict the types of production line failures. Currently, the mainstream network is LSTM. However, the hidden state variables of LSTM need to be calculated one by one, that is, the calculation method of LSTM is serial, and the running speed is slow. For the Transform or Informer network, the time series data is only regarded as one-dimensional information data, without performing Fourier transform to convert the time series data into two-dimensional time processing to obtain more hidden periodic information in the data. And most fault classification networks do not consider that the transition probabilities between fault types are distributed, lacking the problem of distribution limitation for the predicted sequence of fault types.
[0003] In view of the above problems, the present invention proposes an automatic detection method for filling equipment failures based on a TimesNet and CRF fusion network to solve the above problems. Summary of the Invention
[0004] The purpose of the present invention is to provide an automatic detection method for filling equipment failures based on a TimesNet and CRF fusion network to solve the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: An automatic detection method for filling equipment failures based on a TimesNet and CRF fusion network, including the following steps:
[0006] In the first step, extract the characteristic data of M induction equipment in the production line, and use a fixed-length sliding window L for data segmentation to obtain data vectors as the input of the model;
[0007] In the second step, a DataEmbedding layer that fuses different characteristic data, including the information after 1D convolution of the original data vector and the position encoding information;
[0008] In the third step, first perform a fast Fourier transform on the fused characteristic data to obtain the amplitudes and frequencies corresponding to TOP_K, and then reshape the characteristic data according to the TOP_K frequencies, converting from 1D data to K 2D data;
[0009] Fourthly, first pass the K 2D data through the Inception_Block layer, then fuse the features of different-sized convolutions for each 2D data, and finally perform weighted summation on the average features of the K 2D data according to the amplitude size;
[0010] Fifthly, pass the feature data obtained by weighted summation through the fully connected layer and map it to n categories, and the obtained features are represented as h1h2…h L ;
[0011] Sixthly, input the feature data obtained by weighted summation into the CRF layer network for processing. After further constraint by the CRF according to the transition probability matrix, the output prediction results are represented as y1y2…y L .
[0012] Preferably, the dimension of the time series data obtained by the M induction devices is M-dimensional, and the data segmentation is represented as:
[0013] X train ={X (1) ,X (2) ,…,X (L)} (1);
[0014] X (i) ∈R m (2).
[0015] Preferably, the segmented data is normalized. Let the mean vector of the sample features be represented as:
[0016] μ={μ1,μ2...μ j} (3);
[0017] Among them, μ j represents the average value of the features of the jth dimension. Let the standard deviation vector of the sample features be represented as:
[0018] σ={σ1,σ2...σ j} (4);
[0019] Among them, σ j represents the standard deviation of the features of the jth dimension. The normalization calculation is represented as:
[0020]
[0021] Among them, represents the value of the jth feature of the ith sample. Then the data sequence S after normalization processing is represented as:
[0022] S={S1,S2...S L} (6);
[0023] The dimension of the output S sequence is represented as:
[0024] L * M(7).
[0025] Preferably, for the processing of the 1D convolution, the size of the 1D convolution kernel is set to (k, M'), the stride is 1, the padding is 1, and meanwhile, the channel dimension rises from M to M'.
[0026] The output calculation of the processing of the 1D convolution is expressed as:
[0027]
[0028] where, W ij represents the value at each position of the convolution kernel, S ij represents the convolution sequence value, S' represents the convolution output value, b represents the bias, and the dimension of the S' sequence after convolution is expressed as:
[0029] L * M'(9).
[0030] Preferably, for the position encoding processing, position encoding is performed on the dimension of the output S sequence, and the encoding formula is expressed as:
[0031] PE t,2i = sin(t / 10000 2i / d ), PE t,2i+1 = cos(t / 10000 2i / d )(10);
[0032] t ∈ [1, L](11);
[0033] i ∈ [1, M' / / 2](12);
[0034] where, t represents the t-th sample of the sequence, i represents the dimension of the encoding, and the size of the output matrix is expressed as:
[0035] L * M”(13);
[0036] Finally, the S' sequence and the PE matrix are added together to obtain the eigenvector data X, which is used as the input data for the third step.
[0037] Preferably, the fast Fourier transform is expressed as:
[0038] A = Avg(Amp(FFT(X 1D )))(14);
[0039] where, the input eigenvector data is a one-dimensional time series X with a time length of L and a channel number of M' 1D , and then the processed amplitudes are sorted from largest to smallest, and the first K are selected, which is expressed as:
[0040]
[0041] Subsequently, X is reshaped by tensor reshaping 1D The data dimension is transformed into X 2D , with dimensions (L / p i , p i , M'), where i ∈ [1, k];
[0042] If L cannot be divided evenly by p i , then padding with zeros is performed on the original X 1D data. First, calculate the length to be padded based on L and p i , which is expressed as:
[0043] l i = ((L / / p i ) + 1) * p i - L (16);
[0044] i ∈ [1, k] (17);
[0045] Then, a zero tensor of size (l i , M') is generated, and finally, this zero tensor is concatenated to the original X 1D data.
[0046] Preferably, the X 2D data input through the Inception_Block layer has dimensions (L / p i , p i , M'), where i ∈ [1, k]. The Inception_Block layer consists of multiple convolutional network layers with different window sizes, and the convolutional kernel sizes are (1, 1), (3, 3), (5, 5), (7, 7), (9, 9), and (11, 11);
[0047] The processing method of the 2D convolution is to set the size of the 2D convolutional kernel to (k1, k2, M'), where k1, k2 ∈ {1, 3, 5, 7, 9, 11}, the stride is 1, and the channel dimension remains M' unchanged. The processing formula of the 2D convolution is expressed as:
[0048]
[0049] Among them, W ijl represents the value at each position of the convolutional kernel, X ijl represents the convolutional sequence value, X' 2D represents the convolutional output value, b represents the bias, and the dimension of the X' 2D tensor after convolution is (L / p i , p i, M'), where \(i \in [1, k]\). Finally, after the network undergoes multi-scale convolution processing, Dropout is used to achieve regularization.
[0050] Preferably, for \(X'\) 2D The corresponding frequencies \(\{f_1, f_2... f\) k}\) are subjected to softmax calculation, expressed as:
[0051]
[0052] where \(i \in [1, k]\). According to the obtained weights, weighted summation is performed on \(X'\) 2D , and then it is put into the non-linear function GELU for processing. Finally, a residual skip connection is added. The dimension of the processed tensor is \((L / p\) i , p i , M').
[0053] Preferably, for the input 2D tensor with dimension \((L / p\) i , p i , M'), it is reshaped into a 1D tensor with dimension \((L, M')\), and then connected to a fully connected layer and mapped to \(n\) categories.
[0054] Preferably, let the transition score matrix \(T\) be introduced. The matrix element \(T\) ij represents the transition score from fault type \(i\) to fault type \(j\). There are a total of \(n\) fault types, so \(T \in R\) n*n ;
[0055] Let the length of the time series be \(L\), then the score matrix \(P\) of the output layer is in \(R\) L*n , where the matrix element \(P\) ij represents the score at the \(i\)-th moment of the observed time series under the \(j\)-th fault type. For the input \(H=(h_1 h_2... h\) L ), and the output fault sequence is \(Y=(y_1 y_2... y\) L ), then the total score of this fault type is expressed as:
[0056]
[0057] Normalize the sequence path to generate a probability distribution for the output sequence \(Y\), expressed as:
[0058]
[0059] Then maximize the logarithmic probability value of the fault type sequence, expressed as:
[0060]
[0061] Finally, when making a prediction, the Viterbi algorithm is selected to find the optimal fault type sequence.
[0062] Technical effects and advantages of the present invention:
[0063] By means of Fourier transform, the data is raised from one dimension to two dimensions in the present invention, and the hidden periodic information in the data is obtained, thus improving the accuracy of model prediction; at the same time, the backpropagation path of the convolutional network is different from the sequence time, avoiding the problems of gradient explosion and gradient disappearance, and can achieve parallelism to improve the running speed of the model. Finally, according to the problem of transition probability distribution among production line fault types, the CRF algorithm is used for decoding to obtain the optimal fault type sequence, further improving the accuracy of model prediction. Brief description of the drawings
[0064] Figure 1 It is the operation flow chart of the detection method of the present invention.
[0065] Figure 2 It is the architecture diagram of the overall network of the present invention.
[0066] Figure 3 It is the Times Bloc structure diagram of the present invention.
[0067] Figure 4 It is the classification prediction report. Detailed implementation manners
[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0069] The present invention provides an automatic fault detection method for filling equipment based on the TimesNet and CRF fusion network as shown in Figures 1-4 , including the following steps:
[0070] In the first step, extract the characteristic data of M induction equipment in the production line, and use a fixed-length sliding window L for data segmentation to obtain data vectors as the input of the model. There are M induction equipment in total, and the dimension of the time series data obtained is M-dimensional;
[0071] Specifically, the dimension of the time series data obtained by M induction equipment is M-dimensional, and the data segmentation is expressed as:
[0072] X train ={X (1) ,X (2) ,…,X (L)} (1);
[0073] X (i) ∈R m (2).
[0074] Since each sample has M-dimensional features, so X (i) ∈R m .
[0075] After the data is segmented, it is normalized. Let the mean vector of the sample features be expressed as:
[0076] μ = {μ1, μ2... μ j} (3);
[0077] where, μ j represents the average value of the features of the j-th dimension. Let the standard deviation vector of the sample features be expressed as:
[0078] σ = {σ1, σ2... σ j} (4);
[0079] where, σ j represents the standard deviation of the features of the j-th dimension. The normalization calculation is expressed as:
[0080]
[0081] where, represents the value of the j-th feature of the i-th sample. Then the data sequence S after normalization is expressed as:
[0082] S = {S1, S2... S L} (6);
[0083] The dimension of the output S sequence is expressed as:
[0084] L * M (7).
[0085] It should be noted that, as shown in Figure 2 , the overall architecture of the model proposed by the present invention mainly includes a time series input layer, a Data Embedding layer, a Times Block layer, and a CRF layer.
[0086] The experimental data of the present invention comes from a fruit filling production line. The production line operates 8 hours a day, and the equipment acquisition time interval is 1 second. A total of 72,132,81 sample data are collected in a year. Each sample data has 23 features, namely: the number of material pushes, the number of materials to be grabbed, the number of placed containers, the number of container upload detections, the number of filling detections, the fixed state of the filling positioner, the released state of the filling positioner, the number of material grabs, the number of filling rotations, the number of filling descents, the number of fillings, the number of capping detections, the number of capping positions, the number of lid pushes, the number of capping descents, the number of caps, the number of cap tightening detections, the number of cap tightening positions, the number of cap tightening descents, the number of cap tightening rotations, the number of cap tightenings, the number of qualified products, and the number of unqualified products. There are 9 types of equipment failures, namely: material push device failure 1001, material detection device failure 2001, filling device detection failure 4001, filling device positioning failure 4002, filling device filling failure 4003, capping device positioning failure 5001, capping device capping failure 5002, cap tightening device positioning failure 6001, and cap tightening device cap tightening failure 6002. The first 9 months are set as the training set with a total of 5,409,960 samples, and the last 3 months are set as the test set with a total of 1,803,321 samples.
[0087] The second step is the DataEmbedding layer that fuses different feature data, including the information after 1D convolution of the original data vector and the positional encoding information;
[0088] It should be noted that 1D convolution is an operation commonly used in fields such as signal processing, time series analysis, and natural language processing. It is mainly used to extract features from one-dimensional data. It slides a convolution kernel (or filter) over the input data, calculates the weighted sum of the local area, and generates an output feature map.
[0089] Specifically, the processing method of 1D convolution is as follows: set the size of the 1D convolution kernel to (k, M'), the stride to 1, the padding to 1, and at the same time, the channel dimension rises from M to M';
[0090] The processing formula steps of 1D convolution are:
[0091] Let the value of each position of the corresponding convolution kernel be W ij , the corresponding convolution sequence value be S ij , the convolution output value be S', and the bias be b, then:
[0092]
[0093] The dimension of the S' sequence after convolution is:
[0094] L * M' (9).
[0095] The positional encoding process is to perform positional encoding on the dimension of the S sequence output in the first step, that is, to encode L * M, and the encoding formula is expressed as:
[0096] PE t,2i = sin(t / 10000 2i / d ), PE t,2i+1 = cos(t / 10000 2i / d ) (10);
[0097] t ∈ [1, L] (11);
[0098] i ∈ [1, M' / / 2] (12);
[0099] Among them, t represents the t-th sample of the sequence, i represents the dimension of the encoding, and the size of the output matrix is represented as:
[0100] L * M” (13);
[0101] Finally, add the S' sequence and the PE matrix to obtain the eigenvector data X, which is used as the input data for the third step.
[0102] In the third step, first perform a fast Fourier transform on the feature data after fusion processing to obtain the corresponding amplitudes and frequencies of TOP_K, and then reshape the feature data according to the TOP_K frequencies, converting from 1D data to K 2D data;
[0103] Specifically, the fast Fourier transform is expressed as:
[0104] A = Avg(Amp(FFT(X 1D ))) (14);
[0105] Among them, the input feature data is a one-dimensional time series X with a time length of L and a number of channels of M' 1D , and then sort the processed amplitudes from largest to smallest and select the top K, which is expressed as:
[0106]
[0107] After that, reshape X through tensor reshaping 1D The data dimension is transformed into X 2D , with dimensions (L / p i , p i , M'), where i ∈ [1, k];
[0108] If L cannot be divided evenly by p i , then perform a zero-padding operation on the original X 1D data. First, calculate the length that needs to be supplemented according to L and p i , which is expressed as:
[0109] l i = ((L / / p i ) + 1) * pi -L (16);
[0110] i ∈ [1, k] (17);
[0111] Then generate a tensor of all zeros with size (l i , M'), and finally concatenate this all-zero tensor to the original X 1D data, that is, concatenate in the dimension of sequence length.
[0112] In the fourth step, first pass the K 2D data through the Inception_Block layer, then fuse the features of different-sized convolutions for each 2D data, and finally perform weighted summation on the average features of the K 2D data according to the amplitude size;
[0113] Specifically, the input X 2D data passing through the Inception_Block layer has dimensions (L / p i , p i , M'), where i ∈ [1, k]. The Inception_Block layer consists of multiple convolutional network layers with different window sizes, and the convolutional kernel sizes are (1, 1), (3, 3), (5, 5), (7, 7), (9, 9), and (11, 11);
[0114] The processing method of 2D convolution is to set the size of the 2D convolutional kernel to (k1, k2, M'), where k1, k2 ∈ {1, 3, 5, 7, 9, 11}, the stride is 1, and the channel dimension remains M' unchanged. The processing formula of 2D convolution is expressed as:
[0115]
[0116] where, W ijl represents the value at each position of the convolutional kernel, X ijl represents the convolutional sequence value, X' 2D represents the convolutional output value, b represents the bias, and after convolution, the dimension of the X' 2D tensor is (L / p i , p i , M'), where i ∈ [1, k]. Finally, after the network passes through multi-scale convolutional processing, it is all processed by Dropout to achieve regularization.
[0117] Perform softmax calculation on the frequencies {f1, f2... f 2D} corresponding to X', which is expressed as: k}
[0118]
[0119] where \(i\in[1,k]\), weighted summation is performed on \(X'\) according to the obtained weights 2D and then put into the non-linear function GELU for processing, and finally the skip connection of the residual is added to improve the accuracy. The dimension of the processed tensor is \((L / p i ,p i ,M')\), which is used as the feature data to input the next step.
[0120] In the fifth step, the feature data obtained by weighted summation is mapped to \(n\) categories through a fully connected layer, and the obtained feature representations are \(h_1h_2\cdots h L For the input 2D tensor with dimensions \((L / p i ,p i ,M')\), it is reshaped into a 1D tensor with dimensions \((L,M')\), and then connected to a fully connected layer and mapped to \(n\) categories;
[0121] In the sixth step, the feature data obtained by weighted summation is input into the CRF layer network for processing. After further constraint by the CRF according to the transition probability matrix, the output prediction results are expressed as \(y_1y_2\cdots y L .
[0122] Specifically, let the transition score matrix \(T\) be introduced, and the matrix element \(T ij represents the transition score from fault type \(i\) to fault type \(j\). There are a total of \(n\) fault types, so \(T\in R n*n ;
[0123] Let the length of the time series be \(L\), then the score matrix \(P\) of the output layer belongs to \(R L*n , where the matrix element \(P ij represents the score at the \(i\)-th moment of the observed time series under the \(j\)-th fault type. When the input is \(H=(h_1h_2\cdots h L ), and the output fault sequence is \(Y=(y_1y_2\cdots y L ), then the total score of this fault type is expressed as:
[0124]
[0125] Normalize the sequence path to generate the probability distribution about the output sequence \(Y\), which is expressed as:
[0126]
[0127] Then maximize the logarithmic probability value of the processed fault type sequence, which is expressed as:
[0128]
[0129] This formula can enable the model to generate the correct fault type sequence, and select the fault sequence with the highest total score as the optimal sequence in the decoding stage;
[0130]
[0131] Finally, the Viterbi algorithm is selected to find the optimal fault type sequence during prediction.
[0132] The data is transformed from one dimension to two dimensions by Fourier transform to obtain the hidden periodic information in the data, which improves the accuracy of model prediction. At the same time, due to the difference between the backpropagation path of the convolutional network and the sequence time, the problems of gradient explosion and gradient disappearance are avoided, and parallel processing can be achieved to improve the running speed of the model. Finally, it is found that there is a transition probability distribution between the fault types on the production line, and the CRF algorithm is used to decode it to obtain the optimal fault type sequence, further improving the accuracy of model prediction.
[0133] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent substitution on some of the technical features. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An automatic fault detection method for filling equipment based on the TimesNet and CRF fusion network, characterized in that, The detection method includes the following steps: In the first step, extract the characteristic data of M induction devices in the production line, and use a fixed-length sliding window L to segment the data to obtain data vectors as the input of the model; In the second step, a DataEmbedding layer that fuses different characteristic data, including the information after 1D convolution of the original data vector and the position encoding information; In the third step, first perform a fast Fourier transform on the fused characteristic data to obtain the amplitudes and frequencies corresponding to TOP_K, and then reshape the characteristic data according to the TOP_K frequencies, converting from 1D data to K 2D data; In the fourth step, first pass the K 2D data through the Inception_Block layer, then fuse the characteristics of different-sized convolutions for each 2D data, and finally perform a weighted sum of the average characteristics of the K 2D data according to the amplitude size; In the fifth step, the weighted-sum feature data is mapped to n categories through a fully connected layer, and the obtained features are represented as h1h2…h L ; Step 6: Input the feature data obtained by weighted summation into the CRF layer network for processing. After further constraint by the CRF according to the transition probability matrix, the predicted result is output as y1y2…y L .
2. The automatic fault detection method for filling equipment based on the TimesNet and CRF fusion network according to claim 1, characterized in that, The dimension of the time series data obtained by the M induction devices is M-dimensional, and the data segmentation is expressed as: X train = {X (1) , X (2) , …, X (L)}} (1); X (i) ∈R m (2).
3. The automatic fault detection method for filling equipment based on the TimesNet and CRF fusion network according to claim 2, wherein The segmented data is normalized. Let the mean vector of the sample characteristics be expressed as: μ = {μ1, μ2... μ j} (3); Among them, μ j represents the average value of the features in the j-th dimension. Let the standard deviation vector of the sample features be expressed as: σ = {σ1, σ2... σ j} (4); where, σ j represents the standard deviation of the j-th dimensional feature, and the normalization calculation is expressed as: Among them, represents the value of the j-th feature of the i-th sample. Then the standardized data sequence S is expressed as: S = {S1, S2... S L} (6); The dimension of the output S sequence is expressed as: L*M (7).
4. The automatic fault detection method for filling equipment based on the TimesNet and CRF fusion network according to claim 3, characterized in that, The processing of the 1D convolution is to set the size of the 1D convolution kernel to (k, M'), the stride to 1, the padding to 1, and at the same time the channel dimension rises from M to M'; The output calculation of the processing of the 1D convolution is expressed as: Among them, W ij represents the value at each position of the convolution kernel, S ij represents the convolution sequence value, S' represents the convolution output value, b represents the bias, and the dimension of the S' sequence after convolution is expressed as: L*M' (9).
5. The automatic fault detection method for filling equipment based on the TimesNet and CRF fusion network according to claim 3, characterized in that The position encoding processing is to perform position encoding on the dimension of the output S sequence, and the encoding formula is expressed as: PE t,2i = sin(t / 10000 2i / d ), PE t,2i+1 = cos(t / 10000 2i / d ) (10); t∈[1,L] (11); i∈[1,M' / / 2](12); Among them, t represents the t-th sample of the sequence, i represents the encoding dimension, and the size of the output matrix is expressed as: L*M”(13); Finally, add the S' sequence and the PE matrix to obtain the characteristic vector data X as the input data for the third step.
6. The automatic fault detection method for filling equipment based on the TimesNet and CRF fusion network according to claim 5, characterized in that, The fast Fourier transform is expressed as: A = Avg(Amp(FFT(X 1D ))) (14); Among them, the input feature data is a one-dimensional time series X with a time length of L and a channel number of M'. 1D , and then sort the processed amplitudes in descending order, and select the top K, which is expressed as: After that, X is reshaped through tensor reshaping 1D The data dimension is transformed into X 2D , with the dimension of (L / p i , p i , M'), where i ∈ [1, k]; If L does not divide p i , then for the original X 1D data, perform a zero-padding operation. First, calculate the length to be padded based on L and p i and express it as: l i = ((L / / p i ) + 1) * p i - L (16); i∈[1,k](17); Then generate a tensor of all zeros, with size (l i , M'), and finally concatenate this all-zero tensor to the original X 1D data.
7. The automatic fault detection method for filling equipment based on the TimesNet and CRF fusion network according to claim 6, wherein X, the input passing through the Inception_Block layer 2D data, with dimensions (L / p i , p i , M'), where i ∈ [1, k]. The Inception_Block layer is composed of multiple convolutional network layers with different window sizes, and the convolutional kernel sizes are (1, 1), (3, 3), (5, 5), (7, 7), (9, 9) and (11, 11) respectively; The processing method of the 2D convolution is to set the size of the 2D convolution kernel to (k1, k2, M'), where k1, k2 ∈ {1, 3, 5, 7, 9, 11}, the stride is 1, and at the same time the channel dimension remains M' unchanged. The processing formula of the 2D convolution is expressed as: Among them, W ijl represents the value at each position of the convolution kernel, X ijl represents the convolution sequence value, X' 2D represents the convolution output value, b represents the bias, and after convolution, the dimension of the X' 2D tensor is (L / p i , p i , M'), where i ∈ [1, k]. Finally, after the network undergoes multi-scale convolution processing, it is all processed by Dropout to achieve regularization.
8. The automatic fault detection method for filling equipment based on the TimesNet and CRF fusion network according to claim 7, characterized in that, Take X' 2D The corresponding frequencies {f1, f2... f k} are subjected to softmax calculation, expressed as: where \(i\in[1,k]\), weighted summation is performed on \(X'\) according to the obtained weights, then it is put into the non-linear function GELU for processing, and finally a skip connection with residual is added. The dimension of the processed tensor is \((L / p 2D ,p i ,M') i 9. The automatic fault detection method for filling equipment based on the TimesNet and CRF fusion network according to claim 8, characterized in that, For the input 2D tensor with dimensions (L / p i , p i , M'), it is reshaped into a 1D tensor with dimensions (L, M'), and then connected to a fully connected layer to be mapped to n classes.
10. The automatic fault detection method for filling equipment based on the TimesNet and CRF fusion network according to claim 1, wherein Let the transfer score matrix \(T\) be introduced, and the matrix element \(T\) ij represents the transfer score from fault type \(i\) to fault type \(j\). There are a total of \(n\) fault types, so \(T\in R\) n*n ; Let the length of the time series be \(L\), then the score matrix \(P\in\mathbb{R}\) of the output layer L*n , where the matrix element \(P\) ij represents the score at the \(i\)-th moment of the observed time series under the \(j\)-th fault type. Given the input \(H=(h_1 h_2\cdots h\) L ), and the output fault sequence \(Y=(y_1 y_2\cdots y\) L ), then the total score of this fault type is expressed as: Normalize the sequence path to generate a probability distribution about the output sequence Y, which is expressed as: Then maximize the logarithmic probability value of the processed fault type sequence, which is expressed as: Finally, select the Viterbi algorithm to find the optimal fault type sequence during prediction.