Off-grid photovoltaic array fault diagnosis method based on ConvTran

By adopting a ConvTran-based method in photovoltaic array fault diagnosis, using the Transformer model and convolution module combined with tAPE and eRPE position coding, the problem of insufficient accuracy and applicability of photovoltaic array fault diagnosis in the prior art is solved, and more efficient and accurate fault diagnosis is achieved.

CN119988873AActive Publication Date: 2025-05-13CHINA THREE GORGES UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510072503.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The existing photovoltaic array fault diagnosis methods have shortcomings in real-time monitoring and large-scale photovoltaic array fault diagnosis, especially due to the variability and complexity of the operating environment of the photovoltaic array, the mathematical model cannot accurately reflect the actual operating conditions, affecting the accuracy of fault diagnosis.

Method used

Using ConvTran-based off-grid photovoltaic array fault diagnosis method, the length of the time series is shortened and local information is captured by building a Transformer model and using a convolutional module on its architecture. Combining tAPE and eRPE position coding with convolutional input coding improves the position and data embedding of time series data.

Benefits of technology

It improves the applicability and stability of photovoltaic array fault diagnosis, achieves better results for multivariate time series classification tasks, can distinguish slight differences between short-circuit faults and partial shadow occlusion, and significantly improves the accuracy of off-grid photovoltaic array fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988873A_ABST
    Figure CN119988873A_ABST
Patent Text Reader

Abstract

The invention discloses an off-grid photovoltaic array fault diagnosis method based on ConvTran. The off-grid photovoltaic array fault diagnosis method based on ConvTran comprises the following steps: constructing a Transform model; a convolution module is used on a Transform architecture, before an input embedded vector is input into a Transform module, a position embedded vector generated by tAPE is added into the input embedded vector, and after the final output of the Transform module is obtained, global average pooling and a full connection layer are applied for processing to obtain a model with more translation invariance. And finally, applying a Softmax function to obtain a classification prediction result. And collecting multivariate time sequence data of four channels of the photovoltaic array, inputting the multivariate time sequence data into the ConvTran network model, and carrying out fault diagnosis on the photovoltaic array. The method is more suitable for fault type classification of multivariate time series, and shows more excellent applicability and accuracy in off-grid photovoltaic array fault diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of online monitoring and fault diagnosis of electric power equipment, and in particular to a ConvTran-based off-grid photovoltaic array fault diagnosis method. Background Art

[0002] At present, the fault diagnosis methods for photovoltaic arrays can be divided into two categories: visual imaging and electrical characteristic parameters. Visual imaging requires the use of specific detection instruments in a specific environment to diagnose the faults of photovoltaic components. It is also impossible to monitor large photovoltaic arrays in real time. The threshold for use is high and it is difficult to obtain a large number of fault samples. The fault diagnosis methods for photovoltaic arrays with electrical characteristic parameters also include circuit structure method, mathematical model method and machine learning method. The circuit structure method configures the voltage and current sensor embedding method corresponding to the photovoltaic string according to the photovoltaic arrays with different connection structures to locate the approximate range of the fault. However, this method requires a large number of sensors to be arranged at the photovoltaic array, which increases the complexity of wiring and is too costly. It is not suitable for large photovoltaic power stations. The mathematical model method first constructs a mathematical model of the expected photovoltaic array output, and then imports the actual measured voltage and current data into the pre-constructed mathematical model to estimate the working state of the photovoltaic array. However, due to the variability and complexity of the photovoltaic array operating environment, the constructed mathematical model cannot accurately reflect the actual operating status of the photovoltaic array, making the adaptability of this method weak under actual working conditions, thereby affecting the accuracy of fault diagnosis.

[0003] Machine learning is currently the mainstream method for photovoltaic array fault diagnosis. A mapping relationship between fault feature values ​​and fault types in photovoltaic arrays is established through machine learning algorithms to achieve fault diagnosis. Compared with traditional fault diagnosis methods, machine learning shows faster speed and higher accuracy in identifying and classifying photovoltaic array faults.

[0004] Since the output characteristics of the photovoltaic array are related to meteorological conditions such as solar irradiance and temperature, the change of the electrical signal is closely related to time. Therefore, the time-series voltage and current fault diagnosis method is selected, that is, the voltage and current of each branch of the photovoltaic array are measured at the DC junction box of the photovoltaic power generation system. This fault detection method can be used for diagnosis during the operation of the photovoltaic system. Fault identification and location based only on the time-series changes of the photovoltaic array voltage and current is the most efficient fault detection method, which greatly reduces the workload of collecting and preprocessing data sets. Therefore, how to use multivariate time series classification methods to perform fault diagnosis of photovoltaic arrays is an urgent problem to be solved by technicians in this field. Summary of the invention

[0005] In view of this, the present invention provides an off-grid photovoltaic array fault diagnosis method based on ConvTran, which combines the advantages of the ConvTran network model to improve the applicability and stability of photovoltaic array fault diagnosis. The specific method steps are as follows:

[0006] S1: Build the Transformer model, which includes four parts: input module, encoder module, decoder module and output module;

[0007] S2: Use convolutional modules on the Transformer architecture to shorten the length of the time series and capture local information in the original time series;

[0008] S3: The position embedding vector generated by tAPE is added to the input embedding vector before it is fed into the Transformer module so that the model can capture the temporal order of the time series.

[0009] S4: After obtaining the final output of the Transformer module, global average pooling and fully connected layers are applied to obtain a model with more translation invariance. Finally, the Softmax function is applied to obtain the classification prediction result, and finally the ConvTran model is constructed.

[0010] S5: The voltages UA and UB and branch currents IA and IB of the A and B lines on the DC output side of the PV array are collected to form a four-channel multivariate time series data. A time point is collected every 1 second, and a group of samples is packaged every 30 seconds. The samples are input into the ConvTran network model to perform fault diagnosis on the PV array.

[0011] Furthermore, the input of the Transformer architecture is sequence data. First, the input sequence is converted into a feature vector through an embedding algorithm. If the input is a sentence, then word embedding algorithms such as word2vec, GloVe, and one-hot encoding can be used to convert the input sentence into a word vector after word segmentation. After the embedding operation is completed on the input sequence, the feature vector is also positionally encoded. The position encoding has the same dimension as the input embedding, so that the position encoding information and the encoding of the corresponding position sequence vector can be added together, so that the self-attention mechanism can take into account both the order information of the sequence and all the input sequences. For multivariate time series, the Transformer trains all input time series at the same time, and the self-attention layer cannot retain the position information of the time series in the Transformer architecture, so position encoding is needed to help it understand the order of the sequence. Position encoding methods include absolute position encoding and relative position encoding to enhance the time context of the time series input.

[0012] The encoder module consists of multiple encoder stacks with the same structure. Each encoder consists of two sublayers: the self-attention layer and the fully connected feed-forward network. Each sublayer uses a residual connection and then performs layer normalization. It should be noted that although the structure of each encoder is exactly the same, the weight parameters are different, the parameters are trained independently, and the output of the encoder module is input into each decoder to execute the multi-head attention mechanism.

[0013] The decoder module is also composed of multiple decoder stacks with the same structure. In addition to the self-attention layer and the feedforward network, the decoder has a third sublayer, which performs a multi-head attention mechanism on the output of the encoder stack. Unlike the encoder, the decoder's self-attention layer adds masking to ensure that the prediction of position i can only rely on the known output before position i, preventing the current position from being affected by subsequent positions. Like the encoder module, each sublayer of the decoder module also performs residual connection and layer normalization operations. Although the structure of each decoder is exactly the same, the weight parameters are different and the parameters are trained independently.

[0014] The final output module consists of a linear layer and a Softmax layer. The linear layer is a simple fully connected neural network that maps the output vector of the decoding module to a longer vector, the logits vector. The Softmax layer converts the attention score of each sequence segment into a probability distribution between 0 and 1, and selects the result corresponding to the highest probability as the output of this time step.

[0015] The attention mechanism can be described as the process of mapping a query and a set of key-value pairs to an output, where the query, key, value, and output are all vectors.

[0016] For a d x The input sequence x of dimension t , xt={x1,x2,…,x L},in L is the length of the time series, and after self-attention calculation, d z The output sequence of dimension is z t , zt={z1,z2,…,z L},in z i It is calculated by the weighted sum of the input elements, as shown in formula (1).

[0017]

[0018] The weight of each coefficient is α i,j It is calculated by the Softmax function, as shown in formula (2).

[0019]

[0020] Where e ij is the attention weight from position j to position i, which is calculated by the scaled dot product. As shown in formula (3), the higher the correlation between position j and position i, the larger the result after the dot product operation. It is a parameter matrix, which is different for each layer.

[0021]

[0022] Therefore, the attention mechanism essentially calculates the weight coefficient of the corresponding Value through the Query and Key in the input element, and then performs weighted summation of the Value according to the weight coefficient.

[0023] Multi-Head Attention (MHA) replaces the original mode of calculating self-attention only once, and uses h different learned linear transformations to linearly map queries, keys, and values ​​respectively, and then performs scaled dot product attention calculations simultaneously, and finally concatenates the generated output values ​​for linear transformation to produce the final result.

[0024] The core idea of ​​the multi-head attention mechanism is to map the same query, key, and value to different high-dimensional subspaces while maintaining the overall parameter scale unchanged, and then independently perform attention calculations in these subspaces, and finally fuse the attention information in different subspaces. This reduces the dimension of a single vector in each attention calculation, and in a sense prevents overfitting. Since attention has different distribution patterns in different subspaces, multi-head attention is actually exploring the correlation between sequences at different angles, obtaining multiple feature expressions through different self-attention layers, and then splicing the features captured in different subspaces together.

[0025] The initial self-attention considers the absolute position and adds the absolute position embedding P = (p1, ..., p L ), as shown in formula (4).

[0026] x i =x i +p i(4)

[0027] In the formula, the position embedding There are several ways to encode the absolute position, including encoding the fixed position via sine and cosine functions of different frequencies, called Vanilla APE, or learnable encoding via trainable parameters.

[0028] By using sine and cosine functions for fixed position encoding, the position d at the i-th time step is model The dimensional embedding can be expressed as formula (5).

[0029]

[0030] Where, d model is the embedding dimension, the dimension of position embedding is the same as the dimension of sequence feature vector; k is the dimension, k is Within the range of k is the frequency term. Each position in the embedding dimension will get a combination of sine and cosine functions of different periods, thus generating unique texture position information, and ultimately enabling the Transformer model to learn the dependencies and timing characteristics between different positions.

[0031] In addition to absolute position embedding, relative position embedding considers the pairwise relationship between input elements. This method embeds the input element x i and x j The relative distance between them is encoded as a vector The encoding vector is embedded into the self-attention module, and equations (1) and (3) are modified to obtain equations (6) and (7). In this way, the pairwise position relationship is trained during the training process of the Transformer.

[0032]

[0033] Relative position information is provided to the model at two levels: value and key. First, relative position information is incorporated into the model as an additional component of the key. The Softmax operation is shown in Equation (3), which is the same as Vanilla self-attention. Finally, the relative position information is re-provided as a sub-component of the value matrix. In addition, considering that the utility of relative position information will drop significantly after exceeding a certain distance, in order to improve the efficiency and performance of the model, the clip function is introduced to limit and reduce the number of required parameters. When calculating the attention between position i and position j, the distance between them is considered, and the encoding calculation is shown in Equations (8) to (10).

[0034]

[0035] clip(x,k)=max(-k,min(k,x)) (10)

[0036] Where p V and p K are the trainable weights encoding the relative position of values ​​and keys, respectively. in The scalar k is the maximum relative distance.

[0037] From Equation (8), we can see that due to the additional relative position encoding, it requires O(L 2 d) memory. A new method for computing relative position encodings, called the vector method, uses offset operations to reduce its intermediate storage requirements from O(L 2 d) is reduced to O(Ld), abandoning the additional relative position embedding corresponding to the value term and focusing only on the key component. The encoding calculation is shown in Equations (11) and (12). Among them, the Skew program uses padding, reshaping, and slicing to reduce memory requirements.

[0038]

[0039] S rel =Skew(W Q P) (12)

[0040] When applied to time series data, the Transformer model requires effective position encoding to capture the ordering of time series data. This technique applies a new time series absolute position encoding method (tAPE), which incorporates the time series length and input embedding dimension into the absolute position encoding. An efficient relative position encoding method (eRPE) is also applied. These two position encoding methods are simple and effective and can be easily integrated into the Transformer module to improve the generalization ability of time series. The ConvTran network model combines tAPE and eRPE with convolution-based input encoding to improve the position and data embedding of time series data, achieving excellent results in multivariate time series classification tasks.

[0041] Absolute position encoding was originally proposed for language modeling tasks, and usually uses high embedding dimensions such as 512 or 1024 to embed the input of length 512. Higher embedding dimensions can better reflect the similarity between different positions. When using lower embedding dimensions for position encoding, the similarity between two positions calculated by calculating the dot product does not always decrease as the distance between the two positions increases, and the distance perception property disappears.

[0042] Although high embedding dimensions show an ideal monotonically decreasing trend as the distance between two positions increases, they are not suitable for encoding time series datasets because most time series datasets have low data dimensions, and higher embedding dimensions may reduce model throughput due to additional parameters and increase the probability of model overfitting. On the other hand, at low embedding dimensions, the similarity between two random embedding vectors is very high, which is called anisotropy. Therefore, the embedding vector space cannot be fully utilized to distinguish between two positions, and position encoding fails when the embedding dimension is low.

[0043] Therefore, the algorithm requires that the position embedding of the time series is distance-aware and isotropic. In order to incorporate distance awareness, the length of the time series is used in equation (5). In this equation, w k refers to the frequency of the sine and cosine functions that generate the embedding vector. If not modified, as the sequence length L increases, the dot product between positions will become more and more irregular, resulting in a loss of distance perception. After introducing the length parameter in the frequency terms of the sine and cosine functions in Equation (5), the dot product will maintain a monotonic smooth trend.

[0044] As the embedding dimension d model As the value of d increases, the vector embedding is more likely to be sampled from low-frequency sinusoidal functions, resulting in anisotropy. To alleviate this problem, d model The parameters are simultaneously incorporated into the frequency terms of the sine and cosine functions in equation (5). A new time series-based absolute position encoding method (tAPE) is used, where Considering the input embedding dimension d model And the time series length L, as shown in formula (13):

[0045]

[0046] Compared with the Vanilla absolute position encoding, the absolute position encoding method using tAPE shows that as the distance between two positions in the time series increases, the dot product representing the similarity between the two positions has a more stable monotonous downward trend, and the similarity between the tAPE embedding vectors decreases. This is because tAPE can use the embedding space to provide isotropic encoding, maintain the distance perception characteristics, and better use the embedding vector space to distinguish between two positions.

[0047] Input embedding is the basis of all previous relative position encoding methods. The position matrix is ​​added or multiplied with the query, key and value matrices. In this paper, an efficient relative position encoding (eRPE) model independent of input embedding is introduced.

[0048] The calculation formula used in the eRPE model is shown in formula (14).

[0049]

[0050] In the formula, L is the sequence length, e i,j is the attention weight, w i-j is a learnable scalar, representing the relative position weight between position i and position j,

[0051] For each attention module in the multi-head attention mechanism, create a trainable parameter w of size 2L-1, because the maximum distance is 2L-1. For the index i and j of two positions, the corresponding relative scalar is w i-j+L , where the index starts at 1, not 0, and you need to index L from 2L-1 vectors 2 elements.

[0052] First, the relative position embedding w i-j is a static parameter that has nothing to do with the input, and the attention weight e i,j is dynamically determined by the representation of the input sequence. Attention adapts to the input sequence by inputting an adaptive weighting strategy, enabling the model to capture the complex relationships between different time points, which is exactly the most needed feature when extracting high-level concepts from time series, such as seasonal components in time series. However, when the data size is limited, using attention will face a greater risk of overfitting.

[0053] Secondly, the relative position embedding w i-j The relative displacement between positions i and j is considered, rather than their values. This is similar to the translation invariance of convolution, which has been shown to enhance generalization, so w i-jis considered as a scalar rather than a vector in order to achieve translation invariance without increasing the number of parameters. In addition, w for all (i, j) i-j All values ​​can be included in the paired dot product attention function, minimizing the additional computation. This efficient relative position encoding method is eRPE.

[0054] eRPE first applies the Softmax function to the attention matrix and then adds the relative position information to the model. Because the position values ​​that have not passed the Softmax function are clearer, the attention model performs better. Compared with existing models that apply Softmax to relative position embedding, eRPE's clearer position embedding is more conducive to time series classification tasks.

[0055] From the previous analysis, we can see that the complexity of global attention is the quadratic of the sequence length. If the attention proposed in equation (14) is directly applied to the original time series, the calculation speed will be too slow for long time series. Therefore, the ConvTran network model is introduced. First, the convolution module is used to shorten the length of the time series, and the feature map is reduced to a size with lower computational intensity before the new position encoding method is applied.

[0056] Using convolution modules on the Transformer architecture can not only reduce computational intensity and increase network training speed, but convolution operations are also very suitable for capturing local features. Therefore, convolution is used as the first module in the ConvTran model architecture to capture any discriminative local information in the original time series.

[0057] In the convolution module, M time convolution kernels are first applied to the input multivariate time series data, so that the ConvTran network model can extract the time information in the input sequence. Then the output of the time convolution kernel is combined with d model d x ×M-sized spatial convolution kernels are used to perform convolution operations to capture the correlation between variables in the original time series and construct d model This disjoint spatiotemporal convolution first expands the number of input channels and then compresses them. A key reason for this choice is that the feed-forward network (FFN) in the Transformer also expands the size of the input, and then projects the expanded hidden state back to the original size to capture spatial interactions.

[0058] Before the input embedding vector is fed into the Transformer module, the position embedding vector generated by tAPE is added to the input embedding vector so that the model can capture the temporal order of the time series. The size of the tAPE embedding vector is d model, which is the same as the input embedding vector. In multi-head attention, first, a linear layer is used to transform the input of dimension L×d model into a size of L×d z ×3, that is, to obtain the q (query), k (key), and v (value) matrices of size L×d z , where d z represents the dimension of the model and is a custom parameter. These q, k, and v matrices are reshaped into h×L×d z / h to represent the h-th attention head. Each attention head is responsible for capturing different patterns in the time series. For example, one attention head focuses on non-noise data, another attention head focuses on seasonal components, and another attention head focuses on trends. After obtaining the q, k, and v matrices, finally, the attention calculation is performed within the multi-head attention module using Equation (4-14).

[0059] The feed-forward network in the Transformer model is a multi-layer perceptron module, which consists of two linear layers and the Gaussian Error Linear Units (GELUs) activation function. GELUs introduce the idea of stochastic regularization in the activation function. It is a probabilistic description of the neuron input and is a combination of dropout, zoneout, and ReLU, which can improve the generalization ability of the model. Assume the input is X and the mask is m, then m follows a Bernoulli distribution F(x) = P(X < x), where X follows a standard normal distribution, that is, X ~ N(0,1). The mathematical expression of GELUs is shown in Equation (15).

[0060] GELU(x) = xΦ(x) = xP(X < x) (15)

[0061] Similar to the Transformer basic architecture, residual connections and layer normalization are also applied to the multi-head attention layer and the feed-forward network layer in the ConvTran network to obtain the final output of the Transformer module. Then, max pooling and global average pooling (GAP) are applied to the output of the ELU activation function of the last layer to obtain a more translation-invariant model. Finally, the Softmax function is applied to obtain the classification prediction result.

[0062] The present invention can achieve the following beneficial effects:

[0063] Compared with other improved algorithms based on CNN, such as FCN, MC-DCNN, ResNet, MLSTM-FCN, and the improved algorithm based on Transformer, GTN, the ConvTran algorithm has a more balanced ability to identify various faults of photovoltaic arrays, and has the highest fault diagnosis accuracy for off-grid photovoltaic arrays, especially in the case of partial shadow occlusion, the diagnosis accuracy is significantly higher than other algorithms. Combining the improved algorithms of convolution and Transformer, the ConvTran network model improves the location and data embedding of time series data in the Transformer architecture, has better classification effect in multivariate time series classification tasks, can distinguish the slight difference between short circuit faults and partial shadow occlusion, is more suitable for multivariate time series fault type classification, and shows better applicability and accuracy in off-grid photovoltaic array fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 This is the Transformer model architecture diagram of the present invention;

[0065] Figure 2 Schematic diagram of residual connection and layer normalization of the present invention;

[0066] Figure 3 This is a structural diagram of the multi-head attention system of the present invention;

[0067] Figure 4 A diagram of the self-attention module for relative position encoding of the present invention;

[0068] Figure 5 It is the overall architecture diagram of the ConvTran model of the present invention;

[0069] Figure 6 It is the GELU activation function curve diagram of the present invention;

[0070] Figure 7 This is the overall structure diagram of the grid-connected photovoltaic array fault simulation test platform of the present invention;

[0071] Figure 8 This is a confusion matrix diagram of photovoltaic array fault identification based on ConvTran in the present invention;

[0072] Fig. 9 This is a confusion matrix diagram of off-grid photovoltaic array fault identification for the comparison model algorithm of the present invention;

[0073] Fig.10 This is a graph showing the accuracy of the algorithms used in the present invention in identifying off-grid photovoltaic array faults;

[0074] Fig.11 This is a graph showing the total accuracy of the algorithms used in the present invention for off-grid photovoltaic array fault identification. DETAILED DESCRIPTION

[0075] See the instruction manual Figure 1-11 The present invention provides a fault diagnosis method for off-grid photovoltaic array based on ConvTran.

[0076] S1: Build the Transformer model, which includes four parts: input module, encoder module, decoder module and output module;

[0077] S2: Use convolutional modules on the Transformer architecture to shorten the length of the time series and capture local information in the original time series;

[0078] S3: The position embedding vector generated by tAPE is added to the input embedding vector before it is fed into the Transformer module so that the model can capture the temporal order of the time series.

[0079] S4: After obtaining the final output of the Transformer module, global average pooling and fully connected layers are applied to obtain a model with more translation invariance. Finally, the Softmax function is applied to obtain the classification prediction result, and finally the ConvTran model is constructed.

[0080] S5: The voltages UA and UB and branch currents IA and IB of the A and B lines on the DC output side of the PV array are collected to form a four-channel multivariate time series data. A time point is collected every 1 second, and a group of samples is packaged every 30 seconds. The samples are input into the ConvTran network model to perform fault diagnosis on the PV array.

[0081] Furthermore, the input of the Transformer architecture is sequence data. First, the input sequence is converted into a feature vector through an embedding algorithm. If the input is a sentence, then word embedding algorithms such as word2vec, GloVe, and one-hot encoding can be used to convert the input sentence into a word vector after word segmentation. After the embedding operation is completed on the input sequence, the feature vector is also positionally encoded. The position encoding has the same dimension as the input embedding, so that the position encoding information and the encoding of the corresponding position sequence vector can be added together, so that the self-attention mechanism can take into account both the order information of the sequence and all the input sequences. For multivariate time series, the Transformer trains all input time series at the same time, and the self-attention layer cannot retain the position information of the time series in the Transformer architecture, so position encoding is needed to help it understand the order of the sequence. Position encoding methods include absolute position encoding and relative position encoding to enhance the time context of the time series input.

[0082] The encoder module consists of multiple encoder stacks with the same structure. Each encoder consists of two sublayers: the self-attention layer and the fully connected feed-forward network. Each sublayer uses a residual connection and then performs layer normalization. It should be noted that although the structure of each encoder is exactly the same, the weight parameters are different, the parameters are trained independently, and the output of the encoder module is input into each decoder to execute the multi-head attention mechanism.

[0083] The decoder module is also composed of multiple decoder stacks with the same structure. In addition to the self-attention layer and the feedforward network, the decoder has a third sublayer, which performs a multi-head attention mechanism on the output of the encoder stack. Unlike the encoder, the decoder's self-attention layer adds masking to ensure that the prediction of position i can only rely on the known output before position i, preventing the current position from being affected by subsequent positions. Like the encoder module, each sublayer of the decoder module also performs residual connection and layer normalization operations. Although the structure of each decoder is exactly the same, the weight parameters are different and the parameters are trained independently.

[0084] The final output module consists of a linear layer and a Softmax layer. The linear layer is a simple fully connected neural network that maps the output vector of the decoding module to a longer vector, the logits vector. The Softmax layer converts the attention score of each sequence segment into a probability distribution between 0 and 1, and selects the result corresponding to the highest probability as the output of this time step.

[0085] The attention mechanism can be described as the process of mapping a query and a set of key-value pairs to an output, where the query, key, value, and output are all vectors.

[0086] For a d x The input sequence x of dimension t , xt={x1,x2,…,x L},in L is the length of the time series, and after self-attention calculation, d z The output sequence of dimension is z t , zt={z1,z2,…,z L},in z i It is calculated by the weighted sum of the input elements, as shown in formula (1).

[0087]

[0088] The weight of each coefficient is α i,j It is calculated by the Softmax function, as shown in formula (2).

[0089]

[0090] Where e ij is the attention weight from position j to position i, which is calculated by the scaled dot product. As shown in formula (3), the higher the correlation between position j and position i, the larger the result after the dot product operation. It is a parameter matrix, which is different for each layer.

[0091]

[0092] Therefore, the attention mechanism essentially calculates the weight coefficient of the corresponding Value through the Query and Key in the input element, and then performs weighted summation of the Value according to the weight coefficient.

[0093] Multi-Head Attention (MHA) replaces the original mode of calculating self-attention only once, and uses h different learned linear transformations to linearly map queries, keys, and values ​​respectively, and then performs scaled dot product attention calculations simultaneously, and finally concatenates the generated output values ​​for linear transformation to produce the final result.

[0094] The core idea of ​​the multi-head attention mechanism is to map the same query, key, and value to different high-dimensional subspaces while maintaining the overall parameter scale unchanged, and then independently perform attention calculations in these subspaces, and finally fuse the attention information in different subspaces. This reduces the dimension of a single vector in each attention calculation, and in a sense prevents overfitting. Since attention has different distribution patterns in different subspaces, multi-head attention is actually exploring the correlation between sequences at different angles, obtaining multiple feature expressions through different self-attention layers, and then splicing the features captured in different subspaces together.

[0095] The initial self-attention considers the absolute position and adds the absolute position embedding P = (p1, ..., p L ), as shown in formula (4).

[0096] x i =x i +pi (4)

[0097] In the formula, the position embedding There are several ways to encode the absolute position, including encoding the fixed position via sine and cosine functions of different frequencies, called Vanilla APE, or learnable encoding via trainable parameters.

[0098] By using sine and cosine functions for fixed position encoding, the position d at the i-th time step is model The dimensional embedding can be expressed as formula (5).

[0099]

[0100] Where, d model is the embedding dimension, the dimension of position embedding is the same as the dimension of sequence feature vector; k is the dimension, k is Within the range of k is the frequency term. Each position in the embedding dimension will get a combination of sine and cosine functions of different periods, thus generating unique texture position information, and ultimately enabling the Transformer model to learn the dependencies and timing characteristics between different positions.

[0101] In addition to absolute position embedding, relative position embedding considers the pairwise relationship between input elements. This method embeds the input element x i and x j The relative distance between them is encoded as a vector The encoding vector is embedded into the self-attention module, and equations (1) and (3) are modified to obtain equations (6) and (7). In this way, the pairwise position relationship is trained during the training process of the Transformer.

[0102]

[0103] Relative position information is provided to the model at two levels: value and key. First, relative position information is incorporated into the model as an additional component of the key. The Softmax operation is shown in Equation (3), which is the same as Vanilla self-attention. Finally, the relative position information is re-provided as a sub-component of the value matrix. In addition, considering that the utility of relative position information will drop significantly after exceeding a certain distance, in order to improve the efficiency and performance of the model, the clip function is introduced to limit and reduce the number of required parameters. When calculating the attention between position i and position j, the distance between them is considered, and the encoding calculation is shown in Equations (8) to (10).

[0104]

[0105] clip(x,k)=max(-k,min(k,x)) (10)

[0106] Where p V and p K are the trainable weights encoding the relative position of values ​​and keys, respectively. in The scalar k is the maximum relative distance.

[0107] However, it can be seen from Equation (8) that due to the additional relative position encoding, it requires O(L 2 d) memory. A new method for computing relative position encodings, called the vector method, uses offset operations to reduce its intermediate storage requirements from O(L 2 d) is reduced to O(Ld), abandoning the additional relative position embedding corresponding to the value term and focusing only on the key component. The encoding calculation is shown in Equations (11) and (12). Among them, the Skew program uses padding, reshaping, and slicing to reduce memory requirements.

[0108]

[0109] S rel =Skew(W Q P) (12)

[0110] When applied to time series data, the Transformer model requires effective position encoding to capture the ordering of time series data. This technique applies a new time series absolute position encoding method (tAPE), which incorporates the time series length and input embedding dimension into the absolute position encoding. An efficient relative position encoding method (eRPE) is also applied. These two position encoding methods are simple and effective and can be easily integrated into the Transformer module to improve the generalization ability of time series. The ConvTran network model combines tAPE and eRPE with convolution-based input encoding to improve the position and data embedding of time series data, achieving excellent results in multivariate time series classification tasks.

[0111] Absolute position encoding was originally proposed for language modeling tasks, and usually uses high embedding dimensions such as 512 or 1024 to embed the input of length 512. Higher embedding dimensions can better reflect the similarity between different positions. When using lower embedding dimensions for position encoding, the similarity between two positions calculated by calculating the dot product does not always decrease as the distance between the two positions increases, and the distance perception property disappears.

[0112] Although high embedding dimensions show an ideal monotonically decreasing trend as the distance between two positions increases, they are not suitable for encoding time series datasets because most time series datasets have low data dimensions, and higher embedding dimensions may reduce model throughput due to additional parameters and increase the probability of model overfitting. On the other hand, at low embedding dimensions, the similarity between two random embedding vectors is very high, which is called anisotropy. Therefore, the embedding vector space cannot be fully utilized to distinguish between two positions, and position encoding fails when the embedding dimension is low.

[0113] Therefore, the algorithm requires that the position embedding of the time series is distance-aware and isotropic. In order to incorporate distance awareness, the length of the time series is used in equation (5). In this equation, w k refers to the frequency of the sine and cosine functions that generate the embedding vector. If not modified, as the sequence length L increases, the dot product between positions will become more and more irregular, resulting in a loss of distance perception. After introducing the length parameter in the frequency terms of the sine and cosine functions in Equation (5), the dot product will maintain a monotonic smooth trend.

[0114] As the embedding dimension d model As the value of d increases, the vector embedding is more likely to be sampled from low-frequency sinusoidal functions, resulting in anisotropy. To alleviate this problem, d model The parameters are simultaneously incorporated into the frequency terms of the sine and cosine functions in equation (5). A new time series-based absolute position encoding method (tAPE) is used, where Considering the input embedding dimension d model And the time series length L, as shown in formula (13):

[0115]

[0116] Compared with the Vanilla absolute position encoding, the absolute position encoding method using tAPE shows that as the distance between two positions in the time series increases, the dot product representing the similarity between the two positions has a more stable monotonous downward trend, and the similarity between the tAPE embedding vectors decreases. This is because tAPE can use the embedding space to provide isotropic encoding, maintain the distance perception characteristics, and better use the embedding vector space to distinguish between two positions.

[0117] Input embedding is the basis of all previous relative position encoding methods. The position matrix is ​​added or multiplied with the query, key and value matrices. In this paper, an efficient relative position encoding (eRPE) model independent of input embedding is introduced.

[0118] The calculation formula used in the eRPE model is shown in formula (14).

[0119]

[0120] In the formula, L is the sequence length, e i,j is the attention weight, w i-j is a learnable scalar, representing the relative position weight between position i and position j,

[0121] For each attention module in the multi-head attention mechanism, create a trainable parameter w of size 2L-1, because the maximum distance is 2L-1. For the index i and j of two positions, the corresponding relative scalar is w i-j+L , where the index starts at 1, not 0, and you need to index L from 2L-1 vectors 2 elements.

[0122] First, the relative position embedding w i-j is a static parameter that has nothing to do with the input, and the attention weight e i,j is dynamically determined by the representation of the input sequence. Attention adapts to the input sequence by inputting an adaptive weighting strategy, enabling the model to capture the complex relationships between different time points, which is exactly the most needed feature when extracting high-level concepts from time series, such as seasonal components in time series. However, when the data size is limited, using attention will face a greater risk of overfitting.

[0123] Secondly, the relative position embedding w i-j The relative displacement between positions i and j is considered, rather than their values. This is similar to the translation invariance of convolution, which has been shown to enhance generalization, so w i-jis considered as a scalar rather than a vector in order to achieve translation invariance without increasing the number of parameters. In addition, w for all (i, j) i-j All values ​​can be included in the paired dot product attention function, minimizing the additional computation. This efficient relative position encoding method is eRPE.

[0124] eRPE first applies the Softmax function to the attention matrix and then adds the relative position information to the model. Because the position values ​​that have not passed the Softmax function are clearer, the attention model performs better. Compared with existing models that apply Softmax to relative position embedding, eRPE's clearer position embedding is more conducive to time series classification tasks.

[0125] From the previous analysis, we can see that the complexity of global attention is the quadratic of the sequence length. If the attention proposed in equation (14) is directly applied to the original time series, the calculation speed will be too slow for long time series. Therefore, the ConvTran network model is introduced. First, the convolution module is used to shorten the length of the time series, and the feature map is reduced to a size with lower computational intensity before the new position encoding method is applied.

[0126] Using convolution modules on the Transformer architecture can not only reduce computational intensity and increase network training speed, but convolution operations are also very suitable for capturing local features. Therefore, convolution is used as the first module in the ConvTran model architecture to capture any discriminative local information in the original time series.

[0127] In the convolution module, M time convolution kernels are first applied to the input multivariate time series data, so that the ConvTran network model can extract the time information in the input sequence. Then the output of the time convolution kernel is combined with d model d x ×M-sized spatial convolution kernels are used to perform convolution operations to capture the correlation between variables in the original time series and construct d model This disjoint spatiotemporal convolution first expands the number of input channels and then compresses them. A key reason for this choice is that the feed-forward network (FFN) in the Transformer also expands the size of the input, and then projects the expanded hidden state back to the original size to capture spatial interactions.

[0128] Before the input embedding vector is fed into the Transformer module, the position embedding vector generated by tAPE is added to the input embedding vector so that the model can capture the temporal order of the time series. The size of the tAPE embedding vector is d model, which is the same as the input embedding vector. In multi - head attention, first, a linear layer is used to transform the input of dimension \(L\times d\) model to a size of \(L\times d\) z \(\times3\), that is, to obtain the \(q\) (query), \(k\) (key), and \(v\) (value) matrices of size \(L\times d\) z , where \(d\) z represents the dimension of the model and is a custom parameter. These \(q\), \(k\), and \(v\) matrices are reshaped into \(h\times L\times d\) z / \(h\) to represent the \(h\) - th attention head. Each attention head is responsible for capturing different patterns in the time series. For example, one attention head focuses on non - noise data, another attention head focuses on seasonal components, and another attention head focuses on trends. After obtaining the \(q\), \(k\), and \(v\) matrices, finally, the attention calculation is performed within the multi - head attention module using Equation (4 - 14).

[0129] The feed - forward network in the Transformer model is a multi - layer perceptron module, which consists of two linear layers and the Gaussian Error Linear Units (GELUs) activation function. GELUs introduce the idea of stochastic regularization in the activation function. It is a probabilistic description of the neuron input and is a combination of dropout, zoneout, and ReLU, which can improve the generalization ability of the model. Suppose the input is \(X\) and the mask is \(m\), then \(m\) follows a Bernoulli distribution \(F(x)=P(X < x)\), where \(X\) follows a standard normal distribution, that is, \(X\sim N(0,1)\). The mathematical expression of GELUs is shown in Equation (15).

[0130] \(GELU(x)=x\varPhi(x)=xP(X < x)\quad(15)\)

[0131] Similar to the Transformer basic architecture, residual connections and layer normalization are also applied in the multi - head attention layer and the feed - forward network layer in the ConvTran network to obtain the final output of the Transformer module. Then, max - pooling and global average pooling (GAP) are applied to the output of the ELU activation function of the last layer to obtain a more translation - invariant model. Finally, the Softmax function is applied to obtain the classification prediction result.

[0132] Figure 7 An experiment is shown on the grid - connected photovoltaic array fault simulation test platform. The multivariate time - series samples of the output voltage and current of each branch of the off - grid photovoltaic array during normal operation are shown in Table 1, the training set samples when an open - circuit fault occurs in Circuit A are shown in Table 2, the training set samples when a short - circuit fault occurs in Circuit B are shown in Table 3, and the training set samples when shading occurs in Circuit B are shown in Table 4.

[0133] Table 1 Sample examples of multivariate time series training set of off-grid photovoltaic array output during normal operation

[0134]

[0135]

[0136] Table 2 Sample examples of multivariate time series training set of off-grid photovoltaic array output during open circuit fault

[0137]

[0138] Table 3 Sample examples of multivariate time series training set of off-grid photovoltaic array output during short-circuit fault

[0139]

[0140] Table 4 Sample examples of multivariate time series training set of off-grid photovoltaic array output during shadow shading failure

[0141]

[0142] In this embodiment, a total of 1371 sets of multivariate time series sample data were collected in the off-grid photovoltaic array fault simulation test, and the 1371 sets of data were randomly divided into 1034 training sets and 337 test sets. The label of the grid-connected photovoltaic array operation status corresponding to each sample is marked. In order to facilitate the ConvTran network to classify different working conditions of the grid-connected photovoltaic array, the correspondence between the label of the sample set and the operation status of the photovoltaic array is shown in Table 5.

[0143] Table 5 Correspondence between sample set labels and actual operating status of photovoltaic arrays

[0144]

[0145] The training set is input into the ConvTran photovoltaic array fault diagnosis model for training. The training process of the training set is referenced Figure 5 The overall framework of the ConvTran network model is shown in Figure 1. Since the input to the ConvTran network model is four-channel multivariate time series data, namely the branch voltages and branch currents of the two branches A and B of the grid-connected photovoltaic array, Figure 5 where dx is equal to 4 and L is the length of the multivariate time series.

[0146] The parameters of the convolution module in the ConvTran network model are as follows: the number of temporal and spatial convolution kernels is set to 64, and the length of the temporal convolution kernel is set to 8, and the width of the spatial convolution kernel is equal to the input dimension. The parameters of the Transformer module in the ConvTran network model are as follows: 8 attention modules are used in the multi-head attention mechanism to capture the attention changes in the input sequence. Figure 5 The dimensions dmodel and dz of the Transformer encoding are both set to 64. The Transformer feedforward network expands the input size by 4 times and then projects the 4 times wider hidden state back to the original size. In addition, Adam optimization and early stopping based on validation loss are also used.

[0147] After the multivariate time series data of the training set is input into the ConvTran network model, the time convolution kernel in the convolution module is first input to extract the time information in the input sequence, and then the output of the time convolution kernel is convolved with the spatial convolution kernel to capture the correlation between the variables in the original time series and construct d model The input embedding of size L×d is then superimposed on the position embedding vector generated by tAPE and input into the Transformer module, so that the ConvTran network can capture the temporal order of the time series. In the multi-head attention layer, a linear layer is first used to transform L×d model The input is converted into 3 L×d z The 8 attention heads perform attention calculations respectively. The feedforward network in the Transformer module is a multi-layer perceptron module, which consists of two linear layers and a GELU activation function, which improves the generalization ability of the network model. Residual connections and layer normalization are applied to each multi-head attention layer and feedforward network layer to obtain the final output of the Transformer module. Then, maximum pooling and global average pooling (GAP) are applied to the output of the ELU activation function of the last layer to obtain a model with more translation invariance. Finally, the Softmax function is applied to obtain the classification prediction result, which is compared with the label input of the training set, and the parameters of the network model are updated to train the model.

[0148] The test set samples of the multivariate time series of the output voltage and current of each branch of the off-grid photovoltaic array during normal operation are shown in Table 6, the test set samples when an open circuit fault occurs in line A are shown in Table 7, the test set samples when a short circuit fault occurs in line B are shown in Table 8, and the test set samples when shadow occlusion occurs in line B are shown in Table 9.

[0149] Table 6 Sample examples of multivariate time series test set of off-grid PV array output during normal operation

[0150]

[0151] Table 7 Sample examples of multivariate time series test set of off-grid PV array output during open circuit fault

[0152]

[0153] Table 8 Sample examples of multivariate time series test set of off-grid photovoltaic array output during short circuit fault

[0154]

[0155]

[0156] Table 9 Sample examples of multivariate time series test set of off-grid PV array output during shadow shading failure

[0157]

[0158] The test set is input into the trained ConvTran model suitable for photovoltaic array fault diagnosis, and finally the off-grid photovoltaic array fault classification result is output.

[0159] In order to compare and analyze the effectiveness and accuracy of the ConvTran model used in this paper, the same 1034 training sets and 337 test sets were used as data samples, and the improved CNN algorithms FCN, MC-DCNN, ResNet and MLSTM-FCN were used as comparison models of the ConvTran network. In addition, the Gated Transformer Networks (GTN) model based only on the Transformer architecture was supplemented as a comparison model of the ConvTran network.

[0160] The 337 test set samples consist of 189 normal operation samples, 50 open circuit fault samples, 49 short circuit fault samples and 49 shadow occlusion fault samples. The confusion matrix of photovoltaic array fault identification based on ConvTran is as follows: Figure 8 As shown, only 1 sample was misclassified among the 337 test set samples, and the overall classification accuracy was as high as 99.7%.

[0161] Depend on Figure 8 From the confusion matrix, it can be seen that the PV array fault diagnosis model based on ConvTran does not classify the abnormal state of the PV array as a normal operating state. The accuracy of abnormal diagnosis is 100%, which fully plays the role of fault warning. The recognition accuracy of open circuit fault, short circuit fault and shadow occlusion fault is also 100%, achieving accurate identification of faults.

[0162] The confusion matrix of the four CNN-based comparison model algorithms FCN, MC-DCNN, ResNet, MLSTM-FCN and the Transformer-based comparison model algorithm GTN mentioned in the previous section is as follows: Fig. 9 shown.

[0163] Depend on Fig. 9The confusion matrix of each algorithm shows that other improved algorithms based on CNN have misdiagnosed the abnormal state of the photovoltaic array and cannot fully achieve the warning effect. Among them, the FCN algorithm and the ResNet algorithm cannot distinguish between the normal operation state and the shadow occlusion fault state, the MC-DCNN algorithm and the MLSTM-FCN algorithm are easy to confuse the short circuit and the shadow occlusion state, and the GTN algorithm has a certain degree of confusion for the normal operation state, short circuit state and shadow occlusion state, which is not conducive to judging the severity of the fault and affects the subsequent response measures.

[0164] The accuracy of the ConvTran network model and the above methods in off-grid photovoltaic array fault diagnosis is compared, and the results are shown in Table 10. Fig.10 and Fig.11 shown.

[0165] Table 10 Off-grid PV array fault diagnosis accuracy under different algorithms

[0166]

[0167]

[0168] It can be seen from Table 10 that compared with other improved algorithms based on CNN such as FCN, MC-DCNN, ResNet, MLSTM-FCN and Transformer-based improved algorithm GTN, the ConvTran algorithm has a more balanced ability to identify various faults of photovoltaic arrays, and has the highest fault diagnosis accuracy for off-grid photovoltaic arrays, especially in the case of partial shadow obstruction, the diagnosis accuracy is significantly higher than other algorithms.

[0169] This shows that as an improved algorithm that combines convolution and Transformer, the ConvTran network model improves the position and data embedding of time series data in the Transformer architecture, has better classification effect in multivariate time series classification tasks, can distinguish the slight difference between short circuit faults and partial shadowing, is more suitable for multivariate time series fault type classification, and shows better applicability and accuracy in off-grid photovoltaic array fault diagnosis.

[0170] The above descriptions are merely embodiments of the present invention and do not limit the patent scope of the present invention; any equivalent methods or structures made using the contents of the present invention are included in the patent protection scope of the present invention.

Claims

1. A ConvTran-based off-grid photovoltaic array fault diagnosis method, characterized in that: The steps include: S1: Build the Transformer model, which includes four parts: input module, encoder module, decoder module and output module; S2: Use convolutional modules on the Transformer architecture to shorten the length of the time series and capture local information in the original time series; S3: Before the input embedding vector is fed into the Transformer module, the position embedding vector generated by tAPE is added to the input embedding vector so that the model can capture the temporal order of the time series, where tAPE is the time series absolute position encoding method; S4: After obtaining the final output of the Transformer module, global average pooling and fully connected layers are applied to obtain a model with more translation invariance, and finally the Softmax function is applied to obtain the classification prediction result; S5: Collect the multivariate time series data of the four channels of the photovoltaic array, input it into the ConvTran network model, and perform fault diagnosis on the photovoltaic array.

2. The off-grid photovoltaic array fault diagnosis method based on ConvTran as claimed in claim 1, characterized in that: The input of the input module in step S1 is sequence data. The input sequence is converted into a feature vector through an embedding algorithm. After the embedding operation is completed on the input sequence, the feature vector is also positionally encoded, and the position encoding has the same dimension as the input embedding.

3. The off-grid photovoltaic array fault diagnosis method based on ConvTran as claimed in claim 2, characterized in that: The encoder module in step S1 is composed of multiple encoder stacks with the same structure. Each encoder consists of two sublayers: a self-attention layer and a fully connected feedforward network. Each sublayer adopts residual connection and then performs layer normalization.

4. The off-grid photovoltaic array fault diagnosis method based on ConvTran as claimed in claim 3, characterized in that: The decoder module in step S1 consists of multiple decoder stacks with the same structure. In addition to the self-attention layer and the feedforward network, the decoder has a third sublayer that performs a multi-head attention mechanism on the output obtained from the encoder stack.

5. The off-grid photovoltaic array fault diagnosis method based on ConvTran as claimed in claim 4, characterized in that: The output module in step S1 consists of a linear layer and a Softmax layer. The linear layer is a simple fully connected neural network that maps the output vector of the decoding module to a longer vector, namely the logits vector. The Softmax layer converts the attention score of each sequence segment into a probability distribution between 0 and 1, and selects the result corresponding to the highest probability as the output of this time step.

6. The off-grid photovoltaic array fault diagnosis method based on ConvTran as claimed in claim 4, characterized in that: The multi-head attention mechanism uses h different learned linear transformations to linearly map queries, keys, and values ​​respectively, while maintaining the overall parameter scale unchanged, and then performs scaled dot product attention calculations simultaneously, and finally concatenates the generated output values ​​for linear transformation to produce the final result.

7. The off-grid photovoltaic array fault diagnosis method based on ConvTran as claimed in claim 1, characterized in that: The multivariate time series data of the four channels of the photovoltaic array are collected in step S5, specifically, one time point is collected every 1s, and a group of samples is packaged every 30s. A total of 1371 groups of multivariate time series sample data are collected, and the 1371 groups of data are randomly divided into 1034 training sets and 337 test sets.

8. The off-grid photovoltaic array fault diagnosis method based on ConvTran as claimed in claim 7, characterized in that: The multivariate time series data of the four channels of the photovoltaic array are the voltage and current of the two branches A and B on the DC output side.

9. The off-grid photovoltaic array fault diagnosis method based on ConvTran as claimed in claim 8, characterized in that: The parameters of the convolution module in the ConvTran network model are as follows: the number of temporal and spatial convolution kernels is set to 64, and the length of the temporal convolution kernel is set to 8, and the width of the spatial convolution kernel is equal to the dimension of the input; the parameters of the Transformer module in the ConvTran network model are as follows: 8 attention modules are used in the multi-head attention mechanism to capture the attention changes in the input sequence, the dimensions dmodel and dz of the Transformer encoding are both set to 64, and the Transformer feedforward network expands the input size by 4 times, and then projects the 4-times-wide hidden state back to the original size.

10. The off-grid photovoltaic array fault diagnosis method based on ConvTran as claimed in claim 9, characterized in that: The steps of training the model include: after the multivariate time series data of the training set is input into the ConvTran network model, the time convolution kernel in the convolution module is first input to extract the time information in the input sequence, and then the output of the time convolution kernel is convolved with the spatial convolution kernel to capture the correlation between the variables in the original time series, and the input embedding of the size of dmodel is constructed. Then, the position embedding vector generated by tAPE is superimposed on the input embedding vector and input into the Transformer module, so that the ConvTran network captures the time order of the time series. In the multi-head attention layer, a linear layer is first used to convert the input of L×dmodel into 3 matrices of size L×dz. , where L is the length of the multivariate time series, and the eight attention heads perform attention calculations respectively. The feedforward network in the Transformer module is a multilayer perceptron module, which consists of two linear layers and a GELU activation function. Each multi-head attention layer and feedforward network layer applies residual connection and layer normalization to obtain the final output of the Transformer module. After that, the maximum pooling and global average pooling are applied to the output of the ELU activation function of the last layer to obtain a model with more translation invariance. Finally, the Softmax function is applied to obtain the classification prediction result, which is compared with the label input of the training set, and the parameters of the network model are updated to train the model.

Citation Information

Patent Citations

  • Fault diagnosis method based on additive adaptive LSTM-Transformer

    CN118690789A

  • Chemical process fault diagnosis method based on multi-scale spatial-temporal feature fusion

    CN119150065A

  • Multi-mode information fusion method and system for protein representative learning, and terminal and storage medium

    WO2023109714A1

  • Sparse code multiple access encoding and decoding system based on generative adversarial network

    WO2024016424A1